Imaging device, image processing device, and imaging method

The imaging device enhances peripheral vision image quality by using motion compensation with EVS pixel units for event data and grayscale pixel units, addressing the neglect of temporal characteristics in conventional methods and reducing processing load.

WO2026105443A1PCT designated stage Publication Date: 2026-05-21SONY SEMICON SOLUTIONS CORP
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SONY SEMICON SOLUTIONS CORP
Filing Date
2025-09-17
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Conventional imaging technologies fail to consider the temporal characteristics of peripheral vision, neglecting the ability of peripheral vision to recognize movement and changes in objects, leading to suboptimal image quality and increased processing load.

Method used

An imaging device and method that utilizes an EVS pixel unit to capture event data for peripheral vision motion and a grayscale pixel unit for central vision, employing motion compensation to synthesize a motion-compensated image, reducing the need for full peripheral vision imaging and enhancing temporal characteristics.

Benefits of technology

Improves the quality of peripheral vision images by ensuring accurate temporal information while reducing power consumption and processing load, suitable for applications in head-mounted displays and augmented reality devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025032604_21052026_PF_FP_ABST
    Figure JP2025032604_21052026_PF_FP_ABST
Patent Text Reader

Abstract

An imaging device according to the present invention includes a gradation pixel unit and an EVS pixel unit. The gradation pixel unit acquires a partial captured image selectively including central vision. The EVS pixel unit acquires event data indicating movements in peripheral vision.
Need to check novelty before this filing date? Find Prior Art

Description

Imaging device, image processing device, and imaging method

[0001] The present invention relates to an imaging device, an image processing device, and an imaging method.

[0002] A technique called foveal rendering is known, which concentrates the rendering source on the user's area of ​​focus (central vision). By rendering the area of ​​focus at high resolution and the rest of the image (peripheral vision) at low resolution, the amount of computation can be reduced. This technique is used in head-mounted displays and other devices that require real-time processing.

[0003] International Publication No. 2021 / 261248, Japanese Patent Publication No. 2010-028722, Japanese Patent Publication No. 2008-125059

[0004] Human vision differs between central and peripheral vision. Central vision refers to the narrow field of view centered on the point of fixation, and has high spatial characteristics such as resolution, shape, and color recognition. Peripheral vision does not have the same high spatial characteristics as central vision, but it is relatively good at recognizing the movement and changes of objects. Conventionally, it has been proposed to differentiate the image quality of central and peripheral vision based on the difference in spatial characteristics, but the peripheral vision's ability to recognize movement and change (temporal characteristics) has not been taken into consideration at all.

[0005] Therefore, this disclosure proposes an imaging device, an image processing device, and an imaging method that can improve the quality of peripheral vision.

[0006] The present disclosure provides an imaging device having a grayscale pixel unit that acquires a partial image that selectively includes central vision, and an EVS pixel unit that acquires event data indicating movement in peripheral vision. The present disclosure also provides an imaging method in which the information processing of the imaging device is performed by a computer.

[0007] The present disclosure provides an image processing device having an image synthesis unit that acquires a partial image that selectively includes central vision, acquires event data indicating peripheral vision motion, and synthesizes a motion-compensated grayscale image of the peripheral vision with the partial image. The present disclosure also provides an imaging method in which the information processing of the image processing device is performed by a computer.

[0008] This figure shows the characteristics of human vision. This figure outlines a method for improving peripheral vision image quality using motion compensation. This figure shows an example of applying the method of this disclosure. This figure shows another example of the configuration of a composite image. This figure shows an example of the timing of acquiring a partially captured image. This figure shows an example of the configuration of the image display system of this disclosure. This figure shows an example of a method for improving central vision image quality. This figure explains an example of determining a partially captured area based on image analysis of a grayscale image. This figure shows an example of extracting a moving object as the central subject from a grayscale image. This figure shows an example of extracting a pre-set type of object as the central subject from a grayscale image. This figure explains an example of determining a partially captured area based on eye tracking. This figure shows an example of extracting a gazed-on object as the central subject. This figure explains an example of determining a partially captured area based on event features. This figure shows an example of extracting an object in a location with many events as the central subject based on event information. This figure shows an example of extracting an object in a location with a lot of movement as the central subject based on event information. This figure shows an example of extracting a pre-set type of object as the central subject based on event information. This figure shows an example of a processing flow related to the overall processing. This figure shows a modified version of the processing flow related to the overall processing. This figure shows an example of a judgment process regarding the necessity of capturing a partially captured image. This figure explains upsampling of event data. This figure illustrates the correspondence between grayscale pixels and EVS pixels. This figure illustrates an example of a composite image generation method. This figure illustrates an example of calculating motion vectors based on event data and grayscale images. This figure illustrates an image display system according to a modified example. This figure illustrates an example of the hardware configuration of an imaging device and an image processing device. This figure illustrates an example of the appearance of the information processing system of this disclosure. This figure illustrates an example of the appearance of the information processing system of this disclosure. This figure illustrates an example of the hardware configuration of the information processing system.

[0009] Embodiments of the present disclosure will be described in detail below with reference to the drawings. In each of the following embodiments, the same parts will be denoted by the same reference numerals, and redundant descriptions will be omitted.

[0010] The explanation will proceed in the following order: [1. Outline of the Invention] [1-1. Human Visual Characteristics] [1-2. Peripheral Vision Image Quality Improvement Method Using Motion Compensation] [2. Configuration of the Image Display System Disclosed] [3. Central Vision Image Quality Improvement Method] [4. Method for Determining Partial Imaging Areas] [4-1. Determination of Partial Imaging Areas Based on Image Analysis of Grayscale Images] [4-2. Determination of Partial Imaging Areas Based on Eye Tracking] [4-3. Determination of Partial Imaging Areas Based on Event Features] [5. Processing Flow] [5-1. Processing Flow Related to Overall Processing] [5-2. Determination Process Regarding the Necessity of Imaging Partial Images] [6. Upsampling of Event Data] [7. Correspondence between Grayscale Pixels and EVS Pixels] [8. Method for Generating Composite Images] [9. Modified Examples] [10. Others] [11. Hardware Configuration Examples] [12. Effects] [13. XR Application Examples]

[0011] [1. Outline of the Invention] [1-1. Human Visual Characteristics] Figure 1 shows the characteristics of human visual characteristics.

[0012] The human field of vision has an angular range of more than 180° centered on a fixation point. Within this range, the area in which the shape and color of objects can be clearly recognized is called central vision. Central vision has an angular range of about 1° to 2° centered on the fixation point. The field of vision surrounding central vision is called peripheral vision. Peripheral vision has an angular range of about 100° for each of the right and left eyes. Central vision excels in the ability to grasp spatial information such as resolution, shape, and color (spatial characteristics), but is inferior in the ability to grasp temporal information such as movement and change (temporal characteristics). Peripheral vision is inferior in spatial characteristics but superior in temporal characteristics.

[0013] It should be noted that "superior" and "inferior" refer to relative differences in ability within the same type of field of vision (central vision / peripheral vision). For example, while central vision is said to have "inferior" temporal characteristics, this means that the temporal characteristics of central vision are inferior to those of central vision's spatial characteristics, not that the temporal characteristics of central vision are inferior to those of peripheral vision. Even with central vision, minute movements and changes in objects can be sensitively perceived.

[0014] Peripheral vision has inferior spatial characteristics compared to central vision. Therefore, conventional approaches have proposed increasing the image quality of central vision and decreasing the image quality of peripheral vision. This allows for improved user experience while reducing processing load. However, while peripheral vision lacks the high spatial characteristics of central vision, it is relatively easy to perceive the movement and changes of objects. Conventional approaches have completely ignored these temporal characteristics of peripheral vision.

[0015] This disclosure was made in view of the aforementioned issues. In this disclosure, in order to improve the temporal characteristics of peripheral vision, the subject in peripheral vision (peripheral subject) is corrected using a process called motion compensation, which takes into account motion between frames. Motion compensation is performed as a process that predicts the data of the current frame from past frames based on motion between frames. The peripheral subject after motion compensation is combined with the subject in central vision (central subject). This provides high-quality images in both central and peripheral vision. This will be explained in detail below.

[0016] [1-2. Peripheral vision image quality improvement method using motion compensation] Figure 2 is a diagram illustrating the overview of the peripheral vision image quality improvement method using motion compensation.

[0017] First, at a certain time t, the overall image TI and event information are acquired. The overall image TI refers to the grayscale image GI, which includes both central and peripheral vision. The grayscale image GI refers to an image that has grayscale information for each color, such as red (R), green (G), and blue (B). For example, the overall image TI is acquired as the largest possible grayscale image GI using all pixels. Hereafter, the subject SB included in the central vision will be referred to as the central subject SB. C It is stated that the subject SB included in peripheral vision is referred to as peripheral subject SB. S It should be written as follows.

[0018] Event EV refers to the brightness change detected by the EVS (Event-based Vision Sensor). The EVS is a camera that detects brightness changes for each pixel as Event EVs and outputs them at a high speed of approximately 10,000 to 20,000 fps. The EVS independently and asynchronously outputs information (event information) for each pixel, including the coordinates of the pixel where the Event EV occurred, the polarity of the brightness change, and the time when the brightness change occurred. The EVS outputs event information acquired in a time series as Event EV data.

[0019] Event EVs occur at least along the contours of the animal's body. By imaging the location (distribution) of event EVs, an image resembling the extracted contour of the animal can be obtained. The temporal change in the distribution of event EVs indicates the movement of the subject (SB). Motion vectors can be calculated from the event EV data. Motion vectors represent the movement of the image between frames. Motion compensation can be performed using motion vectors. The temporal change in the distribution of event EVs can be caused not only by the movement of the subject (SB) but also by the movement of the camera.

[0020] In this disclosure, specific situations in which foveal rendering should be performed are detected based on event EV, etc. For example, this applies to situations where only a specific subject SB is moving and attention is drawn to that subject SB. In the example in Figure 2, attention is expected to be drawn to the "pedestrian" on the left. In this case, it is conceivable to use foveal rendering to create a difference in image quality between the "pedestrian" and other stationary objects (for example, the "tree" or "stationary person" on the right).

[0021] The analysis target can be either event EVs or grayscale image GIs. If there are differences in the level of attention given to different types of subject SBs, it is also possible to detect specific types of subject SBs from the grayscale image GI and perform foveal rendering. In this case, the analysis target is the grayscale image GI. However, as mentioned above, it is also possible to extract the contour of an animal body based on the distribution of event EVs. Therefore, the type of subject SB may be identified based on that contour. In this case, both the type and movement of the subject SB can be determined based on the analysis results of the event EVs.

[0022] Under certain circumstances, some subjects SB (central subject SB C ) that can be the central view are selectively imaged. The tone image GI of the imaged part of the subject SB is obtained as a partial image PI. In the example of FIG. 2, the "person" at time (t + 1) and the image in its vicinity are obtained as the partial image PI. The partial image PI shows the real-time position, shape, color, etc. of the central subject SB C and includes accurate information about both spatial information and temporal information.

[0023] The tone image GI of the remaining subjects SB (peripheral subjects SB S ) that are peripheral views is obtained by performing motion compensation on the tone image GI obtained most recently. In the example of FIG. 2, the tone image GI of the peripheral subject SB S obtained at time t is corrected based on the motion vector from time t to time (t + 1). The tone image GI (corrected image CR) of the peripheral subject SB obtained by motion compensation S is synthesized with the partial image PI. The tone image GI (synthesized image CI) obtained by synthesis is presented to the user as a display image.

[0024] In the above configuration, the partial image PI corresponds to the central view, and the corrected image CR corresponds to the peripheral view. The corrected image CR is inferior in the accuracy of spatial information such as color and shape compared to the tone image GI obtained in real time. However, the accuracy of temporal information such as motion is ensured by motion compensation. Therefore, a display image that satisfies the temporal characteristics required for peripheral vision can be obtained.

[0025] In addition, in the above configuration, central vision is selectively imaged under specific conditions, and imaging of peripheral vision is omitted. Therefore, the power required for imaging and calculation of peripheral vision is reduced. For example, in conventional Foveated Rendering, imaging of the grayscale image GI of peripheral vision is also performed, and when displaying, the resolution of peripheral vision is reduced to reduce the amount of calculation and power. In contrast, in the method of the present disclosure, the grayscale image GI of peripheral vision is obtained by synthesis without being imaged. Of course, for peripheral vision, events EV are acquired, but the power consumption is smaller than that for imaging the grayscale image GI. Therefore, it is possible to reduce power consumption compared to Foveated Rendering.

[0026] FIG. 3 is a diagram showing an application example of the method of the present disclosure. The method of the present disclosure can be applied to displays such as HMD (Head Mounted Display) and AR (Augmented Reality) glasses. In these displays, composite images CI that are right-eye images and left-eye images are generated respectively for 3D display.

[0027] FIG. 4 is a diagram showing another configuration example of the composite image CI. In the example of FIG. 2, the central subject SB C is only a "person", and the "tree" is displayed as a peripheral subject SB S In the example of FIG. 4, not only the "person" but also the "tree" is displayed as the central subject SB C As in this example, the central subject SB C is not limited to one object. The central subject SB C means a subject SB that is assumed to be the central vision. The central subject SB C may be an object at the actual fixation position detected by eye tracking, or an object at a position with a high probability of being fixated.

[0028] Figure 5 shows an example of the timing for acquiring a partial image PI. In the example in Figure 2, the whole image TI was acquired at time t, and the partial image PI was acquired at the next time (t+1). However, the timing of acquiring the whole image TI and the partial image PI does not need to be consecutive. The acquisition of the partial image PI is performed when a specific situation requiring foveal rendering is detected. If a specific situation occurs continuously, such as when a "person" continues to move, the acquisition of the whole image TI will not be performed, and the acquisition of the partial image PI will continue.

[0029] Motion compensation for peripheral subjects S This is extracted from the most recently acquired overall image TI. In the example in Figure 5, the acquisition of the overall image TI is stopped from time (t+1) to time (t+n), and the partial image PI is continuously acquired. The corrected image CR at each time (t+i) (where i is an integer from 1 to n) during this period is extracted from the peripheral subject SB extracted from the overall image TI at time t. S It is generated by applying motion compensation to the grayscale image GI, corresponding to the amount of movement from time t to time (t+i).

[0030] [2. Configuration of the Image Display System of the Disclosure] The method of the Disclosure described above is implemented by an image display system 1 as shown in Figure 6. Figure 6 is a diagram showing an example of the configuration of the image display system 1 of the Disclosure.

[0031] The image display system 1 includes an imaging device 10 and an image processing device 20. The imaging device 10 is configured as an integrated unit comprising an image sensor that acquires a grayscale image GI and an EVS that acquires event information. The image processing device 20 generates a composite image CI based on the grayscale image GI and event information acquired from the imaging device 10.

[0032] The imaging device 10 includes an EVS pixel unit 11, an EVS pixel control unit 12, and an EVS drive control content calculation unit 13 as functional blocks of the EVS. The imaging device 10 also includes a grayscale pixel unit 14, a grayscale pixel control unit 15, and a grayscale drive control content calculation unit 16 as functional blocks of the image sensor. Hereinafter, for distinction, the pixels of the EVS pixel unit 11 will be referred to as "EVS pixels," and the pixels of the grayscale pixel unit 14 will be referred to as "grayscale pixels."

[0033] The EVS pixel unit 11 incorporates a light receiving sensor for each EVS pixel. The EVS pixel unit 11 acquires the brightness change for each EVS pixel as an event EV. Event EVs have two polarities: a positive polarity where the brightness changes positively, and a negative polarity where the brightness changes negatively. Event EVs often occur around the outlines of animal bodies. Based on the event information, the EVS pixel unit 11 can output an image that appears to have the outline of an animal body extracted.

[0034] By tracing the location where event EVs occur, the movement of subject SB can be detected. The movement of subject SB can be extracted as a motion vector. The motion vector is the surrounding subject SB. S This is used when performing motion compensation. The EVS pixel unit 11 acquires brightness change data for all EVS pixels. The acquired data includes event EV data related to central vision and event EV data related to peripheral vision. The EVS pixel unit 11 can acquire event EV data indicating peripheral vision movement from the pixel region corresponding to peripheral vision.

[0035] The EVS drive control content calculation unit 13 sets the detection timing, detection frequency (fps), and detection threshold for event EVs. The EVS pixel unit 11 determines that an event EV has been detected when a brightness change exceeding the threshold is detected. The EVS pixel control unit 12 controls the driving of the EVS pixel unit 11 based on the information set by the EVS drive control content calculation unit 13.

[0036] The grayscale pixel unit 14 acquires the grayscale image GI of the subject SB. The grayscale drive control content calculation unit 16 sets the imaging timing, imaging frequency (fps), and imaging range of the grayscale image GI. The grayscale pixel control unit 15 controls the driving of the grayscale pixel unit 14 based on the information set by the grayscale drive control content calculation unit 16.

[0037] In this disclosure, under specific circumstances where foveal rendering is required, the imaging range is limited to the central vision range. For example, the grayscale drive control content calculation unit 16 has a whole imaging mode and a partial imaging mode as imaging modes. The whole imaging mode is an imaging mode that acquires a whole image TI including central vision and peripheral vision. The partial imaging mode is an imaging mode that acquires a partial image PI. In response to the detection of a specific situation where foveal rendering is required, the grayscale drive control content calculation unit 16 switches the imaging mode from the whole imaging mode to the partial imaging mode.

[0038] In partial imaging mode, the gradation image GI of the peripheral view is obtained by applying motion compensation to past frames. In partial imaging mode, the gradation drive control content calculation unit 16 calculates the imaging target as a part of the subject SB (central subject SB) C The grayscale pixel unit 14 acquires a grayscale image GI that selectively includes the central view as a partial image PI. The partial image PI is combined with the grayscale image GI of the peripheral view after motion compensation (corrected image CR).

[0039] The image processing device 20 has an image synthesis unit 21. The image synthesis unit 21 acquires a partial image PI from the grayscale pixel unit 14 and acquires event EV data indicating peripheral vision movement from the EVS pixel unit 11. The partial image PI is a partial subject SB that is centrally visible (central subject SB) C The image synthesis unit 21 selectively includes a motion-compensated grayscale image GI (peripheral subject SB) of peripheral vision based on the event EV data. S The corrected image (CR) is combined with the partial image (PI). The image combining unit 21 outputs the resulting grayscale image (GI) (combined image CI) as a display image.

[0040] [3. Image Quality Improvement Techniques for Central Vision] Central vision has high recognition capabilities (spatial characteristics) such as resolution, shape, and color. Therefore, it is desirable to improve the image quality of the central vision area. Various techniques can be employed to improve image quality. Figure 7 shows an example of an image quality improvement technique for central vision.

[0041] In the example in Figure 7, the central subject SB C The current image and past image are combined. The combination is performed by aligning and blending the images. "Past" refers to the time when the overall image TI was acquired. "Present" refers to the time when the foveal rendering was performed. In the example in Figure 7, the time when the overall image TI was acquired is set to "t" and the time when the foveal rendering was performed is set to "t+1", but the time when the foveal rendering was performed is not limited to these.

[0042] For example, the image synthesis unit 21 compensates for motion in the most recently acquired overall image TI based on event EV data. The image synthesis unit 21 then synthesizes the partially captured image PI with the motion-compensated overall image TI. The synthesis process may be performed as a simple average of pixel values, or it may be performed using a DNN. In this method, the central subject SB C The image is acquired as an overlay image (overlay image OI) created by superimposing the current image and the past image after motion compensation. As a result, the consistency between the past and present images is improved, and noise is reduced.

[0043] The methods for improving central vision image quality are not limited to those described above. HDR (High Dynamic Range) can also be applied to a partial image PI using the overall image TI or event EV. For example, the dynamic range of the partial image PI can be increased by changing the exposure time between time t (the time the overall image TI is acquired) and time (t+1) (the time when foveal rendering is performed).

[0044] In addition to the above, there are also known methods that use events to enhance the dynamic range of central vision or to increase the resolution of central vision (see references [1] to [5] below). As a method for increasing the resolution of central vision, there is a known method in which the imaging resolution is lowered at one of the imaging times, the resolution is increased at the other, and super-resolution is performed by combining the images.

[0045] [Reference 1]: Nico Messikommer, Stamatios Georgoulis, Daniel Gehrig, Stepan Tulyakov, Julius Erbach, Alfredo Bochicchio, Yuanyou Li, and David Scaramuzza. “Multi-bracket high dynamic range imaging with event cameras.” In Proc. of Computer Vision and Pattern Recognition, 2022. 3.

[0046] [Reference 2]: Lin Wang, Tae-Kyun Kim, and Kuk-Jin Yoon. “Eventsr: From asynchronous events to image reconstruction, restoration, and super-resolution via end-to-end ” In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 8315- 8325, 2020.

[0047] [Reference 3]: Jin Han, Yixin Yang, Chu Zhou, Chao Xu, and Boxin Shi. “Evintsr-net: Event guided multiple latent frames reconstruction and super-resolution.” In Int. Conf. Compute. Vis. (ICCV), pages 4882-4891, October 2021. 3

[0048] [Reference 4]: B. Wang, J. He, L. Yu, G. -S. Xia, and W. Yang, “Event enhanced high quality image recovery,” European Conference on Computer Vision, 2020

[0049] [Reference 5]: KAI, Dachun, et al. “EvTexture: Event-driven Texture Enhancement for Video Super-Resolution.” arXiv preprint arXiv:2406.13457, 2024

[0050] In the example in Figure 7, the central subject SB C The image was obtained by combining the current image and the past image. However, the central subject SB C The image may be obtained from the current image alone. For example, the image synthesis unit 21 extracts a peripheral vision grayscale image GI from the most recently acquired overall image TI. The image synthesis unit 21 motion-compensates the extracted peripheral vision grayscale image GI based on event EV data. The image synthesis unit 21 then synthesizes the motion-compensated peripheral vision grayscale image GI with the partially captured image PI. Even in this case, a display image with improved peripheral vision temporal characteristics is provided.

[0051] [4. Method for Determining the Partial Imaging Region] The acquisition region (partial imaging region) of the partial imaging image PI is the central subject SB C It is determined as a region that selectively includes the central subject SB. CThis refers to the subject SB that is expected to be the central focus. Central subject SB C This could be a subject SB that is actually being observed, or a subject SB that is highly likely to be observed.

[0052] In the former case, the gaze area can be determined using known methods such as eye tracking. In the latter case, the likelihood of gaze can vary depending on the type of subject SB and the magnitude of its movement. Therefore, the system developer can pre-set which parts should be designated as the partial imaging area. Below, we will describe an example of a method for determining the partial imaging area.

[0053] [4-1. Determination of Partial Imaging Area Based on Image Analysis of Tonal Images] Figure 8 illustrates an example of determining a partial imaging area based on image analysis of a tonal image GI.

[0054] In this example, preliminary analysis identifies the characteristics of locations that are likely to attract the user's attention. Locations that match these characteristics are determined to be locations that the user is highly likely to focus on. The grayscale drive control content calculation unit 16 identifies objects that the user is highly likely to focus on as the subject SB (central subject SBC) of the partial image PI. In the example in Figure 8, locations that the user is likely to focus on are estimated based on the type or movement of the subject SB included in the grayscale image GI.

[0055] For example, the grayscale drive control content calculation unit 16 extracts moving objects or objects of a predetermined type from the overall image TI acquired from the grayscale pixel unit 14. The grayscale drive control content calculation unit 16 then uses the extracted objects as the subject SB (central subject SB) of the partial image PI. C ) to be identified as.

[0056] A "moving object" refers to an object with a movement that is highly likely to attract attention. The type and size of movement that is likely to attract attention is determined based on prior consideration. A specific type of object is the central subject (SB). C When extracting data, the types of objects that are likely to attract attention can be arbitrarily set by the system developer based on prior considerations.

[0057] Figure 9 shows the central subject SB of the grayscale image GI, where a moving object is centrally located. C This figure shows an example of extraction. In the example in Figure 9, the overall image TI contains "pedestrians," "trees," and "a person standing still." The grayscale drive control content calculation unit 16 selects "pedestrians" as the central subject SB. C The system extracts the data and determines that the "pedestrian" and its surrounding area are designated as the partial imaging region.

[0058] Figure 10 shows that pre-defined types of objects are selected from the grayscale image GI, with the central subject SB. C This figure shows an example of extraction. In the example in Figure 10, "person" and "tree" are pre-set as objects to be extracted. Therefore, the grayscale drive control content calculation unit 16 determines "pedestrian," "tree," and "stationary person" and their vicinity as partial imaging areas.

[0059] [4-2. Determination of Partial Imaging Area Based on Eye Tracking] Figure 11 illustrates an example of determining a partial imaging area based on eye tracking.

[0060] The image display system 1 in this example includes a gaze-tracking device 90. The gaze-tracking device 90 detects the user's gaze based on image analysis. Known eye-tracking technology can be used as the gaze detection method. For example, the gaze-tracking device 90 includes an eye-tracking camera 91 and a gaze detection unit 92.

[0061] The eye-tracking camera 91 captures an image of the user's face. The gaze detection unit 92 analyzes the face image to detect the user's gaze. The grayscale drive control content calculation unit 16 identifies the object in the line of sight as the object the user is fixated on. The grayscale drive control content calculation unit 16 determines the object the user is fixated on as the subject SB (center subject SB) of the partially captured image PI. C ) to be identified as.

[0062] Figure 12 shows the object of focus as the central subject SB. C This figure shows an example of extraction. In the example in Figure 12, the user's gaze is directed towards the "pedestrian". Therefore, the grayscale drive control content calculation unit 16 determines the "pedestrian" and the surrounding area as the partial imaging area.

[0063] [4-3. Determination of Partial Imaging Region Based on Event Features] Figure 13 illustrates an example of determining a partial imaging region based on the features of event EV.

[0064] Event EV features refer to characteristics that appear in the distribution of event EVs, or characteristics that appear in the temporal changes of the event EV distribution. For example, from the distribution of event EVs, features related to the number of event EV occurrences at each location (event quantity), the outline of the animal body, and the number of animals can be grasped. From the temporal changes in the distribution of event EVs, features related to the direction and magnitude of the animal body's movement (movement amount) can be grasped.

[0065] In this example, similar to the example in Figure 8, the area most likely to be focused on by the user is determined as the partial imaging area. In the example in Figure 8, the partial imaging area was determined by analyzing the grayscale image GI. In contrast, in this example, the target of analysis is the event EV. The partial imaging area is determined by analyzing the event EV.

[0066] For example, the grayscale drive control content calculation unit 16 compares the event amount or the amount of movement of the subject SB with a preset reference level. The grayscale drive control content calculation unit 16 determines the object included in the region where the event amount has reached the reference level, or the region where the amount of movement, determined based on the time change of the distribution of event EV, has reached the reference level, and determines the subject SB (central subject SB) of the partial image PI. C ) is identified as such. The reference level can be set arbitrarily by the system developer.

[0067] Figure 14 shows the central subject SB of objects in locations with a high number of event EVs based on event information. C This figure shows an example of extraction. The grayscale drive control content calculation unit 16 calculates the event amount for each location based on the distribution of event EV. The grayscale drive control content calculation unit 16 compares the event amount with a threshold that serves as a reference level and determines locations with an event amount equal to or greater than the threshold as partial imaging areas.

[0068] In the example shown in Figure 14, slight movements of the "tree" and the "stationary person" are detected based on event information. Therefore, the grayscale drive control content calculation unit 16 determines not only the "pedestrian" and its vicinity, but also the "tree" and the "stationary person" and their vicinity as partial imaging areas.

[0069] Figure 15 shows the central subject SB of objects in areas with a lot of movement based on event information. C This figure shows an example of extraction. The grayscale drive control content calculation unit 16 calculates the amount of movement for each location based on the distribution of event EVs. The grayscale drive control content calculation unit 16 compares the amount of movement with a threshold that serves as a reference level and determines locations with an amount of movement equal to or greater than the threshold as partial imaging areas. In the example in Figure 15, only the "pedestrian" shows an amount of movement equal to or greater than the threshold. Therefore, the grayscale drive control content calculation unit 16 determines the "pedestrian" and its vicinity as partial imaging areas.

[0070] Figure 16 shows the central subject SB, where objects of a predetermined type are selected based on event information. C This figure shows an example of extraction. The grayscale drive control content calculation unit 16 extracts objects of a preset type based on the distribution of event EVs. Object extraction can be performed, for example, by applying a pattern matching method to the contour of an object grasped from the distribution of event EVs. Objects may also be extracted using a machine learning model that detects objects from the distribution of events. The grayscale drive control content calculation unit 16 places the extracted object into the subject SB (central subject SB) of the partial image PI. C ) is identified as such. In the example in Figure 16, "person" is pre-set as an object that can be the object of attention.

[0071] [5. Processing Flow] [5-1. Processing Flow Related to Overall Processing] Figure 17 shows an example of the processing flow related to overall processing.

[0072] The grayscale drive control content calculation unit 16 determines imaging parameters for the overall image TI (step S1). The imaging parameters include imaging timing, imaging frequency (fps), and imaging range. The EVS pixel unit 11 captures the overall image TI based on the imaging parameters (step S2). The EVS pixel unit 11 outputs the overall image TI to a recording medium such as memory and a display device (step S3).

[0073] The EVS pixel unit 11 acquires event EVs from each EVS pixel as needed (step S4). The grayscale drive control content calculation unit 16 determines whether or not to capture a partial image PI based on the analysis results of the event EVs (step S5). The determination may be made based on the analysis results of the grayscale image GI, or on the analysis results of both the event EVs and the grayscale image GI. The partial image PI is captured when a specific situation is detected in which foveal rendering should be performed, based on the type and movement of the subject SB.

[0074] If partial image capture PI is not performed (Step S5: No), the imaging device 10 determines whether or not to terminate imaging (Step S12). If a termination operation by the user (such as pressing the capture termination button) is detected, imaging is terminated. If a termination operation is detected (Step S12: Yes), the imaging device 10 terminates imaging. If no termination operation is detected (Step S12: No), the process returns to Step S1.

[0075] When capturing a partial image PI (Step S5: Yes), the gradation drive control content calculation unit 16 determines the partial imaging area to be acquired (Step S6). The gradation drive control content calculation unit 16 also determines the imaging timing and imaging frequency (fps) as imaging parameters for the partial image PI (Step S7). The gradation pixel unit 14 captures the partial image PI based on the conditions determined by the gradation drive control content calculation unit 16 (Step S8).

[0076] The image synthesis unit 21 extracts image regions other than the partial image PI from the overall image TI. The image synthesis unit 21 performs motion compensation on the extracted image regions (step S9). The image synthesis unit 21 combines the corrected image CR obtained by motion compensation with the partial image PI to obtain a composite image CI (step S10). The image synthesis unit 21 outputs the composite image CI to a recording unit such as memory and a display device (step S11).

[0077] The imaging device 10 determines whether or not to terminate imaging (step S12). If a termination operation by the user is detected, imaging is terminated. If a termination operation is detected (step S12: Yes), the imaging device 10 terminates imaging. If no termination operation is detected (step S12: No), the process returns to step S1, and the above operations are repeated until a termination operation is detected.

[0078] Figure 18 shows a modified version of the processing flow related to the overall process. The following explanation will focus on the differences from the processing flow in Figure 17.

[0079] In this example, the difference from the example in Figure 17 is that the grayscale drive control content calculation unit 16 sets an upper limit on the number of partial image captures PI that can be captured for one whole image (step S0). In step S5, the grayscale drive control content calculation unit 16 determines whether or not it is necessary to capture partial image captures PI based on both whether or not a specific situation requiring foveal rendering has been detected, and whether or not the number of partial image captures PI has reached the upper limit.

[0080] According to the processing flow in this example, the corrected image CR will not be generated based on excessively old image information. If no upper limit is set on the number of partial image captures PI, the central subject SB C As long as it keeps moving, the same surrounding subject SB S Past images will continue to be used. If excessively old past images are used, the corrected image CR obtained by motion compensation on the past image and the actual current surrounding subject SB will be affected. S Inconsistencies may arise between the two. In this example, major inconsistencies are unlikely because the period for which past images are used is limited.

[0081] [5-2. Determination process regarding the necessity of capturing partial images] Figure 19 shows an example of the determination process regarding the necessity of capturing partial images PI. In this example, the necessity of capturing partial images PI is determined based on the event EV data.

[0082] The grayscale drive control content calculation unit 16 calculates the feature quantities of event EV (step S21). The calculated feature quantities include the event quantity in the entire area within the EVS field of view (total number of EVS pixels where event EV occurred), the amount of movement of all animal bodies, and the number of animal bodies. These feature quantities represent the total amount of movement of the entire subject SB.

[0083] The gradation drive control content calculation unit 16 compares the calculated feature quantity with a preset threshold (reference level) (step S22). For example, if the feature quantity is above the threshold, it is set to a high level, and if the feature quantity is below the threshold, it is set to a low level. If the feature quantity of event EV is at a high level (step S22: Yes), the gradation drive control content calculation unit 16 performs the overall imaging mode (step S23). If the feature quantity of event EV is at a low level (step S22: No), the gradation drive control content calculation unit 16 performs the partial imaging mode (step S24).

[0084] [6. Upsampling of Event Data] Figure 20 is a diagram illustrating the upsampling of event data.

[0085] The EVS pixel section 11 and the grayscale pixel section 14 do not necessarily have the same resolution. The EVS pixel section 11 may have a lower resolution than the grayscale pixel section 14. In this case, it is possible to upsample the event EV data to eliminate the resolution difference.

[0086] For example, the imaging device 10 has a resolution difference calculation unit 31 and a resolution upsampling unit 32. The resolution difference calculation unit 31 calculates the difference in resolution between the grayscale pixel unit 14 and the EVS pixel unit 11. The resolution upsampling unit 32 upsamples the event EV data based on the difference in resolution between the grayscale pixel unit 14 and the EVS pixel unit 11.

[0087] Upsampling can be performed using well-known methods such as bilinear or bicubic sampling. It is also possible to perform upsampling using an arbitrary neural network trained on the data before and after upsampling.

[0088] [7. Correspondence between grayscale pixels and EVS pixels] Figure 21 is a diagram illustrating the correspondence between grayscale pixels and EVS pixels.

[0089] The EVS pixel section 11 and the grayscale pixel section 14 are not necessarily arranged coaxially. If they are not in a coaxial configuration, it may not be possible to properly determine the partial imaging area from the location where the event EV occurs. To avoid this situation, it is conceivable to associate the grayscale pixels with the EVS pixels. For example, the imaging device 10 has feature point extraction sections 33, 34, a feature point matching section 35, and a pixel matching section 36.

[0090] The feature point extraction unit 33 extracts feature points from the grayscale image GI. The feature point extraction unit 34 extracts feature points from the event EV distribution. The feature point mapping unit 35 maps the feature points from the grayscale image GI to the feature points from the event EV distribution. The pixel mapping unit 36 ​​maps the pixels of the EVS pixel unit 11 and the grayscale pixel unit 14 based on the mapping between the feature points from the event EV and the feature points from the grayscale image GI acquired by the grayscale pixel unit 14.

[0091] For extracting feature points of grayscale image GI, known methods such as Harris, AKAZE (Accelerated KAZE), SIFT (Scale-invariant feature transform), and SURF (Speed-Upped Robust Feature) can be used. For extracting feature points of event EV distribution, the method described in reference 6 below can be used.

[0092] [Reference 6] Afshar, S. ; Ralph, N. ; Xu, Y. ; Tapson, J. ; Schaik, A. v. ; Cohen, G. “Event-based feature extraction using adaptive selection thresholds.” Sensors 2020, 20, 1600.

[0093] For mapping feature points, known methods such as K-Nearest Neighbor, brute force, and FLANN (Fast Library for Approximate Nearest Neighbors) can be used.

[0094] [8. Method for Generating Composite Images] Figure 22 illustrates an example of a method for generating composite images (CIs).

[0095] As mentioned above, the image synthesis unit 21 combines the partial captured image PI and the surrounding subject SB after motion compensation. S The image (corrected image CR) is synthesized. Motion compensation is performed based on the motion vector calculated by the motion vector calculation unit 37 of the image processing device 20. For example, the motion vector calculation unit 37 calculates the motion vector based on the time change of the distribution of event EV. The image synthesis unit 21 synthesizes the surrounding subject SB from the overall image TI. S Extract images and perform motion compensation on surrounding subjects based on motion vectors. S The image is combined with the partial image PI.

[0096] In the example shown in Figure 22, the motion vector calculation unit 37 calculates the motion vector based solely on the event EV data. However, the motion vector calculation unit 37 may also calculate the motion vector based on both the event EV data and the grayscale image GI.

[0097] Figure 23 shows an example of calculating motion vectors based on event data and grayscale image GI. In the example in Figure 23, motion vectors are estimated using a neural network. As estimation methods, methods such as Fusion-FlowNet (see reference 7 below) and DCEIFlow (see reference 8 below) can be used.

[0098] [Reference 7] C. Lee, A. K. Kosta, and K. Roy, “Fusion-flownet: Energy-efficient optical flow estimation using sensor fusion and deep fused spiking analog network architectures,” in Int. Conf. on Robotics and Automation (ICRA), 2022, pp. 6504-6510. 3, 4, 8

[0099] [Reference 8] Zhexiong Wan, Yuchao Dai, and Yuxin Mao. “Learning dense and continuous optical flow from an event camera.” IEEE Transactions on Image Processing, 31:7237-7251, 2022. 2, 3

[0100] [9. Modified Example] Figure 24 is a diagram showing an image display system 2 according to a modified example.

[0101] In this modified example, the difference from the image display system 1 in Figure 6 is that the gradation drive control content calculation unit 16 is included in the image processing device 50 instead of the imaging device 40. The imaging device 40 outputs the gradation image GI acquired by the gradation pixel unit 14 to the gradation drive control content calculation unit 16 of the image processing device 50. The gradation drive control content calculation unit 16 outputs control information regarding the imaging timing, imaging frequency (fps), and imaging range of the gradation image GI to the gradation pixel control unit 15 of the imaging device 40. With this configuration, the same processing as the image display system 1 described above is possible.

[0102] The configuration in which the functions of the grayscale drive control content calculation unit 16 are assigned to the image processing device 50 can also be applied to the configurations shown in Figures 11 and 13.

[0103] [10. Others] In the example in Figure 7, the central subject SB CThe current image and past image were combined for the overall image TI. However, the current image and past image may also be combined for the overall image TI. For example, the image combining unit 21 compensates for motion in the most recently acquired overall image TI (time t) based on event EV data. The image combining unit 21 then combines the motion-compensated overall image TI with the overall image TI acquired immediately afterward (time t+1). In this case, an overall image TI with less noise is obtained.

[0104] [11. Hardware Configuration Example] Figure 25 shows an example of the hardware configuration of the imaging devices 10, 40 and the image processing devices 20, 50.

[0105] Information processing for the imaging devices 10, 40 and image processing devices 20, 50 is performed, for example, by a computer 1000. The computer 1000 has a CPU (Central Processing Unit) 1100, RAM (Random Access Memory) 1200, ROM (Read Only Memory) 1300, HDD (Hard Disk Drive) 1400, a communication interface 1500, and an input / output interface 1600. The various parts of the computer 1000 are connected by a bus 1050.

[0106] The CPU 1100 operates based on programs (program data 1450) stored in the ROM 1300 or HDD 1400, and controls each part. For example, the CPU 1100 loads the programs stored in the ROM 1300 or HDD 1400 into the RAM 1200 and executes processing corresponding to various programs.

[0107] ROM 1300 stores boot programs such as the BIOS (Basic Input Output System) executed by the CPU 1100 when the computer 1000 starts up, as well as programs that depend on the computer 1000's hardware.

[0108] The HDD 1400 is a computer-readable non-temporary recording medium that non-temporarily records programs executed by the CPU 1100 and data used by such programs. Specifically, the HDD 1400 is a recording medium that records an information processing program according to the embodiment, which is an example of program data 1450.

[0109] The communication interface 1500 is an interface for the computer 1000 to connect to an external network 1550 (for example, the Internet). For example, the CPU 1100 can receive data from other devices or transmit data it has generated to other devices via the communication interface 1500.

[0110] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from input devices such as a keyboard or mouse via the input / output interface 1600. The CPU 1100 also transmits data to output devices such as a display device, speaker, or printer via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs recorded on a predetermined recording medium (media). Examples of media include optical recording media such as DVDs (Digital Versatile Discs) and PDs (Phase Change Rewritable Discs), magneto-optical recording media such as MOs (Magneto-Optical Discs), tape media, magnetic recording media, or semiconductor memory.

[0111] For example, when the computer 1000 functions as an imaging device 10, 40 and an image processing device 20, 50 according to the embodiment, the CPU 1100 of the computer 1000 realizes the functions of each of the aforementioned parts by executing an information processing program loaded on the RAM 1200. The HDD 1400 stores the information processing program, various models, and various data according to this disclosure. The CPU 1100 reads and executes the program data 1450 from the HDD 1400, but as another example, these programs may be obtained from other devices via an external network 1550.

[0112] [12. Effects] The imaging devices 10 and 40 have a grayscale pixel section 14 and an EVS pixel section 11. The grayscale pixel section 14 acquires a partial image PI that selectively includes the central view. The EVS pixel section 11 acquires event EV data that indicates movement in the peripheral view. In the imaging method of this disclosure, the processing of the imaging devices 10 and 40 is performed by a computer.

[0113] In this configuration, the requirements for the spatial characteristics of the central vision are met by the partially captured image PI. The requirements for the temporal characteristics of the peripheral vision are met by motion compensation using event EV data. As a result, a grayscale image GI that satisfies the required characteristics of both the central and peripheral vision is obtained. Furthermore, the grayscale pixel unit 14 selectively captures the central vision and omits capturing the peripheral vision. Therefore, the power required for capturing and processing the peripheral vision is reduced.

[0114] The imaging devices 10 and 40 each have a grayscale drive control content calculation unit 16. The grayscale drive control content calculation unit 16 switches the imaging mode from whole imaging mode to partial imaging mode in response to the detection of a specific situation in which foveal rendering should be performed. The whole imaging mode is an imaging mode in which a whole image TI including central vision and peripheral vision is acquired. The partial imaging mode is an imaging mode in which a partial image PI is acquired.

[0115] This configuration allows you to choose whether or not to perform foveal rendering depending on the situation.

[0116] The grayscale drive control content calculation unit 16 identifies an object that is highly likely to be focused on by the user as the subject SB of the partially captured image PI.

[0117] This configuration allows for comprehensive image quality enhancement for all objects that could potentially be the object of attention, regardless of whether they are actually being observed or not.

[0118] The gradation drive control content calculation unit 16 extracts moving objects or objects of a predetermined type from the overall image TI. The gradation drive control content calculation unit 16 identifies the extracted objects as subjects SB of the partial image PI.

[0119] With this configuration, a grayscale image (GI) of a moving object or an object of a predetermined type is acquired as the central view.

[0120] The grayscale drive control content calculation unit 16 identifies objects included in the region where the number of event EVs has reached a reference level, or the region where the magnitude of motion, determined based on the time change in the distribution of event EVs, has reached a reference level, as the subject SB of the partial image PI.

[0121] With this configuration, images of areas with a lot of movement are acquired as central vision.

[0122] The grayscale drive control content calculation unit 16 extracts objects of a preset type based on the distribution of event EVs. The grayscale drive control content calculation unit 16 identifies the extracted objects as subjects SB of the partial image PI.

[0123] With this configuration, images of pre-defined types of objects are acquired as the central view.

[0124] The grayscale drive control content calculation unit 16 identifies the object that the user is fixated on as the subject SB of the partially captured image PI.

[0125] In this configuration, the central view captures an image of the object the user is fixated on.

[0126] The grayscale drive control content calculation unit 16 performs the overall imaging mode when the feature quantity of event EV is at a high level. The grayscale drive control content calculation unit 16 performs the partial imaging mode when the feature quantity of event EV is at a low level.

[0127] With this configuration, a partial image PI of the event occurrence region is acquired only when the number of features in the event EV is small. The number of features in the event EV represents the degree of change in the entire image. When the entire image changes significantly (when the number of features is large), it becomes difficult to compensate for motion in a portion of the image and synthesize it. In such cases, acquiring the entire image TI results in better image quality.

[0128] The imaging devices 10 and 40 each have a resolution upsampling unit 32. The resolution upsampling unit 32 upsamples the event EV data based on the difference in resolution between the grayscale pixel unit 14 and the EVS pixel unit 11.

[0129] This configuration allows for obtaining high-resolution event information.

[0130] The imaging devices 10 and 40 each have a pixel mapping unit 36. The pixel mapping unit 36 ​​maps the pixels of the EVS pixel unit 11 and the grayscale pixel unit 14 based on the mapping between the feature points of the event EV distribution and the feature points of the grayscale image GI acquired by the grayscale pixel unit 14.

[0131] This configuration allows for a precise correspondence between grayscale pixels and EVS pixels.

[0132] The image processing devices 20 and 50 each have an image synthesis unit 21. The image synthesis unit 21 acquires a partial image PI that selectively includes the central view. The image synthesis unit 21 acquires event EV data indicating peripheral view motion. Based on the event EV data, the image synthesis unit 21 synthesizes a motion-compensated grayscale image GI of the peripheral view with the partial image PI. In the imaging method of this disclosure, the processing of the image processing devices 20 and 50 is performed by a computer.

[0133] In this configuration, the requirements for the spatial characteristics of central vision are met by the partially captured image PI. The requirements for the temporal characteristics of peripheral vision are met by motion compensation using event EV data. Therefore, a grayscale image GI is obtained that satisfies the required characteristics of both central and peripheral vision.

[0134] The image synthesis unit 21 extracts a peripheral vision grayscale image GI from the overall image TI, which includes the most recently acquired central vision and peripheral vision images. The image synthesis unit 21 compensates for motion in the extracted peripheral vision grayscale image GI based on event EV data. The image synthesis unit 21 then synthesizes the motion-compensated peripheral vision grayscale image GI with the partially captured image PI.

[0135] This configuration allows for the creation of a composite image (CI) that enhances the visual characteristics of both central and peripheral vision.

[0136] The image synthesis unit 21 compensates for motion in the overall image TI, which includes the most recently acquired central and peripheral vision images, based on the event EV data. The image synthesis unit 21 then synthesizes the partial image PI with the motion-compensated overall image TI.

[0137] This configuration yields a composite image (CI) that enhances the visual characteristics of both central and peripheral vision. For some subjects (SB) that are centrally visible, the current subject's tonal image (GI) and the motion-compensated tonal image (GI) of the most recent subject (SB) are combined. As a result, a tonal image (GI) with less noise is obtained for central vision.

[0138] Furthermore, the effects described herein are merely illustrative and not limiting, and other effects may also occur.

[0139] [13. XR Application Examples] Figure 26 is a diagram illustrating an example of the appearance of the information processing system 1001 of the present disclosure. As shown in Figure 26, the information processing system 1001 according to this embodiment is configured as a head-mounted display (HMD). Referring to Figure 26, an example of the appearance of the head-mounted display (HMD) of this embodiment will be described.

[0140] In this example, the HMD 1001 consists of an output mechanism 1011 and a mounting mechanism 1012. The mounting mechanism 1012 includes a mounting band 1013 that wraps around the head when worn by the user to secure the device. However, it does not have to wrap around the head as long as it is secured to the head.

[0141] The output mechanism 1011 includes a housing 1014 shaped to cover the left and right eyes when the HMD 1001 is worn by the user, and has a display panel inside that faces the eyes when worn. The housing 1014 may also be further equipped with a lens that is positioned between the display panel (display unit 2005 (Figure 28)) and the user's eyes when the HMD 1001 is worn, to widen the user's field of view. Stereo images corresponding to the parallax between the two eyes may be displayed in each of the regions formed by dividing the display panel into left and right sections, and stereoscopic vision may be realized by such a display.

[0142] The HMD 1001 may also be equipped with speakers or earphones positioned to correspond to the user's ears when worn. In this example, the HMD 1001 has a camera 1015 on the front of the housing 1014, and captures the surrounding real space as a video in a field of view corresponding to the user's line of sight.

[0143] Camera 1015 includes, for example, an image sensor such as a CCD (Charge Coupled Device) sensor or a CMOS (Complementary Metal Oxide Semiconductor) sensor, a light detection device such as a distance measuring sensor, and an optical system such as an imaging lens. For example, in Figure 26, it is configured as a stereo camera that images the space in front from left and right viewpoints corresponding to the user's left and right eyes. However, camera 1015 is not limited to this and may be a monocular camera or a multi-camera with three or more lenses. Furthermore, it may be a combination of multiple types of sensors. In hand tracking applications, camera 1015 may be configured to image the space below the information processing system. In eye tracking and face tracking applications, camera 1015 may be configured to image the user's eyes or face.

[0144] The HMD 1001 also includes a sensor 2008 (Figure 28). The sensor may include at least one of various sensors for determining the movement, attitude, position, etc., of the HMD 1001, such as an accelerometer, gyroscope, angular velocity sensor, and geomagnetic sensor.

[0145] The HMD1001 may be connected to other processing devices via wireless communication, or it may be connected via a wired connection such as USB (Universal Serial Bus).

[0146] In this case, the HMD 1001 may be configured to run online applications such as games that can be played by multiple users via a network. In this case, the HMD 1001 performs predetermined processing on the image captured by the camera 1015, generates a display image within the field of view of the camera 1015, and displays it.

[0147] The content of the displayed image here is not particularly limited and can vary depending on the functions the user requests from the system and the content of the application launched.

[0148] For example, the HMD 1001 may perform some processing on the image captured by the camera 1015, or superimpose virtual objects that interact with images of real objects. Alternatively, the HMD 1001 may render a virtual world in a field of view corresponding to the user's field of view, based on the captured image or measurements from motion sensors included in the sensor group of the HMD 1001.

[0149] Representative examples of these embodiments include virtual reality (VR), augmented reality (AR), and mixed reality (MR). Furthermore, by using the image captured by the camera 1015 as the display image, a see-through form (VST: VideoSeeThrough) in which the real world can be seen through the screen of the HMD 1001 may be realized.

[0150] Figure 27 is a diagram illustrating an example of the external appearance of the information processing system 1101 of the present disclosure. As shown in Figure 27, the information processing system 1101 according to this embodiment is configured as a glasses-type HMD.

[0151] The HMD body 1111 is worn on the user's head. The HMD body 1111 has a front section 1112, a right temple section 1113 provided on the right side of the front section 1112, a left temple section 1114 provided on the left side of the front section 1112, and a glasses section 1115 attached to the bottom of the front section 1112. In Figure 27, the glasses are a single unit, but there may be two separate glasses for each eye, or the glasses may be configured to cover only one eye.

[0152] The display unit 1103 is a see-through type display unit and is provided on the surface of the glass unit 1115. The display unit 1103 performs AR display of virtual objects in accordance with the control of the processing circuit 2001. The display unit 1103 may also be a non-see-through type display unit. In this case, AR display is performed by displaying an image on the display unit 1103 in which the virtual object is superimposed on the image currently being captured by the camera 1104.

[0153] Camera 1104 includes, for example, an image sensor such as a CCD (Charge Coupled Device) sensor or a CMOS (Complemented Metal Oxide Semiconductor) sensor, a light detection device such as a distance measuring sensor, and an optical system such as an imaging lens. Camera 1104 is provided facing outward on the outer surface of the front part 1112, and captures images of objects in real space and outputs the image information obtained by the capture to the processing circuit 2001. In Figure 27, for example, two cameras 1104 are provided on the front part 1112 with a predetermined distance between them in the lateral direction. Camera 1104 is not limited to this, and may be a monocular camera or a multi-lens camera with three or more lenses. Furthermore, it may be a combination of multiple types of sensors. In hand tracking applications, camera 1104 may be provided to capture images of the space below the information processing system. In eye tracking and face tracking applications, camera 1104 may be provided to capture images of the user's eyes or face.

[0154] The glasses-type HMD 1101 also includes a sensor 2008 (Figure 28). The sensor may include at least one of various sensors for determining the movement, attitude, position, etc., of the HMD 1001, such as an accelerometer, gyroscope, angular velocity sensor, and geomagnetic sensor.

[0155] Next, with reference to Figure 28, an example of the hardware configuration of the information processing system (HMD 1001 or glasses-type HMD 1101) will be described. As shown in Figure 28, the hardware of the information processing system consists of a processing circuit 2001, memory 2002, camera 2003, display unit 2005, input unit 2006, output unit 2007, sensor 2008, communication interface (IF) 2009, external network 2010, and secondary storage device 2011, which are connected to each other via a bus 2012, and can send and receive data and programs.

[0156] The processing circuit 2001 operates based on a program stored in memory 2002 or secondary storage device 2011 and controls the overall operation of the information processing systems 1001 and 1101. The processing circuit is, for example, a processor, and by reading and executing each program from memory 2002, it realizes the functions corresponding to each program read. The processor can include, for example, one or more of the following: a multi-core processor, a controller, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or equivalent discrete logic circuits or integrated logic circuits. The processing circuit may be implemented on multiple chips.

[0157] Memory 2002 can be implemented using semiconductor memory elements such as random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electronically erasable programmable read-only memory (EEPROM), or flash memory, as well as a hard disk, optical disk, etc., and may include any form of memory for storing data and executable software instructions.

[0158] Camera 2003 corresponds to camera 1015 in Figure 26 and camera 1104 in Figure 27, and includes an image sensor such as a CCD (Charge Coupled Device) sensor or a CMOS (Complemented Metal Oxide Semiconductor) sensor, a light detection device such as a distance measuring sensor, and an optical system such as an imaging lens.

[0159] The display unit 2005 is a display panel located inside the housing and consists of a display device such as an LCD (Liquid Crystal Display) or an organic EL (Electroluminescence).

[0160] The input unit 2006, although not shown in Figures 26 and 27, consists of input devices such as a keyboard, mouse, touch panel, microphone, and controller into which the user inputs operation commands, and supplies the various input signals to the processing circuit 2001.

[0161] The output unit 2007 consists of an audio output device such as a speaker, a force feedback device, an odor feedback device, etc., and is controlled by the processing circuit 2001, outputting the processing results as sound, force feedback, or odor.

[0162] Sensor 2008 may include at least one of various sensors for determining the movement, attitude, and position of HMDs 1001 and 1101, such as an acceleration sensor, gyroscope, angular velocity sensor, and geomagnetic sensor. It may also include a biosensor for sensing human biological information and a pressure sensor for sensing input.

[0163] The communication interface 2009 is an interface for the information processing systems 1001 and 1101 to connect to the external network 2010. It communicates with smartphones and other external devices (e.g., PCs (personal computers) and server devices on the network) via wired or wireless connections. For example, the processing circuit 2001 receives data from other devices and transmits data generated by the processing circuit 2001 to other devices via the communication interface 2009.

[0164] The above describes an example of an information processing system to which the technology described herein can be applied. The technology described herein can be applied to the following components of the configuration described above: EVS pixel unit 11, EVS pixel control unit 12, EVS drive control content calculation unit 13, grayscale pixel unit 14, grayscale pixel control unit 15, grayscale drive control content calculation unit 16, image synthesis unit 21, resolution difference calculation unit 31, resolution upsampling unit 32, feature point extraction units 33, 34, feature point mapping unit 35, pixel mapping unit 36, motion vector calculation unit 37, and gaze detection unit 92.

[0165] [Note] The technology can also be configured as follows: (1) An imaging device having a grayscale pixel unit that acquires a partial image that selectively includes the central view, and an EVS pixel unit that acquires event data indicating movement in the peripheral view. (2) The imaging device according to (1) above, having a grayscale drive control content calculation unit that, in response to the detection of a specific situation requiring foveal rendering, switches the imaging mode from a whole imaging mode that acquires a whole image including the central view and the peripheral view to a partial imaging mode that acquires the partial image. (3) The imaging device according to (2) above, wherein the grayscale drive control content calculation unit identifies an object that is highly likely to be gazed upon by the user as the subject of the partial image. (4) The imaging device according to (3) above, wherein the grayscale drive control content calculation unit extracts a moving object or an object of a preset type from the whole image, and identifies the extracted object as the subject of the partial image. (5) The imaging apparatus according to (3) above, wherein the gradation drive control content calculation unit identifies an object included in a region where the number of events has reached a reference level, or in a region where the magnitude of the movement, as determined based on the time change of the distribution of events, has reached a reference level, as the subject of the partial image. (6) The imaging apparatus according to (3) above, wherein the gradation drive control content calculation unit extracts objects of a preset type based on the distribution of events, and identifies the extracted objects as the subject of the partial image. (7) The imaging apparatus according to (2) above, wherein the gradation drive control content calculation unit identifies an object that the user is fixated on as the subject of the partial image. (8) The imaging apparatus according to any one of (2) to (7) above, wherein the gradation drive control content calculation unit performs the overall imaging mode when the feature amount of the events is at a high level, and performs the partial imaging mode when the feature amount of the events is at a low level. (9) The imaging apparatus according to any one of (1) to (8) above, further comprising a resolution upsampling unit that upsamples the event data based on the difference in resolution between the grayscale pixel unit and the EVS pixel unit.(10) An imaging apparatus according to any one of (1) to (9) above, further comprising a pixel mapping unit that maps the pixels of the EVS pixel unit and the grayscale pixel unit together based on the correspondence between the feature points of the distribution of the events and the feature points of the grayscale image acquired by the grayscale pixel unit. (11) An image processing apparatus comprising an image synthesis unit that acquires a partial image that selectively includes the central view, acquires event data indicating movement in the peripheral view, and synthesizes the motion-compensated grayscale image of the peripheral view with the partial image. (12) An image processing apparatus according to (11) above, wherein the image synthesis unit extracts the grayscale image of the peripheral view from a whole image that includes the central view and the peripheral view acquired most recently, compensates the extracted grayscale image of the peripheral view for motion based on the event data, and synthesizes the motion-compensated grayscale image of the peripheral view with the partial image. (13) The image processing apparatus according to (11), wherein the image synthesis unit compensates for motion in the overall image including the most recently acquired central vision and peripheral vision based on the event data, and synthesizes the partial image onto the overall image after motion compensation. (14) An imaging method performed by a computer, comprising acquiring a partial image that selectively includes central vision and acquiring event data indicating motion in the peripheral vision. (15) The imaging method according to (14), comprising synthesizing the gradation image of the peripheral vision that has been motion compensated based on the event data with the partial image.

[0166] 10, 40 Imaging device 11 EVS pixel section 14 Grayscale pixel section 16 Grayscale drive control content calculation section 20, 50 Image processing device 21 Image synthesis section 32 Resolution upsampling section 36 Pixel mapping section EV Event GI Grayscale image PI Partial image SB Subject TI Overall image

Claims

1. An imaging device having a grayscale pixel unit that acquires a partial image that selectively includes central vision, and an EVS pixel unit that acquires event data indicating movement in peripheral vision.

2. The imaging apparatus according to claim 1, further comprising a gradation drive control content calculation unit that, in response to the detection of a specific situation requiring foveal rendering, switches the imaging mode from a whole imaging mode that acquires a whole image including the central view and the peripheral view to a partial imaging mode that acquires the partial image.

3. The imaging apparatus according to claim 2, wherein the grayscale drive control content calculation unit identifies an object that is highly likely to be focused on by the user as the subject of the partial image.

4. The imaging apparatus according to claim 3, wherein the grayscale drive control content calculation unit extracts moving objects or objects of a predetermined type from the overall image and identifies the extracted objects as subjects of the partial image.

5. The imaging apparatus according to claim 3, wherein the grayscale drive control content calculation unit identifies an object included in a region where the number of events has reached a reference level, or in a region where the magnitude of the movement, as determined based on the time change in the distribution of the events, has reached a reference level, as the subject of the partial image.

6. The imaging apparatus according to claim 3, wherein the grayscale drive control content calculation unit extracts objects of a preset type based on the distribution of events and identifies the extracted objects as subjects of the partial image.

7. The imaging apparatus according to claim 2, wherein the grayscale drive control content calculation unit identifies the object that the user is fixated on as the subject of the partial image.

8. The imaging apparatus according to claim 2, wherein the grayscale drive control content calculation unit performs the overall imaging mode when the feature amount of the event is at a high level, and performs the partial imaging mode when the feature amount of the event is at a low level.

9. The imaging apparatus according to claim 1, further comprising a resolution upsampling unit that upsamples the event data based on the difference in resolution between the grayscale pixel unit and the EVS pixel unit.

10. The imaging apparatus according to claim 1, further comprising a pixel mapping unit that maps the pixels of the EVS pixel unit and the grayscale pixel unit together based on the correspondence between the feature points of the distribution of the events and the feature points of the grayscale image acquired by the grayscale pixel unit.

11. An image processing apparatus having an image synthesis unit that acquires a partial image that selectively includes central vision, acquires event data indicating peripheral vision motion, and synthesizes the motion-compensated grayscale image of the peripheral vision with the partial image.

12. The image processing apparatus according to claim 11, wherein the image synthesis unit extracts a grayscale image of the peripheral view from an overall image including the most recently acquired central view and peripheral view, compensates the extracted grayscale image of the peripheral view for motion based on the event data, and synthesizes the grayscale image of the peripheral view after motion compensation with the partially captured image.

13. The image processing apparatus according to claim 11, wherein the image synthesis unit compensates for motion in the overall image, including the most recently acquired central view and peripheral view, based on the event data, and synthesizes the partial image onto the overall image after motion compensation.

14. A computer-based imaging method comprising acquiring a partial image that selectively includes central vision and acquiring data on events indicating peripheral vision movement.

15. The imaging method according to claim 14, comprising combining the motion-compensated grayscale image of peripheral vision, based on the data of the event, with the partial image.