Image processing method and system

By employing foveated rendering based on eye tracking, the method addresses the challenge of increased data and processing demands with improved image quality, achieving efficient image processing and display.

JP2025078023APending Publication Date: 2025-05-19SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024188213
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-03
Filing Date
2024-10-25
Publication Date
2025-05-19

AI Technical Summary

Technical Problem

As image quality improves, the amount of data required to represent the image increases, leading to higher bandwidth requirements for transmission and increased processing time for generation and display.

Method used

The method involves using eye tracking to determine a user's focus point, allowing for foveated rendering where high-quality rendering is applied only to the focused area, reducing data requirements and processing time.

Benefits of technology

This approach reduces the computational cost and bandwidth requirements while maintaining subjective image quality, enhancing the efficiency of image processing and display, particularly in virtual reality applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025078023000001_ABST
    Figure 2025078023000001_ABST
Patent Text Reader

Abstract

To provide an image processing method and an image processing system, which improve image quality of a received image.SOLUTION: An image processing method includes the steps of: receiving an image; receiving sight line data indicating a sight line position of a user relative to the image; executing upscaling processing for at least a part of the received image in order to improve image quality of the at least a part of the received image; and outputting the image subjected to the upscaling processing to a display unit. The step of executing the upscaling processing includes the steps of: using a first kernel size to upscale a first region of the received image corresponding to an attention position of the user relative to the image; and using a second kernel size to upscale a second region of the received image. The first kernel size is larger than the second kernel size.SELECTED DRAWING: Figure 16
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing method and an image processing system.

Background Art

[0002] Providing high-quality image content has been a long-standing issue in the content display context and has been continuously improved. Some of these improvements have been achieved by improved display devices such as televisions with improved resolution that enables more detailed image display and HDR (High Dynamic Range) functions that enable display of a wider brightness range. Also, with the improvement in the processing capabilities available to content providers, for example, with the improvement in the processing capabilities of game consoles, it has become possible to generate more detailed virtual environments.

[0003] Improving image quality is considered particularly important in arrangements aimed at providing high-quality images to users in order to enhance the immersion of virtual reality and augmented reality experiences such as HMD (Head-Mounted Display).

[0004] However, generally, as the image quality improves, the amount of data required to represent the image also increases. For this reason, for example, the bandwidth requirements for transmitting such content increase significantly, which may lead to implementation problems. Similarly, the processing required to generate such content also increases, and the waiting time from when the image is generated until it is displayed may become longer.

[0005] Foveated rendering is an example of a technology proposed to address such problems. Foveated rendering technology uses information regarding the user's line-of-sight direction to determine which parts of the image should be rendered in high quality, allowing areas that the user is not focusing on to be rendered in low quality. This makes it possible to reduce the data size of the entire image without significantly affecting the subjective image quality experienced by the user.

[0006] Similarly, techniques have been proposed to enable the generation of an image by changing the resolution in various regions of the image. In some cases, these techniques can be used to provide a smooth resolution gradient across the entire image. Some of these techniques can provide a hardware-based implementation that varies the quality of the image and can be used in combination with other techniques such as forward rendering. SUMMARY OF THE INVENTION PROBLEMS TO BE SOLVED BY THE INVENTION

[0007] This disclosure arises in the context of the above discussion. MEANS FOR SOLVING THE PROBLEMS

[0008] Various aspects and features of the present invention are defined in the appended claims and the accompanying description, and include at least the following.

[0009] In a first aspect, an image processing method is provided according to claim 1.

[0010] In another aspect, an image processing system is provided according to claim 15. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] A more complete understanding of the present disclosure and many of the attendant advantages thereof will be readily obtained as the same becomes better understood by reference to the following detailed description when considered in connection with the accompanying drawings.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6a

Figure 6b

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13a

Figure 13b

Figure 14

Figure 15

Figure 16

DETAILED DESCRIPTION OF THE INVENTION

[0012] Methods and systems for image processing are disclosed. In the following description, many specific details are set forth in order to provide a thorough understanding of embodiments of the present invention. However, it will be apparent to those skilled in the art that these specific details need not be employed in practicing the present invention. Conversely, specific details known to those skilled in the art are omitted where appropriate for clarity.

[0013] First, a system is described in which eye tracking is used to determine a user's focus point on an HMD. This is an example of a system that can utilize embodiments of the present disclosure, but the embodiments need not be limited to HMDs and can be used in any other implementation of foveated rendering.

[0014] Referring to FIG. 1, user 10 wears an HMD 20 on the user's head 30 (an example of a general head-mountable device, and other examples include audio headphones or a head-mountable light source). The HMD includes, in this example, a frame 40 formed of a rear strap and a top strap, and a display unit 50. As described above, many eye tracking arrangements are considered particularly suitable for use in HMD systems, but such use in an HMD system is not essential.

[0015] The HMD of FIG. 1 may include additional features that are not shown in FIG. 1 for clarity of this initial explanation but will be described later in relation to other drawings.

[0016] The HMD of FIG. 1 completely (or at least substantially completely) blocks the user's field of view of the surrounding environment. All that the user can see, in many embodiments, is a pair of images displayed within the HMD that are supplied from an external processing device such as a gaming console. Of course, in some embodiments, the images may alternatively (or additionally) be generated by a processor or obtained from a memory located within the HMD itself.

[0017] The HMD has associated headphone audio transducers or earpieces 60 that fit over the user's left and right ears 70. The earpieces 60 reproduce an audio signal provided from an external source, which may be the same as the video signal source that provides the video signal for display to the user's eyes.

[0018] The fact that a user can only see what is displayed by the HMD and can only hear what is provided through the earpieces, subject to the noise cancellation or active cancellation characteristics of the earpieces and associated electronics, means that this HMD can be regarded as a so-called "fully immersive" HMD.

[0019] However, note that in some embodiments, the HMD is not a fully immersive HMD and can provide at least some functionality for the user to see and hear their surroundings. This can be achieved by providing some degree of transparency or partial transparency in the placement of the display, and / or projecting an external view (captured, for example, using a camera such as a camera attached to the HMD) through the HMD's display, and / or enabling the transmission of ambient sound through the earpieces, and / or providing a microphone to generate an input sound signal (for transmission to the earpieces) depending on the ambient sound.

[0020] The front camera 122 can capture an image on the front of the HMD during use. Such an image can be used for head tracking purposes in some embodiments and is also suitable for capturing images for an augmented reality (AR)-style experience. The Bluetooth (registered trademark) antenna 124 may be arranged simply as a directional antenna so that it can detect the direction of a nearby Bluetooth transmitter in some cases where it provides a communication function.

[0021] During operation, a video signal is provided for display by the HMD. This may be provided by an external video signal source 80 such as a video game console or a data processing device (such as a personal computer). In that case, the signal may be transmitted to the HMD by a wired or wireless connection. An example of a suitable wireless connection is a Bluetooth (registered trademark) connection. The audio signal of the earpiece 60 can also be transmitted by the same connection. Similarly, the control signal passed from the HMD to the video (audio) signal source can also be transmitted by the same connection. Further, a power source (including one or more batteries and / or connectable to a main power outlet) can be connected to the HMD by a cable. The power source and the video signal source 80 may be separate units or may be embodied as the same physical unit. The cables for power supply and video (and audio) signals may be separate or they may be combined and transmitted by a single cable (for example, using separate conductors like a USB cable or in a manner similar to a "Power over Ethernet" arrangement where data is transmitted as a balanced signal and power as a DC on the same physical wire assembly). The video signal and / or the audio signal can be transmitted, for example, by an optical fiber cable. In other embodiments, at least a part of the functions related to the generation of the image signal and / or the audio signal to be presented to the user may be executed by circuits and / or processing forming part of the HMD itself. The power source may be provided as part of the HMD itself.

[0022] Some embodiments of the present invention are applicable to an HMD having at least one electrical cable and / or optical cable connecting the HMD to another device such as a power source and / or a video (and / or audio) signal source. Thus, embodiments of the present invention can include, for example, the following. (a) An HMD having its own power source (as part of the HMD arrangement) but having a cable connection to a video and / or audio signal source. (b) An HMD having a cable connection to a power source and a video and / or audio signal source, embodied as a single physical cable or multiple physical cables. (c) An HMD having (as part of the HMD's arrangement) its own video signal source and / or audio signal source and a cable connection to a power source. Or, (d) An HMD having a wireless connection to a video and / or audio signal source and a wired connection to a power source.

[0023] When one or more cables are used, the physical location where the cable enters or is connected to the HMD is not particularly important from a technical perspective. Aesthetically, or to prevent the cable from covering the user's face during operation, typically the cable enters or attaches to the HMD at the side or back of the HMD (relative to the orientation of the user's head when worn in normal operation). Thus, the position of the cable with respect to the HMD in FIG. 1 should be treated as merely a schematic representation.

[0024] Thus, the arrangement of FIG. 1 provides an example of a head-mounted display system comprising a frame worn on an observer's head, a frame defining one or two eyeball display positions disposed in front of each of the observer's eyes during use, and display elements mounted for each of the eyeball display positions, the display elements providing a virtual image of a video display of a video signal from a video signal source to that eye of the observer.

[0025] FIG. 1 shows only one example of an HMD. For example, the HMD can use a frame similar to conventional glasses, i.e., having substantially horizontal legs extending from the display portion to the upper rear of the user's ears, and in some cases, curling down behind the ears. In other (non-fully immersive) examples, the user's field of view of the external environment may not actually be completely hidden. An example of such an arrangement will be described below with reference to FIG. 4.

[0026] In the example of FIG. 1, separate displays are provided for each eye of the user. A schematic plan view showing how this is achieved is provided as FIG. 2, which shows the position 100 of the user's eyes and the relative position 110 of the user's nose. The display unit 50 generally includes an external shield 120 that masks ambient light from the user's eyes and an internal shield 130 that prevents one eye from seeing the display intended for the other eye. The combination of the user's face, the external shield 120, and the internal shield 130 forms two compartments 140, one for each eye. Each compartment is provided with a display element 150 and one or more optical elements 160. A method by which the display element and the optical element cooperate to provide a display to the user will be described with reference to FIG. 3.

[0027] Referring to FIG. 3, the display element 150 generates a display image that is refracted by the optical element 160 (schematically shown as a convex lens in this example, but may include a compound lens or other elements) and is larger than the real image generated by the display element 150. This generates a virtual image 170 that appears to the user to be quite far away. As an example, the virtual image may have an apparent image size (image diagonal) of more than 1 m and may be located at a distance of more than 1 m from the user's eyes (or from the frame of the HMD). Generally speaking, depending on the purpose of the HMD, it is desirable to place the virtual image at a position quite far from the user. For example, when the HMD is for viewing movies or the like, it is desirable that the user's eyes be relaxed during viewing, and a distance of at least several meters (distance to the virtual image) is required. In FIG. 3, solid lines (such as line 180) represent real light rays, and dashed lines (such as line 190) represent virtual light rays.

[0028] Figure 4 shows another arrangement. This arrangement can be used when it is desirable that the user's view of the external environment is not completely blocked. However, it is also applicable to an HMD in which the user's external view is completely blocked. In the arrangement of Figure 4, the display element 150 and the optical element 200 cooperate to provide an image projected onto the mirror 210, and the mirror 210 deflects the image towards the position 220 of the user's eyes. The user perceives a virtual image at a position 230 in front of the user and at an appropriate distance from the user.

[0029] In the case of an HMD in which the user's view of the external environment is completely blocked, the mirror 210 can be a substantially 100% reflective mirror. The arrangement of Figure 4 has the advantage that the display element and the optical element can be arranged close to the center of gravity of the user's head and on the side of the user's eyes, making the HMD worn by the user less bulky. Alternatively, if the HMD is designed so as not to completely block the user's view of the external environment, by making the mirror 210 partially reflective, the user can also see the external environment through the mirror 210 with the virtual image superimposed on the actual external environment.

[0030] Also, when separate displays are provided for each of the user's two eyes, it is possible to display a stereoscopic image. An example of a pair of stereoscopic images displayed for the left and right eyes is shown in Figure 5. The images show a lateral displacement with respect to each other, and the displacement of the image features depends on the (real or simulated) lateral separation of the cameras from which the images were taken, the angular convergence of the cameras, and the (real or simulated) distance of each image feature from the camera positions.

[0031] There is a possibility that the drawn left-eye image is actually a right-eye image, or the drawn right-eye image is actually a left-eye image. This is because there are stereoscopic displays in which objects are moved to the right in the right-eye image and to the left in the left-eye image so that the user appears to be looking at the scenery on the other side through the stereoscopic window. However, there are also HMDs that adopt the arrangement shown in Figure 5. Which of these two arrangements to choose is left to the discretion of the system designer.

[0032] Depending on the situation, the HMD may be used simply to watch a movie or the like. In this case, even if the user rotates their head left and right, there is no apparent change in the viewpoint of the displayed image. However, in other applications related to virtual reality (VR) or augmented reality (AR) systems, the user's viewpoint needs to track movements related to the real or virtual space in which the user is located.

[0033] As described above, in some applications of the HMD related to virtual reality (VR) or augmented reality (AR) systems, the user's viewpoint needs to track (track) movements related to the real or virtual space in which the user is located.

[0034] This tracking is performed by detecting the movement of the HMD and changing the apparent viewpoint of the displayed image so that the apparent viewpoint follows the movement. Detection can be performed using any suitable arrangement (or combination of such arrangements). For example, it includes the use of hardware motion detectors (such as accelerometers and gyroscopes), external cameras operable to image the HMD, and outward-facing cameras attached to the HMD.

[0035] Turning to gaze tracking in such an arrangement, FIG. 6 schematically shows two possible arrangements for performing gaze tracking on the HMD. The cameras provided in such an arrangement can be freely selected so as to be able to execute an effective gaze tracking method. In some existing arrangements, visible light cameras are used to capture images of the user's eyes. Alternatively, infrared (IR) cameras are used to reduce interference with the captured signal and interference with the user's vision when a corresponding light source is provided, or to improve performance under low illumination conditions.

[0036] FIG. 6a shows an example of a gaze tracking arrangement in which a camera is disposed within the HMD to capture an image of the user's eyes from a short distance. This may be referred to as near-eye tracking, or head-mounted tracking.

[0037] In this example, the HMD 600 (having the display element 601) includes cameras 610 each arranged to directly capture one or more images of each of the user's eyes using an optical path that does not include the lens 620. This is advantageous in that it can avoid distortion of the captured image due to the optical effect of the lens. Here, as an example of a possible position where a gaze tracking camera may be provided, four cameras 610 are shown, but any number of cameras may be provided at any suitable position so as to effectively image the corresponding eye. For example, only one camera may be provided for one eye, or two or more cameras may be provided for each eye.

[0038] However, in many embodiments, it is considered advantageous to arrange the camera so as to include the lens 620 in the optical path used to image the eye of the user. An example of such a position is shown by the camera 630. As a result, due to the deformation of the captured image by the lens, processing may be required to enable appropriate and accurate tracking, but since the relative positions of the corresponding camera and lens are fixed, it can be performed relatively easily. The advantage of including the lens in the optical path is, for example, to simplify the physical constraints in the design of the HMD.

[0039] FIG. 6b shows an example of a gaze tracking arrangement in which the cameras are arranged to indirectly capture images of the user's eyes. Such an arrangement is particularly suitable for use with IR or other non-visible light sources, as will be apparent from the following description.

[0040] Figure 6b includes a mirror 650 disposed between the display 601 and the viewer's eye (which, of course, can be extended or replicated to the user's other eye as needed). For clarity, additional optical components (such as lenses) are omitted in this figure, but may be present at any suitable location within the depicted arrangement. The mirror 650 in such an arrangement is selected to be partially transmissive. That is, the mirror 650 should be selected such that the camera 640 can acquire an image of the user's eye while the user views the display 601. One way to achieve this is to provide a mirror 650 that is reflective for IR wavelengths but transmissive for visible light. This allows the IR light used for tracking to be reflected from the user's eye towards the camera 640, while the light emitted by the display 601 can pass through the mirror without being blocked.

[0041] Such an arrangement is advantageous in that, for example, it makes it easier to place the camera in a position where it is not visible to the user. Additionally, the fact that the camera captures an image from a position (by reflection) substantially along the axis between the user's eye and the display may result in improved accuracy of eye tracking.

[0042] Of course, the eye tracking arrangement need not be implemented in a head-mounted or other near-eye manner as described above. For example, FIG. 7 schematically shows a system in which a camera is arranged to capture an image of the user from a distance. This distance may vary during tracking and can take on any value depending on the parameters of the tracking system. For example, this distance can be 30 centimeters, 1 meter, 5 meters, 10 meters, or any value as long as the tracking is not performed using a device fixed to the user's head.

[0043] In FIG. 7, an array of cameras 700 is provided that provides multiple views of user 710. These cameras are configured to capture information that identifies at least the direction in which user 710's eyes are focused, using any suitable method. For example, an IR camera can be utilized to identify reflections from user 710's eyes. The array of cameras 700 may be provided to provide multiple views of user 710's eyes at any given time, or may simply be provided such that at least one camera 700 can view user 710's eyes at any given time. Depending on the use case, it may not be necessary to provide such a high level of coverage. Instead, it is clear that only one or two cameras 700 can be used to cover a smaller range of the possible viewing directions of user 710.

[0044] Of course, the technical difficulties associated with such long-distance tracking methods may increase, a higher resolution camera may be required, a more powerful light source for generating IR light may be required, and additional information (such as the orientation of the user's head) may need to be input to determine the focus of the user's line of sight. The details of the arrangement can be determined according to, for example, the required ruggedness, accuracy, size, and / or cost, or other design considerations.

[0045] Despite the technical challenges including those described above, such tracking methods are beneficial in that they enable a wider range of interactions for the user. It is possible to perform eye tracking not only for viewers of an HMD, but also for viewers of, for example, a television.

[0046] The eye-tracking arrangement can vary not only in where the cameras are provided, but also in where the processing of the image data captured to determine the tracking data is performed.

[0047] FIG. 8 schematically shows an environment in which gaze tracking processing can be executed. In this example, user 800 is using HMD 810 associated with a processing unit 830 such as a game machine, and peripheral device 820 allows user 800 to input commands for controlling the processing. HMD 810 may perform gaze tracking along the arrangement exemplified by FIG. 6a or FIG. 6b. That is, HMD 810 may include one or more cameras operable to capture images of one or both of user 800's eyes. Processing unit 830 may be operable to generate content for display on HMD 810, although some (or all) of the content generation may be performed by a processing unit within HMD 810.

[0048] The configuration of FIG. 8 also includes a camera 840 disposed outside of HMD 810 and a display 850. In some cases, camera 840 may be used to perform tracking of user 800 during use of HMD 810. For example, it may be used to identify body movements or head orientations. Camera 840 and display 850 may be provided in the same manner as or instead of HMD 810. For example, they may be used to capture an image of a second user and display the image to that user while a first user 800 is using HMD 810, or a first user 800 may be tracked and content may be displayed on these elements instead of HMD 810. That is, display 850 may be operable to display the generated content provided by processing unit 830, and camera 840 may be operable to capture images of the eyes of one or more users to enable eye tracking to be performed.

[0049] Although the connections shown in FIG. 8 are indicated by lines, this should not be taken to mean that the connections should necessarily be wired. Any suitable connection method, including a wireless network or a wireless connection such as Bluetooth®, is considered appropriate. Similarly, although FIG. 8 shows a dedicated processing unit 830, in some embodiments, it is also conceivable to execute the processing in a distributed manner, such as using a combination of two or more of the HMD 810, one or more processing units, a remote server (cloud processing), or a game console.

[0050] The processing necessary to generate tracking information from the captured image of the user's eyeball 800 may be executed locally by the HMD 810, or the captured image or one or more detection results may be transmitted to an external device (such as the processing unit 830) for processing. In the former case, if such processing is not executed exclusively in the HMD 810, the HMD 810 may output the result of the processing to an external device for use in the image generation processing. In embodiments where the HMD 810 does not exist, the captured image from the camera 840 is output to the processing unit 830 for processing.

[0051] FIG. 9 schematically shows a system for performing one or more gaze tracking processes in an embodiment described with reference to FIG. 8, for example. The system 900 includes a processing device 910, one or more peripheral devices 920, an HMD 930, a camera 940, and a display 950. Of course, in many embodiments, not all elements need to be present within the system 900. For example, if the HMD 930 is present, the camera 940 may be omitted because it is less likely to capture an image of the user's eyes.

[0052] As shown in FIG. 9, the processing device 910 may include one or more of a central processing device (CPU) 911, a graphics processing device (GPU) 912, a storage (such as a hard drive, or any other suitable data storage medium) 913, and an input / output 914. These units may be provided in the form of a personal computer, a game console, or any other suitable processing device.

[0053] For example, the CPU 911 may be configured to generate tracking data from one or more input images of the user's one or more eyes from one or more cameras, or from data indicating the direction of the user's eyes. This may be data obtained by processing an image of the user's eyes, for example, on a remote device. Of course, if the tracking data is generated elsewhere, there is no need for the processing device 910 to perform such processing.

[0054] The GPU 912 may be configured to generate content for display to the user for whom eye tracking is being performed. In some embodiments, the content itself may be changed depending on the obtained tracking data. Examples of this include the generation of content according to the foveated rendering technique. Of course, such a content generation process may be performed elsewhere. For example, the HMD 930 may have an on-board GPU operable to generate content depending on the eye tracking data.

[0055] The storage 913 may be provided to store any suitable information. Examples of such information include program data, content generation data, and eye tracking model data. In some cases, such information may be stored remotely, such as on a server. In such cases, local storage 913 may not be required. Therefore, the discussion of storage 913 should be considered to refer to local (and optionally removable storage media) or remote storage.

[0056] Input / output 914 may be configured to perform any suitable communication appropriate for processing device 910. Examples of such communication include transmitting content to HMD 930 and / or display 950, receiving gaze tracking data and / or images from HMD 930 and / or camera 940, and communicating with one or more remote servers (e.g., via the Internet).

[0057] As described above, peripheral device 920 may be provided so that a user can provide input to processing device 910 to control processing or interact with the generated content in other ways. This may be in the form of pressing a button, or alternatively, in the form of tracked movement to enable gestures to be used as input.

[0058] HMD 930 may include a number of sub-elements omitted from FIG. 9 for clarity. Of course, HMD 930 should include a display unit operable to display images to the user. In addition to this, HMD 930 may include any number of suitable cameras for gaze tracking (as described above), in addition to one or more processing units operable to generate content for display and / or generate gaze tracking data from captured images.

[0059] Camera 940 and display 950 can be configured according to the discussion of the corresponding elements described above with respect to FIG. 8.

[0060] Turning to the image capture process based on gaze tracking, examples of different cameras are discussed. The first is a standard camera that captures a series of images of the eyes that can be processed to determine tracking information. The second is an event camera that generates an output in response to observed changes in brightness.

[0061] In such tracking, it is common to use standard cameras that are widely available and can be manufactured relatively inexpensively. The "standard camera" referred to here means a camera that can capture images of the environment at a predetermined interval and can be combined to generate video content. For example, a typical camera of this type captures 30 images (frames) per second and outputs these images to a processing device to perform feature detection and the like to enable eye tracking.

[0062] Such a camera is composed of a photosensitive array operable to record light information during an exposure time, and the exposure time is controlled by the shutter speed (this speed determines the frequency of image capture). The shutter may be configured, for example, as a rolling shutter (reading the captured information line by line) or a global shutter (simultaneously reading the captured information for the entire frame).

[0063] However, depending on the configuration, it may be advantageous to use an event camera (also called a motion vision sensor) instead. In such a camera, the shutter as described above is not necessary. Instead, each element of the photosensitive array (often called a pixel) is configured to output a signal whenever a threshold luminance change is observed. Therefore, an image in the conventional sense is not output, but an image reconstruction algorithm that can generate an image from the signals output by the event camera can be applied.

[0064] Although the computational complexity increases to generate an image from such data, the output of the event camera can be used for tracking without generating an image. When imaging with infrared light, the pupil of the human eye exhibits a much higher level of brightness than the surrounding features. By selecting an appropriate threshold brightness, the movement of the pupil can trigger an event (and the corresponding output) in the sensor.

[0065] Regardless of the type of camera selected, it may often be advantageous to provide illumination to the eye in order to obtain a suitable image. As an example of this, an IR light source configured to irradiate light in the direction of one or both of the user's eyes can be provided. Subsequently, an IR camera capable of detecting reflections from the user's eyes to generate an image can be provided. Since IR light is invisible to the human eye, it is preferred but not essential as it does not interfere with the user's normal content viewing. In some cases, the illumination may be provided by a light source fixed to the imaging device, but in other embodiments, alternatively, the light source may be arranged away from the imaging device.

[0066] As suggested by the above discussion, the human eyeball does not have a uniform structure. That is, the eyeball is not a perfect sphere, and different parts of the eyeball have different characteristics (such as changes in reflectivity and color). FIG. 10 is a simplified side view of the structure of a typical eyeball 1000. In this figure, features such as the muscles that control eye movement are omitted for clarity.

[0067] The eyeball 1000 is formed in a generally spherical structure filled with an aqueous solution 1010, and a retina 1020 is formed on the rear surface of the eyeball 1000. An optic nerve 1030 is connected to the rear surface of the eyeball 1000. An image is formed on the retina 1020 by light incident on the eyeball 1000, and corresponding signals that transmit visual information are transmitted from the retina 1020 to the brain via the optic nerve 1030.

[0068] Looking towards the front of the eyeball 1000, the sclera 1040 (commonly called the white of the eye) surrounds the iris 1050. The iris 1050 controls the size of the pupil 1060, which is the opening through which light enters the eyeball 1000. The iris 1050 and the pupil 1060 are covered by the cornea 1070, which is a transparent layer that can refract the light entering the eyeball 1000. The eyeball 1000 also includes a lens (not shown) that exists behind the iris 1050 and can be controlled to adjust the focus of the light incident on the eyeball 1000.

[0069] The structure of the eye has a region of high visual acuity (fovea), and on both sides of it, the visual acuity drops sharply. This is shown by the curve 1100 in Figure 11, where the central peak represents the fovea region. Region 1110 is the "blind spot", which corresponds to the region where the optic nerve contacts the retina and thus has no visual acuity. The peripheral part (i.e., the visual field angle farthest from the fovea) is not particularly sensitive to color or details and is instead used for detecting motion.

[0070] Foveal rendering is a rendering technique that takes advantage of the relatively small size of the fovea (about 2.5 degrees) and the sharp drop in visual acuity outside the fovea.

[0071] The eye makes a large number of movements during viewing, and these movements are classified into several categories.

[0072] "Saccade", and the smaller-scale Micro-Saccade, are identified as fast movements in which the eye rapidly moves between different foci (often jerkily). This movement is considered a ballistic movement in that once it starts, it cannot be changed. Saccades are often not conscious eye movements but are made reflexively to survey the environment. Saccades can last up to 200 milliseconds or as short as 20 milliseconds, depending on the distance the eye rotates. The speed of saccades also depends on the total rotation angle, and a typical speed is from 200 degrees to 500 degrees per second.

[0073] "Smooth Pursuit" refers to a movement slower than saccades. Smooth pursuit is generally associated with the viewer consciously tracking the position of the focus and is done to maintain the position of the target within (or at least substantially within) the foveal border region of the viewer's vision. This allows for maintaining a high-quality view of the target of interest despite the movement. If the target's movement is too fast, smooth pursuit may require making saccades several times to catch up, which is because the maximum speed of smooth pursuit is as low as about 30 degrees per second.

[0074] The vestibulo-ocular reflex is a further example of eye movement. The vestibulo-ocular reflex refers to the movement of the eyes that cancels out head movement, that is, the movement of the eyes relative to the head that allows a specific point to be continuously focused on even when the head is moved.

[0075] Another movement is the convergence accommodation reflex. This is a movement that rotates the eyes to converge on a point and correspondingly adjusts the lens within the eyes to focus on that point.

[0076] Further eye movements that may be observed as part of the gaze tracking process are the movements of the eyelids blinking or winking that cover the user's eyes. Such movements can be either reflexive or intentional and often interfere with gaze tracking as they obscure the field of view of the eyes.

[0077] In order to enable a detailed visual analysis of a part of the image displayed by the HMD, eye movements are made by the user wearing the HMD while looking at the image displayed by the HMD. In particular, the eyes are rotated to change the position of the fovea and the pupil, enabling a detailed visual analysis of the part of the image where light is incident on the fovea. Similarly, eye movements are also performed by a user not wearing the HMD while looking at an image displayed by a display unit such as the display unit 850 or 950 described above with reference to FIGS. 8 and 9.

[0078] As described above, foveal rendering is a rendering technique that utilizes the relatively small size of the fovea (about 2.5 degrees) and the sharp decline in visual acuity outside of it. In other words, such a technique renders only a part of the image at the highest level of image quality by taking advantage of the fact that the user is viewing only a very small part of the image in high image quality and the perceived image quality rapidly declines outside of that.

[0079] In conventional forward rendering techniques, typically multiple render passes are required to render an image frame multiple times at different image resolutions and then combine the resulting renderings to achieve regions of different image resolutions within the image frame. Using multiple render passes requires a significant processing overhead and there is a possibility of undesirable image artifacts occurring at the boundaries between regions. Instead, in some cases, hardware can be used that allows for rendering at different resolutions in different parts of the image frame without the need for additional render passes. Thus, while such an implementation with hardware acceleration may be superior in terms of performance, it comes with limitations regarding the smoothness of transitions between regions of different image resolutions within the image frame. In some embodiments, only a limited number of regions can be used and a significant abrupt decrease in image resolution is observed between regions.

[0080] Next, turning to FIG. 12, an embodiment of the present specification relates to an image processing system 1200 that implements a form of forward rendering. In this embodiment, a first region of the image corresponding to the user's gaze position on the image is upscaled using a first, larger, kernel size. And a second region of the image (e.g., the remaining region of the image, or a portion of the image further away from the gaze position) is upscaled using a second, smaller, kernel size. By performing post-processing upscaling on the image, the image can be natively rendered at a lower resolution. As a result, the computational cost is reduced. Thereafter, since the regions of the image can be selectively upscaled to upscale a portion of the image that the user is gazing at (e.g., increase the resolution), the efficiency can be improved. Upscaling provides a computationally efficient technique for improving the quality of the image and makes it possible to shorten the waiting time for outputting the image (which may be natively rendered at a low resolution). As a result, it becomes possible to improve the frame rate of the output content, etc. This approach can provide improved efficiency (and / or an improvement in frame rate) compared to multiple rendering passes and does not require dedicated hardware, while, as will be described later, the use of forward rendering can be made less noticeable to the user.

[0081] By using different kernel sizes to upscale different regions of the image, this approach uses a larger kernel (which requires more computation) for the first focal region to improve the image quality of that region, while using a smaller kernel to still upscale a second further region of the image (in some implementations, the remaining region) (thus reducing the computational cost), making it possible to improve the balance between image quality and computational cost. Accordingly, both the first region and the second region can be natively rendered at a low resolution and efficiently upscaled.

[0082] Also, by upscaling the second region, this approach improves the perceived quality of the image beyond the first region (foveal), while using a larger kernel size, thus reducing the computational cost, and improving the resilience of the foveated rendering process to inaccuracies in the gaze position data (e.g., due to sudden eye movements).

[0083] This approach is particularly applicable to virtual reality applications. Virtual reality presents certain challenges because the viewpoint is constantly changing (since the viewpoint is based on head movement) and foveated rendering is used to distort the displayed image. As a result, the use of foveated rendering may be noticeable to the user and reduce the user's immersion in the content. As described herein, this approach makes it possible to address these challenges by leveraging the foveation effect while reducing the visibility to the user.

[0084] Next, turning to FIGS. 13a, 13b, and 14, in this approach, the first quality of the upscaling of the image 1300 is provided in the first region 1310 corresponding to the user's gaze position (as predicted using, for example, a machine learning model or detected using a detector / gaze tracking device), while the second quality of the upscaling of the image can be provided in the second region 1320 away from the user's gaze position. The first upscaling quality is higher than the second upscaling quality by using different kernel sizes, as described herein.

[0085] The transition from the first upscaled image quality to the second upscaled image quality within the image may be instantaneous at the boundary of the first region as shown in FIG. 13a, or may ramp between the first and second image qualities linearly or non-linearly over a predetermined distance from the first region as shown in FIGS. 13b and 14. In FIG. 13b, the image 1350 includes a first region 1310 and a second region 1370, with a third region 1360 therebetween. As shown in FIG. 14, the slope of the upscaled image quality between the first and second regions through the transition region may be implemented by using a kernel size that gradually decreases as the gaze position moves away, and by selecting an appropriate kernel size for each region, a linear or non-linear slope in the image quality between the regions may be provided. In FIG. 14, the dotted lines A, B, C represent the boundaries between the regions (e.g., A represents the boundary between the first and third regions, B represents the boundary between the third and second regions), and s1, s2, s3 indicate the relative kernel sizes used to upscale the first, second, and third regions, respectively.

[0086] It will be appreciated that by using different kernel sizes, the quality of upscaling of the first and second regions may be different, but the resolution of the upscaled images in the first and second regions may be the same. In one or more examples of the present disclosure, upscaling of an image includes upscaling both the first region and the second region to the same "target" resolution (e.g., 1280x720). By using a larger kernel, the quality of upscaling is higher for the first region than for the second region (thus, for example, artifacts are less likely to occur). However, in contrast to existing techniques, both regions are upscaled to the same resolution. Since there is no variation in the resolution of the image in this way, particularly in embodiments where the second region constitutes the remainder of the image excluding the first region and the user views an output image with a uniform resolution, it is possible to make the present foveated rendering approach less perceptible to the user and in some cases not perceptible at all. This is in contrast to existing foveated rendering techniques where the use of foveated rendering is more prominent for the user because the resolution of the image generally varies across the entire image.

[0087] Alternatively, the first region and the second region can be upscaled to different resolutions, using a high resolution for the first region and a low resolution for the second region.

[0088] Returning to FIG. 12, this shows an example of an image processing system 1200 according to one or more embodiments of the present disclosure.

[0089] The image processing system 1200 includes an input processor 1210, an image upscaling processor 1220, and an output processor 1230. The input processor 1210 receives an image and fixation data indicating the user's fixation position on the image. Next, when the first kernel size for the first region of the received image corresponding to the line-of-sight position and the second kernel size for the second region of the received image are used, and the first kernel size is larger than the second kernel size, the image upscaling processor 1220 performs upscaling processing (e.g., increases its resolution) on at least a part of the image. For example, the image upscaling processor 1220 can increase the resolution of at least a part of the image by interpolating between the pixels of the image using a kernel larger (i.e., based on more adjacent pixels) for the first region than for the second region (e.g., using Lanczos resampling). In this way, the image is upscaled to different qualities in the first region and the third region, and higher quality (e.g., lower possibility of artifacts) is provided in the first region. When the image is upscaled, the output processor 1230 outputs the upscaled image to a display unit (e.g., the display unit 50 of the HMD 20, or a television).

[0090] The image processing system 1200 may be provided as part of a processing device such as the processing device 910, may be provided as part of the HMDs 600, 810, or may be provided as part of a server. Each of the processors 1210, 1220, 1230 may be composed of, for example, a GPU and / or a CPU arranged in a processing device, an HMD, or a server.

[0091] When the image processing system 1200 is provided as part of the processing device 910, the input processor 1210 may receive gaze data from an HMD (such as HMDs 600, 810, etc.) that constitutes a gaze detector or from a detector (such as any one of detectors 610, 630, 640, 700, 840, 940) via a wired or wireless communication (e.g., a Bluetooth (registered trademark) communication link) from the HMD. The output processor 1230 may output an upscaled image for display to the user by transmitting the upscaled image to an HMD or a display unit (such as display unit 950) arranged for the user via a wired or wireless communication. In some examples, the image processing system 1200 may be provided as part of a server, and the input processor 1210 may be configured to receive gaze data from an HMD or a detector (or a processing device such as a personal computer or a game console associated with the HMD or the detector) via a wireless communication, and the output processor 1230 may be configured to output an upscaled image for display to the user by communicating image data corresponding to the upscaled image to an HMD or a display unit (such as display unit 950) arranged for the user.

[0092] Next, the functions of the various processors 1210, 1220, 1230 will be described in more detail.

[0093] First, the input processor 1210 receives an image for output to the user. The image may be, for example, an image frame of a video game. In some cases, the image processing system 1200 may further include a rendering processor configured to render the image and then transmit this image to the input processor 1210.

[0094] The received image may constitute a single image or may be part of an image set (e.g., image frames of a video, etc.). This technology may be applied to upscale each image of the image set in order to provide a video with improved quality for output to the user. The received image may be for a video game. However, this technology can be applied to any type of image.

[0095] Received images are typically of low quality (low resolution such as 720×480 pixels). This allows the images to be rendered efficiently and reduces lag. In this way, for example, a high image frame rate can be achieved.

[0096] As described herein, before outputting the received image, the image is upscaled to increase its quality (e.g., resolution). This makes it possible to provide the user with an improved and more immersive visual experience at reduced computational cost, as upscaling can be more efficient than natively rendering the image at a higher quality.

[0097] The input processor 1210 further receives fixation data indicating the user's fixation position on the image. In other words, the input processor 1210 receives data indicating where the user of the image (i.e., the user to whom the upscaled image will be output) is looking. The gaze data may indicate the detected gaze position of the user and / or the predicted gaze position of the user.

[0098] In consideration of detecting the line-of-sight position, the input processor 1210 may receive line-of-sight data indicating the current line-of-sight position of the user with respect to the image, detected using a detector. The detector may be composed of one or more cameras operable to capture at least one image of the user's eye, and may be configured to detect the user's fixation position. A dedicated detector (e.g., a stand-alone camera) may be placed with respect to the user to detect the user's fixation position. Alternatively, when the user is wearing an HMD, one or more detectors provided as part of the HMD may detect the user's line-of-sight position. Information indicating the user's fixation position may be transmitted from at least one of the HMDs 600, 810 and the detectors 610, 630, 640, 700, 840, 940 to the input processor 1210 via wired or wireless communication.

[0099] In an example where the image is first rendered, the user's line-of-sight position may be detected in parallel with or after the rendering of the image. Thereby, more up-to-date line-of-sight data can be used to select the first and second regions of the image for upscaling, and thus provide an improved alignment between the user's line of sight when viewing the upscaled image and the upscaled regions of the image, providing an improved perceived quality of the image.

[0100] Taking into account the prediction of the gaze position, instead of or in addition to the detected gaze position, the input processor 1210 may receive gaze data indicating the predicted gaze position of the user with respect to the image. The prediction of the gaze position may be determined by a machine learning model. The machine learning model may be trained, for example, to predict the likelihood of the user's gaze position based on the features of the input image. For example, gaze data of users viewing different images may be collected, and the gaze data may be input into the machine learning model together with the corresponding images to train the model. The model may be trained based on this training data to predict the position where the user's gaze is likely to be high with respect to the input image. Then, the input processor 1210 may receive a prediction of the gaze position determined by the machine learning model based on the image (i.e., the upscaled image received by the input processor 1210). If predicted, the fixation data may indicate a plurality of fixation positions (e.g., a plurality of objects of interest within the image) that the user is most likely to fixate on when viewing the image.

[0101] The image and the fixation data may be received by the input processor 1210 from further components of the image processing system 1200 (e.g., a rendering processor or a detector) or from the above-described further device (e.g., an HMD) using any suitable wired or wireless connection.

[0102] The image upscaling processor 1220 implements a form of forward rendering for the received image based on the user's received gaze data. The image upscaling processor 1220 does this by performing an upscaling process on at least a portion of the received image in order to enhance the image quality (e.g., resolution) of at least a portion of the image. The upscaling process is performed using a first kernel size for a first region of the received image corresponding to the user's gaze position on the image and a second kernel size for a second region of the received image (e.g., the remaining region of the received image or the region surrounding the first region). The first kernel size is larger than the second kernel size. Thereby, the quality can be improved more greatly in the first region than in the second region.

[0103] As used herein, the term "kernel" preferably relates to a matrix applied to an image to perform processing of the image. The processing can be performed by determining a convolution between the kernel and the image. In other words, the kernel can define a function for mapping pixels in the input image and neighboring pixels thereof to pixels in the output image. The kernel may sometimes also be referred to as a "convolution matrix" and / or a "mask".

[0104] As used herein, the term "kernel size" preferably relates to the dimensions of the kernel (e.g., the height (i.e., number of rows) and width (i.e., number of columns) of a two-dimensional kernel). The kernel size of a kernel can define the number of pixels of the input image covered / processed by the kernel. The kernel size may be symmetric (e.g., height x width is 3x3, or 5x5), or asymmetric (e.g., height x width is 3x5, or 2x4). In this specification, when it is said that a kernel size (e.g., a first kernel size) is larger than another kernel size (e.g., a second kernel size), preferably, it means that the number of pixels of the input image processed by the kernel having the (e.g., first) kernel size is larger than the number of pixels of the input image processed by the kernel having the other (e.g., second) kernel size. Thus, for example, a kernel size of 5x5 (covering 25 pixels) is considered to be larger than a kernel size of 3x3 (covering 9 pixels) or a kernel size of 6x4 (covering 24 pixels).

[0105] The upscaling process executed by the image upscaling processor 1220 may include upscaling a first region corresponding to the user's line-of-sight position with respect to the image using a first kernel, and upscaling a second region using a second kernel. The number of pixels of the image covered by the first kernel is greater than the number of pixels covered by the second kernel. The first and second kernels can be used as part of a convolution operation (e.g., interpolation or transposed convolution). FIG. 15 shows an example of the upscaling process. In this example, the kernels and kernel sizes used for upscaling the image 1500 are different. In FIG. 15, the grid represents the individual pixels of the image 1500. FIG. 15 shows different exemplary kernels 1520, 1530, 1540 applied to a given pixel P / 1510. Kernel 1520 is symmetric and has a size of 5x5 pixels (i.e., covers 25 pixels centered on pixel 1510). Kernel 1530 is also symmetric but has a smaller size of 3x3 pixels. On the other hand, kernel 1540 is asymmetric and has a size of 7x3 pixels, so it is smaller than kernel 1520 but larger than kernel 1530.

[0106] Thus, for example, when upscaling an input image using interpolation, each output pixel can be interpolated based on the pixels covered by the respective kernels 1520, 1530, 1540 when determining that output pixel, depending on the kernels 1520, 1530, 1540 used. This will be described in more detail later in this specification.

[0107] During upscaling, each of the kernels 1520, 1530, 1540 may be shifted along additional pixels of the input image to determine additional pixels of the output image.

[0108] Also, the kernel size affects the computational cost associated with image processing. Since the number of input image pixels that need to be processed for each output pixel increases, the computational cost may increase as the kernel size increases. At the same time, in various image processing operations such as upscaling, data from more pixels of the input image is considered when determining the output pixel (for example, the output pixel may be interpolated from more adjacent pixels), and the relative increase in image quality may increase as the kernel size increases. For example, in Lanczos interpolation, by increasing the kernel size, a smoother and gentler frequency roll-off can be achieved, and higher-quality anti-aliasing and image quality can be obtained.

[0109] Therefore, increasing the kernel size can improve image quality, but there is a trade-off in that the computational cost increases.

[0110] This disclosure uses a larger kernel to upscale the first region corresponding to the user's fixation position, and thus prioritizes this region in the allocation of computing resources. And a smaller kernel is used to upscale the second region, so that the region is still upscaled, but the computational cost is reduced. This effectively balances this trade-off.

[0111] The upscaling process executed by the image upscaling processor 1220 may use any suitable technique for enhancing the quality of the image. The image quality may be related to any other characteristic of the image that indicates its quality, such as the resolution of the image and / or the degree of aliasing. Therefore, for example, upscaling of the image can increase its resolution and / or reduce the aliasing of the image. In some cases, upscaling of the image may include upsampling of the image.

[0112] Various techniques can be used to upscale an image. For example, the upscaling process may use interpolation (e.g., resampling) and / or deconvolution.

[0113] Considering interpolation, upscaling involves interpolating between the pixels of an image to estimate the values of new pixels. As a result, the total number of pixels increases and the resolution of the image improves. The kernel size used in interpolation can define how many pixels of the original image are considered to generate each pixel of the interpolated upscaled image.

[0114] Examples of suitable interpolation techniques include nearest neighbor interpolation, bilinear interpolation, bicubic interpolation, and / or Lanczos interpolation / resampling. In nearest neighbor interpolation, only one pixel (the nearest pixel) of the original image is used to determine each pixel of the upscaled image. Similarly, in bilinear interpolation and bicubic interpolation, the 2x2 neighborhood and 4x4 neighborhood of the pixels of the original image are used respectively to determine each pixel of the upscaled image, so these methods are considered to have kernel sizes of 2x2 and 4x4 respectively. In Lanczos interpolation, various kernel sizes such as 3x3, 5x5, 7x7, 9x9, etc. can be used. The larger the kernel size, the better the anti-aliasing and the quality of the upscaled image, but the computational cost increases.

[0115] For upscaling the first region and the second region, the same interpolation technique or different interpolation techniques may be used. In any case, the computational resources for upscaling may be mainly allocated to the upscaling of the first region corresponding to the user's line-of-sight position. When using different interpolation techniques, a technique for improving interpolation quality may be used for the upscaling of the first region, and a technique with less computational load may be used for the upscaling of the second region. For example, the first region may be upscaled using bicubic interpolation or Lanczos interpolation that uses a large kernel size to improve the quality of interpolation, and the second region may be upscaled using nearest neighbor interpolation or bilinear interpolation that uses a small kernel size to reduce the computational cost. When the same interpolation technique is used for both the first region and the second region, a larger kernel size can be used for the first region than for the second region. For example, Lanczos interpolation is used for both the first region and the second region, but a larger kernel (e.g., kernel 1520 in FIG. 15) is used for the first region, and a smaller kernel (e.g., kernel 1530 in FIG. 15) is used for the second region.

[0116] In one or more examples, Lanczos interpolation / resampling can be used to upscale an image. Lanczos resampling can provide relatively high-quality interpolation at a relatively low computational cost.

[0117] Any suitable interpolation technique can be used for upscaling an image. For example, an additional interpolation technique that can be used instead of or in addition to the techniques described above is sinc interpolation.

[0118] Considering deconvolution (also known as "transposed convolution"), a transposed convolution kernel can be applied to the input image to expand its spatial dimensions and generate a higher-resolution output image. The weights of the transposed convolution kernel are learned by a neural network. The increase in the resolution of the image (and the associated computational cost) may increase as the size of the transposed convolution kernel increases. Therefore, with respect to interpolation, a larger kernel size can be used for the first region than for the second region.

[0119] In some cases, one or more deep neural network techniques can be used to upscale an image. For example, multiple transposed convolution layers can be arranged in series within a deep neural network to perform progressive upscaling of the image.

[0120] The weights / values of the kernel used for upscaling may be pre-determined. For example, the weights may be determined empirically by an operator. Alternatively, or additionally, the weights may be determined or adjusted using a machine learning model for upscaling of the image during training of the model. For example, a machine learning model (e.g., a neural network) can adjust the kernel weights during training to optimize a cost function such as minimization of the reconstruction error.

[0121] In some cases, instead of or in addition to increasing the resolution of the image, upscaling the image may include performing anti-aliasing on the image. An example of a suitable anti-aliasing technique is Morphological Anti-Aliasing (MLAA). This is a post-processing operation that reduces aliasing (i.e., artifacts that cause edges in the image to appear blocky) by smoothing the image as needed. This is achieved by blending the pixels in the image based on patterns detected within the image. For example, pixels may be detected as belonging to a straight line, and blending may be performed to smooth this line. Since using a larger kernel size may improve the smoothing of the image and further reduce anti-aliasing, a larger kernel may be used for MLAA in the first region than in the second region.

[0122] Upscaling the image (e.g., using interpolation) may be performed as part of a broader upscaling process used to improve the quality of the image, such as FidelityFX Super Resolution (FSR). Any suitable broader upscaling process may be used, including spatial and / or temporal upscaling. Such a broader process may perform further processing on the image, such as sharpening the image, before outputting the image for display.

[0123] The number of pixels covered by the kernel size / kernel for different regions of the image (e.g., the first and second regions) may be predetermined. For example, the kernel size to use for different regions of the image may be determined empirically by an operator.

[0124] Alternatively or additionally, the kernel size may be determined by a machine learning model. For example, if upscaling is performed by a deep learning machine learning model, the machine learning model may adaptively determine the optimal kernel size applied to the upscaling of different image regions during the training of the model.

[0125] Referring back to FIGS. 13a and 13b, these show exemplary regions that can be upscaled in an image using the techniques described herein.

[0126] FIG. 13a shows an exemplary image 1300 including a first region 1310 corresponding to the user's line-of-sight position indicated by the line-of-sight data received by the input processor 1210, and a second region 1320 including the remainder of the image. The first region 1310 may be upscaled using a first, larger kernel size (e.g., 5x5). The second region 1320 may be upscaled using a second, smaller kernel size (e.g., 3x3). Thus, the image 1300 is an example of an image where the entire image is upscaled but different kernel sizes are used for different regions of the image. As described herein, both the first region 1310 and the second region 1320 may be upscaled to the same resolution. Thus, the fact that an output image having a uniform image resolution is provided to the user and a form of foveated rendering is implemented is hidden.

[0127] FIG. 13b shows an exemplary image 1350 including a first region 1310, a second region 1370, and a third region 1360. The first region 1310 corresponds to the user's fixation position indicated by the fixation data.

[0128] In some cases, only a part of the image 1350 may be upscaled. For example, for the image 1350 in FIG. 13b, the first region 1310 is upscaled using a first, larger kernel. The third region 1360 (functioning substantially as the second region in FIG. 13a) is upscaled using a smaller, third kernel. On the other hand, the second region 1370 may not be upscaled. In this way, the first region 1310 corresponding to the user's fixation position and the surrounding third region 1360 may be upscaled. At this time, the computing resources are mainly dedicated to the upscaling of the first region 1310 where the larger kernel is used. In this example, by utilizing foveal perception and not upscaling the second region 1370 that is further away from the user's fixation position, the computational cost of upscaling is reduced.

[0129] Alternatively, for example, as described above with reference to FIG. 13a, or by upscaling the second region 1370 of the image 1350 using a second kernel size, the entire image may be upscaled.

[0130] In some cases, a transition region using a kernel size between the first region and the second region may be provided between the first region and the second region. This can also be explained with reference to the image 1350 shown in FIG. 13b. The transition region of the image 1350 scales up the third region 1360 using a third kernel size that is smaller than the first kernel size (used for upscaling the first region 1310) but larger than the second kernel size (used for upscaling the second region 1370). In this way, the third region 1360 can function as a transition region with an intermediate kernel size and quality (e.g., having intermediate aliasing) between the first region 1310 and the second region 1370. Thereby, a gentle degradation of the image quality can be provided between the first region and the second region. As a result, for a user whose perception of the image quality also degrades depending on the distance from the first focus region, it is possible to make it difficult to perceive the difference in image quality.

[0131] The transition between the first region and the second region, and the ramp of the upscaling quality (e.g., anti-aliasing performance) may be provided using a plurality of transition regions with gradually smaller kernel sizes as the distance from the fixation position increases. The kernel size may vary depending on the distance from the fixation position, and may decrease as the distance from the fixation position increases. The change in the kernel size with respect to the distance from the fixation position may be linear or non-linear.

[0132] Returning to FIG. 14, this shows an exemplary indication kernel size that can be used to implement a transition region between a first region and a second region. In this example, the image consists of a first region between the fixation position and distance A from the fixation position, a third region between distance A and B from the fixation position, and a second region between distance B and C from the fixation position. Beyond distance C from the fixation position, upscaling may not be performed. The first, second, and third regions are upscaled using kernel sizes s1 (e.g., 7x7), s2 (e.g., 5x5), and s3 (e.g., 3x3), respectively. In this way, a gradual degradation of image quality from the first region to the second region is provided, making it difficult for the user to perceive the degradation of image quality.

[0133] In some cases, the drop-off in upscaling quality may gradually become steeper as the distance from the fixation position increases. For example, the difference between the third kernel size s3 and the first kernel size s1 (e.g., the difference in the number of pixels covered by the third kernel and the first kernel) may be smaller than the difference between the third kernel size s3 and the second kernel size s2 (e.g., the difference in the number of pixels covered by the third kernel and the second kernel). In other words, the step size in the variation of the kernel size can increase with the distance from the fixation position. This can improve efficiency as larger step changes in quality further away from the foveal region are less noticeable to the user while allowing for a reduction in computational cost.

[0134] It will be appreciated that the step change in kernel size between regions as shown in FIG. 14 may approximate a ramp of kernel sizes for more regions (e.g., such a ramp may be illustrated considering the line between the kernel sizes at the midpoint of each region). Although FIG. 14 shows only three regions for upscaling, such a ramp approximation will become clearer with a larger number of intervening transition regions, similar to the third region discussed herein.

[0135] Alternatively, instead of or in addition to the kernel size, further parameters of the upscaling process can be varied across regions of the image to obtain a (e.g., linear or non-linear) quality gradient between a first region and a second region. Exemplary relevant parameters may be the target resolution of the upscaling process (i.e., the resolution at which the image is upscaled). For example, considering the image 1350 of FIG. 13b, the target resolution may decrease gradually as the distance from the fixation position increases across the third region 1370 (e.g., by 20 pixels in each dimension for every 20 pixels away from the fixation position), providing a gradual decrease in resolution across the third transition 1360 region.

[0136] In some cases, the kernel size used for upscaling a region (e.g., the first region and / or the second region) may be at least partially modified based on the characteristics of the image and / or the image processing system 1200.

[0137] Considering the characteristics of the image, the kernel size for upscaling a region of the image may be determined depending on the characteristics of the image of that region. Exemplary relevant characteristics may include the orientation of features within the image region (e.g., predominantly vertical or horizontal), and / or the level of detail within the image region.

[0138] Regarding the orientation of features, the orientation of features within the image region may be determined, for example, by extracting features from the image region (e.g., using one or more appropriate feature extraction techniques) and determining the dominant orientation of features within the region (e.g., whether the features within the region are predominantly arranged in a given orientation). In some cases, feature extraction performed at another stage of the upscaling process may be reused for this purpose. For example, features extracted as part of FSR may be analyzed to determine whether they are arranged in any dominant direction.

[0139] In this way, for example, it can be determined that the features in a given region of an image are dominant in the vertical direction (e.g., in the case of an image region showing a fence or grass) or the horizontal direction (e.g., in the case of an image region showing an arrow in the air). The dominance of vertical / horizontal features is determined, for example, based on the number or ratio of vertical features in the image region exceeding a predetermined threshold. The features may be classified as vertical or horizontal, for example, depending on their dominant direction (i.e., the direction in which the features extend) being within a predetermined angle of vertical or horizontal.

[0140] When it is determined that the features in the image region are arranged in a predetermined dominant direction, the relative dimensions of the kernel size can be changed depending on the direction of the features. Thereby, during upscaling, artifacts caused by upscaling (e.g., interpolation) of the image can be reduced by assigning a larger weight to adjacent pixels in the dominant direction of the features. For example, the kernel size used to upscale a given region can be increased depending on the direction of the dominant features. For example, when it is determined that the features in the region of the image are mainly in the vertical direction, instead of a 3x3 kernel size, an asymmetric kernel size (e.g., with a height x width of 5x3) can be used.

[0141] Alternatively, or in addition, the kernel size used for upscaling a given region can be decreased in a direction depending on the non-dominant feature orientation. For example, when it is determined that the features in the region of the image are mainly vertical, instead of a 7x7 kernel size, an asymmetric kernel size (e.g., having a height x width of 7x3) may be used. By reducing the kernel size in the non-dominant feature direction, the computational cost can be efficiently reduced while maintaining the image quality.

[0142] Regarding the level of detail (LOD), the kernel size used for upscaling an area of an image may increase as the level of detail of that area increases, for example as the LOD exceeds a predetermined threshold. Larger kernels, which provide smoother and higher-quality interpolation, may be used for areas with higher LOD where upscaling artifacts become more prominent to the user. On the other hand, smaller kernels may be used for areas with lower LOD in order to reduce the overall computational cost. In some cases, the upscaling technique used may also be changed depending on the LOD of the image area. For example, an upscaling technique with low computational cost but low accuracy (such as bilinear interpolation or nearest neighbor interpolation) may be used for areas with low LOD (e.g., LOD below a first predetermined threshold), and an upscaling technique with high computational cost but high accuracy (such as Lanczos interpolation) may be used for areas with high LOD (e.g., LOD above a second predetermined threshold). This ensures that a computationally less expensive upscaling process (e.g., by using a smaller kernel size and / or a computationally cheaper upscaling technique) is used for areas with lower LOD, where artifacts are less likely to be introduced and / or less likely to be noticed by the user, while a computationally more expensive upscaling process is reserved for areas with higher LOD. As a result, the balance between efficiency and upscaled image quality can be improved, which can help to ensure that these areas are accurately upscaled.

[0143] Modifying the kernel size based on the image characteristics may be done over the overall area as described above (e.g., over the entire first area 1310, second area 1370, and / or third area 1360 of the image 1350 in FIG. 13b), or over sub-areas of those areas, such as sub-areas where a particular feature direction is dominant or where the level of detail is particularly high or low.

[0144] Considering the characteristics of the image processing system 1200, the kernel size for one or more regions of an image can be changed depending on one or more of the frame rate for outputting the upscaled image to a display device, the quality at which at least a portion of the image is upscaled, the available computing resources, or the communication bandwidth. Thereby, the current requirements for outputting the image (e.g., set by the frame rate and the upscaling quality, such as the target upscaling resolution), and / or the currently available resources for performing the upscaling (e.g., set by the available computing (e.g., processing or storage) resources and / or the communication bandwidth) can be used to adjust the computational cost of the upscaling. The kernel size can be decreased with an increase in the output requirements (e.g., an increase in the frame rate or the upscaling quality such as the target upscaling resolution) and / or a decrease in the available resources. This helps to reduce the computational cost of the upscaling and ensures that the output requirements can be met with the currently available resources. The reduction of the kernel size may be determined based on a function determined empirically based on the output requirements and the available resources. How much the kernel size is changed may depend on the distance from the fixation position. For example, to reduce the computational cost, a greater reduction in the kernel size may be performed in a second region than in a first region. Considering the characteristics of the image processing system 1200, the kernel size for one or more regions of an image can be changed depending on one or more of the frame rate for outputting the upscaled image to a display device, the quality at which at least a portion of the image is upscaled, the available computing resources, or the communication bandwidth.Accordingly, the computational cost of upscaling can be adjusted depending on the current requirements for outputting an image (e.g., set by frame rate and upscaling quality), and / or the currently available resources for performing upscaling (e.g., set by available computing (e.g., processing or storage) resources and / or communication bandwidth). The kernel size can be decreased as the output requirements increase (e.g., an increase in frame rate or upscaling quality such as target upscaling resolution) and / or the available resources decrease. This helps to reduce the computational cost of upscaling and ensures that the output requirements can be met with the currently available resources. The reduction in kernel size may be determined based on a function determined empirically based on the output requirements and available resources. How much the kernel size is changed may depend on the distance from the fixation position. For example, a greater reduction in kernel size may be performed in a second region than in a first region to reduce the computational cost.

[0145] The image upscaling processor 1220 upscales a first region of the image corresponding to the fixation position using a first kernel size and upscales a second region using a second kernel size.

[0146] In some cases, in addition to upscaling the first region using the first kernel size, the image upscaling processor 1220 may upscale both the first region and the second region using the second kernel size. In other words, the image upscaling processor 1220 may execute multiple upscaling passes on the image. Thereby, the quality of the image is gradually improved. This is because a computationally inexpensive upscaling pass using the second kernel size has already partially increased the resolution of the first region (e.g., by already adding some of the new pixels using interpolation), so that a computationally expensive upscaling pass using the first kernel size can perform a relatively smaller increase in resolution. Therefore, the efficiency of image upscaling can be further improved. Further, the second upscaling pass can use the result of the first upscaling pass. For example, the second pass can be composed of interpolations between the pixels added by interpolation in the first pass. Thereby, in the second pass, a simpler upscaling technique (e.g., bicubic interpolation instead of Lanczos interpolation) can be used, further improving the efficiency. This approach is also in contrast to performing multiple render passes as in the prior art, where a portion of the image rendered at a resolution lower than the target resolution (e.g., as part of the first render pass or an intermediate render pass) is effectively discarded.

[0147] Upscaling both the first and second regions using the second kernel size can be performed before or after upscaling the first region using the first kernel size. For example, the image upscaling processor 1220 may first upscale the first and second regions (which may together constitute the entire image in some cases) using the second kernel size to a second resolution (e.g., from 720×480p to 1280×720p), and then further upscale the first region to a higher first resolution (e.g., up to 1920×1080p) using the first kernel size. Alternatively, the image upscaling processor 1220 may first upscale the first region to an intermediate resolution (e.g., from 720×480p to 1440×1080p) using the first kernel size, and then upscale the first and second regions (which may together constitute the entire image in some cases) to a second resolution (e.g., 1280×720p) in the second region and a first resolution (e.g., 1920×1080p) in the first region using the second kernel size. Finally, turning to the output processor 1230, the output processor 1230 outputs the image upscaled by the image upscaling processor 1220 to a display device. The display device may be an HMD in some embodiments, but any display device may be used as appropriate to display the image.

[0148] It will be appreciated that the techniques described herein are applicable to VR content. For example, the input processor 1210 may receive a pair of images (e.g., a stereoscopic image pair). The image upscaling processor 1220 may perform upscaling of both images in the pair. The output processor 1230 may output both images to a display device (e.g., an HMD).

[0149] The pair of images received by the input processor 1210 may overlap. The first region and the second region may be disposed in only one of the images, or may extend across both images (for example, in the overlapping region of the pair of images). For example, when the user's line-of-sight position is in the overlapping region that exists in both of the pair of images, the first region and / or the second region of the images may extend across both images. In this case, the first region and the second kernel size may be determined for one of the images and then applied to both images.

[0150] Although the above discussion focuses on the use of an HMD, it can be implemented using any display. For example, a video game displayed on a television may be upscaled to a higher level of quality for a region corresponding to the user's line of sight determined using one or more separate detectors. In such an embodiment, the display of the content and the eye tracking are not performed by an HMD.

[0151] Next, turning to FIG. 16, in a general embodiment of the present invention, the image processing method includes the following steps.

[0152] Step 1610 includes receiving an image, as described elsewhere in this specification.

[0153] Step 1620 includes receiving fixation data indicating the user's fixation position on the image, as described elsewhere in this specification.

[0154] Step 1630 includes, as described elsewhere in this specification, performing an upscaling process on at least a part of the received image in order to enhance the image quality of at least a part of the received image. The step of performing the upscaling process includes, as described elsewhere in this specification, upscaling a first region of the received image corresponding to the user's line-of-sight position with respect to the image using a first kernel size, and upscaling a second region of the received image using a second kernel size. The first kernel size is larger than the second kernel size.

[0155] Step 1640 includes, as described elsewhere in this specification, outputting the upscaled image to a display unit.

[0156] It will be apparent to those skilled in the art that various modifications of the above methods corresponding to the operations of the various embodiments of the methods and / or apparatuses as described and claimed herein are contemplated within the scope of the present disclosure, but are not limited thereto.

[0157] The step 1630 of performing the upscaling process includes, as described elsewhere in this specification, upscaling a first region corresponding to the user's line-of-sight position with respect to the image using a first kernel, and upscaling a second region of the image using a second kernel. The number of pixels of the image covered by the first kernel is larger than the number of pixels covered by the second kernel.

[0158] The step of performing the upscaling process 1630 includes, as described elsewhere in this specification, upscaling a first region of the received image to a first quality (e.g., resolution, aliasing, or occurrence / likelihood of artifacts), and upscaling a second region of the received image to a second quality. The second quality is lower than the first quality.

[0159] The degree of upscaling increases as the kernel size increases such that, as described elsewhere in this specification, upscaling using a larger kernel size improves the image quality.

[0160] The kernel size is proportional to the degree of upscaling such that, as described elsewhere in this specification, a larger upscaling is performed on the region associated with the larger kernel size.

[0161] The step of performing the upscaling process 1630 includes the step of increasing the resolution of at least a part of the received image, as described elsewhere in this specification.

[0162] In this case, optionally, as described elsewhere in this specification, the resolution is increased to the same resolution in both the first region and the second region of the received image.

[0163] The step of performing the upscaling process 1630 includes the step of interpolating between at least some of the pixels of the received image, as described elsewhere in this specification.

[0164] In this case, optionally, the step of performing the upscaling process 1630 includes using a first kernel size for interpolating between pixels within the first region of the received image and using a second kernel size for interpolating between pixels within the second region of the received image, as described elsewhere in this specification.

[0165] In this case, optionally, the interpolation is performed using Lanczos resampling, as described elsewhere in this specification.

[0166] The step of performing the upscaling process 1630 includes upscaling both the first region and the second region using a second kernel size and upscaling the first region using a first kernel size, as described elsewhere in this specification.

[0167] As described elsewhere in this specification, the second region includes the remaining portion of the received image excluding the first region.

[0168] The step of performing the upscaling process 1630 includes upscaling a third region of the received image disposed between the first region and the second region using a third kernel size, as described elsewhere in this specification. The third kernel size is larger than the second kernel size and smaller than the first kernel size.

[0169] Optionally, in this case, the difference between the third kernel size and the first kernel size is smaller than the difference between the third kernel size and the second kernel size, as described elsewhere in this specification.

[0170] As described elsewhere in this specification, the method further includes the step of modifying the kernel size for upscaling at least one region of the received image depending on one or more characteristics of the received image in at least one region.

[0171] Optionally, in this case, one or more characteristics of the received image include the direction of features in the received image, as described elsewhere in this specification.

[0172] Optionally, here, the relative dimensions of the kernel size are changed depending on the orientation of the features, as described elsewhere in this specification.

[0173] Optionally, in this case, one or more features of the received image constitute the level of detail of the received image, as described elsewhere in this specification.

[0174] Here, as an option, as described elsewhere in this specification, the kernel size increases as the level of detail increases.

[0175] As described elsewhere in this specification, further comprising the step of changing the kernel size for upscaling at least one region of the received image, depending on the frame rate for outputting the upscaled image to a display device, the quality at which at least a portion of the image is upscaled, the available computing resources or communication bandwidth.

[0176] Further comprising the step of detecting the user's line-of-sight position using a detector. The detector comprises one or more cameras operable to capture at least one image of the user's eye, as described elsewhere in this specification.

[0177] The display device is a head-mountable display, as described elsewhere in this specification.

[0178] Further comprising rendering the image at a first low resolution. Here, the upscaling process increases the resolution of at least a portion of the image to a second high resolution, as described elsewhere in this specification.

[0179] At least a portion of the image to be upscaled constitutes the entire received image, as described elsewhere in this specification.

[0180] The step of performing the upscaling process 1630 includes the step of performing transposed convolution of the image, as described elsewhere in this specification.

[0181] The image is part of a video game, as described elsewhere in this specification.

[0182] The above method may be implemented on conventional hardware that has been appropriately adapted as applicable, either by software instructions or by the inclusion or replacement of dedicated hardware.

[0183] Accordingly, the necessary adaptation to existing portions of conventional equivalent devices may be implemented in the form of a computer program product containing processor-implementable instructions stored on a non-transitory machine-readable medium such as a floppy disk, optical disk, hard disk, solid state disk, PROM, RAM, flash memory, or any combination thereof or other storage medium, or may be implemented in hardware as an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array) or other configurable circuit suitable for use in adapting conventional equivalent devices. Alternatively, such a computer program may be transmitted via a data signal on a network such as Ethernet (registered trademark), a wireless network, the Internet, or a combination thereof or other network.

[0184] Therefore, referring back to FIG. 12, in an exemplary embodiment of the present invention, the image processing system 1200 may be configured as follows.

[0185] As described elsewhere in this specification, an input processor 1210 (e.g., a processing device, an HMD, or a server CPU) configured to receive an image and receive fixation data indicating a user's fixation position on the image (e.g., by appropriate software instructions).

[0186] An image upscaling processor 1220 (e.g., a CPU of a processing device, HMD, or server), configured (e.g., by appropriate software instructions) to perform upscaling processing on at least a part of a received image in order to improve the image quality of at least a part of the received image, wherein the step of performing upscaling processing includes, as described elsewhere herein, upscaling a first region of the received image corresponding to a user's line-of-sight position with respect to the image using a first kernel size, and upscaling a second region of the received image using a second kernel size. The first kernel size is larger than the second kernel size.

[0187] An output processor 1230 (e.g., a CPU of a processing device, HMD, or server), configured (e.g., by appropriate software instructions) to output the upscaled image to a display device, as described elsewhere herein.

[0188] The above system 1200 operating under suitable software instructions may implement the methods and techniques described herein.

[0189] It goes without saying that the functionality of these processors need not require a one-to-one mapping between functionality and device or processor, and may be implemented by any suitable number of processors arranged in any suitable number of devices.

[0190] The foregoing discussion discloses and describes merely exemplary embodiments of the invention. As will be understood by those skilled in the art, the invention can be embodied in other specific forms without departing from its spirit or essential characteristics. Accordingly, the disclosure of the invention is intended to be illustrative rather than limiting the scope of the invention, as is the case with the scope of other claims. The disclosure, including readily distinguishable variations of the teachings herein, partially defines the scope of the terms of the foregoing claims so that the inventive subject matter is not dedicated to the public.

Claims

1. An image processing method comprising: receiving an image; receiving gaze data indicative of a user's gaze position relative to the image; performing an upscaling process on at least a portion of the received image to enhance an image quality of at least a portion of the received image; outputting the upscaled image to a display unit; Including, The step of performing the upscaling process includes: upscaling a first region of the received image corresponding to a user's gaze position relative to the image using a first kernel size; and upscaling a second region of the received image using a second kernel size; 13. The image processing method according to claim 12, wherein the first kernel size is greater than the second kernel size.

2. 2. The method of claim 1, wherein the step of performing an upscaling process comprises increasing a resolution of at least a portion of the received image.

3. 3. The method of claim 2, wherein the resolution is increased to the same resolution in both the first region and the second region.

4. 4. The image processing method according to claim 1, wherein the step of performing the upscaling process includes a step of interpolating between at least some pixels of the received image.

5. 5. The method of claim 4, wherein the step of interpolating is performed using Lanczos resampling.

6. The step of performing the upscaling process includes: upscaling both the first region and the second region using the second kernel size; upscaling the first region using the first kernel size; 4. The image processing method according to claim 1, further comprising:

7. 7. The image processing method according to claim 1, wherein the second region includes a remainder excluding the first region.

8. performing the upscaling process includes upscaling a third region of the received image, the third region being located between the first region and the second region, using a third kernel size; 8. The image processing method according to claim 1, wherein the third kernel size is larger than the second kernel size and smaller than the first kernel size.

9. 9. An image processing method according to any one of claims 1 to 8, further comprising the step of modifying a kernel size for upscaling at least one region in dependence on one or more characteristics of the received image in said at least one region.

10. the one or more characteristics of the received image include an orientation of a feature in the received image; 10. The method of claim 9, wherein the relative dimensions of the kernel sizes are altered depending on the orientation of the features.

11. The one or more characteristics of the received image include a level of detail of the received image; 11. The method of claim 9, wherein the kernel size increases as the level of detail increases.

12. 12. The image processing method according to claim 1, further comprising the step of modifying a kernel size for upscaling at least one region of the received image depending on a frame rate for outputting the upscaled image to a display device, the quality to which at least a portion of the image is upscaled, available computing resources or communication bandwidth.

13. 13. The image processing method according to claim 1, wherein the display unit is a head-mountable display.

14. A computer program comprising computer executable instructions to cause a computer system to carry out a method according to any one of claims 1 to 13.

15. An image processing system comprising: An input processor; an image upscaling processor; an output processor; Preparation, The input processor includes: receiving an image; and receiving gaze data indicative of a user's gaze position relative to the image; The image upscaling processor includes: performing an upscaling process on at least a portion of the received image to enhance image quality of at least a portion of the received image; The upscaling step includes the steps of: upscaling a first region of the received image corresponding to a user's gaze position relative to the image using a first kernel size; and upscaling a second region of the received image using a second kernel size; the output processor performs the steps of outputting the upscaled image to a display unit; 13. An image processing system, comprising: a first kernel size that is greater than a second kernel size;