Video Processing and Playback System and Method

The use of HMDs with eye and head tracking, along with foveal rendering, addresses the discomfort in VR/AR game streaming by allowing viewers to explore the virtual environment freely and maintain high-quality graphics, enhancing immersion and reducing computational overhead.

JP7855444B2Active Publication Date: 2026-05-08SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
SONY INTERACTIVE ENTERTAINMENT LLC
Filing Date
2022-07-19
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Conventional video game streaming systems, particularly for VR or AR games, cause discomfort and frustration for viewers due to the passive nature of watching a recorded stream where the perspective follows the distributor's head and/or eye movements, rather than the viewer's.

Method used

A video recording and playback system that utilizes head-mounted displays (HMDs) with eye and head tracking, combined with foveal rendering techniques, to generate high-resolution images based on the viewer's gaze and head position, allowing viewers to freely explore the virtual environment beyond the original user's field of view.

Benefits of technology

Enables a more immersive viewing experience by allowing viewers to look in different directions and maintain high-quality graphics, reducing computational overhead and enabling richer graphics with the same resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007855444000001
    Figure 0007855444000001
  • Figure 0007855444000002
    Figure 0007855444000002
  • Figure 0007855444000003
    Figure 0007855444000003
Patent Text Reader

Abstract

To provide video processing and playback systems and methods.SOLUTION: A video processing method is for processing a circular panoramic recording video including an original field of view region at a first resolution and a further peripheral region outside the original field of view region at a second resolution lower than the first resolution, and includes a step of performing spatial upscaling of the further peripheral region to a resolution higher than the second resolution.SELECTED DRAWING: Figure 15
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to video processing and playback systems and methods.

Background Art

[0002] With conventional video game streaming systems such as Twitch (registered trademark) and video hosting platforms such as YouTube (registered trademark) and Facebook (registered trademark), players of video games have been able to widely distribute the playing of these games to viewers.

[0003] A major difference between playing a video game and watching a video recording of the game play lies in the passive nature of the experience. This is true both at the points of decision during the game and in terms of the player's perspective (which is determined, for example, by the player's input).

[0004] When the game is a VR or AR game, the latter problem is more serious. In this case, usually the player of the game determines the perspective at least partly based on the movement of their own head or eyes. Therefore, when watching a live or recorded stream of such a VR or AR game, the recorded image will follow the movement of the head and / or eyes of the distributor rather than the viewer. This can make the viewer feel uncomfortable and also lead to frustration for viewers who want to look in a different direction than the distributor.

Summary of the Invention

Problems to be Solved by the Invention

[0005] The present disclosure aims to alleviate or reduce such problems.

Means for Solving the Problems

[0006] Various aspects and features of the present invention are defined in the context of the appended claims and specification. The present invention includes, in at least a first aspect, a video recording method, in another aspect, a video recording distribution method, in yet another aspect, a video recording viewing method, in yet another aspect, a video recording system, and in yet another aspect, a video playback system.

[0007] It should be understood that the general description above and the detailed description below are illustrative of the invention and not limiting. [Brief explanation of the drawing]

[0008] A complete understanding of this disclosure and its many advantages can be obtained by referring to the attached drawings and reading the following detailed description. [Figure 1] This is a schematic diagram of an HMD (Head-Mounted Display) worn by a user. [Figure 2] This is a schematic plan view of the HMD. [Figure 3] This is a schematic diagram illustrating the formation of a virtual image using an HMD (Head-Mounted Display). [Figure 4] This is a schematic diagram of another type of display used in HMDs. [Figure 5] This is a schematic diagram of a pair of three-dimensional images. [Figure 6a] This is a schematic plan view of the HMD. [Figure 6b] This is a schematic diagram of a near-eye tracking configuration. [Figure 7] This is a schematic diagram of a remote tracking configuration. [Figure 8] This is a schematic diagram of an eye-tracking environment. [Figure 9] This is a schematic diagram of an eye-tracking system. [Figure 10] This is a schematic diagram of the human eye. [Figure 11] This is a schematic diagram of the graph of human visual acuity. [Figure 12a] This is a schematic diagram of foveal rendering. [Figure 12b] This is a schematic diagram of foveal rendering. [Figure 13a] It is a schematic diagram showing the change in resolution. [Figure 13b] It is a schematic diagram showing the change in resolution. [Figure 14a] It is a schematic diagram of an extended rendering scheme according to an embodiment of the present invention. [Figure 14b] It is a schematic diagram of an extended rendering scheme according to an embodiment of the present invention. [Figure 15] It is a flowchart of a video processing method according to an embodiment of the present invention. [Figure 16] It is a flowchart of a video playback method according to an embodiment of the present invention. [[ID=​​​​​​​​​​​​​​​​​​During operation, the video signal for the display is provided by the HMD. This may be provided by an external video signal source 80 (e.g., a video game console or a data processing device such as a personal computer). In this case, the signal may be transmitted to the HMD by a wired or wireless connection 82. A suitable example of a wireless connection is a Bluetooth® connection. The audio signal for the earpiece 60 may be transmitted by the same connection. Similarly, any control signals sent from the HMD to the video (audio) signal source may be transmitted by the same connection. Furthermore, a power supply 83 (which may include one or more batteries and / or be connected to a mains outlet) may be connected to the HMD by a cable 84.

[0013] Thus, the configuration in Figure 1 provides an example of a head-mountable display system comprising a frame mounted on the viewer's head and a display element mounted relative to the line of sight. The frame defines one or two line of sight positions, which are positioned in front of the viewer's eyes during use. The display element provides a virtual image of the video display signal from a video signal source to the viewer's eyes. Figure 1 is merely one example of an HMD, and other forms are possible. For example, the HMD may use a frame similar to conventional eyeglasses.

[0014] In the example in Figure 1, separate displays are provided for the user's left and right eyes. Figure 2 is a schematic plan view of how this is achieved. Figure 2 shows the position of the user's eyes 100 and the relative position 110 of the user's nose. The display portion 50 schematically comprises an external shield 120 to block ambient light from the user's eyes and an internal shield 130 to prevent the other eye from seeing the display viewed by one eye. With respect to the user's face, the external shield 120 and the internal shield 130 form two compartments 140 for each eye. Within each compartment are provided a display element 150 and one or more optical elements 160. Figure 3 shows the optical path formed by the display element and optical elements (which provides the user with a display).

[0015] Referring to FIG. 3, display element 150 generates a display image. (In this example) the display image is refracted by optical element 160 (shown schematically as a single convex lens, but may be a compound lens or the like). As a result, a virtual image 170 is generated. To the user, virtual image 170 appears larger and much farther away than the real image generated by display element 150. In FIG. 3, solid lines (e.g., line 180) represent actual light rays, and dotted lines (e.g., line 190) represent virtual light rays.

[0016] FIG. 4 shows an alternative configuration. Here, display element 150 and optical element 200 cooperate to provide an image projected onto mirror 210. Mirror 210 reflects the image towards the position 220 of the user's eyes. The user perceives the virtual image to be at a forward position 230 in front of the user and moderately distant from the user.

[0017] When separate displays are provided for each of the user's left and right eyes, a stereoscopic image can be displayed. FIG. 5 shows an example of a pair of stereoscopic images for display to the left and right eyes.

[0018] When an HMD is used in a virtual reality (VR) system or the like, the user's viewpoint needs to track movements with respect to the space where the user is located.

[0019] For tracking, head tracking and / or eye tracking may be used. Head tracking is performed by detecting the movement of the HMD and changing the apparent viewpoint of the displayed image. As a result, the apparent viewpoint tracks the movement. For tracking the movement, any suitable configuration including a hardware motion detector (e.g., an accelerometer or gyroscope, etc.), an external camera capable of photographing the HMD, and an outward-facing camera attached to the HMD may be used.

[0020] Regarding eye tracking, FIGS. 6a and 6b show two possible configurations.

[0021] Figure 6a shows an example of an eye-tracking configuration. In this configuration, cameras are placed within the HMD. This captures images of the user's eyes from a close distance. This is sometimes called near-eye tracking or head-mounted tracking. In this example, the HMD 400 (along with the display element 601) is given cameras 610. Each of these cameras is positioned to directly capture one or more images of each respective eye. The figure shows four cameras 610 as an example of a possible arrangement of eye-tracking cameras. However, typically, it is desirable to have one camera per eye. Optionally, if eye movement is constant as usual, only one eye may be tracked. One or more such cameras may be arranged such that a lens 620 is included in the optical path for capturing images of the eyes. An example of such an arrangement using camera 630 is illustrated. One advantage of including the lens in the optical path is that it simplifies the physical constraints on the HMD design.

[0022] Figure 6b shows an example of an eye-tracking configuration. Here, the camera is positioned to indirectly capture an image of the user's eye. Figure 6b includes a mirror 650 positioned between the display 601 and the viewer's eye. For clarity, all additional optical elements such as lenses are omitted in this figure. In such a configuration, the mirror 650 is selected to be partially light-transmitting. That is, the mirror 650 is selected so that when the user looks at the display 601, the camera 640 can capture an image of the user's eye. One way to achieve this is to use a mirror 650 that reflects IR wavelength light but transmits visible light. This allows the IR light used for tracking to be reflected from the user's eye towards the camera 640, while the light emitted from the display 601 passes through the mirror without interference. One advantage of such a configuration is that the camera can be easily positioned outside the user's field of view. Furthermore, the accuracy of eye tracking is improved because (thanks to the reflection) the camera captures the image from a position substantially along the axis between the user's eye and the display.

[0023] Alternatively, the eye-tracking configuration does not have to be the head-mounted or near-eye type described above. For example, Figure 7 is a schematic diagram of a system in which cameras are positioned to capture images of the user from a distance. In Figure 7, an array of cameras 700 is given, which provides multiple images of the user 710. These cameras are positioned in a preferred manner to capture information to determine at least the direction in which the user 710's eyes are focusing.

[0024] Figure 8 is a schematic diagram of the environment in which the eye-tracking process takes place. In this example, user 800 is using an HMD 810 associated with a processing unit 830 (e.g., a game console) and a peripheral device 820 for inputting commands to control the process. The HMD 810 may perform eye tracking according to the configuration illustrated in Figure 6a or Figure 6b. That is, the HMD 810 may have one or more cameras for capturing images of one or both of user 800's eyes. The processing unit 830 may generate content to display on the HMD 810. However, some (or all) of the display content may be generated by the processing unit within the HMD 810.

[0025] The configuration shown in Figure 8 includes a camera 840 positioned outside the HMD 810 and a display 850. In some cases, the HMD 810 may be used, for example, to identify body movements or head orientation, and the camera 840 may be used to track the user 800. In an alternative configuration, the camera 840 may be mounted outward on the HMD to determine the movement of the HMD based on the motion in the captured video.

[0026] The processing required to generate tracking information from the captured image of the user 800's eyes may be performed in-situ by the HMD 810. Alternatively, the captured image or one or more detection results may be sent to an external device for processing (e.g., a processing unit 830). In the former case, the HMD 810 may output the processing results to the external device.

[0027] Figure 9 is a schematic diagram of a system that performs one or more eye-tracking and head-tracking processes. This system performs, for example, the processes described in Figure 8. System 900 comprises a processing device 910, one or more peripheral devices 920, an HMD 930, a camera 940, and a display 950.

[0028] As shown in Figure 9, the processing device 910 comprises one or more central processing units (CPUs) 911, a graphics processing unit (GPU) 912, storage (hard drives or other suitable storage media) 913, and input / output 914. These units may be provided in the form of a personal computer or in the form of any other suitable processing device.

[0029] For example, the CPU 911 may be configured to generate tracking data from input images of one or more user eyes obtained from one or more cameras, or from data representing the direction of the user's gaze. This data may be obtained, for example, from processed images of the user's eyes by a remote device. Needless to say, if the tracking data is generated elsewhere, the processing device 910 does not need to perform this processing.

[0030] Alternatively or additionally, one or more cameras (other than eye-tracking cameras) may be used to track head movements as described above, or any suitable motion tracker, such as an accelerometer within the HMD, may be used.

[0031] A GPU may be deployed to generate content to be displayed to users who are subject to eye tracking or head tracking.

[0032] Depending on the tracking data acquired, the display content itself may be improved. One example of this is the generation of display content using foveal rendering technology. Of course, the generation process of such display content may be carried out by other methods. For example, the HMD930 may have an onboard GPU that generates display content using eye tracking and / or head motion data.

[0033] A storage device 913 may be provided to store any suitable information. Examples of such information include program data, display content generation data, and eye-tracking and / or head-tracking model data. This information may also be stored on a remote server. That is, the storage device 913 may be local, remote, or a combination of both.

[0034] Such storage may be used to record the generated display content.

[0035] Input / output 914 may be provided to enable communication suitable for the processing device 910. Examples of such communication include transmitting display content to the HMD 930 and / or display 950, recognizing eye tracking data, head motion data and / or images from the HMD 930 and camera 94, and communicating with one or more remote devices (e.g., via the Internet).

[0036] A peripheral device 920 may be provided. This allows the user to provide input to the processing unit 910 in order to control processing or to interact with the generated display content. The peripheral device 920 may be a button or the like, or it may be via a motion track that enables gestures that can be used as input.

[0037] The HMD930 may be configured similarly to the corresponding elements in Figure 2. The camera 940 and the display 950 may be configured similarly to the corresponding elements in Figure 8.

[0038] Referring to Figure 10, it can be seen that the structure of the human eye is not uniform. In other words, the eye is not a perfect sphere. Different parts of the eye have different characteristics (for example, different refractive indices and colors). Figure 10 is a simplified side view of the structure of a typical eye. For clarity, features such as the muscles that control eye movement have been omitted from this figure.

[0039] The eye 1000 is formed with a nearly spherical structure and filled with an aqueous solution 1010. The retina 1020 is formed in front of the eye 1000. The optic nerve 1030 is connected to the posterior part of the eye 1000. An image is formed on the retina by light incident on the eye 1000. Signals transmitting visual information are sent from the retina 1020 to the brain via the optic nerve 1030.

[0040] Referring to the front of the eye 1000, the sclera 1040 (usually called the white of the eye) surrounds the iris 1050. This iris 1050 controls the size of the pupil 1060, which is the opening when light enters the eye 1000. The iris 1050 and pupil 1060 are covered by the cornea 1070, a transparent layer that refracts light entering the eye 1000. The eye 1000 also has a lens (not shown) located behind the iris 1050. This lens is controlled to adjust the focus of light entering the eye 1000.

[0041] The eye has a region of high visual acuity (the fovea), and visual acuity rapidly decreases towards both sides of this fovea. Figure 11 illustrates this with curve 1100. The peak near the center of Figure 11 corresponds to the foveal region. Region 1110 is the "blind spot." The blind spot is the region where visual acuity is lost. This is because the optic nerve connects to the retina in this region. The peripheral area (i.e., the region where the visual angle is far from the fovea) is not very sensitive to color or detail and is used to detect movement.

[0042] As described above, foveal rendering (or foveal adaptive rendering) is effective in a relatively small area near the fovea (approximately 2.5 to 5 degrees), and visual acuity rapidly deteriorates outside this area.

[0043] Conventional foveal rendering techniques typically require multiple render passes. This is to allow rendering of image frames multiple times at different resolutions. The rendering results are then combined, creating regions of different resolutions within a single image frame. Using multiple render passes introduces significant processing overhead and can result in undesirable image artifacts at the boundaries between regions.

[0044] Alternatively, hardware that can render parts of an image with different resolutions within a single image may be available (so-called flexible-scale rasterization). In this case, no additional render passes are required. If such hardware is available, these hardware-accelerated implementations offer performance advantages.

[0045] Figure 12a is a schematic diagram of foveal rendering for the displayed scene 1200. The user directs their gaze in the direction of the region of interest. As described above, the direction of the gaze is tracked. For clarity, in this example, the direction of the gaze is directed towards the center of the displayed field of view. Therefore, the region 1210, which roughly coincides with the user's high-resolution foveal region, is rendered in high resolution. On the other hand, the peripheral region 1220 is rendered in low resolution. Through gaze tracking, the high-resolution areas of the image are projected onto the foveal region of the user's eye where visual acuity is high, while the low-resolution areas of the image are projected onto the areas of the user's eye where visual acuity is low. By continuously tracking and rendering the user's gaze, the user is created to perceive the entire image as a high-resolution image because the image always appears in the high-resolution portion of the user's own field of view. However, in reality, typically the majority of the image is rendered in low resolution. This significantly reduces the computational overhead required to render the entire image.

[0046] This offers several advantages. Firstly, it allows users to receive richer, more complex, and / or more detailed graphics than before, using the same computing resources. Furthermore, it allows rendering two images (e.g., left and right images of a stereoscopic image displayed on a head-mounted display) instead of a single image (e.g., an image displayed on a television) using the same computing resources. Secondly, it reduces the amount of data transmitted to displays such as HMDs. More selectively, it reduces the computing cost of image preprocessing (e.g., reprojection) on HMDs.

[0047] Referring to Figure 12b, selectively, foveal rendering can vary the resolution between the foveal and peripheral regions of an image in multi-step or stepwise manner. This is due to the smooth decrease in visual acuity from the foveal to the peripheral region of the eye, as shown in Figure 11.

[0048] Therefore, in scene 1200' shown in the modified example, the foveal region 1210 is surrounded by a transitional region 1230 positioned between the foveal region and the reduced peripheral region 1220'.

[0049] The transition region may be rendered at an intermediate resolution between the resolution of the foveal region and the resolution of the peripheral region.

[0050] Refer to Figures 13a and 13b. Alternatively, this may be rendered as a function of distance from the estimated line of sight position. For example, this may be done using a pixel mask and pixels that gradually become sparser with distance. This represents that the corresponding image pixels are rendered first, and the remaining pixels are mixed in according to the nearby rendered colors. Alternatively, this may be done using a flexible-scale rasterization system with a suitable resolution distribution curve. Figure 13a shows a linear transition of resolution. Figure 13b shows a nonlinear transition of resolution, reflecting the nonlinear attenuation of visual acuity as the user moves away from the fovea of ​​the eye. In the second approach, the resolution attenuates faster, so the computational overhead can be reduced more efficiently.

[0051] In this way, eye-tracking is possible (e.g., by using one or more eye-tracking cameras, and subsequently calculating the user's gaze and gaze position on the virtual image). Selectively, foveal rendering may be applied to maintain a high-resolution illusion. In this case, the quality of the resulting image can be improved, at least in the foveal region, while reducing the computational overhead associated with image generation. And / or, when generating two normal images, a second viewpoint can be provided at less than twice the cost (e.g., by generating a pair of stereoscopic images).

[0052] Furthermore, when the HMD is worn, if the gaze region 1210 is the display region of the gaze-based maximum area of ​​interest, the entire rendered scene is the display region of the head position-based general area of ​​interest. That is, the displayed field of view 1200 reflects the position of the user's head when wearing the HMD. In contrast, the foveal rendering within that region reflects the user's gaze position.

[0053] In reality, the peripheral area of ​​the displayed field of view 1200 can be considered a special case where the area is rendered with zero resolution (i.e., not actually rendered), because the user cannot see outside the displayed field of view.

[0054] However, if a second user wants to view the recorded gameplay of the original user while wearing their own HMD, the above does not apply (even if they are viewing the same content as the original user). In the embodiment described below, according to Figure 14a, the principle of foveal rendering can be extended to an area beyond the field of view 1200 displayed to the original user. This means rendering peripheral areas further outside the original user's field of view at an even lower resolution. These lower-resolution areas are usually not visible to the original user (because they are only displayed along with the current field of view 1200). However, these can be rendered as part of the same rendering pipeline using the same technique as foveal rendering within the current field of view.

[0055] In this embodiment, the game console or other rendering source renders a higher set of the displayed image 1200. Selectively, the high-resolution foveal region 1210 is rendered first. Then, selectively, the peripheral region 1220, which is within the user's field of view, is rendered along with the transition region 1230 (not shown in Figure 14a). Subsequently, a further peripheral region 1240, i.e., outside the user's field of view, is rendered. In the context of this specification, “rendering” means generating image data that is displayable (and / or recordable), immediately ready, or output in some visible form.

[0056] This further peripheral region is typically a sphere with the user's head as its virtual center (or, more precisely, a sphere is formed). This further peripheral region is rendered at a lower resolution than the internal region of the user's field of view, which is displayed to the user.

[0057] Selectively referring to Figure 14b, a transition region 1250 may be created around the user's field of view in a similar manner to the transition region shown in Figure 12b. In this case, the resolution of the peripheral region 1220 within the user's field of view is reduced to a lower resolution for the further spherical peripheral region. Again, this may be an intermediate resolution or a linear or nonlinear reduction. The relative size of the transition region may be a matter of design choice or may be determined experimentally. For example, viewers of a recording of the original user who want to track the original user's head movements (typically because the original user is tracking an object of interest or event of interest in the game) may not need to track the displayed field of view completely because their reaction time is limited. Therefore, the size of the transition region may be chosen based on the relative time lag when tracking the field of view displayed to the user as it moves around the periphery of a virtual sphere. This time lag may be a function of the size and velocity of the field of view. Therefore, for example, if the original user moves their head quickly and / or over a long distance, the transition region 1250 is long in time, and its size is a function of velocity and / or distance, and selectively a function of the overall computing resources (in this case, selectively, the resolution of the rest of the additional spherical region may be reduced in time to conserve overall computing resources). Conversely, if the original user's field of view is relatively fixed, the transition region may be relatively small. For example, large enough to accommodate the small head movements of the second user, or large enough to accommodate the different (and possibly larger) field of view of the next head-mounted display (for example, if recording was made using a first-generation head-mounted display with a 110° field of view, the transition region may be expanded to 120° in anticipation of a second-generation head-mounted display with a wider field of view).

[0058] The rendering of the spherical image may be performed within the rendering pipe, for example, as a cubemap, or using other suitable spherical rendering techniques.

[0059] As described above, the original user sees only the displayed field of view 1200. Selectively, the displayed field of view 1200 itself comprises a high-resolution foveal region, a selective transition region, and a peripheral region. Alternatively, where the head-mounted display does not perform eye tracking, the displayed field of view has a predetermined resolution. The rest of the rendered spherical image is not seen by the original user and is rendered at a lower resolution. Selectively, a transition region exists between the displayed field of view and the rest of the sphere.

[0060] Therefore, in this scheme, the displayed field of view can be considered a head-based foveal rendering scheme, rather than a gaze-based foveal rendering scheme. In this scheme, when the user moves their head, the relatively high-resolution displayed field of view moves around the periphery of the entire rendered sphere. On the other hand, selectively, when the user moves their gaze at the same time, a higher-resolution area moves around within the displayed field of view. The original user sees only the displayed field of view. However, viewers who subsequently watch a recording of the rendered image can potentially access the entire sphere, regardless of the original user's field of view within the sphere.

[0061] Therefore, while viewers typically try to track the original user's field of view, if their current field of view differs from the original user's, they may look at other parts of the spherical image to enjoy the periphery, look at areas the original user was not interested in, or simply gain a greater sense of immersion.

[0062] In the same way that conventional images are recorded in a game console's circular buffer, for example, the entire image (a spherical upper set of images displayed to the original user) may be recorded in a circular buffer. The game console's hard disk, solid disk, and / or RAM may be used to record footage of 1 minute, 5 minutes, 15 minutes, 30 minutes, or 60 minutes of the entire image. Unless the user specifically wishes to save / archive the recorded material (individual files may be copied to the hard disk or solid disk, or uploaded to a server if desired), the oldest footage may be overwritten with new footage. Similarly, the entire image may be uploaded to a distribution server and streamed live, streamed or uploaded from the circular buffer, or uploaded to a distribution server or VOD server and streamed later.

[0063] As a result, when the original user wearing the HMD moves their head, a spherical image is generated in which the field of view displayed to that original user becomes a high-resolution area. Selectively, within this high-resolution area, an even higher-resolution area corresponding to the line of sight position within the field of view is generated.

[0064] Selectively, metadata may be recorded along with the spherical image. The metadata may be part of the video recording or in an associated file. The metadata indicates where the displayed field of view is located within the spherical image. This may be used, for example, to assist a second user if they become confused or lose track of the original user's field of view (for example, if the original user shoots a spaceship out of their field of view while watching a space battle, the second user will lose their viewpoint to track the spaceship and the original user's field of view). In this case, navigation tools such as an arrow indicating the current direction of the original user's displayed field of view, or a bright spot at the edge of the periphery of the second user's field of view, would be helpful in guiding them back to the highest resolution area in the recorded image.

[0065] In this way, even if the second user changes their gaze and moves to a different location, they can be sure to return to the original user's field of view.

[0066] A second user might look around the scene if another event occurs or if other objects exist within the virtual environment. These might be of no interest to the original user, but more interesting to the second user.

[0067] Therefore, selectively, a game console (or game or other application) may maintain lists, tables, or other relevant data. This data may represent the level of interest in specific objects (such as non-player characters) or environmental factors, and / or the level of interest in specific events (such as the appearance of objects or characters, or explosions), or some script events tagged as of high interest.

[0068] In such cases, if such objects or events occur within a spherical image outside the field of view displayed to the original user, the region of the sphere corresponding to such objects or events may be rendered at a relatively high resolution (e.g., a resolution corresponding to the middle of the transition region 1250 or the peripheral region 1220 that was initially displayed). Selectively, to conserve overall computing resources, other parts of the spherical image may be rendered at a lower resolution. Selectively, the resolution may be increased depending on the level of interest of the object or event (e.g., resolution increases of 0, 1, and 2 for objects or events of 0, low, and high interest, respectively).

[0069] These objects or events may have a transition region similar to the transition region 1230 or 1250 around them, allowing the image to transition smoothly into the surrounding area. This allows a second user to see objects or events that the original user did not see. The resolution in this case is higher than the resolution of the less interesting parts of the spherical image.

[0070] Consider cases where the principle of foveal rendering is extended beyond the original user's field of view, adding areas that generate further peripheral or spherical regions (or annular or cylindrical regions), or where foveal rendering is not actually used (e.g., because eye tracking is not present) and the principle of foveal rendering is applied outside the original user's field of view. In such cases, the above scheme may optionally be started or stopped by one or more users, the operating system of the application game console that generates the rendered environment, or a helper application (e.g., an application for distribution / streaming or uploading).

[0071] For example, the above scheme may be turned off by default because it incurs computing overhead and is unnecessary when the running game is not being streamed or distributed. Therefore, the above scheme may be given to the user as an option, such as being turned on when the game is being streamed or distributed, or being turned on in response to a command to start streaming or uploading.

[0072] Consider a scenario where viewers want to see a different viewpoint than the original user. In this case, too, the application generating the game or rendered environment may trigger the aforementioned event, for example, in response to a game event or a specific level or cutscene.

[0073] [frame rate] The above scheme increases computational overhead because it requires rendering more scenes, even lower-resolution scenes within the view displayed to the original user.

[0074] To mitigate this, portions rendered outside the field of view displayed to the original user (or selectively outside the transition region 1250, which is the boundary with the field of view displayed to the original user) may be rendered at a lower frame rate than portions rendered within the field of view (or selectively within the transition region).

[0075] Therefore, for example, the field of view may be rendered at 60 frames per second (fps), the rest of the sphere at 30 fps, and selectively, rendering at a higher resolution than 60 fps is possible if computing resources allow.

[0076] Selectively, to restore a frame rate of 60fps, the upload server of the recorded image may insert the remaining frames of the spherical image.

[0077] More generally, the rest of the sphere (including selectively the transitional portion around the original user's field of view) is rendered at a fraction of the frame rate of the field of view displayed to the original user (typically 1 / 2 or 1 / 4). This portion of the image is then frame-inserted by the game console or the server to which the recorded image is transmitted.

[0078] [Upscale] As an alternative to or addition to temporal / frame interpolation to compensate for reduced frame rates, spatial upscaling may be used to compensate for reduced image resolution within a sphere. This may be achieved through offline processing (e.g., on the game console or server mentioned above) or using the client device of the next user of the content.

[0079] Suitable upscaling techniques are known and include bilinear and bicubic interpolation algorithms, sinc and Lanczos resampling algorithms, etc.

[0080] Alternatively or additionally, machine learning (e.g., neural) rendering or inpainting techniques (e.g., convolutional neural networks trained to upscale images) may be used. In this embodiment, a machine learning system can be trained to upscale an image using the resolution difference between the foveal region (or visual field region) and a lower resolution (depending on the application, the resolution of the peripheral region, further peripheral region, or transition region). Selectively, each machine learning system can be trained for its respective upscaling rate.

[0081] Such a machine learning system is trained using a target image with full resolution and input images with reduced resolution (e.g., an image generated by downscaling the target image, or a target image re-rendered at a lower resolution / quality). In embodiments of the present invention, the training set may include a rendered target image (corresponding to the foveal region, or the visual field region if the foveal region is absent) and a corresponding input image (corresponding to one or more other regions). Typically, the machine learning system is not trained on the entire image, but on fixed-size tiles extracted from the image. For example, the tiles may be 16x16 pixels, 32x32 pixels, or 64x64 pixels. The target may be a corresponding tile of the same size, but the target represents a higher-resolution image. Therefore, this target tile may correspond only to a subset of the image found within the input tile. For example, if the input resolution is 640x480 and the target resolution is 1920x1080, a 32x32 input tile corresponds to an image region in the image that is approximately 6.75 times the size of the 32x32 output tile. This allows the machine learning system to use the surrounding pixels of the input image. This can contribute to upscaling the portion corresponding to the input tile by using information from repeating patterns or textures in the input, or the slope or curve of chrominance or brightness can contribute to a better evaluation.

[0082] The output tile does not have to be the same size as the input tile; it may be enlarged to a size equivalent to the image area of ​​the input tile. On the other hand, the input tile may represent any portion of the image (up to the entire image), as long as the machine learning system (and the equipment on which the machine learning system is used) allows.

[0083] Using surrounding pixels of an input image tile contributes to the upscaling of the portion corresponding to the output tile, but is not limited to this; it may also be used when upscaling using the techniques described above, and is not limited to machine learning.

[0084] While training images can be any images, machine learning systems perform better when trained on the same game (and / or past games in a series with the same look) as the upscaled footage.

[0085] Selectively, these arbitrary interpolation techniques may use additional information from other image frames (e.g., past and / or future image frames). This additional information determines other complementary information.

[0086] In some embodiments of this application, when the viewpoint moves around the scene, image information from the original field of view may be required. This may provide higher-resolution reference pixels, which may substantially replace the processing of portions rendered at lower resolutions. For example, when the user's head moves to the left, the current central portion of the scene may pan to the right and be rendered at a lower resolution. However, the high-resolution data for that portion of the scene is obtained from an earlier frame, when it was at the center of the field of view.

[0087] Selectively, frames may include metadata indicating the direction of the center of the field of view. When upscaling peripheral or further peripheral regions of a frame, the system may determine whether (when) these regions were last at the center of the field of view and may retrieve high-resolution pixels from the frame.

[0088] Alternatively or additionally, the system may generate a spherical reference image using pixel data provided from the last frame (in which pixels are rendered at high resolution). In this case, the foveal field of view is treated like a brush, leaving a trail of high-resolution pixels from its trailing edge in each frame. As the user looks around the environment, this brush paints a high-resolution image of the current field of view. The peripheral area (the field of view if no foveal area exists) can also be treated like a brush (its values ​​are prioritized by the foveal pixels). This allows the largest surface area of ​​the sphere to be painted with these higher-resolution pixels. The same approach can be used for any further transitional areas (if any). In summary, the most recent highest-resolution pixel value with respect to a given location on the reference sphere is stored and updated each time the user looks around. These values ​​can also be deleted, for example, after a predetermined amount of time has elapsed (or if the user moves more than a predetermined amount, or if the game environment changes more than a predetermined degree).

[0089] These pixels may then be used directly to fill pixels for the current upscaling of peripheral or further peripheral regions, or they may be used as additional input data for any of the techniques described above. For example, when upscaling peripheral and further peripheral regions of the current frame, the spherical reference image may include (e.g.) 40% high-resolution pixels of the sphere, because if the user has recently turned backward, 40% of the spherical field of view will be within the foveal resolution (or field of view resolution) region over 20 or 30 consecutive frames. Thus, the upscaler can use high-resolution data (e.g., of a size equivalent to, or somewhat larger than, a high-resolution target tile) as input, in conjunction with the low-resolution data of the current frame being upscaled.

[0090] Typically, a neural network trained on both current low-resolution and the associated high-resolution inputs will perform better on high-resolution targets. In this situation, the neural network may be trained on inputs with multiple resolutions (e.g., foveal, peripheral, and further peripheral resolutions) to accommodate the case where the user's gaze is distributed relatively randomly (this determines which parts of the reference spherical image are filled with higher-resolution information). As an improvement on this approach, the probabilities of the user's gaze direction during gameplay can be estimated, and the neural network can be trained using resolutions selected at frequencies corresponding to these probabilities. For example, since it is rare for a user to look directly behind them, inputs at such times are selected as having the lowest resolution during training (but different from the current input because they are made from older frame data and are still complementary). On the other hand, the left and right of the forward field of view are more likely to yield high-quality data and are selected as having the highest resolution during training.

[0091] Alternatively or additionally, selectively, a machine learning system may be trained to upscale videos having walkthroughs of game environments trained at lower and higher resolutions. This game environment may be created by a developer who has, for example, experienced the environment and rendered a screen sphere at the target resolution (this may be independent of the resulting frame rate / elapsed time, as these are not for the purpose of gameplay). Thus, the machine learning system is trained particularly on unsolved games and uses complete target and input data (complete resolution information for the complete sphere, and their downsampled versions, or lower-resolution renderings generated using a script that generates the same in-game progression for both versions of the video, for example). Again typically, these are presented to the upscaler in a tiled format.

[0092] Alternative strategies may be used to improve the reliability of the upscaling process. For example, when rendering a sphere using a cubemap, each machine learning system may be trained on each facet of the cubemap. This would train the systems specifically for the front, back, top, bottom, left, and right views within the sphere. This allows the machine learning systems to be tuned to match the typical resolution data and content obtained (e.g., different for top and bottom). Selectively, the machine learning systems for top and back (assuming the reliability of these parts of the sphere is less critical than for other parts) may be smaller and simpler.

[0093] In principle, recorded video containing the remaining sphere has a spatially and / or temporally reduced resolution. Therefore, in order to interpolate and / or upscale the frames, these resolutions are compensated, at least in part, by parallel and / or subsequent processing by the game console and / or storage / distribution server.

[0094] The server may then deliver the (spatially and / or temporally) upscaled video image (or, if the above modifications are not applied, the originally upscaled video image) to one or more viewers (or further servers with such capabilities).

[0095] Viewers can then watch the video using an application on their client device. Alternatively, they can track the original user's viewpoint or freely look around the scene. In this case, the resolution outside the original user's field of view / foveal region is improved compared to the original recording.

[0096] [Summary of the Embodiment] Referring to Figure 15, the video processing method according to an embodiment of the present disclosure is a method for processing annular panoramic recorded video comprising an original field of view ("FoV") having a first resolution and a further peripheral region outside the original field of view having a second resolution lower than the first resolution. The method includes step S1510 of spatially upscaling the further peripheral region to a resolution higher than the second resolution. As described above, the upscaled resolution may be the resolution of the transition region, the original field of view (FoV), or the foveal region. On the other hand, in view of the purpose of the original field of view (FoV) or the foveal region, it is desirable to use a lower resolution (e.g., the resolution of the transition region) or to base it on the visual heatmap of the original user or a viewer of similar material before, especially in areas of relatively little interest to the user (e.g., the sky in most games).

[0097] It will be apparent to those skilled in the art that various aspects of the methods corresponding to the operation of the embodiments of the apparatus described herein and in the claims are within the scope of the present invention. These methods include the following: -In one embodiment, the original field of view region comprises a foveal region having a third resolution higher than a first resolution. The method includes the step of spatially upscaling the original field of view region to a resolution substantially equal to the third resolution. -In one embodiment, the annular panoramic recorded video comprises a first transition region between the foveal region and the original field of view region, and a second transition region between the original field of view region and a further peripheral region. The first transition region has a resolution intermediate between a third resolution and the first resolution. The second transition region has a resolution intermediate between the first resolution and the second resolution. -In one embodiment, the spatial upscaling step is performed by a machine learning system, which is trained with input image data at a lower input resolution within the recording resolution and with corresponding target image data at a higher input resolution within the recording resolution. The upscaled resolution is one or more of the resolutions of the transition region, the original field of view (FoV), or the foveal region. -In one embodiment, the method includes the steps of storing the locations of at least a subset of image data having a resolution higher than a second resolution within each frame with respect to a predetermined number of preceding frames, and when upscaling a predetermined portion of the current frame of an annular panoramic recording video, using image data of one or more preceding frames having a higher resolution at the location of the predetermined portion of the current frame as input. -In a similar embodiment, the method includes the steps of: storing the positions of image data having the third resolution within each frame with respect to a predetermined number of preceding frames, wherein the original field of view region comprises a foveal region having a third resolution higher than a first resolution; and using image data from one or more preceding frames having the third resolution at the positions of the predetermined portion of the current frame as input when upscaling a predetermined portion of the current frame of an annular panoramic recorded video. -In one embodiment, the method includes the steps of generating a reference annular panoramic image using at least a subset of image data having a resolution higher than a second resolution in each of a predetermined number of preceding frames, and using image data from a corresponding portion of the reference annular panoramic image as input when upscaling a predetermined portion of the current frame of an annular panoramic recorded video. The annular panoramic image stores the most recently rendered higher resolution pixels in each direction on the reference annular panoramic image (selectively, using the most recent second resolution data if no other data is available). -In this case, selectively, pixel data relating to higher resolution regions of a given image frame is stored in preference to pixel data relating to lower resolution regions, based on the reference annular panoramic image. -Similarly in this example, the selective, spatially upscaling step is performed by a machine learning system, which is trained with input image data and corresponding input data from a reference annular panoramic image at lower input resolutions within the recording resolution, and then trained with corresponding target image data at higher input resolutions within the recording resolution. -In one embodiment, an annular panoramic image is rendered using a cubemap, and the step of spatially upscaling is performed by multiple machine learning systems trained on one or more facets of the cubemap. -In one embodiment, the annular panoramic image is cylindrical or spherical.

[0098] Referring to Figure 16, one embodiment of the present disclosure is a video output method which includes the following:

[0099] The first step S1610 is to obtain a spatially upscaled annular panoramic recording video according to the method described above. This video may be obtained from the device performing the upscaling, from the server to which the video was uploaded, or alternatively, by performing the upscaling (e.g., on a distribution server or client device).

[0100] A second step S1620 outputs a circular panoramic recorded video for display to the user. Typically, this is output to a port of the video signal source 80 (e.g., the user's client device) for viewing by an HMD (or potentially the client device itself, or mounted on the HMD frame, if the client device is a mobile phone or handheld console).

[0101] In one embodiment, the annular panoramic video recording selectively includes the original field of view for each frame. If the user's field of view moves a predetermined amount outside the original field of view during playback, a visual indicator is displayed showing where the original field of view is located in the annular panoramic video recording (e.g., an arrow pointing to the viewpoint or a bright spot at the edge of the periphery of the current image).

[0102] It will be understood that the above methods can be executed using ordinary hardware to which suitable software instructions can be applied, or (in addition to or instead of) dedicated hardware.

[0103] Implementation using existing parts of a typical equivalent device is possible in the form of a computer program product with a processor capable of executing instructions recorded on a non-temporary computer-readable medium (e.g., floppy disk®, optical disk, hard disk, solid disk, PROM, RAM, flash memory, or a combination thereof), or it can be done using hardware (e.g., ASIC (application specific integrated circuit), FPGA (field programmable gate array), or other configurable circuits suitable for a typical device). Such computer programs may be transmitted via data signals over a network (e.g., Ethernet®, wireless network, internet, or a preferred combination thereof).

[0104] In the summary of this disclosure, the video processing system (for example, processing system 910, i.e., a video game console such as PlayStation 5®, typically combined with a head-mounted display 810) is, A video processor for performing spatial upscaling of an annular panoramic recorded video comprising an original field of view having a first resolution and a further peripheral region outside the original field of view having a second resolution lower than the first resolution, wherein the video processor comprises a spatial upscaling processor for spatially upscaling the further peripheral region to a resolution higher than the second resolution.

[0105] It will be apparent to those skilled in the art that the embodiments of the video processing systems described herein and in the claims relating to the methods and techniques described herein fall within the scope of the present invention.

[0106] Similarly, a video processing system (e.g., video processing system 910, a video game console such as PlayStation 5®, typically combined with a head-mounted display 810) comprises a playback processor (e.g., GPU 911 and / or CPU 912) that acquires spatially upscaled annular panoramic recorded video according to the method described above (e.g., by preferred software instructions), and a display processor (e.g., GPU 911 and / or CPU 912) that outputs the video for display to the user (e.g., by preferred software instructions).

[0107] Again, it will be apparent to those skilled in the art that the embodiments of the video processing systems described herein and in the claims, corresponding to the methods and techniques described herein and in the claims, are within the scope of the present invention.

[0108] The above discussion merely discloses and describes examples of embodiments of the present invention. Those skilled in the art will understand that the present invention can be realized in other specific forms without departing from the spirit and essential features of the invention. Accordingly, the disclosure of the present invention is for illustrative purposes only and is not intended to limit the scope of the invention or the claims. This disclosure, including any identifiable modifications of the above teachings, partially defines the scope of the terms of the claims. The subject matter of the invention is not dedicated to the public.

Claims

1. A video processing method for processing an annular panoramic recorded video comprising an original field of view having a first resolution and a further peripheral region outside the original field of view having a second resolution lower than the first resolution, The original field of view region comprises a foveal region having a third resolution higher than the first resolution, The steps include spatially upscaling the further peripheral region to a resolution higher than the second resolution, A method characterized by comprising the step of spatially upscaling the original field of view to a resolution substantially equal to the third resolution.

2. The method according to claim 1, characterized in that the spatial upscaling step is a step of upscaling the further peripheral region to a resolution substantially equal to the first resolution.

3. The annular panoramic video recording comprises a first transition region between the foveal region and the original field of view region, and a second transition region between the original field of view region and the further peripheral region. The first transition region has a resolution intermediate between the third resolution and the first resolution. The method according to claim 1, characterized in that the second transition region has a resolution intermediate between the first resolution and the second resolution.

4. The aforementioned spatial upscaling step is performed by a machine learning system. The method according to claim 1, characterized in that the machine learning system is trained with input image data at a lower input resolution within the recording resolution and with corresponding target image data at a higher input resolution within the recording resolution.

5. With respect to a predetermined number of preceding frames, the steps include storing the positions of at least a subset of image data having a resolution higher than the second resolution within each frame, The method according to claim 1, characterized in that, when upscaling a predetermined portion of the current frame of the annular panoramic recorded video, the method includes the step of using image data of one or more preceding frames having a higher resolution at the position of the predetermined portion of the current frame as input.

6. The original field of view region comprises a foveal region having a third resolution higher than the first resolution, A step of storing the position of the image data having the third resolution within each frame with respect to a predetermined number of preceding frames, The method according to claim 1, characterized in that, when upscaling a predetermined portion of the current frame of the annular panoramic video recording, the method includes the step of using image data of one or more preceding frames having the third resolution at the position of the predetermined portion of the current frame as input.

7. A step of generating a reference annular panoramic image using at least a subset of image data having a resolution higher than the second resolution in each of a predetermined number of preceding frames, The step of upscaling a predetermined portion of the current frame of the aforementioned annular panoramic recorded video, using image data from a corresponding portion of a reference annular panoramic image as input, The method according to claim 1, characterized in that the annular panoramic image stores pixels with higher resolution that have been most recently rendered in each direction on the reference annular panoramic image.

8. The method according to 7, characterized in that, with respect to the reference annular panoramic image, pixel data relating to a higher resolution region of a predetermined image frame is stored in preference to pixel data relating to a lower resolution region.

9. The aforementioned spatial upscaling step is performed by a machine learning system. The method according to 7, characterized in that the machine learning system is trained with input image data and corresponding input data from the reference annular panoramic image at a lower input resolution within the recording resolution, and is trained with corresponding target image data at a higher input resolution within the recording resolution.

10. The aforementioned annular panoramic video is rendered using a cubemap. The method according to claim 1, characterized in that the spatial upscaling step is performed by a plurality of machine learning systems trained on one or more facets of the cubemap.

11. The method according to claim 1, characterized in that the annular panoramic video recording is cylindrical or spherical.

12. A video output method, The steps of obtaining a spatially upscaled annular panoramic video recording according to the method of claim 1, A method characterized by comprising the step of outputting the annular panoramic recorded video for display to a user.

13. The aforementioned annular panoramic video recording includes the original field of view area for each frame, The method according to 12, characterized in that if the user's field of view deviates by a predetermined amount from the original field of view during playback, a visual display indicating where the original field of view is located in the annular panoramic recorded video is displayed.

14. A computer program characterized by causing a computer to execute the method described in claim 1.

15. A video processor for performing spatial upscaling of an annular panoramic recorded video comprising an original field of view having a first resolution and a further peripheral region outside the original field of view having a second resolution lower than the first resolution, The original field of view region comprises a foveal region having a third resolution higher than the first resolution, The further peripheral region is spatially upscaled to a resolution higher than the second resolution, A video processor further comprising a spatial upscaling processor that spatially upscales the original field of view to a resolution substantially equal to the third resolution.

16. A playback processor that acquires a spatially upscaled annular panoramic recorded video according to the method of claim 1, A video playback device comprising a graphics processor that outputs the annular panoramic recorded video for display to a user.

17. The method according to claim 1, characterized in that the step of substantially upscaling the further peripheral region to the third resolution includes the step of spatially upscaling the further peripheral region to a resolution higher than the second resolution.

Citation Information

Patent Citations

  • Video streaming method

    JP2016165105A

  • Image transmission / reception system, image transmission device, image reception device, image transmission / reception method and program

    WO2020179205A1