Video recording and playback system and method
Patent Information
- Application Number
- JP2022108929
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-07-16
- Filing Date
- 2022-07-06
- Publication Date
- 2025-06-23
- Estimated Expiration
- 2042-07-06
AI Technical Summary
Traditional video game streaming systems fail to account for the passive nature of viewer experience in VR or AR games, where the viewer's point of view is determined by the broadcaster's head or eye movements, leading to frustration and discomfort.
A method and system that uses head and eye tracking to dynamically adjust the video playback based on the viewer's movements, employing foveal rendering techniques to maintain high resolution in the user's gaze area while rendering peripheral areas at lower resolution, allowing for a more immersive experience.
Enables viewers to freely explore the virtual environment while maintaining high resolution in their line of sight, reducing computational overhead and providing a more engaging and immersive experience compared to conventional systems.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] TECHNICAL FIELD This disclosure relates to video recording and playback systems and methods. [Background technology]
[0002] Traditional video game streaming systems such as Twitch® and video hosting platforms such as YouTube® and Facebook® have enabled video game players to broadcast their gameplay to a wide audience.
[0003] The key difference between playing video games and watching video recordings of these gameplays is the passive nature of the experience, both in terms of the decisions made during the game and the player's perspective (which is determined, for example, by the player's input).
[0004] This latter issue is exacerbated when the game is a VR or AR game, where players typically base their viewpoint, at least in part, on their own head or eye movements. Thus, when watching a live or recorded stream of such a VR or AR game, the recorded image will likely track the head and / or eye movements of the streamer, not the viewer. This can be nauseating for viewers and can be frustrating for viewers who want to look in a different direction from the streamer. Summary of the Invention [Problem to be solved by the invention]
[0005] The present disclosure aims to alleviate or mitigate these problems. [Means for solving the problem]
[0006] Various aspects and features of the present invention are defined within the context of the accompanying claims and specification, and include in at least a first aspect a method for recording video, in another aspect a method for distributing video recordings, in yet another aspect a method for viewing video recordings, in yet another aspect a video recording system, and in yet another aspect a video playback system.
[0007] It is to be understood that both the foregoing general description and the following detailed description are exemplary, but are not restrictive, of the invention. [Brief explanation of the drawings]
[0008] A more complete understanding of the present disclosure and its many advantages will be obtained by reading the following detailed description in conjunction with the accompanying drawings. [Figure 1] FIG. 1 is a schematic diagram of an HMD worn by a user. [Figure 2] FIG. 2 is a schematic plan view of an HMD. [Figure 3] FIG. 1 is a schematic diagram showing the formation of a virtual image by an HMD. [Figure 4] Schematic diagram of another type of display used in an HMD. [Figure 5] FIG. 1 is a schematic diagram of a stereo image pair. [Figure 6a] FIG. 2 is a schematic plan view of an HMD. [Figure 6b] FIG. 1 is a schematic diagram of a near-eye tracking configuration. [Figure 7] FIG. 1 is a schematic diagram of a remote tracking configuration. [Figure 8] FIG. 1 is a schematic diagram of an eye-tracking environment. [Figure 9] FIG. 1 is a schematic diagram of an eye-tracking system. [Figure 10] 1 is a schematic diagram of the human eye. [Figure 11] 1 is a schematic diagram of a graph of human visual acuity. [Figure 12a] FIG. 1 is a schematic diagram of foveated rendering. [Figure 12b] FIG. 1 is a schematic diagram of foveated rendering. [Figure 13a] FIG. 10 is a schematic diagram showing a change in resolution. [Figure 13b] FIG. 10 is a schematic diagram showing a change in resolution. [Figure 14a] FIG. 2 is a schematic diagram of an enhanced rendering scheme according to an embodiment of the present invention; [Figure 14b] FIG. 2 is a schematic diagram of an enhanced rendering scheme according to an embodiment of the present invention; [Figure 15] 1 is a flow diagram of a video recording method according to an embodiment of the present invention; [Figure 16] 2 is a flow diagram of a video playback method according to an embodiment of the present invention; DETAILED DESCRIPTION OF THE INVENTION
[0009] This specification discloses a video recording and playback system and method thereof. In the following description, some specific details are set forth in order to provide a thorough understanding of embodiments of the present invention. However, it will be apparent to those skilled in the art that the use of these specific details is not essential to the practice of the present invention. Conversely, for the sake of clarity, certain details known to those skilled in the art may be omitted, where necessary.
[0010] In the following, in the drawings, the same or similar components are designated by the same reference numerals. In Fig. 1, a user 10 wears an HMD 20 (which in one example is a conventional head-mountable device, but in other examples may include audio headphones or a head-mountable light source) on the user's head 30. The HMD includes a frame 40 (which in this example is formed by a rear strap and a top strap) and a display portion 50.
[0011] Optionally, the HMD has associated headphone transducers or earpieces 60 that fit over the user's left and right ears 70. The earpieces 60 reproduce an audio signal provided from an external audio source (which may be the same as the video signal source that provides the video signal to the display).
[0012] In operation, a video signal for display is provided by the HMD. This may be provided by an external video signal source 80 (e.g., a data processing device such as a video game console or a personal computer). In that case, the signal may be transmitted to the HMD by a wired or wireless connection 82. An example of a suitable wireless connection includes a Bluetooth® connection. Audio signals for the earpieces 60 may be carried by the same connection. Similarly, any control signals sent from the HMD to the video (audio) signal source may be carried by the same connection. Additionally, a power source 83 (which may include one or more batteries and / or be connected to a mains power outlet) may be connected to the HMD by a cable 84.
[0013] Thus, the configuration of Figure 1 provides an example of a head-mountable display system that includes a frame mounted on a viewer's head and a display element mounted relative to a gaze display position. The frame defines one or two gaze display positions. The gaze display positions are positioned in front of the viewer's eyes during use. The display element presents a virtual image of a video display signal from a video signal source to the viewer's eyes. Figure 1 shows only one example of an HMD; other configurations are possible. For example, an HMD may use a frame similar to traditional eyeglasses.
[0014] In the example of FIG. 1, separate displays are provided for each of the user's eyes. FIG. 2 is a schematic plan view of how this is achieved. FIG. 2 shows the positions of the user's eyes 100 and the relative position of the user's nose 110. The display portion 50 generally comprises an outer shield 120 for blocking ambient light from the user's eyes and an inner shield 130 for preventing one eye from seeing the display seen by the other eye. With respect to the user's face, the outer shield 120 and inner shield 130 form two compartments 140, one for each eye. Within each compartment, a display element 150 and one or more optical elements 160 are provided. FIG. 3 shows the optical paths formed by the display element and optical elements (which provide a display to the user).
[0015] Referring to Figure 3, display element 150 generates a display image. The display image is refracted (in this example) by optical element 160 (shown schematically as a single convex lens, but could also be a compound lens, etc.). As a result, virtual image 170 is generated. To a user, virtual image 170 appears to be larger and much farther away than the real image generated by display element 150. In Figure 3, solid lines (e.g., line 180) represent actual light rays, and dotted lines (e.g., line 190) represent virtual light rays.
[0016] An alternative configuration is shown in Figure 4, where display element 150 and optical element 200 cooperate to provide an image that is projected onto mirror 210. Mirror 210 reflects the image towards the user's eye position 220. The user perceives the virtual image as being in front of the user at a position 230, moderately far away from the user.
[0017] When a user is given a separate display for each eye, stereoscopic images can be displayed. Figure 5 shows an example pair of stereoscopic images for display to each eye.
[0018] When an HMD is used in a virtual reality (VR) system, the user's viewpoint needs to track movement relative to the space in which the user is located.
[0019] Tracking may involve head tracking and / or eye tracking. Head tracking is achieved by detecting movement of the HMD and changing the apparent viewpoint of the displayed image, so that the apparent viewpoint tracks the movement. Movement tracking may involve any suitable configuration, including hardware motion detectors (such as accelerometers or gyroscopes), external cameras capable of capturing images of the HMD, and outward-facing cameras attached to the HMD.
[0020] For eye tracking, two possible configurations are shown in Figures 6a and 6b.
[0021] FIG. 6a shows an example of an eye-tracking configuration. In this configuration, cameras are placed within the HMD to capture images of the user's eyes from a close distance. This is sometimes referred to as near-eye tracking or head-mounted tracking. In this example, the HMD 400 (along with the display element 601) is provided with cameras 610. Each of these cameras is positioned to directly capture one or more respective images. Four cameras 610 are shown as an example of a possible arrangement of eye-tracking cameras. However, one camera per eye is typically desirable. Optionally, only one eye may be tracked if eye movement is typically constant. One or more such cameras may be positioned with a lens 620 in the optical path for capturing images of the eyes. An example of such an arrangement is shown using camera 630. One advantage of including a lens in the optical path is that it simplifies the physical constraints on the HMD design.
[0022] FIG. 6b shows an example of an eye-tracking configuration, in which a camera is positioned to indirectly capture an image of a user's eyes. FIG. 6b includes a mirror 650 positioned between the display 601 and the viewer's eyes. For clarity, this diagram omits all additional optical elements, such as lenses. In such a configuration, the mirror 650 is selected to be partially transparent. That is, the mirror 650 is selected so that the camera 640 can capture an image of the user's eyes when the user looks at the display 601. One way to achieve this is to employ a mirror 650 that reflects light in the IR wavelengths but transmits visible light. In this way, the IR light used for tracking is reflected from the user's eyes toward the camera 640, while the light emitted by the display 601 passes through the mirror without interference. One advantage of such a configuration is that the camera can be easily positioned outside the user's field of view. Furthermore, the accuracy of eye tracking is improved because (thanks to reflection) the camera captures images from a position substantially along the axis between the user's eyes and the display.
[0023] Alternatively, the eye tracking configuration need not be head-mounted or near-eye as described above. For example, Figure 7 is a schematic diagram of a system in which cameras are positioned to capture images of a user from a distance. In Figure 7, an array of cameras 700 is provided, providing multiple images of a user 710. The cameras are positioned to capture information to identify, using a suitable method, at least the direction in which the user's 710 eyes are focused.
[0024] FIG. 8 is a schematic diagram of an environment in which the eye tracking process takes place. In this example, a user 800 is using an HMD 810 associated with a processing unit 830 (e.g., a game console) and a peripheral device 820 to input commands for controlling the processing. The HMD 810 may perform eye tracking according to the configuration illustrated in FIG. 6a or 6b. That is, the HMD 810 may include one or more cameras for capturing images of one or both eyes of the user 800. The processing unit 830 may generate content to be displayed on the HMD 810. However, some (or all) of the display content may be generated by a processing unit within the HMD 810.
[0025] The configuration of Figure 8 includes a camera 840 positioned external to the HMD 810 and a display 850. In some cases, the HMD 810 may be used to determine, for example, body movements and head orientation, and the camera 840 may be used to track the user 800. In an alternative configuration, the camera 840 may be mounted on the HMD facing outward to determine the movement of the HMD based on movements in captured video.
[0026] The processing required to generate tracking information from the captured images of the user's 800 eyes may be performed locally by the HMD 810. Alternatively, the captured images or one or more detection results may be transmitted to an external device (e.g., processing unit 830) for processing. In the former case, the HMD 810 may output the processed results to the external device.
[0027] Figure 9 is a schematic diagram of a system that performs one or more eye tracking and head tracking processes, such as those described in Figure 8. The system 900 includes a processing device 910, one or more peripherals 920, an HMD 930, a camera 940, and a display 950.
[0028] 9, processing device 910 includes one or more central processing units (CPUs) 911, graphics processing units (GPUs) 912, storage (such as a hard drive or any other suitable storage media) 913, and input / output 914. These units may be provided in the form of a personal computer or any other suitable processing device.
[0029] For example, the CPU 911 may be configured to generate the tracking data from one or more input images of the user's eyes obtained from one or more cameras, or from data representing the user's gaze direction. This may be data obtained, for example, from processed images of the user's eyes by a remote device. Of course, if the tracking data is generated elsewhere, the processing device 910 does not need to perform such processing.
[0030] Alternatively or additionally, one or more cameras (other than eye-tracking cameras) may be used to track head movements as described above, or any suitable motion tracker, such as an accelerometer in the HMD, may be used.
[0031] A GPU may be arranged to generate content for display to an eye-tracked or head-tracked user.
[0032] The display content itself may be improved in response to the acquired tracking data. One example is generating the display content using foveated rendering techniques. Of course, the display content generation process may be performed in other ways. For example, the HMD 930 may have an on-board GPU that uses eye tracking and / or head motion data to generate the display content.
[0033] Storage 913 may be provided for storing any suitable information. By way of example, such information may include program data, display content generation data, eye-tracking and / or head-tracking model data. Such information may be stored on a remote server. That is, storage 913 may be local, remote, or a combination thereof.
[0034] Such storage may be used to record the generated display content.
[0035] Input / output 914 may be arranged for appropriate communications with processing device 910. By way of example, such communications may include transmitting display content to HMD 930 and / or display 950, eye tracking data, head motion data and / or recognition of images from HMD 930 or camera 94, and communications with one or more remote devices (e.g., via the internet).
[0036] Peripherals 920 may be provided to allow a user to provide input to the processing unit 910 to control processing or to interact with the generated display content. Peripherals 920 may be buttons or the like, or via a motion track that enables gestures to be used as input.
[0037] The HMD 930 may be configured similarly to the corresponding elements in Figure 2. The camera 940 and display 950 may be configured similarly to the corresponding elements in Figure 8.
[0038] Referring to Figure 10, it can be seen that the structure of the human eye is not uniform; that is, the eye is not a perfect sphere. Different parts of the eye have different characteristics (e.g., different refractive indices and colors). Figure 10 is a simplified side view of the structure of a typical eye 1000. For clarity, this view omits features such as the muscles that control eye movement.
[0039] The eye 1000 is formed with a nearly spherical structure and is filled with an aqueous solution 1010. The retina 1020 is formed at the front of the eye 1000. The optic nerve 1030 connects to the back of the eye 1000. Light entering the eye 1000 forms an image on the retina. Signals conveying visual information are sent from the retina 1020 to the brain via the optic nerve 1030.
[0040] Referring to the front of the eye 1000, the sclera 1040 (commonly called the white of the eye) surrounds the iris 1050. This iris 1050 controls the size of the pupil 1060. The pupil 1060 is the opening through which light enters the eye 1000. The iris 1050 and pupil 1060 are covered by the cornea 1070. The cornea 1070 is a transparent layer that refracts light entering the eye 1000. The eye 1000 also includes a lens (not shown) located behind the iris 1050. This lens is controlled to adjust the focus of the light entering the eye 1000.
[0041] The eye has an area of high visual acuity (the fovea), with visual acuity rapidly decreasing on either side of the fovea. Figure 11 illustrates this with curve 1100. The peak near the center of Figure 11 corresponds to the foveal region. Region 1110 is the "blind spot." This is the area where vision is lost because this is where the optic nerve connects to the retina. The periphery (i.e., the area of vision far away from the fovea) is less sensitive to color and detail and is used to detect motion.
[0042] As mentioned above, foveal rendering (or foveal-adaptive rendering) is effective in a relatively small region near the fovea (approximately 2.5 to 5 degrees), and visual acuity deteriorates rapidly outside this region.
[0043] Conventional foveated rendering techniques typically require multiple render passes, allowing an image frame to be rendered multiple times at different resolutions. The rendered results are then composited to create regions of different resolutions within a single image frame. Using multiple render passes requires significant processing overhead and can introduce undesirable image artifacts at the boundaries between regions.
[0044] Alternatively, hardware may be available that can render parts of different resolutions into a single image (so-called flexible scale rasterization), in which case no additional render passes are required. If such hardware is available, such hardware-accelerated implementations can also be advantageous in terms of performance.
[0045] FIG. 12a is a schematic diagram of foveated rendering for a displayed scene 1200. The user directs their gaze toward a region of interest. As described above, the gaze direction is tracked. For clarity, in this example, the gaze direction is directed toward the center of the displayed field of view. Thus, a region 1210 that roughly corresponds to the user's high-resolution foveal region is rendered at a high resolution, while a peripheral region 1220 is rendered at a low resolution. Through gaze tracking, the high-resolution region of the image is projected onto the foveal region, where the user's eyes have higher visual acuity, while the low-resolution region of the image is projected onto the user's eyes with lower visual acuity. By continuously tracking and rendering the user's gaze, the user is given the illusion that the entire image is a high-resolution image because it always appears in the high-resolution portion of the user's field of view. However, in reality, the majority of the image is typically rendered at a low resolution. This significantly reduces the computational overhead of rendering the entire image.
[0046] This is advantageous in several ways. First, the same computational resources can be used to present the user with richer, more complex, and / or more detailed graphics than previously possible. Furthermore, the same computational resources can be used to render two images (e.g., left and right images of a stereoscopic image displayed on a head-mounted display) instead of a single image (e.g., an image displayed on a television). Second, the amount of data sent to a display such as an HMD can be reduced. Optionally, the computational cost of pre-processing (e.g., re-projection) of images in the HMD can be reduced.
[0047] Referring to Figure 12b, optionally, foveal rendering can change resolution in multiple steps or in stages between the foveal and peripheral regions of the image, due to the smooth decrease in visual acuity from the fovea to the peripheral regions of the eye, as shown in Figure 11.
[0048] Thus, in the scene 1200' displayed in the variant, the foveal region 1210 is surrounded by a transition region 1230 located between the foveal region and the reduced peripheral region 1220'.
[0049] The transition region may be rendered at a resolution intermediate between the resolution of the foveal region and the resolution of the peripheral region.
[0050] See Figures 13a and 13b. Alternatively, this may be rendered as a function of distance from the estimated gaze position. For example, this may be done using pixels that become increasingly sparse with distance and a pixel mask. This means that the corresponding image pixel is rendered first, and the remaining pixels are blended according to the colors of nearby rendered pixels. Alternatively, this may be done by a flexible scale rasterization system using a suitable resolution distribution curve. Figure 13a shows a linear resolution transition. Figure 13b shows a non-linear resolution transition, reflecting the non-linear attenuation of visual acuity as one moves away from the fovea of the user's eye. In the second approach, the resolution decays more quickly, which is more efficient and reduces computational overhead.
[0051] In this way, eye tracking is possible (e.g., by using one or more eye-tracking cameras and then calculating the user's gaze and gaze position on the virtual image). Optionally, foveated rendering may be applied to maintain the illusion of high resolution. This can improve the quality of the resulting image, at least in the foveal region, while reducing the computational overhead associated with image generation. And / or provide a second viewpoint (e.g., generate a stereoscopic image pair) at twice the cost of generating two regular images.
[0052] Furthermore, when wearing an HMD, if the gaze region 1210 is the viewing region of the gaze-based maximum region of interest, then the entire rendered scene is the viewing region of the head-position-based general region of interest. That is, the displayed field of view 1200 reflects the user's head position when wearing the HMD, whereas the foveated rendering within that region reflects the user's gaze position.
[0053] In fact, the peripheral area of the displayed field of view 1200 can be thought of as a special case of an area that is rendered at zero resolution (i.e., not actually rendered) because the user cannot see outside the displayed field of view.
[0054] However, this does not apply if a second user wants to view the recorded gameplay of the original user while wearing their own HMD (even if they are viewing the same content as the original user). In the embodiment described below, in accordance with FIG. 14a, the principles of foveated rendering can be extended to areas beyond the field of view 1200 displayed to the original user. This means that peripheral areas further outside the original user's field of view are rendered at a lower resolution. These lower resolution areas are typically not visible to the original user (because they are only displayed in conjunction with the current field of view 1200). However, they can be rendered as part of the same rendering pipeline using the same techniques as foveated rendering within the current field of view.
[0055] In this embodiment, a game console or other rendering source renders a superset of the displayed image 1200. Optionally, a high-resolution foveal region 1210 is rendered first. Optionally, a peripheral region 1220 within the field of view displayed to the user is then rendered, along with a transition region 1230 (not shown in FIG. 14a). Then, a further peripheral region 1240, i.e., outside the field of view displayed to the user, is rendered. Note that in this context, "rendering" means generating image data that can be displayed (and / or recorded), or that can be immediately prepared, or that can be output in some visible form.
[0056] This further peripheral region is typically a sphere (or more precisely, a sphere is formed) virtually centered at the user's head, and is rendered at a lower resolution than the internal region 1220 of the field of view displayed to the user.
[0057] Referring optionally to FIG. 14b, a transition region 1250 may be created around the periphery of the field of view displayed to the user in a manner similar to the transition region shown in FIG. 12b. In this case, the resolution of the peripheral region 1220 in the field of view displayed to the user is reduced to a lower resolution due to the spherical additional peripheral region. Again, this may be an intermediate resolution or a linear or non-linear drop-off. The relative size of the transition region may be a matter of design choice or may be determined experimentally. For example, a viewer of the original user's recording who wishes to track the original user's head movements (typically because the original user is tracking an object or event of interest in a game) may have limited reaction time and therefore may not need to perfectly track the displayed field of view. Thus, the size of the transition region may be chosen based on the relative time lag in tracking the field of view displayed to the user as it moves around the virtual sphere. This time lag may be a function of the size and speed of the field of view. Thus, for example, if the original user moves their head quickly and / or a long distance, transition region 1250 may be long in time, with its size being a function of speed and / or distance, and optionally also a function of overall computational resources (in which case, optionally, the resolution of the remainder of the further spherical region may be reduced in time to conserve overall computational resources). Conversely, if the original user's field of view is relatively fixed, the transition region may be relatively small, e.g., large enough to accommodate the small head movements of the second user, or large enough to accommodate the different (possibly larger) field of view of the next head-mounted display (e.g., if recorded using a first-generation head-mounted display with a 110° field of view, the transition region may be expanded to 120° in anticipation of a second-generation head-mounted display with a wider field of view).
[0058] The rendering of the spherical image may be done in the rendering pipe, for example as a cube map, or using any other suitable spherical rendering technique.
[0059] As described above, the original user sees only the displayed field of view 1200. Optionally, the displayed field of view 1200 itself comprises a high-resolution foveal region, an optional transition region, and a peripheral region. Alternatively, where the head-mounted display does not provide eye tracking, the displayed field of view has a predetermined resolution. The remainder of the rendered spherical image is not seen by the original user and is rendered at a lower resolution. Optionally, there is a transition region between the displayed field of view and the remainder of the sphere.
[0060] Thus, in this scheme, the displayed field of view can be thought of as a head-based foveated rendering scheme, rather than gaze-based foveated rendering, in which as the user moves their head, a relatively high-resolution displayed field of view moves around the entire rendered sphere, while optionally, as the user simultaneously moves their gaze, higher-resolution regions move around within the displayed field of view. The original user sees only the displayed field of view, but a viewer who subsequently watches a recording of the rendered image potentially has access to the entire sphere, regardless of where they are within the sphere of the original user's field of view.
[0061] Thus, while the viewer will typically attempt to follow the original user's field of view, when their current field of view differs from that of the original user, they may choose to look elsewhere in the spherical image to enjoy their surroundings, to look at areas that the original user was not interested in, or simply to achieve a greater sense of immersion.
[0062] For example, an entire image (a spherical superset of the image originally displayed to the user) may be stored in the circular buffer in the same way that a traditional image is stored in a circular buffer on a gaming console. For example, the hard disk, solid-state disk, and / or RAM of the gaming console may be used to store 1, 5, 15, 30, or 60 minutes of footage of the entire image. Unless the user specifically desires to save / archive the recorded material (in which case individual files may be copied to a hard disk or solid-state disk or uploaded to a server), the oldest footage may be overwritten with new footage. Similarly, the entire image may be uploaded to a distribution server and streamed live, streamed or uploaded from the circular buffer, or uploaded to a distribution or VOD server for later streaming.
[0063] The result is a spherical image in which the field of view displayed to the original user when the original user wearing the HMD moves their head is a high-resolution region, and optionally, within this high-resolution region, an even higher-resolution region is generated that corresponds to the gaze position within the field of view.
[0064] Optionally, metadata may be recorded along with the spherical image. The metadata may be part of the video recording or an associated file. The metadata describes where the displayed field of view is located within the spherical image. This may be used, for example, to assist a second user if he or she becomes disoriented or loses track of the original user's field of view (e.g., while watching a space battle, if the original user shoots a spaceship out of view, the second user will lose sight of the spaceship and the original user's field of view). In this case, navigation tools such as an arrow indicating the current direction of the original user's displayed field of view or a bright spot at the peripheral edge of the second user's field of view may help guide the second user back to the highest resolution area within the recorded image.
[0065] In this way, even if the second user changes their line of sight and moves to another location, they can be sure to return to the original user's field of view.
[0066] The second user may look around the scene if other events occur or if there are other objects in the virtual environment that are of no interest or interest to the original user but are more interesting to the second user.
[0067] Thus, optionally, a game console (or game or other application) may maintain lists, tables, or other related data representing the interest of particular objects (e.g., non-player characters) or environmental factors, and / or the interest of particular events (e.g., object or character appearances, explosions, etc.), or portions of scripted events that have been tagged as being of interest.
[0068] In such cases, if such objects or events occur within the spherical image outside the field of view displayed to the original user, the areas within the sphere corresponding to such objects or events may be rendered at a relatively high resolution (e.g., a resolution corresponding to the middle of the transition region 1250 or the initially displayed peripheral region 1220). Optionally, other portions of the spherical image may be rendered at a lower resolution to conserve overall computational resources. Optionally, the resolution may be increased depending on the interest level of the object or event (e.g., resolution increases of 0, 1, and 2 for objects or events with an interest level of 0, low, and high, respectively).
[0069] Such objects or events may have transition regions around them similar to transition regions 1230 or 1250 to allow the image to transition smoothly to the surrounding area, thereby allowing the second user to see objects or events not seen by the original user, at a higher resolution than the resolution of portions of the spherical image that are of less interest.
[0070] Consider the case where the principles of foveated rendering are extended beyond the original user's field of view to add additional peripheral or spherical regions (or toroidal or cylindrical regions), or where foveated rendering is not actually used (e.g., because there is no eye tracking) and the principles of foveated rendering are applied outside the original user's field of view. In such cases, the above scheme may optionally be activated or deactivated by one or more users, the application generating the rendered environment, the game console's operating system, or a helper application (e.g., an application for broadcasting / streaming or uploading).
[0071] For example, the above scheme may be off by default because it represents a computational overhead that is not needed while the game is running and is not being streamed or distributed, and therefore may be provided to the user as an option to be turned on when the game is being streamed or distributed, or to be turned on in response to an instruction to begin a stream or upload.
[0072] Consider the case where a viewer wishes to see a different gaze direction than the original user. Again, such events may be initiated by the application generating the game or rendered environment, for example, in response to a game event or a particular level or cutscene.
[0073] [Variations] The above scheme increases computational overhead because it requires rendering more of the scene, even if it is a lower resolution scene within the field of view displayed to the original user.
[0074] To mitigate this, portions rendered outside the field of view displayed to the original user (or optionally outside the transition region 1250 that borders the field of view displayed to the original user) may be rendered at a lower frame rate than within the field of view (or optionally the transition region).
[0075] Thus, for example, the field of view may be rendered at 60 frames per second (fps), while the rest of the sphere may be rendered at 30 fps, or optionally at a higher resolution than 60 fps if computational resources allow.
[0076] Optionally, the video upload server may interpolate the remaining frames of the sphere to restore the 60 fps frame rate.
[0077] More generally, the remainder of the sphere (optionally including a transitional portion around the periphery of the original user's field of view) is rendered at a frame rate that is a fraction (typically 1 / 2 or 1 / 4) of the frame rate of the field of view displayed to the original user, and this portion of the image is then frame-inserted by the game console or a server to which the recorded image is sent.
[0078] Alternatively or additionally to independent temporal frame interpolation to compensate for reduced frame rate, independent spatial interpolation may be performed to compensate for reduced image resolution. This may be performed using offline processing (e.g., by the game console or server described above) and (optionally) using information from successive frames and / or information from the original view moving around the scene (to provide higher resolution reference pixels, which may be replaced with processing or information from portions rendered at lower resolution). Optionally, a machine learning system may be trained to scale up the video (specifically, video trained on lower or higher resolution walkthroughs of the game environment, which may be generated, for example, by a developer moving through the environment, or a spherical image may be rendered at a target resolution, where the frame rate / runtime may be independent of the resulting frames, since this is not intended for gameplay).
[0079] In principle, the recorded video containing the remaining part of the sphere may be shortened in time or space, at least partially compensated for by parallel or sequential processing on the game console and / or storage and distribution servers.
[0080] The server can then distribute the scaled-up (temporally and / or spatially) video image (or the original uploaded video image, if the above variant does not apply) to one or more viewers (or to further servers with such functionality).
[0081] This viewer can then watch the video using an application on their client device, or they can follow the original user's viewpoint and look around the scene freely.
[0082] [Outline of the embodiment] Referring to FIG. 15, the method of the embodiment of the present invention generally includes the following steps:
[0083] A first step S1510 renders a field of view of the virtual environment for a first user wearing a head-mounted display at a first resolution, as described above.
[0084] A second step S1520 renders the virtual environment outside the field of view of the first user at a second resolution lower than the first resolution, as described above.
[0085] A third step S1530 outputs the rendered view for display to the first user, as described above.
[0086] A fourth step S1540 is recording the combined rendering as a video for subsequent viewing by a second user, as described above.
[0087] In principle, the third and fourth steps are interchangeable. Further alternatively, the third step may precede or be performed in parallel with the second step. Similarly, the fourth step may be performed sequentially based on the rendering of the first and second steps.
[0088] Those skilled in the art will appreciate that variations of the above methods that are equivalent to the operation of various embodiments of the methods and / or apparatus of the present invention are within the scope of this disclosure. These variations include, but are not limited to: The method includes rendering a foveal region of the virtual environment within the field of view at a third resolution higher than the first resolution based on eye tracking of the first user. The step of rendering an image of the virtual environment outside the field of view of the first user generates a spherical or cylindrical image of the virtual environment. The method includes rendering a transition region outside the first user's field of view, the transition region being a boundary of the rendered field of view and having a resolution intermediate the first resolution and the second resolution. The size of the transition region depends on one or more of the following: the difference in the rendered field of view position between successive image frames; the rate of change of the rendered field of view position between successive image frames; the fluctuation in the rendered field of view position caused by small movements of the first user's head; and the expected value of the fluctuation in position caused by small movements of the second user's head. The method includes rendering objects or events of interest in the rendered image of the virtual environment that are outside the field of view of the first user at a resolution higher than the second resolution, while maintaining the second resolution for the remainder of the rendered image of the virtual environment that is outside the field of view of the first user. - The step of rendering an image of the virtual environment outside the field of view of the first user depends on one or more items: the first user, an application that renders the virtual environment, an operating system of the device that renders the virtual environment, and a helper application of the operating system of the device that renders the virtual environment. The method includes the steps of: acquiring a recorded video including a first user's field of view of a virtual environment on a head-mounted display rendered at a first resolution and an image of the virtual environment outside the first user's field of view rendered at a second resolution lower than the first resolution; and increasing either or both the spatial resolution in recording the image of the virtual environment outside the first user's field of view and the temporal resolution in recording the image of the virtual environment outside the first user's field of view.
[0089] On the other hand, a method of delivering recorded video using the method described in the summary of the embodiments, wherein the recorded video includes a first user's field of view of a virtual environment in a head-mounted display rendered at a first resolution and an image of the virtual environment outside the first user's field of view rendered at a second resolution lower than the first resolution, includes steps of receiving a request to download or stream the recorded video to a second user, and downloading or streaming the recorded video to the second user.
[0090] The method optionally includes increasing either or both the spatial resolution in recording images of the virtual environment outside the first user's field of view and the temporal resolution in recording images of the virtual environment outside the first user's field of view prior to distribution.
[0091] Referring to FIG. 16, a method for viewing a video recording (including a first user's field of view of a virtual environment in a head-mounted display rendered at a first resolution and an image of the virtual environment outside the first user's field of view rendered at a second resolution lower than the first resolution) described in the summary of an embodiment of the present invention includes the following steps:
[0092] The first step S1610 requests a download or stream of video from a remote source (eg, a distribution or streaming server, or potentially a peer-to-peer server or game console).
[0093] A second step S1620 receives a video download or stream from a remote source.
[0094] A third step S1630 outputs at least a portion of the stream or video for display to a second user wearing a head-mounted display.
[0095] Here, the step of outputting the video stream includes the following substeps.
[0096] The first substep S1632 detects the field of view of the second user wearing the head-mounted display, as described above.
[0097] A second substep S1634 provides the corresponding portion of the recorded stream or video for output to the second user's head-mounted display, as described above.
[0098] Again, those skilled in the art will appreciate that variations of the above methods that are equivalent to the operation of various embodiments of the methods and / or apparatus of the present invention are within the scope of this disclosure. These variations include, but are not limited to: -The method includes a step of calculating a relative position of the second user's current field of view to the first user's field of view based on data related to the recorded video, and if the relative position deviates by more than a predetermined threshold, a step of calculating a correction direction for the second user to move their field of view towards the corresponding field of view of the first user, and a step of displaying the correction direction in the field of view shown to the second user.
[0099] It will be appreciated that the above methods may be implemented using conventional hardware or (in addition to or instead of) specialized hardware to which suitable software instructions can be applied.
[0100] Implementation using existing parts of conventional equivalent devices can be in the form of a computer program product having a processor capable of executing instructions recorded on a non-transitory computer-readable medium (e.g., floppy disk, optical disk, hard disk, solid state disk, PROM, RAM, flash memory, or a combination of these recording media), or can be in hardware (e.g., ASIC (application specific integrated circuit), FPGA (field programmable gate array), or other configurable circuitry suitable for conventional devices). Such a computer program can be transmitted via a data signal over a network (e.g., Ethernet, a wireless network, the Internet, or a suitable combination of these networks).
[0101] In summary of this disclosure, a video recording system 910 (e.g., a video game console such as a PlayStation 5®, typically combined with a head-mounted display 810) includes a rendering processor (e.g., a GPU 912 and / or a CPU 911) configured to render (e.g., by suitable software instructions) a first user's field of view of a virtual environment on the head-mounted display at a first resolution. The rendering processor is also configured to render (again, by suitable software instructions, for example) images of the virtual environment outside the first user's field of view at a second resolution lower than the first resolution. These two steps of rendering may be performed sequentially or in parallel, or may be part of a single process that dynamically changes resolution during rendering (e.g., flexible scale rasterization). An output processor (e.g., a GPU 912, a CPU 911, and / or an input / output 914) is configured to output (again, by suitable software instructions, for example) the rendered field of view for display to the first user. Further, as a video that can be subsequently viewed by a second user, a storage unit (not shown, e.g., a hard disk, solid state drive or ROM, typically associated with the CPU 911 and / or GPU 912) is configured to record (again, e.g., by suitable software instructions) the combined rendering as a video for subsequent viewing by the second user.
[0102] Summary examples of these embodiments that implement the methods and techniques herein (e.g., by suitable software instructions) include, but are not limited to, systems within the scope of this application that include a server configured to use suitable software instructions to increase the spatial and / or temporal resolution of renderings prior to distribution.
[0103] Similarly, a video playback system 910 (e.g., a video game console such as a PlayStation 5®, typically combined with a head-mounted display 810) is adapted to play back recorded video (including a first user's field of view of a virtual environment in the head-mounted display rendered at a first resolution and an image of the virtual environment outside the first user's field of view rendered at a second resolution lower than the first resolution). The video playback system includes:
[0104] A transmitter (e.g., input / output 914, optionally in conjunction with CPU 911) configured to send (e.g., by suitable software instructions) a request to download or stream video from a remote source (e.g., a distribution or streaming server hosting or relaying the video).
[0105] A receiver configured (eg, by suitable software instructions) to receive a video download or stream from a remote source.
[0106] A graphics processor (e.g., GPU 912 and / or CPU 911) configured to output (e.g., by suitable software instructions) at least a portion of the stream or video for display to a second user wearing a head-mounted display.
[0107] The graphics processor is further configured (e.g., by suitable software instructions) to detect the field of view of a second user wearing a head-mounted display (e.g., using any head movement tracking technology) and provide a corresponding portion of the recorded stream or video for output to the second user's head-mounted display.
[0108] Summary examples of these embodiments that implement (e.g., by suitable software instructions) the methods and techniques herein include, but are not limited to, graphics processors (e.g., GPU 912 and / or CPU 911) within the scope of this application. Such graphics processors are configured to calculate a relative position of the second user's current field of view to the first user's field of view based on data associated with the recorded video (e.g., metadata within the recorded video or associated file that defines the original first user's field of view). The graphics processor is further configured to calculate a corrective direction for the second user to move their field of view toward the corresponding first user's field of view if the relative positions of these fields of view deviate by more than a predetermined threshold, and to display the corrective direction within the current field of view shown to the second user.
[0109] The foregoing discussion discloses and describes merely exemplary embodiments of the present invention. Those skilled in the art will recognize that the present invention can be embodied in other specific forms without departing from the spirit or essential characteristics thereof. Accordingly, the disclosure of the present invention is intended to be illustrative and not limiting of the scope of the present invention and the claims that follow. The present disclosure, including any identifiable variations of the above teachings, defines in part the scope of the claim terms. The subject matter of the invention is not dedicated to the public.
Claims
1. A video recording method, comprising: rendering a field of view of a virtual environment of a first user of a head-mounted display at a first resolution; rendering an image of the virtual environment outside the field of view of the first user at a second resolution lower than the first resolution; outputting the rendered field of view for display to the first user; and recording, as a video, the combined rendering for subsequent viewing by a second user.
2. The method according to claim 1, further comprising rendering a foveal region of the virtual environment at a third resolution higher than the first resolution within the field of view based on the gaze tracking of the first user.
3. The method according to claim 1 or 2, wherein the step of rendering an image of the virtual environment outside the field of view of the first user comprises generating a spherical or cylindrical image of the virtual environment.
4. The method according to claim 1, further comprising rendering a transition region, which is a boundary of the rendered field of view outside the field of view of the first user and has a resolution intermediate between the first resolution and the second resolution.
5. The size of the transition region depends on i. the difference in the positions of the rendered field of view between consecutive image frames, ii. the rate of change of the positions of the rendered field of view between consecutive image frames, iii. the fluctuation of the positions of the rendered field of view caused by minute movements of the head of the first user, iv. the expected value of the position fluctuations caused by minute movements of the head of the second user, and depends on one or more of the above items.
6. rendering, at a resolution higher than the second resolution, an object or event of interest within a rendered image of a virtual environment that is outside the field of view of the first user, The method according to claim 1, characterized in that the remaining part of the rendered image of the virtual environment that is outside the field of view of the first user maintains the second resolution. **Claim 7** The step of rendering an image of a virtual environment outside the field of view of the first user, i. the first user, ii. an application for rendering a virtual environment, iii. an operating system of a device for rendering a virtual environment, iv. a helper application of an operating system of a device for rendering a virtual environment, The method according to claim 1, characterized in that it depends on one or more of the above items. **Claim 8** obtaining a recorded video including a field of view of a virtual environment of a first user of a head-mounted display rendered at a first resolution and an image of a virtual environment outside the field of view of the first user rendered at a second resolution lower than the first resolution; i. the spatial resolution in recording an image of a virtual environment outside the field of view of the first user, ii. the temporal resolution in recording an image of a virtual environment outside the field of view of the first user, The method according to claim 1, characterized in that it includes the step of raising either or both of them. **Claim 9** A method for distributing a video recorded using the method according to claim 1, wherein the recorded video includes a field of view of a virtual environment of a first user of a head-mounted display rendered at a first resolution and an image of a virtual environment outside the field of view of the first user rendered at a second resolution lower than the first resolution, Receiving a request to download or stream the recorded video to a second user; A method comprising: downloading or streaming the recorded video to the second user. **Claim 10** Before delivering, i. The spatial resolution in recording an image of a virtual environment outside the field of view of the first user, ii. The temporal resolution in recording an image of a virtual environment outside the field of view of the first user, The method according to claim 9, comprising the step of increasing either or both of them. **Claim 11** A method of viewing a recorded video including a field of view of a virtual environment of a first user rendered at a first resolution and an image of a virtual environment outside the field of view of the first user rendered at a second resolution lower than the first resolution, Requesting a download or stream of a video from a remote source; Receiving a download or stream of the video from the remote source; Outputting at least a part of the stream or video for display to a second user wearing a head-mounted display, The step of outputting, Detecting the field of view of a second user wearing a head-mounted display; Supplying a corresponding part of the recorded stream or video for output to the head-mounted display of the second user. **Claim 12** Calculating a relative position of the current field of view of the second user with respect to the field of view of the first user based on data related to the recorded video, When the relative position deviates more than a predetermined threshold, Calculating a correction direction for the second user to move the field of view towards the corresponding first user's field of view; displaying the correction direction within the field of view presented to the second user, the method of claim 11, comprising. **Claim 13** A computer program, characterized in that the method according to claim 1 is executed by a computer. **Claim 14** A video recording system comprising: a rendering processor configured to render the field of view of the virtual environment of the first user of the head-mounted display at a first resolution; an output processor configured to output for display to the first user the rendered field of view; a storage unit configured to record as a video the combined rendering for subsequent viewing by a second user, a video recording system, comprising. The rendering processor renders an image of the virtual environment outside the field of view of the first user at a second resolution lower than the first resolution, the video recording system. **Claim 15** A video playback system for playing a recorded video including the field of view of the virtual environment of the first user of the head-mounted display rendered at a first resolution and an image of the virtual environment outside the field of view of the first user rendered at a second resolution lower than the first resolution, comprising: a transmitter configured to send a request to download or stream a video from a remote source; a receiver configured to receive the download or stream of the video from the remote source; a graphics processor configured to stream or output at least a portion of the video for display to a second user wearing the head-mounted display, comprising. The video playback system is characterized in that the graphic processor is configured to detect the field of view of a second user wearing a head-mounted display and supply a corresponding portion of a recorded stream or video for output to the head-mounted display of the second user.