Techniques for enabling high-fidelity upscaling of video
By generating and combining videos with varying fields of view or focal lengths from the same camera position and time, high-fidelity zooming is achieved without pixelation, addressing the inefficiencies of current video zooming methods.
Patent Information
- Application Number
- JP2024539958
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-01-07
- Filing Date
- 2022-12-29
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2042-12-29
AI Technical Summary
Zooming a video at high levels results in pixelation, which can be mitigated by increasing resolution but at the cost of excessive storage and bandwidth consumption, lacking an adequate solution in current computer-related technologies.
Generating multiple videos from the same camera position and time with varying field of views or focal lengths, combining them to simulate zooming without loss of fidelity, using a processor to present these videos on a display.
Achieves high-fidelity zooming without pixelation by seamlessly transitioning between videos with different fields of view or focal lengths, optimizing storage and bandwidth usage.
Smart Images

Figure 0007752776000001 
Figure 0007752776000002 
Figure 0007752776000003
Abstract
Description
[Technical Field]
[0001] FIELD OF THE INVENTION This application relates generally to techniques for zooming video without loss of resolution. [Background technology]
[0002] As recognized herein, when zooming a video, at high levels of zoom, the image becomes pixelated. This can be mitigated by providing video with very high resolution, but such video consumes excessive storage and bandwidth. Currently, there is no adequate solution to the above computer-related technical problem. Summary of the Invention
[0003] Thus, in one aspect, at least one storage device that is not a transient signal includes instructions executable by at least one processor, the instructions causing the processor to present a first video on a display, combine a second video with the first video in response to a zoom command, and present the second video combined with the first video on the display. The first video and the second video are generated from substantially the same camera position, at substantially the same time, and with substantially the same resolution as each other. However, to achieve the appearance of zooming without loss of fidelity, the second video is generated by a physical or virtual lens having a smaller field of view (FOV) than the physical or virtual lens used to generate the first video. Alternatively, the second video may be generated by a camera having a shorter focal length than the first video.
[0004] The zoom command may be a first zoom command, and the instructions may be executable to present only the second video on the display in response to continued input of the first zoom command or input of a second zoom command. In some examples, the instructions may be executable to combine the second video with a third video and present the third video combined with the second video on the display in response to continued input of the first zoom command or input of a third zoom command after the second zoom command, where the first, second, and third videos may be generated from substantially the same camera position, at substantially the same time, and at substantially the same resolution as each other, but the third video is generated by a physical or virtual lens having a smaller FOV than the FOV of the physical or virtual lens used to generate the second video.
[0005] In effect, the processor may access a fourth and fifth video, each having a successively smaller FOV than the immediately preceding video, for use in the successive input of zoom commands.
[0006] The display may be a head-mounted display (HMD), such as a virtual reality (VR), three-dimensional (3D) computer game display, or the like.
[0007] In another aspect, a method includes presenting a first video on a display in a wide-angle mode. The method includes, in response to a zoom-in command, presenting the first video in a standard-angle mode, and in response to a continued zoom-in command, presenting the first video in a telephoto mode. In response to the continued zoom-in command, the method includes presenting a second video on the display in the wide-angle mode.
[0008] In another aspect, an apparatus includes at least a processor programmed to present a first video on a display and, in response to a zoom command, present a second video on the display, the second video being generated by a physical or virtual lens having a field of view (FOV) smaller than the FOV of the physical or virtual lens used to generate the first video and / or based on a focal length shorter than the focal length at which the first video is presented.
[0009] The details of the present application, both as to its structure and operation, can best be understood in reference to the accompanying drawings, in which like reference numerals refer to like parts and in which: [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a block diagram of an exemplary system in accordance with the present principles; [Figure 2] Exemplary logic consistent with the present principles is illustrated in exemplary flow chart form. [Figure 3] Illustrates a user zooming by stepping along the Z axis. [Figure 4] 1 illustrates a schematic diagram of zooming. [Figure 5] 10 shows a schematic representation of an offset between videos. [Figure 5A] FIG. 2 is a block diagram of an example rendering module and decoding module. [Figure 6] Illustrates the view from five cameras. [Figure 7] Illustrates multi-FOV and multi-position content capture. DETAILED DESCRIPTION OF THE INVENTION
[0011] The present disclosure generally relates to computer ecosystems, including aspects of consumer electronics (CE) device networks, including, but not limited to, computer gaming networks, including wireless networks operating on 5G or ATSC 3.0. The systems herein may include a server component and a client component, which may be connected via a network such that data may be exchanged between the client and server components. The client component may include one or more computing devices, including game consoles such as Sony PlayStation®, Microsoft, Nintendo, or other game consoles, virtual reality (VR) headsets, augmented reality (AR) headsets, portable televisions (e.g., smart TVs, Internet-enabled televisions), portable computers such as laptops and tablet computers, and other mobile devices including smartphones, as well as additional examples described below. These client devices may operate in a variety of operating environments. For example, some client computers may employ, by way of example, the Linux operating system, a Microsoft operating system, a Unix operating system, or an operating system manufactured by Apple or Google. These operating environments may be used to run one or more browsing programs, such as browsers manufactured by Microsoft, Google, or Mozilla, or other browser programs capable of accessing websites hosted by Internet servers, as described below. Additionally, an operating environment according to the present principles may be used to run one or more computer game programs.
[0012] Servers and / or gateways may be used, which may include one or more processors that execute instructions that configure the server to receive and transmit data over a network such as the Internet. Alternatively, the clients and servers may be connected via a local intranet or virtual private network. The servers or controllers may be exemplified by game consoles such as the Sony PlayStation™, personal computers, etc.
[0013] Information may be exchanged between the client and the server over a network. For this purpose and for security, the server and / or client may include firewalls, load balancers, temporary storage, and proxies, as well as other network infrastructure for reliability and security. One or more servers may form an apparatus that implements a method for providing a secure community, such as an online social website or a gamer network, to network members.
[0014] The processor may be a single-chip or multi-chip processor capable of implementing logic through various wiring such as address lines, data lines, and control lines, and registers and shift registers.
[0015] Components included in one embodiment may be used in other embodiments in any suitable combination. For example, any of the various components described herein and / or shown in the drawings may be combined, interchanged, or excluded from other embodiments.
[0016] "A system having at least one of A, B, and C" (and similarly "a system having at least one of A, B, or C" and "a system having at least one of A, B, and C") includes systems having A only, B only, C only, A and B together, A and C together, B and C together, and / or A, B, and C together.
[0017] 1 , an exemplary system 10 is shown, which may include one or more of the exemplary devices described above and further described below in accordance with the present principles. A first exemplary device included in system 10 is a consumer electronics (CE) device, such as, but not limited to, an audio-video device (AVD) 12, such as an Internet-enabled TV having a TV tuner (equivalently, a set-top box that controls the TV). AVD 12 may alternatively be a computerized Internet-enabled (“smart”) phone, a tablet computer, a notebook computer, a head-mounted device (HMD) and / or headset, such as smart glasses or a VR headset, another wearable computerized device, a computerized Internet-enabled music player, computerized Internet-enabled headphones, a computerized Internet-enabled implantable device, such as an implantable skin device, or the like. It should be understood that AVD 12 is nevertheless configured to implement the present principles (e.g., to communicate with other CE devices to implement the present principles, to execute the logic described herein, and to perform any other functions and / or operations described herein).
[0018] Accordingly, to implement these principles, AVD 12 may be established by some or all of the components shown in Figure 1. For example, AVD 12 may include one or more touch-enabled displays 14, which may be implemented by high-definition or ultra-high-definition "4K" or higher flat screens. Touch-enabled display(s) 14 may include, for example, a capacitive or resistive touch-sensing layer having a grid of electrodes for touch sensing consistent with the present principles.
[0019] The AVD 12 may also include one or more speakers 16 for outputting audio in accordance with the present principles and at least one additional input device 18, such as an audio receiver / microphone, for inputting audible commands to the AVD 12 to control it. The exemplary AVD 12 may also include one or more network interfaces 20 for communicating over at least one network 22, such as the Internet, a WAN, or a LAN, under the control of one or more processors 24. Thus, the interface 20 may be a Wi-Fi transceiver, which is an example of a wireless computer network interface, such as, but not limited to, a mesh network transceiver. It should be understood that the processor 24 controls the AVD 12 to implement the present principles, including other elements of the AVD 12 described herein, such as controlling the display 14 to present images thereon and receiving input therefrom. It should further be noted that the network interface 20 may be a wired or wireless modem or router, or a wireless telephone transceiver, or other suitable interface, such as the Wi-Fi transceiver discussed above.
[0020] In addition to the above, AVD 12 may also include one or more input and / or output ports 26, such as a High-Definition Multimedia Interface (HDMI®) port or a Universal Serial Bus (USB) port for physically connecting to another CE device, and / or a headphone port for connecting headphones to AVD 12 to present audio from AVD 12 to a user through the headphones. For example, input port 26 may be connected via wire or wireless to a cable or satellite audio-video content source 26a. Thus, source 26a may be a separate or integrated set-top box or a satellite receiver. Alternatively, source 26a may be a game console or disc player containing content. When implemented as a game console, source 26a may include some or all of the components described below in connection with CE device 48.
[0021] AVD 12 may further include one or more computer memory / computer-readable storage media 28, such as non-transitory disk-based or solid-state storage, possibly embodied within the AVD chassis as a standalone device, or as a personal video recording device (PVR) or video disc player either internal or external to the AVD chassis for playing AV programs, or as removable memory media or a server as described below. In some embodiments, AVD 12 may also include a location or position receiver, such as, but not limited to, a cellular receiver, a GPS receiver, and / or an altimeter 30, configured to receive geographic location information from a satellite or cellular tower and provide that information to processor 24 and / or to determine the altitude at which AVD 12 is disposed in conjunction with processor 24. Component 30 may also be implemented by an inertial measurement unit (IMU), typically including a combination of accelerometers, gyroscopes, and magnetometers, or by event-based sensors to determine the position and orientation of AVD 12 in three dimensions.
[0022] Continuing with the description of AVD 12, in some embodiments, AVD 12 may include one or more cameras 32, such as an infrared camera, a digital camera such as a webcam, an event-based sensor, and / or a camera integrated into AVD 12 and controllable by processor 24, capable of collecting photos / images and / or video in accordance with the present principles. AVD 12 may also include a Bluetooth® transceiver 34 and other near field communication (NFC) elements 36 for communicating with other devices using Bluetooth® and / or NFC technology, respectively. An exemplary NFC element can be a radio frequency identification (RFID) element.
[0023] Continuing further, AVD 12 may include one or more auxiliary sensors 38 (e.g., pressure-sensitive sensors, motion sensors such as accelerometers, gyroscopes, cyclometers, or magnetic sensors, infrared (IR) sensors, optical sensors, speed and / or cadence sensors, event-based sensors, gesture sensors (e.g., for sensing gesture commands)) that provide input to processor 24. For example, one or more of auxiliary sensors 38 may include one or more pressure sensors forming a layer of touch-enabled display 14 itself, and may be, without limitation, piezoelectric pressure sensors, capacitive pressure sensors, piezoresistive strain gauges, optical pressure sensors, electromagnetic pressure sensors, etc.
[0024] The AVD 12 may also include an over-the-air television broadcast port 40 for receiving terrestrial television broadcasts, which provides input to the processor 24. In addition to the above, it should be noted that the AVD 12 may also include an infrared (IR) transmitter and / or an IR receiver and / or an IR transceiver 42, such as an IR Data Association (IRDA) device. A battery (not shown) may be provided to power the AVD 12, and may also be a kinetic energy harvester that can convert kinetic energy into power to charge the battery and / or power the AVD 12. A graphics processing unit (GPU) 44 and a field programmable gate array 46 may be included. One or more haptic / vibration generators 47 may be provided to generate haptic signals that can be sensed by a person holding or in contact with the device. The haptic generator 47 may vibrate all or a portion of the AVD 12 using an electric motor connected to a centered and / or off-balanced weight via a rotatable shaft of the motor, such that the shaft can rotate under the control of the motor (and which may in turn be controlled by a processor such as processor 24) to generate vibrations of various frequencies and / or amplitudes, and simulated forces in various directions.
[0025] 1 , in addition to AVD 12, system 10 may include one or more other CE device types. In one embodiment, first CE device 48 may be a computer game console that can be used to transmit computer game audio and video to AVD 12 via commands sent directly to AVD 12 and / or through a server, as described below, while second CE device 50 may include similar components as first CE device 48. In the illustrated embodiment, second CE device 50 may be configured as a computer game controller operated by a player or a head-mounted display (HMD) worn by a player. The HMD may include a head-up transparent or non-transparent display for presenting AR / MR content or VR content, respectively.
[0026] In the illustrated embodiment, only two CE devices are shown, but it should be understood that fewer or more devices may be used. The devices herein may implement some or all of the components shown for AVD 12. Any of the components shown in the following figures may incorporate some or all of the components shown for AVD 12.
[0027] Referring now to the aforementioned at least one server 52, it includes at least one server processor 54, at least one tangible computer-readable storage medium 56, such as disk-based or solid-state storage, and at least one network interface 58 that, under the control of the server processor 54, enables communication with other devices of Figure 1 via network 22, and may indeed facilitate communication between the server and client devices in accordance with the present principles. It should be noted that the network interface 58 may be, for example, a wired or wireless modem or router, a Wi-Fi transceiver, or other suitable interface, such as, for example, a wireless telephone transceiver.
[0028] Thus, in some embodiments, server 52 may be an entire internet server or server "farm" and may include and perform "cloud" functionality such that, for example, in an exemplary embodiment for a network gaming application, devices of system 10 may access the "cloud" environment via server 52. Alternatively, server 52 may be implemented by one or more game consoles or other computers located in or near the same room as the other devices shown in FIG.
[0029] The components shown in the following figures may include some or all of the components shown in Figure 1. Any user interfaces (UIs) described herein may be integrated and / or extended, and UI elements may be mixed and matched between UIs.
[0030] 2 illustrates that, in an embodiment, "N" videos are generated by each virtual or physical camera and associated physical or virtual lens. N can be an integer greater than or equal to 2. In one embodiment, N is equal to 5.
[0031] In one embodiment, each of the N videos has the same resolution, such as, but not limited to, 4K. However, in other embodiments, the N videos may not all have the same resolution.
[0032] In either case, in one embodiment, the videos may be taken from the same or substantially the same location and at the same or substantially the same time. "Substantially the same location" means, for example, within the constraints of physically locating two cameras in the same place; the cameras may be closely positioned side-by-side, yet separated by the width of the camera housings. "Substantially the same time" means at the same real or virtual time, or within a few seconds of each other.
[0033] However, the first video may be generated using a physical or virtual lens having a first field of view (FOV), the second video may be generated using a physical or virtual lens having a second FOV smaller than the first FOV, and each successive video may be generated with a successively smaller FOV than the preceding video in the chain. However, each FOV may be centered at the same location or point or center. Note that in addition to or instead of successively smaller FOVs, the physical or virtual cameras may have successively shorter focal lengths.
[0034] Moving to block 202, the videos are synchronized with each other, for example, by aligning key frames in each video with each other, and in a particular embodiment, encoding the videos as H264. Alignment is described further below.
[0035] When a user wishes to play a video, the video is presented in block 204 using a first video, i.e., the video with the widest FOV. In block 206, as the user zooms in using an input device or by moving their head along the Z axis while wearing the HMD presenting the video, a video with the next smaller FOV is combined with the first video, eventually replacing it. Continued zooming results in successive videos presented with successively smaller FOVs, such that zooming is emulated without loss of fidelity. Thus, during playback, content from the telephoto camera is inserted into content from the wide-angle camera according to a pre-calculated alignment metric to create the perception of watching a single video. With accurate alignment, the presence of the inner video displayed within the outer video is not apparent to the viewer.
[0036] FIG. 3 illustrates a user 300 wearing an HMD 302 that zooms by moving along a Z axis 304 .
[0037] FIG. 4 continues to further illustrate. Note that FIG. 4 illustrates an implementation in which, in addition to using different FOVs, scenes are captured at different positions, while FIG. 6, described below, illustrates a case in which three or more videos are captured from the same position. More specifically, in the example shown in FIG. 4, multiple (e.g., three) lenses are used with different FOVs to capture three videos from the same physical or virtual camera arrangement, and the same three lenses, each with a different FOV, are used to capture three videos from a second arrangement. Thus, after recording, six videos are captured simultaneously.
[0038] A first video 400 is shown in its widest mode 402. As the user zooms in, the video is shown in its standard mode 404, and finally, under continued zoom, in its telephoto mode 406, with each mode taking up space on the display. It should be understood that the transition between the three modes shown is continuous and gradual as the user zooms, but for simplicity's sake, only three common modes are shown.
[0039] When the zoom in the telephoto mode 406 of the first video reaches a threshold limit, further zooming results in combining the first video with the second video 408 in its maximum wide-angle mode 410. It should be understood that the second video 408 may eventually or immediately replace the first video entirely as the zoom progresses from the telephoto mode 406 of the first video to the wide-angle mode 410 of the second video 408.
[0040] As the user continues to zoom in, the second video 408 is shown in its standard angle mode 412 and eventually, under continued zoom, in its telephoto mode 414, with each mode taking up space on the display.
[0041] Continued zooming from the telephoto mode 414 of the second video results in the second video being combined with the third video 416 having its maximum wide-angle mode 418. It should be understood that the third video 414 may eventually or immediately replace the second video entirely as the zoom progresses from the telephoto mode 414 of the second video to the wide-angle mode 418 of the third video 416.
[0042] As the user continues to zoom in, the third video 416 is shown in its standard angle mode 420, and eventually, under continued zoom, in its telephoto mode 422, with each mode occupying the display. Note that if the scene is captured from only a single position, steps 408-422 are not available.
[0043] While FIG. 4 illustrates the use of three videos generated by physical or virtual lenses, each with a successively smaller FOV, it should be understood that, consistent with the principles of FIG. 4, only two videos need be used, or four or more videos may be used.
[0044] Note that multiple videos, each with a progressively smaller FOV, can be generated for multiple likely regions of user focus. A central focus can be used as a baseline, and then offsets from that point in terms of distance and direction can be used and transmitted as metadata indicating when the user is focusing on a point that is the offset away from the central focus. For each offset, a series of nested videos can be pre-computed or computed on-the-fly for a particular focus as the user focuses on that point. If the user happens to focus on a point for which there are no nested videos with progressively smaller FOVs, traditional magnification techniques can be used.
[0045] A heatmap of previous user focus for each scene can be used to determine which points in the scene should have a series of nested videos generated for them: only the video of the area on which the user focused can be decoded.
[0046] Reference is now made to FIG. 5 for a description of a match metric that may be determined before or during capture.
[0047] The interpolation ratio (R) can be determined to be the ratio of the number of pixels in the outer video (wider FOV) to the number of pixels in the inner video (narrower FOV) in a single dimension. In Figure 5, W0 is the width in pixels of the outer video, W1 is the width in pixels of the inner video, and after alignment, R = W0 / W1. The interpolation ratio depends on the focal lengths of the two cameras and the resolution of the camera sensors.
[0048] The horizontal offset (Oh) is shown in Figure 5 and is the horizontal offset of the inner video or ROI measured from the center of the frame of the outer video. Similarly, the vertical offset (Ov) is the vertical offset of the inner video or ROI measured from the center of the frame of the outer video.
[0049] Figure 5 illustrates that frames of video with wider and narrower FOVs are aligned during display using the above offsets with the alignment metric. Specifically, a camera placement is determined, with two cameras with different FOVs capturing the same scene simultaneously. In the simplest case, Oh = Ov = 0, and the ROI is the center of the video frame. An interpolation ratio of 2 could be achieved by using an FOV of 60 for the wide-angle lens and an FOV of approximately 32.2 for the telephoto lens. Disabling automatic features such as camera auto-exposure facilitates blending of the two frames during display.
[0050] See FIG. 5A. The raw video from the two cameras, labeled 500 and 502, is synchronized and encoded as two separate bitstreams. In this case, two decoders are used to decode both bitstreams simultaneously. In other embodiments using one decoder, the video data from each camera may be compressed as a single bitstream but may be independently decodable, for example, as HEVC tiles. In either case, the video player 504, which generates output pixels for the display 506, includes a decoding module (DM) 508 and a rendering module (RM) 510. The DM, in turn, includes one or more decoders 512 capable of decoding the compressed bitstream(s). The RM includes GPU shaders that can sample the video texture and render it to the display.
[0051] The match metric may be fixed or may change over time. In the fixed case, the match metric may be sent to the DM and / or RM only once. In the case of a dynamic match metric, the DM and / or RM may be updated with each change in the metric. One way to achieve this is to pass the match metric as metadata in the compressed bitstream. In other embodiments, the match metric may be calculated automatically using motion estimation and image matching algorithms.
[0052] A video player that renders the decoded video data on a display accepts magnification control from the user using a device such as a mouse or video game controller. The magnification level (ML) selected by the user is used to determine the portions of the outer and inner videos that will be visible on the display. The system can set upper and lower bounds on ML to avoid magnification levels that result in image quality degradation. When the user zooms in, the value of ML increases, and when the user zooms out, ML decreases. As ML increases, the number of visible pixels in the outer video decreases and the number of visible pixels in the inner video increases. The RM's GPU shader uses the ML value, a matching metric, and the frame number of each bitstream for synchronization to create the perception of watching a single video rather than two separate videos. In other embodiments, an additional "feathering" step can be performed by the shader to mask the boundary at the junction of the inner and outer videos.
[0053] When the ML is small and the number of visible pixels of the inner video is small, rendering of the inner video can be skipped without noticeable difference in the image quality of the displayed video. When decoded video data of the inner video is not displayed, decoding of the non-displayed video data can be omitted, thereby improving system performance and efficiency. One way this can be achieved is to utilize the ML to determine which video bitstreams need to be decoded and render only frames from the bitstreams that are actively being decoded. When a decoder is in an active state, access units (AUs) of the bitstreams are decoded successfully, and the decoded video data is sent to the RM for rendering to the display. When a decoder is in an inactive state, decoding of the AUs can be partially or completely skipped, and video data for bitstreams corresponding to the inactive decoders is not rendered to the display.
[0054] When the ML changes, a decoder in an active state may become inactive, and vice versa. Switching a decoder from an active state to an inactive state can be done immediately, but switching from an inactive state to an active state may not be done immediately. The reason for this is that the current AU may depend on the previous AU, and if decoding of the previous AU is skipped when the decoder is in the inactive state, the current AU may have an error when decoded. To avoid this problem, switching from an inactive state to an active state may be performed only when the current AU is a key frame (IDR frame). To support this, a seek state may be used, in which when the ML exceeds a threshold, an inactive decoder first switches to a seek state where the decoder waits for an IDR. When the current AU is an IDR, the decoder switches from the seek state to the active state. The DM passes the bitstream ID of the active decoder to the RM and passes invalid IDs for decoders in the seek state or in the inactive state to the RM. The RM uses these IDs to render only valid pixels to the display.
[0055] For applications requiring high magnification levels or a smoother transition from a zoomed-out view to a zoomed-in view, more than two camera views may be required. For such use cases, three or more cameras with varying focal lengths or FOVs may be used. As before, the same scene is captured simultaneously from a single location using these cameras.
[0056] An example of fields of view that can be captured using five cameras is shown in FIG. 6 (the five fields of view are labeled "Wide 1," "Wide 2," "Telephoto 1," "Telephoto 2," and "Standard").
[0057] The video data from each camera in Figure 6 can be synchronized and compressed as individual bitstreams or independently decodable substreams. While all of these streams can be simultaneously decoded and selectively rendered according to the desired ML, a more efficient approach is to decode only the streams that will ultimately be displayed. The number of decoders required in the DM can be equal to the maximum number of video streams being simultaneously rendered at any given moment. The number of decoders required for the setup shown in Figure 5 of one outer video and one inner video can be limited to two, even if more than two video streams are used. This is achieved using a "stream switching" strategy described below.
[0058] The stream processed by each decoder is determined by the value of ML. When the application starts, the first decoder (D1) can process the widest bitstream (B1), and the second decoder (D2) can process the second bitstream (B2) with a lower FOV. As the user increases ML, there comes a point where the pixels of B1 are no longer rendered on the display. D1 then transitions to a seek state and prepares to decode the next bitstream (B3) in the viewing list. The RM uses the bitstream ID and alignment metric passed from the decoder to display the decoded pixels of each bitstream at the appropriate magnification. When the RM detects a change in the bitstream ID, it updates the rendering process to use the correct texture and sampling coordinates.
[0059] In other embodiments, to facilitate smooth stream switching, the following steps can be performed during the encoding process.
[0060] First, the bitstreams use similar coding settings so that the same instance of the decoder can process AUs from multiple bitstreams without requiring additional memory. The IDRs of different bitstreams can be aligned and evenly spaced according to the rate at which the user can increase or decrease the ML. Second, the IDR positions and AU offsets for each bitstream can be pre-calculated to avoid performing the same calculations in the DM.
[0061] In further embodiments, the DM may include one or more additional decoders to predict the next bitstream to be processed based on ML and decode these streams before the decoded pixels are made visible on the display. This strategy can help increase the rate of change of ML. An alternative approach to achieve this is to encode the bitstream using only IDR.
[0062] Referring now to FIG. 7 , an alternative technique for applications requiring high magnification levels is multi-location content capture instead of multi-FOV content capture. Instead of capturing a scene from one location using cameras with different FOVs, the scene can be captured at locations 700, 702 using the same FOV but with different scene capture directions. In other embodiments, as shown in FIG. 7 , both multi-FOV and multi-location content capture can be used together. In other embodiments, the RM can include a stage for distortion correction between multi-location content or multi-view content. In other embodiments, audio is also captured from different locations, and the audio streams are also switched according to the ML for a more immersive experience.
[0063] Although particular embodiments are shown and described in detail herein, it should be understood that the subject matter encompassed by the present invention is limited only by the claims.
Claims
1. A device, at least one storage device containing instructions executable by at least one processor that are not transitory signals, said instructions causing said processor to: causing the first video to be presented on the display; in response to a zoom command, combine a second video with the first video and present the second video combined with the first video on the display, wherein the first video and the second video are generated from substantially the same camera position as each other and at substantially the same resolution, and the second video is generated by a physical or virtual lens having a field of view (FOV) that is smaller than the FOV of a physical or virtual lens used to generate the first video to create the appearance of zooming without loss of resolution; the zoom command is a first zoom command, and the instructions are executable to present only the second video on the display in response to continued input of the first zoom command or input of a second zoom command; the instructions are executable to, in response to continuing to input the first zoom command or inputting a third zoom command after the second zoom command, combine the second video with a third video and present the third video combined with the second video on the display, wherein the first, second, and third videos are generated from substantially the same camera position, at substantially the same time, and at substantially the same resolution as each other, and the third video is generated by a physical or virtual lens having a FOV smaller than the FOV of a physical or virtual lens used to generate the second video. device.
2. The device of claim 1 , wherein the first video and the second video are generated by a virtual lens.
3. The device of claim 1 , wherein the first video and the second video are generated by a physical lens.
4. 10. The device of claim 1, wherein the processor has access to a fourth video and a fifth video, each having a successively smaller FOV than the immediately preceding video, for use with successive input of zoom commands.
5. The device of claim 1 comprising the display.
6. The device of claim 5 , wherein the display comprises a head-mounted display (HMD).
7. 1. A method comprising: presenting the first video on a display; and presenting a second video on the display in response to a zoom-in command; The method includes presenting the first video in a wide-angle mode and then, in response to a first zoom command, presenting the video in a standard-angle mode.
8. The method of claim 7 , including presenting the video in telephoto mode after a second zoom command that follows the first zoom command.
9. The method of claim 8 , wherein the zoom command that causes the presentation of the second video is a third zoom command following the second zoom command, and the second video is presented in a wide-angle mode.
10. 8. The method of claim 7, wherein the first video is generated using a lens having a first field of view (FOV) and the second video is generated using a lens having a second FOV that is smaller than the first FOV.
11. The method of claim 10 , wherein the first video and the second video are generated to be from cameras located at the same location and at the same time.
12. 8. The method of claim 7, wherein the first video and the second video are captured using the same camera field of view from respective first and second physical or virtual camera positions.
13. 1. An apparatus comprising: The system includes at least a processor, the processor: presenting the first video on a display; and programmed to present a second video on the display in response to a zoom command, the second video being generated by a physical or virtual lens having a field of view (FOV) smaller than the FOV of a physical or virtual lens used to generate the first video, and / or the second video being generated based on a focal length shorter than the focal length at which the first video is presented; the zoom command is a first zoom command, and the processor is programmed to present only the second video on the display in response to continued input of the first zoom command or input of a second zoom command; the processor is programmed to combine the second video with a third video and present the third video combined with the second video on the display in response to continued input of the first zoom command or input of a third zoom command after the second zoom command, wherein the first, second, and third videos are generated from substantially the same camera position, at substantially the same time, and at substantially the same resolution as each other, and the third video is generated by a physical or virtual lens having a FOV smaller than the FOV of a physical or virtual lens used to generate the second video.
14. The apparatus of claim 13 , wherein the first video and the second video have substantially the same resolution.
15. The apparatus of claim 13 , wherein the first video and the second video are generated by a virtual lens.
16. The apparatus of claim 13 , wherein the first video and the second video are generated by a physical lens.
17. 14. The apparatus of claim 13, wherein the processor has access to a fourth video and a fifth video, each having a successively smaller FOV than the immediately preceding video, for use with successive input of zoom commands.
18. The apparatus of claim 13 , wherein the first video and the second video are generated from substantially the same camera position and at substantially the same time as each other.
19. The apparatus of claim 13 , wherein the first video and the second video are generated from respective first and second camera positions that are spaced apart from one another.
20. A method comprising: presenting the first video on a display; and presenting a second video on the display in response to a zoom-in command; A method comprising reducing required computational power by selectively decoding only content in a video represented by at least one bitstream that is made visible.
21. A method comprising: presenting the first video on a display; and presenting a second video on the display in response to a zoom-in command; A method comprising smoothly switching between respective bitstreams of camera views at least in part by using similar encoding settings so that the same instance of a decoder can process bitstream access units (AUs) from multiple bitstreams without requiring additional memory, wherein keyframes of the different bitstreams are aligned and evenly spaced according to a rate at which a user can increase or decrease the magnification level, and wherein keyframe positions and AU offsets for each bitstream are pre-calculated to avoid performing the same calculations in the decoding module.
22. A method comprising: presenting the first video on a display; and presenting a second video on the display in response to a zoom-in command; 1. A method comprising: executing a decoding module (DM) including one or more additional decoders to predict a next bitstream to be processed; and decoding the next bitstream before decoded pixels from the next bitstream are visible on the display.
23. A method comprising: presenting the first video on a display; and presenting a second video on the display in response to a zoom-in command; 10. A method comprising: encoding a bitstream associated with said video using only keyframes.
24. A method comprising: presenting the first video on a display; and presenting a second video on the display in response to a zoom-in command; A method including switching audio streams in conjunction with switching video when zooming.
Citation Information
Patent Citations
A control framework with a zoomable graphical user interface for organizing, selecting and launching media items
JP2007516496A
Image distribution apparatus
JP2018026603A
Multi-camera video stabilization
US11190689B1
Imaging systems and methods for immersive surveillance
US20150271453A1
Composite Image Associated with a Head-Mountable Device
US20160027210A1