Content distribution method, image capture and processing system, playback system, method of operating a playback system, and computer readable medium
By capturing images at high resolution and low frame rate in a virtual reality system and using interpolation technology, the problems of high cost and bandwidth limitations are solved, achieving a high-quality virtual reality experience while reducing artifact risks and device complexity.
Patent Information
- Application Number
- CN202411158622.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-04-24
- Filing Date
- 2020-04-24
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2040-04-24
AI Technical Summary
High frame rate and high resolution image capture and transmission are costly and bandwidth-limited in virtual reality systems, especially in 3D virtual reality systems, where the difference between left and right eye image data makes it difficult to convey depth and 3D properties.
Images are captured at high resolution but below the playback frame rate. Interpolation techniques are used to generate interpolated frames, which are then selectively updated for high-motion segments. Interpolated frame data is transmitted using limited bandwidth, reducing the risk of artifacts and optimizing processing and bandwidth complexity.
By effectively utilizing bandwidth, the risk of motion artifacts in virtual reality devices is reduced, image detail preservation and 3D experience quality are improved, and reliance on high-cost equipment is reduced.
Smart Images

Figure CN119031110B_ABST
Abstract
Description
[0001] Related patent applications
[0002] This application is a divisional application of the invention patent application with international application number PCT / US2020 / 029958, international application date of April 24, 2020, entry into the Chinese national phase date of December 23, 2021, Chinese national application number 202080046243.2, and invention title "Content Distribution Method, Image Capture and Processing System, Playback System, Method for Operating the Playback System, and Computer-Readable Medium".
[0003] This application claims the benefit of U.S. Provisional Patent Application Serial No. 62 / 838,300, filed on April 24, 2019, which is expressly incorporated herein by reference in its entirety. Technical Field
[0004] The present invention relates to methods and apparatus for encoding and transmitting images with motion for virtual reality systems (e.g., systems that support playback on virtual reality devices, such as those capable of rendering 2D or 3D image content). Background Technology
[0005] High-quality image capture can be important for practical virtual reality applications. Unfortunately, capturing high-resolution images at high frame rates can involve using extremely expensive high-speed, high-resolution cameras, which may be impractical for many applications due to cost considerations.
[0006] While capturing high-resolution images to support virtual reality is desirable, the overall transmission of a large number of high-resolution images may present communication problems, given the bandwidth constraints of the communication networks available for transmitting images to playback devices.
[0007] The problem of image capture and data transmission constraints becomes particularly acute in the context of non-stereoscopic systems, especially in 3D video and 3D virtual reality systems, where separate left-eye and right-eye image data often need to be transmitted.
[0008] In the context of 3D virtual reality systems, the difference between the image content for the left and right eyes is typically what conveys the sense of depth and 3D nature of the viewed object to the human observer. Therefore, for a truly immersive depth experience, it is generally important to capture and preserve image details.
[0009] In light of the foregoing, it should be understood that methods and apparatus are needed to support the use of high-resolution images in virtual reality applications. It would be desirable if the image capture rate did not need to match the image playback rate to allow the use of cameras that might not support the required playback rate when capturing images at a high level of detail. Furthermore, methods and apparatus are needed that support playback and allow users of the playback system sufficient detail to enjoy the experience. It is desirable that at least some of the methods and / or apparatus address the problem of delivering image content to playback devices with limited bandwidth for delivering such content, and it is desirable that at least some of the methods are available where the content may include moving objects (such as those anticipated in a sporting event where a ball or other objects may be moving). Not all embodiments need to address all of the aforementioned problems, but embodiments and features that address one or more of the problems discussed are useful and desirable. Summary of the Invention
[0010] This invention describes methods and apparatus for capturing, transmitting, and using image data to support virtual reality experiences. Images, such as frames, are captured at a high resolution but at a frame rate lower than the frame rate required to support playback. Interpolation is applied to the captured frames. The captured frames, along with interpolated frame information, are transmitted to a playback device. The combination of the captured frames and the interpolated frames corresponds to a second frame playback rate higher than the image capture rate. By operating the camera at a high image resolution but a slower frame rate than when the same camera can capture images at a lower resolution, detail is preserved. Interpolation is performed before delivery to the user device, wherein the segments to be interpolated are selected based on motion and / or lens field of view (FOV) information. A relatively small amount of interpolated frame data is transmitted compared to the captured frame data to efficiently utilize bandwidth.
[0011] While various features are discussed in this invention description, not all embodiments need to include all features discussed in the invention description. Therefore, the discussion or reference to one or more features in this invention description is not intended to imply that the one or more features discussed are necessary or essential in all embodiments.
[0012] The described method and apparatus offer several advantages over other methods. This method addresses and avoids some more severe motion artifacts that can occur when playing immersive content on certain HMDs (Head-Mounted Displays), which may occur when the content is limited to updating at the capture frame rate. By understanding the nature of the artifacts—for example, that they are created due to the motion and / or shape of the lens portions used to capture the image—this method mitigates the risk of such artifacts by selectively updating high-motion segments and / or segments corresponding to the portions of camera lenses subjected to high distortion, without requiring the configuration of multiple VR (Virtual Reality) systems.
[0013] Compared to other systems, this method also reduces the processing and bandwidth complexity required for content delivery, with fewer artifacts than some known systems, because it effectively utilizes the limited bandwidth available for transmitting image data to the playback device. One or more aspects of this method can be implemented on an embedded device at the final point in the content delivery pipeline.
[0014] An exemplary content distribution method relates to a content distribution method comprising: storing an image captured at a first frame rate; performing interpolation to generate interpolated frame data to support a second frame rate higher than the first frame rate; and transmitting the captured frame data and the interpolated frame data to at least one playback device.
[0015] The present invention also relates to various playback system features, methods, and implementation schemes.
[0016] An exemplary method implementation relates to a method of operating a playback system, the method comprising: receiving capture frame data and interpolated frame data; recovering capture frames from the received capture frame data; generating one or more interpolated frames from the received interpolated frame data; rendering a video sequence comprising one or more capture frames and at least one interpolated frame; and outputting one or more rendered images to a display device.
[0017] Many variations of the above-described methods and apparatus are possible. Attached Figure Description
[0018] Figure 1 An exemplary system is shown, implemented according to some embodiments of the present invention, which can be used to capture content, stream content, and output content to one or more users.
[0019] Figure 2 An exemplary content delivery system with encoding capabilities that can be used for encoding and streaming content, according to features of the present invention, is shown.
[0020] Figure 3 It shows the functions that can be used to receive, decode, and display data. Figure 2 An exemplary content playback system for streaming content.
[0021] Figure 4 A camera setup is shown, comprising multiple pairs of cameras for capturing left-eye and right-eye images corresponding to different 120-degree sectors of a 360-degree field of view, and one or more cameras pointed to the sky to capture a view of the sky.
[0022] Figure 5This demonstrates how five different environment meshes corresponding to different camera views can be combined to create a complete spherical view / environment, onto which a background image can be applied as part of the playback operation.
[0023] Figure 6 The diagram shows the complete assembly of 5 grids to create a spherical simulation environment.
[0024] Figure 7 An environment mesh model corresponding to a sector of a camera setup is shown, in which an image from an image is applied (e.g., projected) onto the environment mesh to generate a background image.
[0025] Figure 8 The application of images captured by cameras corresponding to each sector, as well as sky and ground cameras equipped with cameras, is shown to simulate a complete 3D environment in the form of a sphere, which can be used as a background for foreground objects to be applied.
[0026] Figure 9 A method for processing and capturing image content according to the present invention is shown, which can be and sometimes is performed by... Figure 1 The content delivery system shown and / or Figure 1 The content delivery system shown is implemented using image processing, calibration, and encoding equipment.
[0027] Figure 10 This is an illustration showing how the transmission frame rate generated by the content delivery system exceeds the capture frame rate in various embodiments of the invention.
[0028] Figure 11 An exemplary overall transmission frame sequence is shown, wherein the CF start is used to indicate, for example, a capture frame transmitted as part of a base video layer, and the IF is used to indicate interpolated frame information transmitted as an enhancement layer in some embodiments, wherein each set of interpolated frame information includes, for example, one or more segments generated by using motion interpolation.
[0029] Figure 12 A capture frame is shown as a base frame, where a circle corresponds to a portion of the frame that includes image content corresponding to a fisheye lens. In some embodiments, the fisheye lens is used to capture the edge portions outside the circle that are not of interest, as they will not be used as textures and correspond to areas outside the captured area of interest.
[0030] Figure 13 This demonstrates how information used to reconstruct a specific segment can be sent as interpolated frame data.
[0031] Figure 14 This demonstrates how velocity vectors and various rates, along with positions relative to the fisheye lens capture area, can be considered to determine the frame rate that should be supported by interpolation.
[0032] Figure 15 It is a graph showing the various interpolation frame times and which segments will be interpolated and transmitted based on the selected frame rate that a particular segment is to support.
[0033] Figure 16 This is a flowchart of an exemplary playback method according to an exemplary implementation. Detailed Implementation
[0034] Figure 1 An exemplary system 100 implemented according to some embodiments of the present invention is illustrated. System 100 supports the delivery of content (e.g., imaging content delivery) to one or more client devices (e.g., playback devices / content players) located at client locations. System 100 includes an exemplary image capture device 102, a content delivery system 104, a communication network 105, and multiple client locations 106, ..., 110. Image capture device 102 supports the capture of stereoscopic images. Image capture device 102 captures and processes imaging content according to features of the present invention. Communication network 105 may be, for example, a hybrid fiber-coaxial (HFC) network, a satellite network, and / or the Internet.
[0035] Content delivery system 104 includes image processing, calibration, and encoding apparatus 112 and content delivery device 114, such as streaming server 114. Image processing, calibration, and encoding apparatus 112 is responsible for performing various functions, including performing camera calibration based on one or more target images and / or grid patterns captured during the camera calibration process; generating a distortion correction or compensation grid that a playback device can use to compensate for distortion introduced by the calibrating camera; processing, for example, cropping and encoding the captured images; and providing calibration and / or environmental information to content delivery device 114, which can be provided to the playback device and used in the rendering / image playback process. Content delivery device 114 can be implemented as a server, as will be discussed below, responding to content requests using image calibration information, optional environmental information, and one or more images captured by camera apparatus 102 that can be used to simulate a 3D environment. The streaming of images and / or content can be, and sometimes is, a function of feedback information, such as the viewer's head position and / or user selection corresponding to the location of an event at camera apparatus 102, which will be the source of the images. For example, a user can select or switch between images from a camera positioned at the centerline and images from a camera positioned at the target location on site. The simulated 3D environment and streaming images are then changed to correspond to those images from the camera equipment selected by the user. Therefore, it should be understood that... Figure 1A single camera setup 102 is shown, but multiple camera setups may exist within the system and be located at different physical locations during sports or other events, where a user can switch between different locations and select which playback device 122 is used to deliver content to the content server 114. While separate devices 112, 114 are shown in the image processing and content delivery system 104, it should be understood that the system can be implemented as a single device comprising separate hardware for performing various functions, or having different functions controlled by different software or hardware modules but implemented in or on a single processor.
[0036] Encoding device 112 may, and in some embodiments does, include one or more encoders for encoding image data according to the invention. Encoders can be used in parallel to encode different portions of the scene and / or to encode a given portion of the scene to produce encoded versions with different data rates. Using multiple encoders in parallel can be particularly useful when supporting real-time or near-real-time streaming.
[0037] Content streaming device 114 is configured to stream, for example, via communication network 105, encoded content to deliver encoded image content to one or more client devices. Via network 105, content delivery system 104 can send information and / or exchange information with devices located at client locations 106, 110, as shown by link 120 traversing communication network 105.
[0038] Although the encoding device 112 and the content delivery server 114 are in Figure 1 The examples show them as separate physical devices, but in some implementations, they are implemented as a single device that encodes and streams content. The encoding process can be a 3D (e.g., stereoscopic) image encoding process, in which information corresponding to the left-eye and right-eye views of a scene portion is encoded and included in the encoded image data, enabling 3D image viewing. The specific encoding method used is not critical to this application, and a wide range of encoders can be used as or to implement the encoding device 112.
[0039] Each customer location 106, 110 may include multiple playback systems, such as devices / players, for example, means for decoding and playing back / displaying imaging content streamed by content streaming device 114. Customer location 1 106 includes playback system 101, which includes a decoding device / playback device 122 coupled to display device 124. Customer location N 110 includes playback system 111, which includes a decoding device / playback device 126 coupled to display device 128. In some embodiments, display devices 124, 128 are head-mounted stereoscopic display devices. In various embodiments, playback system 101 is a head-mounted system supported by a strap worn around the user's head. Therefore, in some embodiments, customer location 1 106 includes a playback system 1 101, which includes a decoding device / playback device 122 coupled to a display 124 (e.g., a head-mounted stereoscopic display), and customer location N 110 includes a playback system N 111, which includes a decoding device / playback device 126 coupled to a display 128 (e.g., a head-mounted stereoscopic display).
[0040] In various implementations, decoding devices 122, 126 present imaging content on corresponding display devices 124, 128. Decoding devices / players 122, 126 may be devices capable of decoding imaging content received from content delivery system 104, generating imaging content using the decoded content, and rendering the imaging content (e.g., 3D image content) on display devices 124, 128. Either decoding device / playback device 122, 126 can be used as... Figure 3 The decoding device / playback device 800 shown. System / playback device (such as...) Figure 3 The system / playback device shown can be used as any one of the decoding device / playback device 122, 126.
[0041] Figure 3 An exemplary content delivery system 700 with encoding capabilities for encoding and streaming content, according to features of the present invention, is shown.
[0042] This system can be used to perform object detection, encoding, storage, and transmission and / or content output according to the features of the present invention. The content delivery system 700 can be used as... Figure 1 System 104. Although Figure 3 The system shown is used for the encoding, processing and streaming of content, but it should be understood that system 700 may also include, for example, the ability to decode and display processed and / or encoded image data to an operator.
[0043] System 700 includes a display 702, an input device 704, an input / output (I / O) interface 706, a processor 708, a network interface 710, and a memory 712. The various components of system 700 are coupled together via a bus 709, which allows data to be transferred between the components of system 700.
[0044] The memory 712 includes various modules, such as routines, which, when executed by the processor 708, control the system 700 to perform partitioning, encoding, storage, and streaming / transmission and / or output operations according to the invention.
[0045] The memory 712 includes various modules, such as routines, which, when executed by the processor 707, control the computer system 700 to implement the immersive stereoscopic video acquisition, encoding, storage, transmission, and / or output method according to the present invention. The memory 712 includes a control routine 714, a partitioning module 706, an encoder 718, a detection module 719, a streaming controller 720, received input images 732 (e.g., 360-degree stereoscopic video of a scene), an encoded scene portion 734, timing information 736, an environment mesh model 738, a UV map 740, and multiple sets of correction mesh information, including first correction mesh information 742, second correction mesh information 744, third correction mesh information 746, fourth correction mesh information 748, fifth correction mesh information 750, and sixth correction mesh information 752. In some embodiments, the modules are implemented as software modules. In other embodiments, the modules are implemented in hardware, for example, as a single circuit, wherein each module is implemented as a circuit for performing the function corresponding to the module. In other embodiments, the modules are implemented using a combination of software and hardware. The memory 712 also includes a frame buffer 715 for storing captured images (e.g., left-eye and right-eye images captured by the left and right cameras of a camera pair) and / or interpolated frame data representing interpolated frames generated by interpolating a sequence of buffered frames between captured frames. In some embodiments, fragment definition and position change information is stored in a portion 717 of the memory 712. The definition information may, and sometimes does, define a set of blocks corresponding to objects interpreted as image fragments.
[0046] As will be discussed below, interpolated frames may and sometimes do include one or more interpolated segments that are combined with data from a previous or subsequent frame to form a complete interpolated frame. In various embodiments, an interpolated frame is represented using a portion of the data used to represent the captured frame, for example, 1 / 20 or less, and in some cases 1 / 100—the amount of data used to represent the captured frame for intercoding. Therefore, when transmitting encoded frames and interpolated frames, relatively little data can be used to transmit the interpolated frame. In some cases, the decoder of playback device 122 uses image content from previous frames to perform a default padding operation to fill in any portions of the interpolated frame that were not transmitted to the playback device. The segments of the interpolated frame generated by interpolation may be transmitted to playback device 122 using motion vectors as inter-coded data or as inter-coded data, depending on the embodiment.
[0047] Control routine 714 includes device control routines and communication routines to control the operation of system 700. Dividing module 716 is configured to divide the received stereoscopic 360-degree version of the scene into N scene parts according to the features of the invention.
[0048] Encoder 718 may include, and in some embodiments does include, multiple encoders configured to encode received image content (e.g., a 360-degree version of a scene and / or one or more scene portions) according to features of the invention. In some embodiments, the encoder includes multiple encoders, each configured to encode a stereoscopic scene and / or partitioned scene portions to support a given bitrate stream. Thus, in some embodiments, multiple encoders may be used to encode each scene portion to support multiple different bitrate streams for each scene. The output of encoder 718 is encoded scene portion 734 stored in memory for streaming to a client device (e.g., a playback device). The encoded content may be streamed to one or more different devices via network interface 710.
[0049] The detection module 719 is configured to detect network-controlled switching of streaming content from the current camera pair (e.g., the first stereo camera pair) to another camera pair (e.g., the second or third stereo camera pair). Specifically, the detection module 719 detects whether the system 700 has switched from a streaming content stream generated using images captured by a given stereo camera pair (e.g., the first stereo camera pair) to a streaming content stream generated using images captured by another camera pair. In some embodiments, the detection module is also configured to detect user-controlled changes from receiving a first content stream including content from the first stereo camera pair to receiving a second content stream including content from the second stereo camera pair; for example, detecting a signal from a user playback device indicating that the playback device is attached to a content stream different from its previously attached content. The streaming controller 720 is configured to control the streaming of encoded content to deliver encoded image content to one or more client devices, for example, via the communication network 105.
[0050] The streaming controller 720 includes a request processing module 722, a data rate determination module 724, a current head position determination module 726, a selection module 728, and a streaming control module 730. The request processing module 722 is configured to process received requests for imaging content from a client playback device. In various embodiments, the request for content is received via a receiver in network interface 710. In some embodiments, the request for content includes information indicating the identity of the requesting playback device. In some embodiments, the request for content may include a data rate supported by the client playback device, the user's current head position, such as the position of a head-mounted display. The request processing module 722 processes the received request and provides the retrieved information to other elements of the streaming controller 720 for further action. While the request for content may include data rate information and current head position information, in various embodiments, the data rate supported by the playback device may be determined based on network testing and other network information exchange between system 700 and the playback device.
[0051] The data rate determination module 724 is configured to determine the available data rate that can be used to stream imaging content to a client device. For example, since multiple encoded scene portions are supported, the content delivery system 700 can support streaming content to the client device at multiple data rates. The data rate determination module 724 is further configured to determine the data rate supported by a playback device that requests content from the system 700. In some embodiments, the data rate determination module 724 is configured to determine the available data rate for delivering image content based on network measurements.
[0052] The current head position determination module 726 is configured to determine the user's current viewpoint and / or current head position, such as the position of a head-mounted display, based on information received from the playback device. In some embodiments, the playback device periodically sends current head position information to the system 700, wherein the current head position determination module 726 receives and processes the information to determine the current viewpoint and / or current head position.
[0053] Selection module 728 is configured to determine which parts of the 360-degree scene to stream to the playback device based on the user's current viewpoint / head position information. Selection module 728 is also configured to select the encoded version of the determined scene parts based on the available data rate to support content streaming.
[0054] The streaming control module 730 is configured, according to features of the invention, to control the streaming of image content (e.g., multiple portions of a 360-degree stereoscopic scene) at various supported data rates. In some embodiments, the streaming control module 730 is configured to control the streaming of N portions of the 360-degree stereoscopic scene to a playback device requesting the content, to initialize scene memory in the playback device. In various embodiments, the streaming control module 730 is configured to, for example, periodically send selected encoded versions of determined scene portions at a determined rate. In some embodiments, the streaming control module 730 is further configured to send 360-degree scene updates to the playback device according to time intervals (e.g., once per minute). In some embodiments, sending 360-degree scene updates includes sending N or Nx scene portions of the complete 360-degree stereoscopic scene, where N is the total number of portions into which the complete 360-degree stereoscopic scene has been divided, and X represents the selected scene portion most recently sent to the playback device. In some embodiments, the streaming control module 730 waits for a predetermined time after initially sending N scene segments for initialization before sending 360-degree scene updates. In some embodiments, timing information for controlling the sending of 360-degree scene updates is included in timing information 736. In some embodiments, the streaming control module 730 is further configured to identify scene segments that have not yet been transmitted to the playback device during the refresh interval; and to transmit an updated version of the identified scene segments that have not been transmitted to the playback device during the refresh interval.
[0055] In various implementations, the streaming control module 730 is configured to periodically deliver at least a sufficient number of N parts to the playback device to allow the playback device to fully refresh a 360-degree version of the scene at least once during each refresh cycle.
[0056] In some implementations, the streaming controller 720 is configured to control the system 700 to transmit stereoscopic content streams (e.g., encoded content streams 734) via a transmitter in a network interface 710, the stereoscopic content streams including those from one or more cameras (e.g., a stereoscopic camera pair, such as...). Figure 4 The image content captured (as shown) generates an encoded image. In some embodiments, the streaming controller 720 is configured to control the system 700 to transmit an environment mesh model 738 to one or more playback devices for rendering the image content. In some embodiments, the streaming controller 720 is further configured to transmit a first UV map to the playback device as part of an image rendering operation, used to map portions of the image captured by the first stereo camera to a portion of the environment mesh model.
[0057] In various embodiments, the streaming controller 720 is further configured to provide (e.g., via a transmitter in network interface 710) one or more sets of correction grid information to the playback device, such as first correction grid information, second correction grid information, third correction grid information, fourth correction grid information, fifth correction grid information, and sixth correction grid information. In some embodiments, the first correction grid information is used to render image content captured by the first camera of the first stereo camera pair, the second correction grid information is used to render image content captured by the second camera of the first stereo camera pair, the third correction grid information is used to render image content captured by the first camera of the second stereo camera pair, the fourth correction grid information is used to render image content captured by the second camera of the second stereo camera pair, the fifth correction grid information is used to render image content captured by the first camera of the third stereo camera pair, and the sixth correction grid information is used to render image content captured by the second camera of the third stereo camera pair. In some embodiments, the streaming controller 720 is further configured to, for example, instruct the playback device to use the third and fourth correction grid information when content captured by the second stereo camera pair, rather than content from the first stereo camera pair, is streamed to the playback device by sending control signals. In some implementations, the streaming controller 720 is further configured to, in response to the detection module 719 detecting i) a network control switch of streaming content from the first stereo camera pair to the second stereo camera pair, or ii) a user-controlled change from receiving a first content stream including content from the first stereo camera pair to receiving a second content stream including encoded content from the second stereo camera pair, indicate to the playback device that third and fourth correction grid information should be used.
[0058] The memory 712 also includes an environment mesh model 738, a UV map 740, and a correction mesh information set, which includes first correction mesh information 742, second correction mesh information 744, third correction mesh information 746, fourth correction mesh information 748, fifth correction mesh information 750, and sixth correction mesh information 752. The system provides the environment mesh model 738 to one or more playback devices for rendering image content. The UV map 740 includes at least a first UV map used to map portions of the image captured by the first stereo camera pair onto a portion of the environment mesh model 738 as part of the image rendering operation. The first correction mesh information 742 includes information generated based on measurements of one or more optical characteristics of the first lens of the first camera of the first stereo camera pair, and the second correction mesh includes information generated based on measurements of one or more optical characteristics of the second lens of the second camera of the first stereo camera pair. In some embodiments, the first stereo camera pair and the second stereo camera pair correspond to a forward viewing direction but to different locations at the area or event location where content is being captured for streaming.
[0059] In some embodiments, processor 708 is configured to perform various functions corresponding to the steps discussed in flowcharts 600 and / or 2300. In some embodiments, the processor uses routines and information stored in memory to perform various functions and control system 700 to operate according to the method of the invention. In one embodiment, processor 708 is configured to control system to provide playback device with first correction mesh information and second correction mesh information, the first correction mesh information being used to render image content captured by a first camera and the second correction mesh information being used to render image content captured by a second camera. In some embodiments, the first stereo camera pair corresponds to a first direction, and the processor is further configured to control system 700 to transmit a stereo content stream comprising an encoded image generated from image content captured by the first and second cameras. In some embodiments, processor 708 is further configured to transmit to playback device an environment mesh model to be used for rendering image content. In some embodiments, processor 708 is further configured to transmit to playback device a first UV map to be used for mapping portions of the images captured by the first stereo camera pair to a portion of the environment mesh model as part of an image rendering operation. In some embodiments, processor 708 is further configured to control system 700 to provide playback device with third and fourth correction grid information, the third correction grid information being used to render image content captured by the first camera of the second stereo camera pair, and the fourth correction grid information being used to render image content captured by the second camera of the second stereo camera pair. In some embodiments, processor 708 is further configured to control system 700 to instruct playback device (e.g., via network interface 710) that the third and fourth correction grid information should be used when content captured by the second camera pair, rather than content from the first camera pair, is streamed to playback device. In some embodiments, processor 708 is further configured to, in response to system detection of: i) a network control switch of streaming content from the first stereo camera pair to the second stereo camera pair, or ii) a user-controlled change from receiving a first content stream including content from the first stereo camera pair to receiving a second content stream including encoded content from the second stereo camera pair, control system 700 to instruct playback device that the third and fourth correction grid information should be used. In some implementations, the processor 708 is further configured to control the system 700 to provide the playback device with fifth correction grid information and sixth correction grid information, the fifth correction grid information being used to render image content captured by the first camera of the third stereo camera pair, and the sixth correction grid information being used to render image content captured by the second camera of the third stereo camera pair.
[0060] Figure 3A playback system 300 implemented according to an exemplary embodiment of the present invention is shown. The playback system 300 is, for example... Figure 1 The playback system 101 or playback system 111. An exemplary playback system 300 includes a computer system / playback device 800 coupled to a display 805 (e.g., a head-mounted stereoscopic display). The computer system / playback device 800 implemented according to the present invention can be used to receive, decode, store, and display content from a content delivery system (such as...). Figure 1 and Figure 2 The content delivery system shown receives the imaged content. The playback device can be used with a 3D head-mounted display such as the OCULUS RIFT™ VR (virtual reality) headset, which can be a head-mounted display 805. Device 800 includes the ability to decode the received encoded image data and generate 3D image content to display to a client. In some embodiments, the playback device is located at a client's location, such as a residence or office, but may also be located at an image capture site. Device 800 can perform signal reception, decoding, display, and / or other operations according to the invention.
[0061] Device 800 includes a display 802, a display device interface 803, an input device 804, a microphone (mic) 807, an input / output (I / O) interface 806, a processor 808, a network interface 810, and a memory 812. The various components of playback device 800 are coupled together via a bus 809, which allows data transfer between components of system 800. While in some embodiments, display 802 is included as an optional element as shown using dashed boxes, in some embodiments, an external display device 805 (e.g., a head-mounted stereoscopic display device) may be coupled to the playback device via display device interface 803.
[0062] Through I / O interface 806, system 800 can be coupled to external devices to exchange signals and / or information with other devices. In some embodiments, through I / O interface 806, system 800 can receive information and / or images from external devices and output information and / or images to external devices. In some embodiments, through interface 806, system 800 can be coupled to an external controller, such as a handheld controller.
[0063] Processor 808 (e.g., CPU) executes routines 814 and modules in memory 812 and uses the stored information to control system 800 to operate according to the invention. Processor 808 is responsible for the overall general operation of system 800. In various embodiments, processor 808 is configured to perform functions discussed to be performed by playback system 800.
[0064] Via network interface 810, system 800 transmits signals and / or information (e.g., including encoded images and / or video content corresponding to a scene) to and / or receives signals and / or information from various external devices via a communication network (e.g., such as communication network 105). In some embodiments, the system receives one or more content streams, including encoded images captured by one or more different cameras, from content delivery system 700 via network interface 810. The received content streams may be stored as received encoded data, such as encoded image 824. In some embodiments, interface 810 is configured to receive a first encoded image including image content captured by a first camera and a second encoded image corresponding to a second camera. Network interface 810 includes a receiver and a transmitter, through which receiving and transmitting operations are performed. In some embodiments, interface 810 is configured to receive correction mesh information corresponding to multiple different cameras, including first correction mesh information 842, second correction mesh information 844, third correction mesh information 846, fourth correction mesh information 848, fifth correction mesh information 850, and sixth correction mesh information 852, which are then stored in memory 812. Furthermore, in some embodiments, via interface 810, the system receives one or more masks 832, an environment mesh model 838, and a UV map 840, which are then stored in memory 812.
[0065] The memory 812 includes various modules, such as routines, which, when executed by the processor 808, control the playback device 800 to decode and output operations according to the present invention. The memory 812 includes a control routine 814, a request to the content generation module 816, a head position and / or viewpoint determination module 818, a decoder module 820, a stereoscopic image rendering engine 822 (also referred to as a 3D image generation module and determination module), and data / information including received encoded image content 824, decoded image content 826, a 360-degree decoded scene buffer 828, generated stereoscopic content 830, a mask 832, an environment mesh model 838, a UV map 840, and multiple received sets of correction mesh information, including a first correction mesh information 842, a second correction mesh information 844, a third correction mesh information 846, a fourth correction mesh information 848, a fifth correction mesh information 850, and a sixth correction mesh information 852.
[0066] Control routine 814 includes device control routines and communication routines to control the operation of device 800. Request generation module 816 is configured to generate a request for content to be sent to the content delivery system to provide the content. In various embodiments, the request for content is sent via network interface 810. Head position and / or viewing angle determination module 818 is configured to determine the user's current viewing angle and / or current head position, such as the position of a head-mounted display, and report the determined position and / or viewing angle information to content delivery system 700. In some embodiments, playback device 800 periodically sends current head position information to system 700.
[0067] Decoder module 820 is configured to decode encoded image content 824 received from content delivery system 700 to generate decoded image data, such as decoded image 826. Decoded image data 826 may include a decoded stereoscopic scene and / or a decoded scene portion. In some embodiments, decoder 820 is configured to decode a first encoded image to generate a first decoded image, and to decode a received second encoded image to generate a second decoded image. The decoded first and second images are included in the stored decoded image 826.
[0068] The 3D image rendering engine 822 performs rendering operations (e.g., using content and information received and / or stored in memory 812, such as decoded image 826, environment mesh model 838, UV map 840, mask 832, and mesh correction information) and generates a 3D image according to the features of the invention for display to a user on display 802 and / or display device 805. The generated stereoscopic image content 830 is the output of the 3D image generation engine 822. In various embodiments, the rendering engine 822 is configured to perform a first rendering operation using first correction information 842, a first decoded image, and environment mesh model 838 to generate a first image for display. In various embodiments, the rendering engine 822 is further configured to perform a second rendering operation using second correction information 844, a second decoded image, and environment mesh model 838 to generate a second image for display. In some such embodiments, the rendering engine 822 is further configured to perform the first and second rendering operations using a first UV map (included in the received UV map 840). When a first rendering operation is performed to compensate for distortion introduced into the first image by the lens of the first camera, first correction information provides information about the corrections to be made to the node positions in the first UV map. Similarly, when a second rendering operation is performed to compensate for distortion introduced into the second image by the lens of the second camera, second correction information provides information about the corrections to be made to the node positions in the first UV map. In some embodiments, the rendering engine 822 is further configured to use a first mask (included in mask 832) to determine how portions of the first image are combined with portions of the first image corresponding to different fields of view when a portion of the first image is applied to the surface of the environment mesh model as part of the first rendering operation. In some embodiments, the rendering engine 822 is further configured to use the first mask to determine how portions of the second image are combined with portions of the second image corresponding to different fields of view when a portion of the second image is applied to the surface of the environment mesh model as part of the second rendering operation. The generated stereoscopic image content 830 includes a first image and a second image (e.g., corresponding to a left-eye view and a right-eye view) generated as a result of the first and second rendering operations. In some embodiments, a portion of the first image corresponding to a different field of view corresponds to either a sky or a ground field of view. In some embodiments, the first image is a left-eye image corresponding to the forward field of view, and the first image corresponding to a different field of view is a left-eye image captured by a third camera corresponding to a side field of view adjacent to the forward field of view. In some embodiments, the second image is a right-eye image corresponding to the forward field of view, and the second image corresponding to a different field of view is a right-eye image captured by a fourth camera corresponding to a side field of view adjacent to the forward field of view. Therefore, the rendering engine 822 renders the 3D image content 830 to the display.In some implementations, the operator of the playback device 800 can control one or more parameters via the input device 804 and / or the selection operation to be performed (e.g., selecting to display a 3D scene).
[0069] Network interface 810 allows the playback device to receive content and / or transmit information from streaming device 114, such as viewing head position and / or position (camera setup) selection indicating the selection of a specific viewing position at the event. In some embodiments, decoder 820 is implemented as a module. In such embodiments, when executed, decoder module 820 causes the received images to be decoded, while 3D image rendering engine 822 causes the images to be further processed according to the invention and optionally stitched together as part of the rendering process.
[0070] In some embodiments, interface 810 is further configured to receive additional mesh correction information corresponding to multiple different cameras, such as third mesh correction information, fourth mesh correction information, fifth mesh correction information, and sixth mesh correction information. In some embodiments, rendering engine 822 is also configured to use mesh correction information corresponding to a fourth camera (e.g., fourth mesh correction information 848) when rendering an image corresponding to a fourth camera, which is one of multiple different cameras. Determination module 823 is configured to determine which mesh correction information the rendering engine 822 should use based on which camera's image content is used for the rendering operation or based on an instruction from a server indicating which mesh correction information should be used when rendering an image corresponding to the received content stream. In some embodiments, determination module 823 may be implemented as part of rendering engine 822.
[0071] In some implementation schemes, Figure 2 The memory 712 and Figure 3 The modules and / or elements shown in memory 812 are implemented as software modules. In other embodiments, although the modules and / or elements are shown as being included in memory, they are implemented in hardware as, for example, separate circuits, where each element is implemented as circuitry for performing a function corresponding to that element. In other embodiments, the modules and / or elements are implemented using a combination of software and hardware.
[0072] Although Figure 2 and Figure 3Elements shown as included in memory, but shown as included in systems 700 and 800, can and in some embodiments are implemented entirely in hardware within the processor, for example as separate circuitry of the corresponding device, such as within processor 708 in the case of a content delivery system, and within processor 808 in the case of a playback system 800. In other embodiments, some elements are implemented, for example as circuitry within the corresponding processors 708 and 808, while other elements are implemented, for example as circuitry outside the processor and coupled to the processor. It should be understood that the level of integration of modules on the processor and / or the level of integration of some modules outside the processor can be one of the design choices. Alternatively, instead of being implemented as circuitry, all or some of the elements can be implemented in software and stored in memory, wherein the software module controls the operation of the corresponding systems 700 and 800 to perform the functions corresponding to the module when the module is executed by its corresponding processor (e.g., processors 708 and 808). In other embodiments, various elements are implemented as a combination of hardware and software, for example, wherein circuitry outside the processor provides input to the processor, and that input then operates under software control to perform a portion of the module's functionality.
[0073] Although Figure 2 and Figure 3 Each of the embodiments is shown as a single processor, such as a computer, but it should be understood that each of processors 708 and 808 may be implemented as one or more processors, such as a computer. When one or more elements of memories 712 and 812 are implemented as software modules, the modules include code that, when executed by the processors of the corresponding system (e.g., processors 708 and 808), configures the processor to implement the functions corresponding to the module. Figure 7 and Figure 8 In the embodiments shown, the various modules are stored in a memory, which is a computer program product including a computer-readable medium comprising code for causing at least one computer (e.g., a processor) to perform the functions corresponding to the modules, such as separate code for each module.
[0074] Modules that are entirely hardware-based or entirely software-based can be used. However, it should be understood that any combination of software and hardware (e.g., circuit-implemented modules) can be used to implement these functions. Figure 2 The illustrated module controls and / or configures the system 700 or its components, such as the processor 708, to perform the functions of corresponding steps of the method of the present invention, such as those shown and / or described in the flowchart. Similarly, Figure 3 The module control and / or configuration system 800, or its components, such as processor 808, shown herein, functions to perform the corresponding steps of the method of the present invention, such as those shown and / or described in the flowchart.
[0075] To facilitate understanding of the image capture process, reference will now be made. Figure 4 The exemplary camera equipment shown. Camera equipment 1300 can be used as... Figure 1 The system is equipped 102 and includes multiple stereo camera pairs, each corresponding to a different of the three sectors. A first stereo camera pair 1301 includes a left-eye camera 1302 (e.g., a first camera) and a right camera 1304 (e.g., a second camera), designed to capture images corresponding to those seen by the left and right eyes of a person positioned at the location of the first camera pair. A second stereo camera pair 1305 corresponds to a second sector and includes a left camera 1306 and a right camera 1308, while a third stereo camera pair 1309 corresponds to a third sector and includes a left camera 1310 and a right camera 1312. Each camera is mounted in a fixed position within a support structure 1318. An upward-facing camera 1314 is also included. Figure 4 An invisible, downward-facing camera may be included below camera 1314. Stereo camera pairs are used in some embodiments to capture paired upward and downward images; however, in other embodiments, a single upward camera and a single downward camera are used. In other embodiments, a downward image is captured before the equipment is placed and used as a static ground image for the duration of the event. Given that the ground view tends not to change significantly during the event, this approach is often satisfactory for many applications. The output of the camera equipped with 1300 is captured and processed.
[0076] When using Figure 4 When the camera is equipped, each sector corresponds to a known 120-degree viewing area relative to the camera's position, where captured images from different sector pairs are stitched together based on known images mapped to the simulated 3D environment. While typically a 120-degree portion of each image captured by the sector camera is used, the camera captures a wider image corresponding to an approximately 180-degree viewing area. Therefore, the captured images can be masked in the playback device as part of the 3D environment simulation. Figure 5 This is a composite diagram 1400 illustrating how an environment mesh portion can be used to simulate a 3D spherical environment, the environment mesh portion corresponding to different camera pairs of the apparatus 102. Note that one mesh portion of each sector in the apparatus 102 is shown, where the sky mesh is used relative to the top camera view, and the ground mesh is used for the ground image captured by the downward-facing camera. Although the masks used for the top and bottom images are essentially circular, the masks applied to the sector images are truncated so that the top and bottom portions of the scene area will be provided by the top and bottom cameras, respectively.
[0077] When combined, the overall grid corresponding to different cameras generates a spherical grid, such as... Figure 6 As shown. Note that the grid is shown for monocular images, but in the case of capturing stereo image pairs, it is used for both left-eye and right-eye images.
[0078] Figure 5 The type of mesh and mask information shown can and sometimes is transmitted to the playback device. The transmitted information will vary depending on the equipment configuration. For example, if a larger number of sectors are used, the mask corresponding to each sector will correspond to an observation area of less than 120 degrees, where more than 3 environment meshes are needed to cover the diameter of the sphere.
[0079] Environmental calibration information is shown as optionally being transmitted to the playback device in step 1132. It should be understood that environmental calibration information is optional, as the environment can be assumed to be within a default size range if such information is not transmitted. In cases where multiple different default sphere sizes are supported, instructions regarding which sphere size to use can and sometimes are transmitted to the playback device.
[0080] Image capture operations can be performed continuously during the event, specifically with respect to each of the three sectors that can be captured by camera equipment 102.
[0081] It should be noted that although multiple camera views corresponding to different sectors are captured, the image capture rate does not need to be the same for all sectors. For example, the front-facing sector corresponding to, for example, the main playback field can capture images at a faster frame rate than the cameras corresponding to other sectors and / or the top (sky) view and the bottom (ground) view.
[0082] Figure 7 The mapping from the image portion corresponding to the first sector to the corresponding 120-degree portion of the sphere representing the 3D viewing environment is shown.
[0083] Images corresponding to different parts of the 360-degree environment are combined to the extent necessary to provide the observer with a continuous viewing area, for example, depending on head position. For instance, if the viewer is viewing the intersection of two 120-degree sectors, the portions of the image corresponding to each sector will be stitched together and presented to the viewer based on the known angles and positions of each image within the entire 3D environment being simulated. Image stitching and generation will be performed for each of the left and right eye views, such that, in the case of a stereoscopic implementation, two separate images are generated, one for each eye.
[0084] Figure 8This demonstrates how, and sometimes how, multiple decoded, corrected, and cropped images can be calibrated and stitched together to create a 360-degree viewing environment that can be used as a background for a foreground image of an object represented by point cloud data.
[0085] Figure 9 Methods for capturing, processing, and delivering captured image content, as well as interpolated content, are illustrated. Figure 9 The method shown can be implemented by combining a stereoscopic image capture system 102 with a content delivery system 104. The content delivery system in... Figure 1 The image processing device 112 and the content delivery device 114 are shown as a combination of separate image processing devices 112 and 114, but in some embodiments, these components are implemented as a single device, such as... Figure 2 The content delivery system 700 shown can and sometimes is used as Figure 1 The content delivery system 104. In some embodiments, the processor 708 in the content delivery system communicates with the cameras of a camera pair (e.g., left-eye camera 1302 and right-eye camera 1304) and controls one or more cameras in the camera pair to capture images of the environment at a first frame rate and provide the captured images to the content delivery system 104 for processing. In some embodiments, high-resolution images are captured at the first frame rate, but the image content is provided to playback devices 122, 126 to support a second, higher frame rate, wherein interpolation is performed by the content delivery system to generate image data to support the higher frame rate. This allows the cameras to operate in a high-resolution operating mode, where a high level of detail is captured, but could potentially be at a slower frame rate if the cameras instead capture lower-resolution images. The capture of high-resolution content provides the detail that allows for a high-quality virtual reality experience, while interpolation facilitates the frame rate, which allows for a realistic motion experience when supporting 3D.
[0086] In some implementation schemes, the use of Figure 1 The system shown implements Figure 9 Method 900 shown, wherein Figure 2 Content delivery system 700 in Figure 1 The system is used as a content delivery system 104, and images are captured by a camera in camera equipment 102, which sometimes uses... Figure 4 This can be achieved, for example, under the control of a content delivery system and / or a system operator.
[0087] In start step 902, method 900 begins with powering on components of system 100, such as content delivery system 104 and camera assembly 102, and initiating operation. Operation proceeds from start step 902 to image capture step 904, where the left camera 1302 and right camera 1304 of the camera pair of camera assembly 102 are used to capture images, such as left-eye and right-eye images, respectively, at a first image capture rate and a first resolution. In some embodiments, the first image resolution is a first maximum resolution supported by cameras 1302, 1304, and the first image (e.g., frame) capture rate is lower than the second image (e.g., frame) capture supported when the first and second cameras operate at a lower resolution. Thus, in at least some embodiments, the cameras are operated to maximize the level of detail captured, although at the tradeoff of a lower image capture rate than that achievable with a lower image resolution.
[0088] In step 908, the image 908 captured by the camera in step 902 is stored in a frame buffer 715, which may include, and sometimes includes, in the memory 712 of the content delivery system. Capture and storage can and sometimes are performed on a continuous basis when capturing and processing images of ongoing events or sporting events.
[0089] The high-resolution image, captured at a first frame rate and usable for processing, proceeds to step 910, where motion analysis is performed on the frames (e.g., captured frames in a video sequence, such as a left-eye frame and / or right-eye frame sequence) to detect motion and identify moving segments. In at least one such embodiment, a block is a rectangular portion of a frame, and a segment is a group of one or more moving blocks. A segment may correspond to, and sometimes does indeed correspond to, an object, such as a ball, whose position changes over time within the environment of the captured image, such as a stadium or sports field.
[0090] The motion analysis performed in step 910 allows for the identification of image (e.g., frame) fragments that, if not updated at a rate higher than the first frame rate during playback, may, and in some cases will, adversely affect the 3D playback experience. The processing in step 910 and other steps in flowchart 900 can, and in some embodiments, be implemented by and / or under the control of processor 708 of the content delivery system, which controls the components of system 100 to achieve… Figure 9 The steps of method 900 are shown.
[0091] In some embodiments, step 910 includes one, more, or all of steps 912, 914, and 916. In step 912, the images of the captured video sequence are analyzed to identify matching groups of blocks, such as pixel blocks corresponding to objects appearing in consecutive frames of the image sequence (e.g., a sequence of left-eye and / or right-eye image frames). Then, in step 914, the identified matching groups of blocks are defined as segments. In some such embodiments, each segment is a set of physically adjacent blocks, such as a set of blocks corresponding to a physical object moving within the environment of the captured images. In at least some embodiments, information regarding the position and size defining the segment (e.g., an object image) is stored in step 914, where the position information is stored on a per-frame basis. In step 916, the positional change of the segment from one frame time to the next frame time in the captured frame is determined, and positional change information is stored on a per-segment basis (e.g., an object). In some embodiments, the change information is stored as a motion vector indicating how the segment moves position from the previous frame (e.g., the immediately preceding frame in a frame sequence). By analyzing the positional changes of segments corresponding to objects such as a ball in a frame over time, the movement of the segment from one frame time to the next frame time, as well as the ball's position in the environment and its position relative to the camera lens used to capture the image of the environment, can be determined. In step 916, the segment definition and positional change information are stored in portion 717 of memory 712 in some embodiments. The motion vectors determined for one or more segments in step 916 can be used, and sometimes used, to generate interpolated frames.
[0092] With image segments defined and motion vectors for segments indicating segment motion between capture frames generated, the operation proceeds from step 910 to step 918.
[0093] In step 918, the motion vector corresponding to the segment is considered, and the segment velocity is determined. It is also determined which part of the camera's fisheye lens the segment corresponds to for capturing it. Determining the part of the fisheye lens used to capture the image segment and how much the segment will move is likely important because some parts of the fisheye lens are more curved than others. Image segments that may be captured by the same part of the lens over a prolonged time period may suffer less lens-introduced distortion between frame times than segments that will move more in terms of curvature in the lens region used to capture the image segment. Therefore, step 918 allows the effects of motion to be considered in conjunction with lens curvature when determining the rate at which the segment should be presented to the user and, therefore, the interpolation rate (if any) that should be supported for a particular segment.
[0094] With the effect of the fisheye lens FOV now considered and quantified in terms of segment motion relative to the lens FOV, the operation proceeds from step 918 to step 920. In step 920, information defining each captured frame is stored, including segment information and image pixel values. Also stored as part of the frame information are pixel values representing the captured image. The frame information and data may be stored in the segment portion of the frame buffer 715 and / or memory 717.
[0095] The operation proceeds from step 920 to step 922. In step 922, the frame rate of at least some image segments is determined based on the amount of motion of one or more segments. In some embodiments, the frame rate determination is performed on a per-segment basis. In some cases, step 922 includes step 923, which involves determining the segment frame rate based not only on the amount or rate of motion of the segment but also on the position of the segment in the field of view during the time period corresponding to the frame rate determination. In some cases, different frame rates are determined for segments of the same size corresponding to different portions of the field of view. In some embodiments, segments that change position relative to the camera field of view corresponding to a more curved lens portion are given a higher frame rate than segments that change position in the FOV portion corresponding to a less curved lens portion. In this way, in at least some embodiments, the effect of lens appearance distortion can be reduced by supporting higher frame rates for segments that are more affected by the combination of motion and lens shape than for segments that are less affected by the combination of motion and lens shape. In at least some cases, the frame rate of segments undergoing high motion is selected to be greater than the actual captured image frame rate. Segments that are not subjected to motion or are subjected to motion rates below a threshold (e.g., a predetermined or variable segment rate motion threshold used to control whether interpolation is performed on a segment) are determined to have a selected frame rate equal to the image capture rate (e.g., a first frame rate) and therefore will not be interpolated. The interpolation segment motion threshold can and sometimes varies depending on the amount of data available to be transmitted to the playback device (e.g., the data rate). Thus, for a first low level of data rate used to transmit data to the playback device, in some embodiments, a first higher threshold is required for the segment interpolation to be performed, and therefore a higher motion rate is required, while for a second higher level of data rate used to transmit data to the playback device, in some embodiments, a second lower threshold is required for the segment interpolation to be performed, and therefore a lower motion rate is required.
[0096] With the frame rate determined for the identified image segments in steps 922 and / or 923, the operation proceeds to step 924. In step 924, segments are interpolated between captured frames to generate a set of interpolated segments for each non-captured frame. For non-captured frames, frame information is passed in the transport stream, wherein different numbers of interpolated segments are generated and passed for at least some different interpolated (e.g., non-captured) frames.
[0097] By making decisions about which segments to interpolate before transmission to the playback device, information typically unavailable to the playback device, such as the shape of the camera lens used to capture the image and the position of the segments relative to the camera lens's field of view (FOV), can be used to determine the frame rate to support and which segments to interpolate for a specific frame time. Furthermore, the relatively powerful processing resources are often available in the distribution system compared to the playback device. This is because the distribution system can perform content interpolation and encoding for multiple devices, allowing the use of relatively powerful and expensive processors or processor sets to support distribution. Given that playback devices are typically owned by individual customers, their cost and corresponding processing power are usually much smaller than those of the distribution system. In fact, in many cases, the playback device can discover the normal decoding and display of left-eye and right-eye image content to support the computational classification of the 3D virtual reality experience, even if it does not support interpolation between captured frames. Therefore, for various reasons, it is beneficial to shift interpolation-related decisions and processing to the distribution side rather than the playback side. Furthermore, by using the interpolation according to the invention, the distribution system can deliver frames at a higher frame rate than the image capture rate, thereby allowing a more enjoyable and realistic viewing experience than reducing the resulting image capture and losing the details needed for an actual 3D experience or making the changes between frames too large due to motion.
[0098] In step 924, the segments are interpolated to support the frame rate selected for the segments in step 922. Therefore, segments for which different frame rates are selected (i.e., determined) in step 922 and determined (i.e., selected) in step 920 will be interpolated with their different corresponding frame rates. Segments for which an image capture rate is selected will not undergo interpolation. The interpolated image data is used as interpolated frame information. The interpolated frame information is typically much smaller, for example, less than 1 / 20th of the amount of data used to represent the captured image frame, or in some cases less than 1 / 200th. Therefore, in at least some cases, the interpolated frame information can be passed to the playback device without significantly affecting the total amount of data transmitted. The interpolated frame information may include interpolated segments passed as intra-frame coded image data and / or inter-frame coded image data. The interpolated frame information generated in step 924 may, and sometimes may, also include padding information that provides instructions to the playback device, wherein missing frame information is obtained and / or how to pad a portion of the interpolated frame, which, in the absence of a padding operation, may have gaps due to the movement of the segments relative to the previous frame. In various implementations, the padding image data is obtained from either the previous or next frame. The padding portion may be part of an environment occluded by a stationary but moving object. Such padding data may be obtained from a frame that is one or more capture frame times away from the interpolation frame time.
[0099] If interpolated fragments and / or other interpolated frame data have been generated in step 924, the operation proceeds to encoding step 925, in an implementation where frame encoding is performed, or directly to delivery step 926. In step 926, the captured frame data and interpolated frame data are subsequently delivered to at least one playback device, for example, after being stored in memory 700 of the content delivery system.
[0100] In step 925, when used, the captured frames and / or interpolated frames are encoded. MPEG or other image encoding content can be used to encode the captured image frames. In some embodiments, left-eye and right-eye image data corresponding to frame time are incorporated into a single frame for encoding purposes, while in other embodiments, the left-eye and right-eye image data are encoded as separate image streams and / or interleaved image streams.
[0101] Intra-frame and / or intra-frame coding techniques can be used and sometimes used to encode captured frames and interpolated frames.
[0102] It should be understood that capture frames and interpolated frames can be delivered to the playback device in encoded or unencoded form. In most embodiments, the encoding of the captured image frames, which can be considered as the base view layer, is implemented together with the encoding of the interpolated frames. In step 925, the encoding of the frame information to be stored and transmitted to the playback device is implemented. Then, in step 926, the frame data is transmitted to the playback device. In some embodiments, step 926 includes steps 927 and 928. Capture frame data (e.g., capture frames) is transmitted to the playback device in encoded form if encoding is used, and in unencoded form if encoding is not used in step 925. In some embodiments, the capture frames form the base video layer transmitted in step 927. In step 928, interpolated frame information is transmitted to the playback device. The interpolated frame information may include encoded interpolated frames that convey interpolated segments and / or padding information that can be used to construct a complete interpolated frame. In some cases, the motion vectors included in the interpolation frame information indicate where in the captured frame, and in other cases, in another interpolation frame, the fragment content to be included in the interpolation frame can be found, as well as the distance the content will be moved from its original position in the source frame to form part of the interpolation frame. Therefore, the motion vectors corresponding to the fragments in the interpolation frame can, and sometimes do, indicate the source of the fragment image content to be included in the interpolation frame generated using the fragment motion vectors, and the position of the fragment within the frame where it will be located.
[0103] In other implementations, interpolation frame information is passed as a segment motion vector and / or other information about how the remainder of the interpolation frame is constructed from one or more other frames, such as from which padding content is obtained and placed at a position in the generated interpolation frame, where the placement position is specified by information in the interpolation frame information.
[0104] While the various accompanying figures illustrate a transmission order that matches the expected frame display order, it should be understood that the actual transmission order can vary as transmitted frames are buffered and reordered by both the content delivery system and the playback system. In such cases, if the content delivery system reorders captured and interpolated frames for transmission, the playback device will overlay and restore the frames to the appropriate display order before displaying them.
[0105] Although flowchart 900 ends at step 928, it should be understood that the capture, processing, and transmission of content can be performed on a continuous basis or over a period of time, for example, corresponding to an event. Therefore, the steps of method 900 are repeatedly performed as new content (e.g., an image of the environment) is captured, processed, and stored for future transmission or delivery to one or more playback systems.
[0106] It should be understood that, according to Figure 9The method, which includes capturing images (e.g., frames) and interpolating frame data (e.g., interpolated frames), in some embodiments, passes the content to one or more playback systems 101, 111 for decoding and display, for example as part of a virtual reality experience, in which separate left-eye and right-eye images are presented to the user to provide a 3D experience. Figure 10 The illustration 1000 shows how the Transmit Frame Rate (TFR) 1002 generated by the Content Delivery System 104 exceeds the Capture Frame Rate (CFR) in various embodiments of the present invention.
[0107] Figure 11 An exemplary overall transmission frame sequence 1100 is illustrated, where CF is used to indicate capture frames, for example, transmitted as part of a base video layer, and IF is used to indicate interpolated frame information transmitted as an enhancement layer in some embodiments, wherein each set of interpolated frame information includes one or more segments, for example, generated by using motion interpolation. Line 1106 is used to indicate transmission frame sequence numbers. Exemplary frame group 1102 includes a first capture frame, which is capture frame 1 (CF1 1108), followed by a plurality of interpolated frames 1104 (interpolated frame 1 (IF1) 1110, interpolated frame 2 (IF2) 1112, interpolated frame 2 (IF3) 1114...), followed by a second capture frame (CF2) 1116. CF1 1008 corresponds to transmission frame number 1; IF1 1110 corresponds to transmission frame number 2; IF2 1112 corresponds to transmission frame number 2; IF3 1114 corresponds to transmission frame number 4; CF2 1116 corresponds to transmission frame number Y, for example, where Y = the number of interpolated frames in frame 1104 + 2.
[0108] Figure 12 Figure 1200 shows capture frames (capture frame 1 (CF1) 1202, capture frame 2 (CF2) 1204, capture frame 3 (CF3) 1206, ..., capture frame N (CFN) 1208) used as base frames, having circles (1203, 1205, 1207, ..., 1209), respectively corresponding to portions of frames (CF1 1202, CF2 1204, CF3, 1206, ..., CFN 1208), which include image content corresponding to the fisheye lens used to capture the image, and edge portions ((1210, 1212, 1214, 1215), (1218, 1220, 1222, 1224), (1226, 1228, 1230, 1232), ..., (1234, 1236, ...)) used as base frames. The regions (1238, 1240) are outside the circles (1203, 1205, 1207, …, 1209) respectively, and are not of interest in some implementations because they will not be used as textures and correspond to areas outside the region of interest of the captured scene.
[0109] Figure 13 Figure 13000 illustrates how information used to reconstruct a specific segment can be transmitted as interpolated frame data. It should be noted that information for different segments can be transmitted at different frame rates, and some interpolated frame times and updated segment information for some segments of the frame set are omitted. When determining the frame rate of an individual segment, the segment's position relative to the fisheye lens of the captured image and the direction of motion of objects within the segment (as indicated by directional arrows) can and sometimes is considered. Segment information is transmitted as supplementary information and is sometimes referred to as side information because it is information transmitted in addition to the underlying video frame layer information. In some implementations, the segment information for interpolated frames is generated by using motion interpolation before encoding and transmission to the playback device.
[0110] Figure 13000 includes a set of frames, including capture frame 1 (CF1) 13002, six interpolated frames (IF1 13004, IF2 13006, IF3 13008, IF4 13010, IF5 13012, IF6 13014) and capture frame 2 (CF1) 13016. Five exemplary segments of interest (segment 1 (S1) 13050, segment 1 (S2) 13052, segment 3 (S3) 13054, segment 4 (S4) 13056, segment 5 (S5) 13058) are identified within the image capture area 13003 of capture frame 1 (CF1) 13002. The direction of motion of the objects in each segment (S1 13050, S2 13052, S3 13054, S4 13056, S5 13058) is represented by directional arrows (13051, 13053, 13055, 13057, 13059), respectively.
[0111] Five exemplary segments (segment 1 (S1) 13050', segment 1 (S2) 13052', segment 3 (S3) 13054', segment 4 (S4) 13056', and segment 5 (S5) 13058') are identified within the image capture area 13017 of capture frame 2 (CF2) 13016. Exemplary segment 1 (S1) 13050' of CF2 13016 includes the same object as segment 1 (S1) 13050 of CF1 13002. Exemplary segment 2 (S2) 13052' of CF2 13016 includes the same object as segment 2 (S2) 13052 of CF1 13002. Exemplary segment 3 (S3) 13054' of CF2 13016 includes the same object as segment 3 (S3) 13054 of CF1 13002. Exemplary fragment 4 (S4) 13056' of CF213016 includes the same object as fragment 4 (S4) 13056 of CF1 13002. Exemplary fragment 5 (S5) 13058' of CF213016 includes the same object as fragment 5 (S5) 13058 of CF1 13002.
[0112] The direction of motion of the objects in each segment (S1 13050', S2 13052', S3 13054', S4 13056', S5 13058') is represented by directional arrows (13051', 13053', 13055', 13057', 13059').
[0113] Figure 14 The graph 14000 shows how velocity vectors and various rates, as well as the position relative to the fisheye lens capture area (13003), can be considered to determine the frame rate that should be supported by interpolation. Figure 14 The identified motion segments (S1 13050, S2 13052, S3 13054, S4 13056, S5 13058) are shown, each with a defined corresponding velocity vector (13051, 13053, 13055, 13057, 13059). The selected rates (R1 14002, R2 14004, R3 14006, R4 14008, R5 14010) used for interpolation are based on the determined velocity vectors (13051, 13053, 13055, 13057, 13059), respectively, and are based on the positions of segments (S1 13050, S2 13052, S3 13054, S4 13056, S5 13058) in the field of view 13003 of the capture frame 13002.
[0114] Figure 15Figure (1500) shows the various interpolation frame times and which segments will be interpolated and passed based on the selected frame rates to be supported by a particular segment. Note that different frame rates will be supported for different segments, where Y indicates that information for that segment will be included, and N indicates that for the given frame times listed in the top row, no interpolation information will be generated and passed as side information. Note that a quantity subscript is used after IF to indicate the interpolated frame associated with the frame information. For example, IF1 corresponds to interpolated frame 1, and IF2 corresponds to interpolated frame 2.
[0115] The first column 1502 includes information about the motion segment in each row of the identification table. The second column 1504 includes information identifying the selected rate used for interpolation of each motion segment. The third column 1508 includes information indicating whether interpolation information for the motion segment should be generated, included, and transmitted for interpolation frame IF1. The fourth column 1510 includes information indicating whether interpolation information for the motion segment should be generated, included, and transmitted for interpolation frame IF2. The fifth column 1512 includes information indicating whether interpolation information for the motion segment should be generated, included, and transmitted for interpolation frame IF3. The fifth column 1514 includes information indicating whether interpolation information for the motion segment should be generated, included, and transmitted for interpolation frame IF5. The sixth column 1516 includes information indicating whether interpolation information for the motion segment should be generated, included, and transmitted for interpolation frame IF6.
[0116] Line 1518 includes information identifying that for motion segment S1, the selected rate is R1, and generating, including, and transmitting interpolation information for motion segment S1 for each of interpolation frames IF1, IF2, IF3, IF4, IF5, and IF6.
[0117] Line 1520 includes information identifying that for motion segment S2, the selected rate is R2, and interpolation information to be generated, included, and transmitted for each of interpolation frames IF2, IF3, IF4, and IF5 but not for interpolation frames IF1 and IF6.
[0118] Line 1522 includes information identifying that for motion segment S3, the selected rate is R3, and generating, including, and transmitting interpolation information for motion segment S3 for each of interpolation frames IF3, IF4, and IF5 but not for interpolation frames IF1, IF2, and IF6.
[0119] Line 1524 includes information identifying that for motion segment S4, the selected rate is R4, and generating, including, and transmitting interpolation information for motion segment S4 for each of interpolation frames IF3 and IF4 but not for interpolation frames IF1, IF2, IF5, and IF6.
[0120] Line 1526 includes information identifying that for motion segment S5, the selected rate is R5, and interpolation information for motion segment S4 to be generated, included, and transmitted for interpolation frames IF3 but not for interpolation frames IF1, IF2, IF4, IF5, and IF6.
[0121] Below is a list of exemplary embodiments using the numbering. The number in each list is used to refer to an embodiment included in the list using the numbering.
[0122] First list of implementation schemes for numbering method
[0123] Method Implementation Scheme 1 A content distribution method comprising: storing (908) an image (751) captured at a first frame rate; performing interpolation (924) to generate interpolated frame data to support a second frame rate higher than the first frame rate; and transmitting (926) the captured frame data and the interpolated frame data to at least one playback device.
[0124] Method Implementation Scheme 2 is based on the method described in Method Implementation Scheme 1, wherein the first frame rate is the highest frame rate supported by the camera used to capture images when the camera is operating at the maximum image capture resolution.
[0125] Method implementation scheme 3, according to the method implementation scheme 1, further includes: operating the camera of the first camera in the camera pair to capture the image at a first rate.
[0126] Method implementation scheme 4 is based on the method implementation scheme 3, wherein when capturing images at a lower resolution, the first camera supports a higher frame rate than the first frame rate.
[0127] Method implementation scheme 5 is based on the method implementation scheme 1, wherein transmitting capture frame data includes transmitting the capture frame at a first data rate corresponding to the image capture rate.
[0128] Method implementation scheme 6 is based on the method described in method implementation scheme 5, wherein the combination of captured frame data and interpolated frame data corresponds to a second frame rate that the playback device will use to display the image.
[0129] Method Implementation Scheme 7 is based on the method described in Method Implementation Scheme 6, wherein the second frame rate is a stereoscopic frame rate at which left-eye and right-eye images are displayed to the user of the playback device to support a 3D viewing experience.
[0130] Method Implementation Scheme 7A is based on the method described in Method Implementation Scheme 7, wherein the interpolated frame data transmission interpolated frame, which supplements the captured frame, thereby increasing the frame rate from the captured frame rate to a second frame rate.
[0131] Method implementation scheme 8, according to the method implementation scheme 1, further includes: performing (910) motion analysis on blocks of captured frames to detect motion and identify moving segments; selecting (922) a frame rate for at least a first segment based on the amount of motion of the first segment from the current captured frame and one or more other captured frames (e.g., at least the next captured frame); and wherein performing interpolation (924) to generate interpolated frame data includes interpolating segments between captured frames to generate interpolated frame information.
[0132] Method Implementation Scheme 9 is based on the method described in Method Implementation Scheme 8, wherein the at least one segment is an image segment of a moving object.
[0133] Method implementation scheme 9A is based on the method described in method implementation scheme 8, wherein the moving object is a ball.
[0134] Method implementation scheme 10 is based on the method implementation scheme 8, wherein the interpolated segment includes a version of the first segment interpolated to support the selected frame rate.
[0135] Method implementation scheme 11, according to the method implementation scheme 8, further includes: transmitting (926) the captured frame as part of the base video layer to the playback device; and transmitting (928) information corresponding to one or more interpolated segments corresponding to the interpolated frame time as part of additional video information.
[0136] Method implementation scheme 11A is the method according to method implementation scheme 8, wherein the base video layer includes 3D video information in the form of multiple left-eye capture frames and right-eye capture frames.
[0137] Method implementation scheme 11B is based on the method described in method implementation scheme 11A, wherein transmitting (926) capture frame data includes transmitting a plurality of capture frames in coded form in the underlying video layer, the video underlying layer having a first frame rate.
[0138] Method implementation scheme 11C is based on the method described in method implementation scheme 11B, wherein the 3D video information includes left-eye and right-eye information corresponding to each frame time of the base video layer.
[0139] The first numbered list of system implementation schemes
[0140] System Implementation 1 An image capture and processing system includes: at least a first camera (1302) for capturing (904) images at a first frame rate; and a content delivery system (700) including: a memory (712); and a processor (708) coupled to the memory, the processor being configured to control the content delivery system to: perform interpolation (924) on the captured frames to generate interpolated frame data, thereby supporting a second frame rate higher than the first frame rate; and deliver (926) the captured frame data and the interpolated frame data to at least one playback device.
[0141] System Implementation Scheme 2 is based on the image capture and processing system described in System Implementation Scheme 1, wherein the first frame rate is the highest frame rate supported by the camera used to capture images when the camera is operating at the maximum image capture resolution.
[0142] System Implementation Scheme 3 is based on the image capture and processing system described in System Implementation Scheme 1, wherein the first camera captures images at a first frame rate performed by the first camera of the camera pair.
[0143] System Implementation Scheme 4 is an image capture and processing system according to System Implementation Scheme 3, wherein when capturing an image at a lower resolution, the first camera supports a higher frame rate than the first frame rate.
[0144] System Implementation Scheme 5 is based on the image capture and processing system described in System Implementation Scheme 1, wherein, as part of transmitting capture frame data, the processor controls the content delivery system to transmit the capture frame at a first data rate corresponding to the image capture rate.
[0145] System Implementation Scheme 6 is based on the image capture and processing system described in System Implementation Scheme 5, wherein the combination of captured frame data and interpolated frame data corresponds to a second frame rate for the playback device to display the image.
[0146] System Implementation Scheme 7 is based on the image capture and processing system described in System Implementation Scheme 6, wherein the second frame rate is a stereoscopic frame rate at which the left-eye and right-eye images are displayed to the user of the playback device to support a 3D viewing experience.
[0147] System Implementation Scheme 7A: According to the image capture and processing system described in System Implementation Scheme 7, interpolated frame data is transmitted as an interpolated frame, which supplements the captured frame, thereby increasing the frame rate from the captured frame rate to a second frame rate.
[0148] System Implementation Scheme 8 is an image capture and processing system according to System Implementation Scheme 1, wherein the processor (708) is further configured to control the content delivery system: performing (910) motion analysis on blocks of captured frames to detect motion and identify moving segments; selecting (922) a frame rate for at least a first segment based on the amount of motion of the first segment from the current captured frame and one or more other captured frames (e.g., at least the next captured frame); and wherein interpolation (924) is performed to generate interpolated frame data including interpolating segments between captured frames to generate interpolated frame information.
[0149] System Implementation Scheme 9 is an image capture and processing system according to System Implementation Scheme 8, wherein the at least one segment is an image segment of a moving object.
[0150] System Implementation Scheme 9A is the image capture and processing system according to System Implementation Scheme 8, wherein the moving object is a ball.
[0151] System implementation scheme 10 is an image capture and processing system according to system implementation scheme 8, wherein the interpolated segment includes a version of the first segment interpolated to support the selected frame rate.
[0152] System Implementation Scheme 11 is an image capture and processing system according to System Implementation Scheme 8, wherein the processor (708) is further configured to control a content delivery system as part of passing (926) capture frame data and interpolated frame data to at least one playback device: passing (926) capture frames to the playback device as part of a base video layer; and passing (928) information corresponding to one or more interpolated segments corresponding to the interpolated frame time as part of additional video information.
[0153] System Implementation Scheme 11A is an image capture and processing system according to System Implementation Scheme 8, wherein the basic video layer includes 3D video information in the form of multiple left-eye capture frames and right-eye capture frames.
[0154] System implementation scheme 11B is based on the image capture and processing system described in system implementation scheme 11A, wherein transmitting (926) capture frame data includes transmitting a plurality of capture frames in coded form in the underlying video layer, the video underlying layer having a first frame rate.
[0155] System implementation scheme 11C is the image capture and processing system according to system implementation scheme 11B, wherein the 3D video information includes left-eye information and right-eye information corresponding to each frame time of the basic video layer.
[0156] Computer-readable first numbered list
[0157] Media Implementation Plan
[0158] Computer-readable medium embodiment 1: A non-transitory computer-readable medium comprising computer-executable instructions, which, when executed by a processor (706) of a content delivery system (700), cause the processor to control the content delivery system (700) to: access an image (751) stored in a memory (712) of the content delivery system (700), the image having been captured by a camera (1302) at a first frame rate; perform interpolation (924) to generate interpolated frame data, thereby supporting a second frame rate higher than the first frame rate; and transmit (926) the captured frame data and the interpolated frame data to at least one playback device.
[0159] Second list of implementation schemes for numbering method
[0160] Method Implementation Scheme 1. A method for operating a playback system (300), the method comprising: receiving (1603) capture frame data and interpolated frame data; recovering (1607) capture frames from the received capture frame data; generating (1610) one or more interpolated frames from the received interpolated frame data; rendering (1618) a video sequence comprising one or more capture frames and at least one interpolated frame; and
[0161] Output one or more rendered images (1620) to a display device (702).
[0162] Method Implementation Scheme 2. The method according to Method Implementation Scheme 1, wherein the captured frame corresponds to a first frame rate; and the generated video sequence comprising the captured frame and one or more interpolated frames has a second frame rate higher than the first frame rate.
[0163] Method Implementation Scheme 3. The method according to Method Implementation Scheme 2, wherein outputting one or more rendered images (1620) to a display device (702) includes outputting the rendered images at a second frame rate.
[0164] Method Implementation Scheme 4. The method according to Implementation Scheme 3, wherein the second frame rate is matched with the display refresh rate supported by the display device (702); and wherein the display device is a head-mounted display device.
[0165] Method Implementation Scheme 5. The method according to Method Implementation Scheme 1, wherein the second frame rate is a frame rate higher than the image capture rate used to capture images received by the playback system.
[0166] Method Implementation Scheme 6. The method according to Implementation Scheme 5, wherein the playback system does not interpolate between received frames to increase the frame output rate beyond the received frame rate.
[0167] Method Implementation Scheme 7. The method according to Method Implementation Scheme 2, wherein generating (1610) one or more interpolation frames from received interpolation frame data comprises: generating a portion of a first interpolation frame from a first capture frame using (1612) motion vectors; and filling a region of the first interpolation frame for which no interpolation fragment data is provided using (1614) content from one or more previous frames.
[0168] Second List of System Implementation Plans
[0169] System Implementation Scheme 1. A playback system (300), the method comprising: a display device (805); a network interface (810) receiving (1603) capture frame data and interpolated frame data; and a processor (808) configured to control the playback system to: recover (1607) capture frames from the received capture frame data; generate (1610) one or more interpolated frames from the received interpolated frame data; render (1618) a video sequence comprising one or more capture frames and at least one interpolated frame; and output (1620) one or more rendered images to the display device (702).
[0170] System Implementation Scheme 2. The playback system (300) according to System Implementation Scheme 1, wherein the captured frame corresponds to a first frame rate; and wherein the generated video sequence includes the captured frame and one or more interpolated frames and has a second frame rate higher than the first frame rate.
[0171] System Implementation Scheme 3. The playback system (300) according to System Implementation Scheme 2, wherein outputting one or more rendered images (1620) to a display device (702) includes outputting the rendered images at a second frame rate.
[0172] System Implementation Scheme 4. The playback system (300) according to System Implementation Scheme 3, wherein the second frame rate matches the display refresh rate supported by the display device (702); and wherein the display device is a head-mounted display device.
[0173] System Implementation Scheme 5. The playback system (300) according to System Implementation Scheme 1, wherein the second frame rate is a frame rate higher than the image capture rate used to capture images received by the playback system.
[0174] System Implementation Scheme 6. The playback system (300) according to System Implementation Scheme 5, wherein the playback system does not interpolate between received frames to increase the frame output rate beyond the received frame rate.
[0175] System Implementation Scheme 7. According to the playback system (300) of System Implementation Scheme 2, wherein as part of generating (1610) one or more interpolation frames from received interpolation frame data, the process is configured to: generate a portion of a first interpolation frame from a first capture frame using (1612) motion vectors; and fill regions of the first interpolation frame for which no interpolation fragment data is provided using (1614) content from one or more previous frames.
[0176] Second numbered list of computers
[0177] Readable media implementation scheme
[0178] Computer-readable medium implementation scheme 1. A non-transitory computer-readable medium comprising computer-executable instructions (814) that, when executed by a processor (808) of a content playback system (300), cause the processor to control a content delivery system (700) to: receive (1603) capture frame data and interpolated frame data; recover (1607) capture frames from the received capture frame data; generate (1610) one or more interpolated frames from the received interpolated frame data; render (1618) a video sequence comprising one or more capture frames and at least one interpolated frame; and output (1620) one or more rendered images to a display device (702).
[0179] One goal of immersive VR experiences is to deliver realistic experiences. To achieve this, the content consumed by users must adhere to very high-quality specifications. Specifically, high-motion content such as sports is extremely sensitive to motion fidelity. Any issues regarding motion capture, processing, transmission, playback, and rendering must be addressed appropriately. Various factors, such as the content capture system, encoding and playback, and the display refresh rate, affect motion fidelity. Specifically, under certain conditions, when the capture and rendering frame rates are lower than the device / HMD display refresh rate, high-resolution, fast-motion immersive content results in motion blur, flicker, and clipping artifacts.
[0180] Various implementations relate to methods and / or apparatuses that take into account one, more, or all of the factors mentioned above and support dynamic processing of content to achieve good fidelity of content delivered and played back to users of virtual reality devices. Various features involve motion segmentation, depth analysis, and estimation performed on captured images, taking into account the amount of motion in each time period, for example, a set of 3D image content corresponding to the time period, and allocating a frame rate to provide smooth motion sensing given the amount of motion of an object detected in a captured frame, the amount of motion corresponding to a time period comprising multiple frames (e.g., 2, 15, 30, or more sequentially captured frames).
[0181] In various implementations, the use of fisheye lens-based image capture is considered. The motion trajectory of an object or portion of content is taken into account relative to the curvature of the fisheye lens. The mapping function takes into account the rate of change of the real-world display size and the movement perceived by the fisheye lens in the 3D spherical domain for capturing, for example, on-site images (such as the scene of a sporting event).
[0182] In 3D implementations, this method processes and manipulates stereo images (e.g., pairs of left-eye and right-eye images). Considering the stereo nature and the capture of both left and right eye images, choosing a desired frame rate different from the capture rate based on motion and stereo-related issues can lead to reduced computation, bandwidth, and processing compared to systems that capture, encode, and transmit images at a fixed rate regardless of motion or 3D problems. In various implementations, image capture, processing, encoding, and communication are performed in real time, for example, while motion or other events are still occurring during image capture.
[0183] In various implementations, images are captured at a rate lower than the rate at which the playback device generates and displays images to the user at the playback device. The playback device can be, and sometimes is, a virtual reality system including a processor and a head-mounted display, which in some implementations displays different image content to the user's left and right eyes. In at least some implementations, the captured image content is analyzed, and a frame rate for transmitting frames to the playback device is determined. In many cases, the determined frame rate, intended to give the user a sense of smooth motion, is typically higher than the image capture rate. Interpolation of the captured images is used to generate a complete set of frames corresponding to the determined frame rate. The rate at which a particular segment (sometimes corresponding to a moving object) is transmitted to the playback device is selected based on the amount of motion exhibited by the segment / object and / or the portion of the segment captured by the fisheye lens of the camera device. In various implementations, segments subjected to high motion / change rates are encoded and transmitted at a rate different from that of relatively static segments.
[0184] In various implementations, images are captured at high resolution but at a lower frame rate than that achieved by a virtual reality playback device. Frames are interpolated between captured frames to account for object motion. The interpolation process can utilize the entire set of captured image content because it is performed before lossy encoding and transmission to the playback device. Therefore, the system processing the captured images has more information available to it than a playback device that typically receives compressed image data and images that have been degraded by compression. It should be understood that the system generating the content to be transmitted to the playback device may also have one or more processors that can perform the interpolation operation before encoding. Encoding may take into account the differences between the encoded set of frames to be transmitted to the playback device and the movement or changes from one frame to the next.
[0185] In some systems using this invention, a fisheye lens is used to capture image content to be transmitted to a playback device. Fisheye lenses tend to distort the image being captured. In some, but not necessarily all, embodiments, for coding purposes, such as when determining which portions of a captured frame or an interpolated frame should be transmitted to the playback device, the motion position relative to which portion of the fisheye lens is used to capture the image content corresponding to the motion is considered.
[0186] Figure 16 A playback method 1600 is shown, which can be, and in some embodiments, by Figure 1 An implementation of a playback system. In some implementations, when system 300 is used... Figure 1 In the system, this method is by Figure 3 The playback system 300 shown is implemented.
[0187] Method 1600 begins at start step 1602, wherein the playback system 300 is powered on, and the processor 808, under the control of routine 814 stored in memory 800, begins to control the playback system 300 to achieve... Figure 16 The method.
[0188] The operation proceeds from start step 1602 to receive frame information step 1603, where capture frames and interpolated frame data, such as interpolated frames, are received. In some embodiments, step 1603 includes step 1604 of receiving base frames (e.g., coded capture frames corresponding to a first frame rate) and step 1606 of receiving interpolated frame data (e.g., data representing interpolated frames). Capture frame and interpolated frame data have already been discussed with respect to content delivery system 104, which provides such data to the playback system, and therefore will not be described in detail again.
[0189] If both capture frame data and interpolated frame data have been received, and if the received capture frame is encoded, the operation proceeds from step 1603 to step 1608. In step 1608, before proceeding to step 1610, the received capture frame (e.g., a base frame) is decoded to produce an unencoded capture frame. If the capture frame is received in unencoded form in step 1603, the decoding step 1608 is skipped, and the operation proceeds directly from step 1603 to step 1610.
[0190] In step 1610, one or more interpolated frames are generated from the received interpolated frame data. In some embodiments, step 1610 includes steps 1612 and 1614. In step 1612, motion vectors and / or intra-coded segments are used to generate portions of the interpolated frame that differ, for example, in position or content from one or more previous or subsequent frames. In step 1614, content from one or more previous or subsequent frames is used to fill one or more regions of the interpolated frame for which no interpolated segment data is provided. At the end of step 1610, the playback device has both capture frames and interpolated frames, which can be used to generate a sequence of video frames with a second frame rate higher than the image capture rate. The second frame rate can and sometimes is equal to the display rate and / or reference rate implemented by the playback device. Therefore, by using interpolated frames, a refresh rate higher than the image capture rate can be supported in the playback device, while avoiding the need for the playback device to interpolate frames to support a higher playback rate. Therefore, in some embodiments, the playback device supports playback or frame refresh rates higher than the image capture rate without performing interpolation between frames to support the playback rate. This is because the content delivery system performs any necessary frame interpolation to generate interpolated frames before the frames are provided to the playback device.
[0191] The operation proceeds from step 1610 to step 1610, where a frame sequence is generated, for example, by interpolating and capturing frames in a desired display sequence. In some embodiments that support stereoscopic images, separate left-eye and right-eye image sequences are generated in step 1616, each image sequence having a frame rate corresponding to a second frame rate (i.e., the supported frame rate).
[0192] The operation proceeds from step 1616 to step 1618, where, for example, the left-eye and right-eye images are rendered by applying the images as textures to a mesh model, thereby rendering images from one or more video sequences generated in step 1616.
[0193] The operation proceeds from step 1618 to step 1620, where at least a portion of the rendered frame is output to the display. The size of the portion of the output rendered frame (e.g., an image) depends on the size of the display device used and / or the user's field of view. As part of step 1620, in a stereoscopic implementation, different images may be displayed to the user's left and right eyes, for example, using different portions of a head-mounted display.
[0194] The operation of proceeding from step 1620 back to step 1603 is shown to indicate that the process can be repeated over time, wherein the playback device receives, processes and displays sets of images corresponding to different parts of the video sequence at different points in time.
[0195] While in some implementations the display outputs images at a second frame rate, this method does not preclude the playback device from performing further interpolation to further increase the refresh rate supported by the playback device. However, significantly, performing at least some frame interpolation outside the playback device to support a frame rate higher than the capture frame rate eliminates the need for the playback device to perform interpolation to achieve a second frame rate higher than the frame capture rate. Because frames are captured and transmitted at a high level of detail, a practical 3D experience can be achieved using a less expensive camera than might be necessary in the case of a camera capable of supporting a second frame rate, where the same level of detail is used to capture the image.
[0196] In some implementations, motion segmentation is used to segment the captured image, such as frames, for processing by including encoding for transmission purposes in some implementations. In some implementations, the image is segmented or processed on a block-by-block basis during the encoding process or as a separate pre-encoding process, wherein the image comprising a set of blocks (e.g., rectangular portions of a frame) is analyzed to detect objects moving (i.e., in motion) from one frame to the next. In some implementations, the system analyzes the temporal motion of multiple blocks, for example, in a variable-sized look-forward buffer comprising the content of the captured image (e.g., blocks of previously captured frames). The data is processed to form, i.e., to identify meaningful motion segments in the scene, such as a set of blocks as units moving from one frame to another, as perceived from observing multiple frames. This set of blocks may correspond to, and sometimes will correspond to, movable moving objects such as a ball or other objects, while other objects in the scene area, such as the sports field of a sporting event, remain fixed. In some implementations, taking into account the angular velocity of the object when captured in the fisheye-based video, the identified image segments subjected to motion are trimmed based on spatial texture information. That is, some image segments subjected to motion are excluded from further consideration to obtain the final motion segments to be considered for additional processing. The groups of segments considered for further processing are then ordered in some, but not all, implementations in order of relative complexity, where complexity can and sometimes does depend on the level of detail and / or the number of different colors in each segment. Segments with higher levels of detail and more colors or brightness levels are considered more complex than segments with less detail, fewer colors, and / or a lower number of brightness levels.
[0197] In some implementations, depth analysis is used to determine depth, such as the distance to an object in a segment undergoing motion. Therefore, in at least some implementations supporting 3D content, for example, when presenting different left-eye and right-eye images to a user during playback, the depth of the object and thus the depth associated with the image segment representing such an object are considered. Motion in 2D left-eye and right-eye video segments (e.g., images) is considered, and its impact on the perception of objects in 3D (e.g., stereo domain) is taken into account when determining the frequency at which frames should be encoded and transmitted to support smooth motion of 3D objects. Taking into account the 3D effect and depth based on distance to cameras (e.g., left-eye and right-eye cameras), and capturing left-eye and right-eye images respectively, in some implementations, this facilitates and is used to determine an appropriate motion interpolation rate based on the distance of the segment from the camera origin. Therefore, in many cases, object distance may, and sometimes does, affect the rate at which the object will be interpolated at a higher rate of reciprocity, thus providing a better 3D effect. The interpolation rate associated with a segment corresponding to a single object can, and sometimes does, be adjusted based on the display refresh rate of one or more playback devices supported or used. It should be understood that, since different objects are at different distances, different interpolation rates can and sometimes need to be selected for different objects due to differences in depth. In some implementations, objects that are closer to the camera and therefore perceived as closer to the user during playback are given a higher priority for interpolation than objects that are farther away (e.g., objects at greater depth). This is partly because closer objects tend to be more prominent, and users expect such objects to have a higher level of detail than more distant objects.
[0198] In some implementations, motion estimation and representation are performed on each segment, where the segment may correspond to an object or set of objects in an image that may and sometimes does move in position over time, for example, where the segment changes position from frame to frame. The change in position may be, and sometimes is, due to the movement of objects corresponding to the segment's movement, such as a ball, while the camera and other objects remain stationary. In some implementations, the frame rate for each segment is calculated based on spatiotemporal and depth data associated with the segment. In some implementations, motion interpolation is performed individually for each segment, and motion vectors and optimal block information for each segment are calculated. Motion vectors and other information (e.g., padding information for image portions that cannot be accurately represented by motion vectors) are passed as side information and / or metadata and may be embedded in the stream along with information representing the captured frame. Thus, portions of frames determined to correspond to some segments may be refreshed more frequently by the passed data than other portions of the captured frame. If content is delivered to multiple sets / versions of display devices, such as head-mounted displays (HMDs), multiple versions of motion segment side information based on device specifications may exist.
[0199] It should be understood that various features involve image capture, processing, and content transmission in a manner that allows efficient use of the limited data transmission capacity of the communication channel to the user's playback device. By transmitting information corresponding to segments rather than the entire frame for at least some frame time, and wherein segments are selected based on which part of the motion and / or fisheye lens was used to capture the portion of the image corresponding to the segment, segment updates can occur at a higher rate than the rate of content updates for the entire frame, where segment updates contribute to a satisfactory 3D playback experience. During playback, received segments corresponding to motion can be decoded and displayed, allowing the detected motion portions of the scene area to be updated more frequently at the playback device than other portions of the scene area corresponding to the environment of the captured image. Entire frames at a frame rate lower than the desired playback frame rate can and sometimes are transmitted as part of the video basestream. In at least some of these embodiments, information corresponding to one or more segments is transmitted as enhancement information, which provides information corresponding to one or more segments of a time period (e.g., frame time) occurring between frames in the base video stream. Enhancement information can be, and sometimes is combined with, the base stream information at the playback device to generate a video output stream with left-eye and right-eye images, which has a higher frame rate than the base video stream.
[0200] Playback on client devices is supported by various methods and apparatuses. The playback process on a client device includes an initial step where an underlying video stream (e.g., one or more captured frames) is decoded, followed by the extraction of side information passed through the underlying stream. This side information allows some segments of the displayed image to be updated at a faster rate than other segments of the displayed image. Interpolated blocks in frames with side information are reconstructed from one or more reference images (e.g., captured images) using motion compensation, such that the playback frame rate matches the supported display refresh rate of the playback device, which in some, but not necessarily all, embodiments exceeds the frame rate of the captured image.
[0201] In some implementations, when generating an interpolated frame, to fill gaps in pixels where gaps have been created due to motion and no motion information is provided to fill gaps from another portion of the image, the playback device uses content from a juxtaposed block that is positionally attached to the location where the gap occurs. In the absence of motion indication, the motion-free portion of the interpolated frame is generated by acquiring content from a previous or subsequent frame, where the content is taken from a position in the previous or subsequent frame corresponding to the location of one or more motion-free blocks. For example, the content of a block from another frame with the same x, y coordinates as a block without specified motion estimation content is sometimes acquired from the temporally nearest reference frame in the generated interpolated frame. In cases where the motion-free blocks exhibit drastic brightness and / or color variations, a more optimized bidirectional prediction / averaging can be, and is sometimes explicitly signaled in the information passed to the playback device, in which case the brightness and / or chromaticity may depend on the content of multiple frames.
[0202] Motion representation :
[0203] In some implementations, motion-compensated interpolation involves motion estimation to produce satisfactory or optimal motion vectors for each variable-size block shape of an object or segment. For pixels located outside the selected motion segment in the frame, blocks are co-located by the playback device at the interpolation frame time using a reference image. In some implementations, this rule for specifying references for “motionless” pixels / segments is implicitly computed based on the nearest temporal neighborhood. Where bidirectional prediction / averaging is preferred or optimal for “motionless” pixels, it is transmitted as part of the interpolation frame data, which is passed in the form of side information in addition to the coded frame, which typically includes the base video layer of the capture frame. In some implementations, the motion vectors of the selected motion segments are explicitly specified in the transmitted side information, in addition to the base layer video data and source block coordinates, to form a complete motion segment.
[0204] A variety of feature structures can be used to provide a general solution for motion fidelity, which can be scaled and customized for any number of unique HMDs and their variants. It can also be extended to AR / MR / XR applications.
[0205] The features of this method naturally align with any field-of-view (FOV) based streaming and display method, where only the motion segments involved and associated side information can be streamed along with the FoV.
[0206] Although the steps are shown in an exemplary order, it should be understood that in many cases the order of the steps can be changed without adversely affecting the operation. Therefore, the order of steps is considered exemplary and not restrictive unless proper operation requires the exemplary order of steps.
[0207] Some embodiments involve a non-transitory computer-readable medium embodying a set of software instructions, such as computer-executable instructions, for controlling a computer or other device to encode and compress stereoscopic video. Other embodiments involve a computer-readable medium embodying a set of software instructions, such as computer-executable instructions, for controlling a computer or other device to decode and decompress video at a player end. While encoding and compression are mentioned as possible separate operations, it should be understood that encoding can be used to perform compression, and therefore encoding may include compression in some cases. Similarly, decoding may involve decompression.
[0208] The techniques of various embodiments can be implemented using software, hardware, and / or a combination of software and hardware. Various embodiments relate to apparatus, such as an image data processing system. Various embodiments also relate to methods, such as a method for processing image data. Various embodiments further relate to non-transitory machines, such as computer-readable media, such as ROM, RAM, CD, hard disk, etc., which include machine-readable instructions for controlling the machine to implement one or more steps of the method.
[0209] Various features of the present invention are implemented using modules. Such modules can be implemented as software modules, and in some embodiments are implemented as software modules. In other embodiments, the modules are implemented in hardware. In other embodiments, the modules are implemented using a combination of software and hardware. In some embodiments, the modules are implemented as separate circuits, wherein each module is implemented as a circuit for performing the function corresponding to the module. A wide variety of embodiments are contemplated, including some in which different modules are implemented in different ways, such as some in hardware, some in software, and some using a combination of hardware and software. It should also be noted that some of the routines and / or subroutines, or steps performed by such routines, can be implemented in dedicated hardware, in contrast to software executed on a general-purpose processor. Such embodiments are still within the scope of the present invention. Many of the methods or method steps described above can be implemented using machine-executable instructions (such as software) contained in a machine-readable medium (such as a storage device, e.g., RAM, floppy disk, etc.) to control a machine (e.g., a general-purpose computer with or without additional hardware) to implement all or part of the methods described above. Therefore, among other things, the present invention relates to a machine-readable medium comprising machine-executable instructions for causing a machine (e.g., a processor and associated hardware) to perform one or more steps of one or more of the methods described above.
[0210] Based on the above description, numerous other variations of the methods and apparatus of the various embodiments described above will be apparent to those skilled in the art. Such variations should be considered within the scope.
Claims
1. A method of generating a video sequence, the method comprising: acquiring a stereoscopic frame set having a first frame rate, wherein the stereoscopic frame set comprises a series of image pairs; for each of a series of left eye images and a series of right eye images in the series of image pairs, generating one or more intermediate frames to obtain one or more left eye intermediate frames and one or more right eye intermediate frames, including: acquiring motion vector data indicative of motion of a portion of the stereoscopic frame set, acquiring depth data for one or more portions of the series of left eye images and the series of right eye images of the stereoscopic frame set, and generating one or more intermediate frames using the stereoscopic frame set, the motion vector data, and the depth data such that objects at lesser depths are given a higher rate of interpolation than objects at greater depths; and generating a left eye video sequence comprising the series of left eye images and the one or more left eye intermediate frames, and a right eye video sequence comprising the series of right eye images and the one or more right eye intermediate frames, wherein the left eye video sequence and the right eye video sequence are associated with a second frame rate that is greater than the first frame rate.
2. The method of claim 1, wherein generating the one or more intermediate frames comprises: generating a portion of a first intermediate frame from a first acquired frame based on the motion vector data; and filling areas of the first intermediate frame for which no interpolation segment data is provided using content from one or more previous frames of the stereoscopic frame set.
3. The method of claim 1, wherein the left eye video sequence has a second frame rate that is higher than the first frame rate.
4. The method of claim 3, further comprising: outputting the left eye video sequence to a display device, wherein the second frame rate corresponds to a display refresh rate supported by the display device.
5. The method of claim 4, wherein the display device is included in a head mounted device.
6. The method of claim 4, wherein the display device supports a 3D viewing experience.
7. A non-transitory computer readable medium comprising computer readable code executable by one or more processors for: acquiring a stereoscopic frame set having a first frame rate, wherein the stereoscopic frame set comprises a series of image pairs; acquiring motion vector data indicative of motion of at least a portion of the stereoscopic frame set, acquiring depth data for one or more portions of a series of left eye images and a series of right eye images of the stereoscopic frame set, and generating one or more left eye intermediate frames using the stereoscopic frame set, the motion vector data, and the depth data such that objects at lesser depths are given a higher rate of interpolation than objects at greater depths; and generating a left eye video sequence comprising the series of left eye images and the one or more left eye intermediate frames, and a right eye video sequence comprising the series of right eye images and the one or more right eye intermediate frames of the stereoscopic frame set, wherein the left eye video sequence and the right eye video sequence are associated with a second frame rate that is greater than the first frame rate. 8. The non-transitory computer-readable medium of claim 7, wherein the computer- readable code for generating the one or more left-eye intermediate frames and the one or more right-eye intermediate frames comprises computer-readable code for: generating a portion of a first intermediate frame from a first acquired frame based on the motion vector data; and filling areas of the first intermediate frame for which no interpolation segment data is provided using content from one or more previous frames of the stereoscopic frame set.
9. The non-transitory computer-readable medium of claim 7, wherein the left-eye video sequence has a second frame rate that is higher than the first frame rate.
10. The non-transitory computer-readable medium of claim 9, further comprising computer-readable code for: outputting the left-eye video sequence to a display device, wherein the second frame rate corresponds to a display refresh rate supported by the display device.
11. The non-transitory computer-readable medium of claim 10, wherein the display device is included in a head-mounted device.
12. The non-transitory computer-readable medium of claim 10, wherein the display device supports a 3D viewing experience.
13. A system for generating a video sequence, comprising: one or more processors; and one or more computer-readable media comprising computer-readable code executable by the one or more processors for: acquiring a stereoscopic frame set having a first frame rate, wherein the stereoscopic frame set comprises a series of image pairs; acquiring motion vector data indicative of motion of a portion of the stereoscopic frame set, acquiring depth data for one or more portions of a series of left-eye images and a series of right-eye images of the stereoscopic frame set, and generating one or more left-eye intermediate frames using the stereoscopic frame set, the motion vector data, and the depth data such that objects at lesser depths are given a higher interpolation rate than objects at greater depths; and generating a left-eye video sequence comprising the series of left-eye images and the one or more left-eye intermediate frames, and a right-eye video sequence comprising a series of right-eye images and the one or more right-eye intermediate frames, wherein the left-eye video sequence and the right-eye video sequence are associated with a second frame rate that is greater than the first frame rate.
14. The system of claim 13, wherein the computer-readable code for generating the one or more left-eye intermediate frames and the one or more right-eye intermediate frames comprises computer-readable code for: generating a portion of a first intermediate frame from a first acquired frame based on the motion vector data; and filling areas of the first intermediate frame for which no interpolation segment data is provided using content from one or more previous frames of the stereoscopic frame set.
15. The system of claim 13, wherein the left-eye video sequence has a second frame rate that is higher than the first frame rate.
16. The system of claim 15, further comprising computer-readable code for: outputting the left-eye video sequence to a display device, wherein the second frame rate corresponds to a display refresh rate supported by the display device.
17. The system of claim 16, wherein the display device is included in a head-mounted device.
18. The system of claim 16, wherein the display device supports a 3D viewing experience.
Citation Information
Patent Citations
Method and system for processing 2d / 3d video
US20110058016A1
Video motion processing including static scene determination, occlusion detection, frame rate conversion, and adjusting compression ratio
US20180288433A1