Reference frame reprojection for improved video coding
By projecting the reconstructed reference frames in the video codec, generating the reprojected reference frames, and using the reference frames for motion estimation and compensation during the encoding and decoding process, the efficiency problems of the prior art when processing complex motion videos such as rotation and zoom are solved, and higher compression efficiency and video quality are achieved.
Patent Information
- Application Number
- CN201810876368.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-08-03
- Filing Date
- 2018-08-03
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2038-08-03
AI Technical Summary
Existing video codecs have low compression efficiency, video quality, and computational efficiency when processing video content with rotation, zoom, and other complex motion.
By using scene pose difference data for projection transformation, a reconstructed reference frame after reprojection is generated, and the reference frame is used for motion estimation and motion compensation during encoding and decoding.
Improves the compression efficiency and video quality of video codecs, and reduces computational complexity, especially when dealing with video content with complex scene pose changes.
Smart Images

Figure CN109391815B_ABST
Abstract
Description
Background Art
[0001] Existing video codecs (e.g., H.264 or MPEG-4 Part 10, Advanced Video Coding (AVC) codec, H.265 High Efficiency Video Coding (HEVC) codec, etc.) operate using the principle of motion compensation prediction performed on blocks of variable partition size. The motion estimation and compensation can use block-based search for the block of the current frame to find the best matching block in one or more reference frames. The best matching block is referenced using a reference index of the reference frame and a motion vector indicating the motion between the current frame block and the best matching block in the reference frame. The reference index and motion vector found via motion estimation at the encoder are encoded into the bitstream and sent to the decoder. Both the encoder and the decoder use the reference index and motion vector in motion compensation to reconstruct (at the decoder side) a frame for further use as a reference frame and for final presentation. These techniques can be most efficient when the encoded video content is generated based on a single camera model (which can pan but has minimal rotation, zoom, distortion, etc.). However, content including higher levels of rotation, zoom magnification or reduction, distortion, etc. may provide difficulty.
[0002] Therefore, it may be advantageous to increase the compression efficiency, video quality, and computational efficiency of a codec system for processing video content with rotation, zoom, and other effects. It is with respect to these and other considerations that the present improvements are needed. BRIEF DESCRIPTION OF THE DRAWINGS
[0003] The materials described herein are illustrated in the accompanying drawings by way of example and not by way of limitation. For simplicity and clarity of illustration, the elements shown in the drawings are not necessarily drawn to scale. For example, the dimensions of some elements may be exaggerated relative to other elements for clarity. In addition, where considered appropriate, reference numerals have been repeated between the drawings to indicate corresponding or similar elements. In the drawings:
[0004] Figure 1 is an illustrative diagram of an example context for video coding using a reprojected reconstructed reference frame;
[0005] Figure 2 is an illustrative diagram of an example encoder for video encoding using a reprojected reconstructed reference frame;
[0006] Figure 3 A block diagram illustrating an example decoder for video decoding using a reprojected reconstructed reference frame;
[0007] Figure 4is a flow chart illustrating an example process for encoding a video using a reprojected reconstructed reference frame;
[0008] Figure 5 is a flow chart illustrating an example process for conditionally applying frame reprojection based on evaluating scene pose difference data;
[0009] Figure 6 An example of multiple reprojected reconstructed reference frames used in video encoding is shown;
[0010] Figure 7 shows example post-processing of a reconstructed reference frame after reprojection following a zoom up operation;
[0011] Figure 8 shows example post-processing of a reconstructed reference frame after reprojection following a zoom-out operation;
[0012] Fig. 9 An example projection transformation is shown applied only to the region of interest;
[0013] Fig.10 is a flow chart illustrating an example process for video encoding using a reprojected reconstructed reference frame;
[0014] Fig.11 is an illustrative diagram of an example system for video encoding using a reprojected reconstructed reference frame;
[0015] Fig.12 is an illustrative diagram of an example system; and
[0016] Fig.13 Example small form factor devices are shown, each arranged in accordance with at least some implementations of the present disclosure. DETAILED DESCRIPTION
[0017] One or more embodiments or implementations are now described with reference to the accompanying drawings. Although specific configurations and arrangements have been discussed, it should be understood that this is done for illustrative purposes only. Those skilled in the art will recognize that other configurations and arrangements may be used without departing from the spirit and scope of the specification. It is apparent to those skilled in the art that the techniques and / or arrangements described herein may also be used in various other systems and applications other than those described herein.
[0018] Although the following description sets forth various implementations that may, for example, appear in an architecture (e.g., a system on chip (SoC) architecture), the implementations of the techniques and / or arrangements described herein are not limited to a particular architecture and / or computing system, and may be implemented by any architecture and / or computing system for similar purposes. For example, the techniques and / or arrangements described herein may be implemented using various architectures such as multiple integrated circuit (IC) chips and / or packages, and / or various computing devices and / or consumer electronics (CE) devices (e.g., set-top boxes, smart phones, etc.). In addition, although the following description may set forth a large number of specific details (e.g., the logical implementations, types and interrelationships of system components, logical partitioning / integration selection, etc.), the required subject matter may be practiced without these specific details. In other instances, in order not to obscure the material disclosed herein, some materials (e.g., control structures and complete software instruction sequences) may not be shown in detail.
[0019] The materials disclosed herein may be implemented in hardware, firmware, software, or any combination thereof. The materials disclosed herein may also be implemented as instructions stored on a machine-readable medium, which may be read and executed by one or more processors. A machine-readable medium may include any medium and / or mechanism for storing or transmitting information in a form readable by a machine (e.g., a computing device). For example, a machine-readable medium may include: a read-only memory (ROM); a random access memory (RAM); a magnetic disk storage medium; an optical storage medium; a flash memory device; an electrical, optical, acoustic, or other form of propagated signal (e.g., a carrier wave, an infrared signal, a digital signal, etc.), and other media.
[0020] References in the specification to "one implementation," "implementation," "example implementation," etc. indicate that the implementation being described may include a particular feature, structure, or characteristic, but every embodiment may not necessarily include the particular feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same implementation. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an embodiment, it is considered that it is within the knowledge of those skilled in the art to implement such features, structures, or characteristics in relation to other implementations, whether or not explicitly described herein.
[0021] Methods, devices, apparatus, computing platforms, and articles related to video encoding, and in particular, to reprojecting reconstructed video frames to provide reprojected reconstructed reference frames for motion estimation and motion compensation, are described herein.
[0022] As discussed, current motion estimation and motion compensation techniques use reference frames to search for blocks of the current frame for the best match. In order to improve coding efficiency, such motion estimation and motion compensation compress information that is redundant in time. However, the use of these techniques may be limited when the encoded video content includes higher levels of rotation, zooming in or out, distortion, etc. For example, in the context of virtual reality devices, augmented reality devices, portable devices, etc., the device may move frequently in various directions (e.g., with 6 degrees of freedom: up / down, front / back, left / right, roll, yaw, pitch). In these contexts, video capture and / or video generation can provide a sequence of frames with complex motion (e.g., panning, rotation, zooming, distortion, etc.) between frames.
[0023] In some embodiments discussed herein, a reconstructed reference frame corresponding to a first scene pose (e.g., a view of a scene at or about the time of capturing / generating a reference frame) may be transformed by a projective transformation based on scene pose difference data indicating a change in scene pose from a first scene pose to a second scene pose subsequent to the first scene pose. The second scene pose corresponds to a view of a scene at or about the time of capturing a current frame or at or about the time of generating or rendering a current frame. The scene pose difference data provides data (e.g., metadata) outside the encoded frame indicating a change in scene pose between the reference frame and the encoded current frame. As used herein, the term scene pose is used to indicate the pose of a scene relative to a viewpoint or viewport of a captured scene. In the context of virtual reality, a scene pose indicates a pose or view of a generated scene relative to a viewpoint of a user of a virtual reality (VR) device (e.g., a user wearing a VR head accessory). In the context of an image capture device (e.g., a camera of a handheld device, a head-mounted device, etc.), a scene pose indicates a pose or view of a scene captured by the image capture device. In the context of augmented reality (AR), scene pose indicates the pose or view of the scene analyzed (eg, by an image capture device) and the pose of any information generated about the scene (eg, overlay information, images, etc.).
[0024] Scene pose difference data can be utilized (leverage) by applying projection transformation to the reconstruction reference frame to generate the reconstruction reference frame after reprojection.As used herein, the term projection transformation indicates that the transformation of parallelism, length and angle is not necessarily maintained between input frame and output frame.These projection transformations can be contrasted with the affine transformation that maintains parallelism, length and angle, and are therefore more limited in capturing complex scene pose changes.Scene pose difference data can be any suitable data or information (for example, 6 degrees of freedom (6-DOF) difference or increment information, transformation or transformation matrix, motion vector field, etc.) indicating scene pose changes.In addition, depending on the format of scene pose difference data, projection transformation can be performed using any one or more suitable techniques.As further discussed herein, after applying projection transformation, other techniques can be used to adapt the obtained frame to the size and shape of the reconstruction reference frame for motion estimation and / or motion compensation.
[0025] The reprojected reconstructed reference frame is then used in motion estimation (at the encoder) and / or motion compensation (at the encoder or decoder) to generate motion information corresponding to the encoded current frame (e.g., motion vectors generated at the encoder) and / or a reconstructed current frame. For example, the reconstructed current frame is generated in a loop at the encoder (e.g., in a local decoding loop at the encoder) for motion estimation / compensation of subsequent frames. The reconstructed current frame is also generated at the decoder for motion compensation of subsequent frames and / or for final presentation to the user. These techniques provide a closer match between the encoded current frame and the reference frame used for motion estimation / compensation (e.g., the reprojected reconstructed reference frame), which improves coding efficiency. These and other advantages will be apparent to those skilled in the art based on the discussion herein. In addition, the discussed techniques can be used in any suitable coding context, such as, for example, in a codec based on the H.264 / MPEG-4 Advanced Video Coding (AVC) standard, a codec based on the High Efficiency Video Coding (H.265 / HEVC) standard, the proposed video coding (H.266) codec, a codec based on the Alliance for Open Media (AOM) standard (e.g., the AV1 standard), a codec based on an MPEG standard (e.g., the MPEG-4 standard), a codec based on the VP9 standard, or any other suitable codec or extension or profile thereof.
[0026] Figure 1 is an illustrative diagram of an example context 130 for video encoding using a reprojected reconstructed reference frame, arranged in accordance with at least some implementations of the present disclosure. Figure 1As shown, context 130 may include system 100 and system 110, which are communicatively coupled via communication link 131. In an embodiment, context 130 is a virtual reality context, which includes system 100 as a host system, which generates virtual reality frames (e.g., of game content, entertainment content, etc.), which are encoded and sent to system 110. In this context, system 110 may be characterized as a sink, etc., and system 110 may be a head mounted display (HMD) including optical devices (not shown) that provide a 3-dimensional (3D) effect to a user when viewing frames presented via display 115.
[0027] It should be understood that in these contexts, the system 110 may often be moving in 3-dimensional (3D) space as the user moves to view different parts of the virtual scene, interact with virtual content, etc. Thus, the system 110 can move through the 3D space in motion characterized as 6-DOF motion 135, indicating that the system 110 can move in the following manner: translation: forward / backward (e.g., in the x-direction), up / down (e.g., in the y-direction), left / right (e.g., in the z-direction); and rotation: rotation by yaw (e.g., angle α around the z-axis), roll (e.g., angle β around the y-axis), and pitch (e.g., angle γ around the y-axis). It should be understood that 3D content (e.g., VR frames) can be generated based on a hypothetical line of sight of a user of the system 110 (e.g., in a forward direction along the x-axis).
[0028] Although discussed with respect to VR video frames and content, the context 130, system 100, system 110, and other systems discussed herein may operate on frames or pictures including any suitable content. In some embodiments, the technology discussed may be applied to wireless virtual reality (VR), augmented reality (AR), mixed reality (MR), and the like using outside-inside or inside-outside 6-DOF data or information. In some embodiments, the technology discussed may be applied to a camera, a smartphone with a camera, and the like, such that the camera or smartphone includes an integrated inertial measurement unit (IMU) that provides 3-DOF data or information. In some embodiments, the technology discussed may be applied to security cameras that have the ability to provide pan-tilt-zoom (PTZ) control data or information. In some embodiments, the technology discussed may be applied to cloud gaming or entertainment services with virtual camera orientation that has the ability to provide translation data or information, 6-DOF data or information, and the like.
[0029] As shown, system 110 may include a scene pose tracking module 111, a transceiver 112, a decoder 113, a rendering module 114, and a display 115. In addition, system 100 may include an application module 101, a rendering module 102, an encoder 103, and a transceiver 104. Continuing the discussion, without limitation to the VR context, a user wears system 110 as a head mounted display. As the user moves, the user's scene perspective (e.g., view pose) changes as the user moves. In addition, system 110 and system 100 are communicatively coupled via a communication link 131, which may be a wireless link (e.g., WiFi, WiGiG, etc.) or a wired connection (e.g., a universal serial bus coupling, a transport agnostic display coupling, etc.). As the user moves, scene pose tracking module 111 tracks the position and orientation of system 110. As shown, such position and orientation data 121 (via communication link 131 between transceivers 104, 112) is provided to system 100. Such position and orientation data 121 may be provided in any suitable format (e.g., 6-DOF data (e.g., x, y, z, α, β, and γ values relative to an initialized zero position), 6-DOF difference data (e.g., Δx, Δy, Δz, Δα, Δβ, and Δγ values) relative to a previously known 6-DOF position, 3-DOF data (e.g., x, y, z values relative to an initialized zero position), 3-DOF difference data (e.g., Δx, Δy, Δz values) relative to a previously known 6-DOF position, etc.). Furthermore, while discussed with respect to 6-DOF and 3-DOF position and orientation information, any number of degrees of freedom in any combination may be implemented.
[0030] As shown, application module 101 (which may be running on a central processing unit of system 100) receives position and orientation data 121 as metadata. Application module 101 uses the latest metadata (e.g., P) indicating the current scene pose. curr ) to generate rendering data 123, which will be used to render the current frame. For example, the application module 101 may be running a game application, an entertainment application, etc. that responds to the position and orientation data 121 when generating the rendering data 123. For example, as the user of the system 110 moves and / or interacts with the virtual scene, the position and orientation data 121 and / or other input data are used to generate the next view of the user's virtual scene. As shown in the figure, the rendering module 102 uses the rendering data 123 to generate a rendering frame 124. For example, the rendering module 102 can implement a rendering pipeline via a graphics processing unit, etc. to generate a rendering frame 124 based on the rendering data 123.
[0031] In addition, as shown, the system 100 generates pose difference data 122 via the application module 101 or another module or component thereof. As discussed, in some examples, the pose difference data 122 can be provided via the position and orientation data 121. In any case, the pose difference data 122 indicates the scene pose difference between the scene pose corresponding to a previously encoded reconstructed reference frame (e.g., a frame before the rendered frame 124) and the scene pose corresponding to the rendered frame 124 (e.g., the current frame). As discussed, the pose difference data 122 may include any suitable data or information indicating the scene pose change between frames, time instances, etc. In an embodiment, the pose difference data 122 is indicated as scene pose difference metadata. For example, for the scene pose P corresponding to the rendered frame 124 (e.g., the current frame), curr and the scene pose P corresponding to the reconstructed reference frame in the encoder 103 (eg, in the frame buffer of the encoder 103 ) ref , the posture difference data 122 provides the difference between the postures of each scene: ΔP = P ref –P curr For example, ΔP may provide 6-DOF difference data (eg, Δx, Δy, Δz, Δα, Δβ, and Δγ values) with respect to a previously known 6-DOF position.
[0032] As further discussed herein, a transform (e.g., a projective transform) is applied to a reconstructed reference frame (generated by the encoder 103 as further discussed herein) based on the pose difference data 122 (along with other techniques as needed) to generate a reprojected reconstructed reference frame. As used herein, the term reprojection in the context of a reprojected frame indicates that the frame (e.g., having a first scene pose or projection) has been transformed to another scene pose or projection using the pose difference data 122. The reprojected reconstructed reference frame is then used as a motion estimation and motion compensation reference frame for encoding a rendered frame 124 to generate at least a portion of the bitstream 125. For example, the motion estimation performed by the encoder 103 may include searching blocks of the rendered frame 124 on a block basis by searching the reprojected reconstructed reference frame. The motion vectors (e.g., motion vector fields) generated by the motion estimation are then used to reconstruct the rendered frame 124 for use as a reference frame by the encoder 103. In addition, motion vectors and transformed and quantized prediction residuals (e.g., the residual between the original block of the rendered frame 124 and the reference block of the reprojected reconstructed reference frame) are encoded into a bitstream 125, which is sent to the system 110 via the communication link 131.
[0033] In some embodiments, the relative timing of rendering the rendered frame 124 from the rendered data 123 and transforming the reconstructed reference frame based on the pose difference data 122 provides that the rendered frame 124 is rendered at least partially simultaneously and the reconstructed reference frame is transformed so that at least some of the operations are performed simultaneously. Such simultaneous rendering and transformation can provide reduced latency in the processing pipeline of the system 100.
[0034] As shown, at the system 110, the pose difference data 122 is provided to the decoder 113. As shown, in some embodiments, the pose difference data 122 is received from the system 100 via the communication link 131. In some embodiments, the pose difference data 122 is provided from the scene pose tracking module 111 through the position and orientation data 121. In some embodiments, the pose difference data 122 is provided separately from the position and orientation data 121 from the scene pose tracking module 111 or another module or component of the system 110. In an embodiment, the pose difference data 122 can be provided via a bitstream 125. For example, the pose difference data 122 can be standardized as metadata describing the reprojection of the reconstructed reference frame and included in the bitstream 125. The pose difference data 122 (e.g., metadata) can have any suitable format and can be further compressed for inclusion in the bitstream 125. In some embodiments, the system 100 can modify the pose difference data 122 before providing the pose difference data 122 to the system 110. In any case, the pose difference data 122 provided to the encoder 103 and the decoder 113 must be the same (or at least provide the same implementation of the reprojection of the reconstructed reference frame), and it must be applied to the same reconstructed reference frame in the same way so that the encoder 103 and the decoder 113 generate the same reprojected reconstructed reference frame. Otherwise, during motion compensation, the encoder 103 and the decoder 113 will reference different frames and the encoding will be corrupted.
[0035] The decoder 113 applies the pose difference data 122 to the reconstructed reference frame to generate a reprojected reconstructed reference frame (e.g., the reprojected reconstructed reference frame discussed above with respect to the encoder 103). The system 110 receives the bitstream 125 via the communication link 131, and the decoder 113 decodes the bitstream 125 to determine the motion vectors and the transformed and quantized prediction residuals corresponding to the rendered frame 124, as described above. The decoder 113 then inverse quantizes and inverse transforms the transformed and quantized prediction residuals, and uses the motion vectors to determine the reference blocks of the reprojected reconstructed reference frame. The reconstructed (e.g., inverse quantized and transformed) prediction residuals and the corresponding reference blocks are then added to form a reconstructed block, which can be combined with other reconstructed blocks and optionally intra-frame prediction reconstructed blocks to provide a reconstructed frame, which can optionally be deblocked filtered to generate a reconstructed frame 126, which is a reconstruction of the rendered frame 124. The reconstructed frame 126 is provided along with pose difference (PD) data 132 to a rendering module 114 that may be implemented by a graphics processing unit of the system 110, thereby providing even more up-to-date scene pose information, so that the rendering module 114 reprojects or warps the reconstructed frame based on the pose difference data 132 to provide a final frame 127 for display via the display 115 and presentation to a user.
[0036] Thus, context 130 provides a head mounted display (e.g., system 110) communicatively coupled to a host (e.g., system 100) via communication link 131. System 110 sends tracking information (e.g., position and orientation data 121) from scene pose tracking module 111 (e.g., 6-DOF system), where position and orientation data 121 may include position and orientation data of system 110. Applications (e.g., games, entertainment applications, etc.) running on application module 101 of system 100 receive the tracking information as metadata. The latest metadata (scene pose information, P curr ) is used to render the rendering frame (e.g., by providing the rendering data 123 to the rendering module 102 to generate the rendering frame 124) and reproject the rendering frame with the previous scene pose information P ref For example, the scene pose difference or scene pose difference data (eg, ΔP = P ref –P curr ) is used by encoder 103 (or another module or component of system 100) to reproject and reconstruct a reference frame.
[0037] These pose difference data 122 (received from system 100 or generated at system 110) are also used by decoder 113 to reproject the reconstructed reference frame. Encoder 103 encodes rendered frame 124 using the reprojected reconstructed reference frame to generate bitstream 125 that is sent to system 110. The reprojected reconstructed reference frame at decoder 113 and information from bitstream 125 are used to decode reconstructed frame 126 (corresponding to rendered frame 124). Reconstructed frame 126 is provided to rendering module 114 along with pose difference data 132, and final reprojection, lens distortion and correction (if applicable) based on the latest head pose is performed, and the resulting frame 127 is sent for display.
[0038] As discussed, in some embodiments, the tracked position and orientation data 121 is used to generate pose difference data 122. In an embodiment, scene pose prediction can be used to generate subsequent scene pose data and / or pose difference data 122 (e.g., for frame rendering). For example, the delay between measuring position and orientation data 121, rendering a rendered frame 124, and displaying a frame 127 may provide an undesirable user interface (e.g., artifacts, lag time, etc.), which can be at least partially addressed using scene pose prediction. The scene pose prediction can be performed using any one or more suitable techniques. For example, the latest or subsequent scene pose data (or scene pose difference data) can be generated based on extrapolating from a previous scene pose using a previously known scene pose difference. In an embodiment, the scene pose data corresponding to the reconstructed reference frame and the subsequent scene pose data (e.g., from the scene pose tracking module 111) can be used to extrapolate scene pose data that is still after the measured scene pose data. For example, a processing time may be determined that includes a sum of a time for generating rendering data 123, a rendering time (e.g., rendering complexity) for rendering rendering frame 124, an encoding time for generating bitstream 125, a data transmission time over communication link 131 for delivering bitstream 125, and / or a decoding time for generating reconstructed frame 126. The processing time may be approximated or determined by measuring these times during operation of system 100 and / or system 110.
[0039] Using the pose difference between a known first scene pose at a first time instance and a known second scene pose at a second time instance after the first time instance, an extrapolation technique can be used to determine a pose instance at a time instance of the second time instance plus the processing time (e.g., at a third time instance). For example, the scene pose difference between the first time instance and the second time instance can be linearly extrapolated to the scene pose difference at the third time instance. For example, the extrapolated scene pose can be provided as: P3 = P2 + (P2-P1) * (t3-t2) / (t2-t1), where P3 is the extrapolated scene pose at time t3, P2 is the scene pose at time t2, and P1 is the scene pose at time t1. In an embodiment, in order to reduce the possibility of overshoot of the pose difference when extrapolating from the second time instance to the third time instance, the extrapolation can be multiplied by a predetermined factor (e.g., 2 / 3 or 1 / 2, etc.) to reduce the linear extrapolation (e.g., P3=P2+k*(P2-P1)*(t3-t2) / (t2-t1), where k is a predetermined factor). The predicted or extrapolated scene pose (e.g., P3) and / or scene pose difference data (e.g., P3-P1) can then be used throughout the processing pipeline discussed (e.g., at the application module 101, the rendering module 102, the encoder 103, and the decoder 113), as discussed above.
[0040] Furthermore, as discussed with respect to rendering module 114, reprojection of reconstructed frame 126 may be performed at a final frame buffer (not shown) immediately prior to presentation via display 115. This reprojection may be based on the difference between the predicted or extrapolated scene pose head pose used for rendering and the latest available scene pose at the time of display from scene pose tracking module 111, and such reprojection may further mitigate the undesirable user interface effects discussed.
[0041] Figure 2 2 is an illustrative diagram of an example encoder 200 for video encoding using a reprojected reconstructed reference frame, arranged in accordance with at least some implementations of the present disclosure. For example, the encoder 200 may be implemented as the encoder 103 in the system 100. Figure 2 As shown, the encoder 200 may include a projection transformation module 213, a difference finder 212, an intra prediction module 201, a motion estimation module 202, a difference finder 203, a transformation module 204, a quantization module 205, an entropy encoder 214, an inverse quantization module 206, an inverse transformation module 207, an adder 208, a motion compensation module 209, an intra decoding module 210, switches 215, 216, and a deblocking filter module 211. The encoder 200 may include additional modules and / or interconnections that are not shown for clarity of presentation.
[0042] As shown, encoder 200 receives input frame 224 having current scene pose 221 corresponding to input frame 224, and encoder 200 has previously generated reconstructed reference frame 225 corresponding to reference scene pose 223, so that current scene pose 221 is behind in time relative to reference scene pose 223. As discussed, current scene pose 221 can be a scene pose measured (e.g., by scene pose tracking module 111) or a predicted scene pose (e.g., predicted using extrapolation, etc.). It should be understood that although current scene pose 221 corresponds to input frame 224 and reference scene pose 223 corresponds to reconstructed reference frame 225, the timing or time instance (e.g., measurement time) of these scene poses and the timing or time instance (e.g., the time to present them) of the frame can be the same or different. Input frame 224 (or multiple input frames) can include frames or pictures of a video sequence in any suitable format. For example, input frame 224 can be a frame in a video sequence of any number of video frames. These frames can be in any suitable format and can include any suitable content (e.g., VR frames or content, AR frames or content, MR frames or content, image frames captured (e.g., via a mobile camera device, security camera, etc.), etc.). Frames can be divided into or include segments or planes that allow parallel processing of video data and / or separation of it into different color components. For example, a frame of color video data may include a luma plane or component and two chroma planes or components at the same or different resolutions relative to the luma plane. The input frame 224 can be divided into blocks of any size, which contain data corresponding to, for example, MxN pixel blocks. These blocks may include data from one or more planes or color channels of pixel data. As used herein, the term block may include macroblocks, coding units, etc. of any suitable size. It should be understood that these blocks may also be divided into sub-blocks for prediction, transformation, etc.
[0043] As shown, the difference between the current scene pose 221 and the reference scene pose 223 can be determined by a difference finder 212 to generate scene pose difference data, which in the context of the encoder 200 is provided by a transformation matrix 222. For example, the 6-DOF scene pose difference data (e.g., Δx, Δy, Δz, Δα, Δβ, and Δγ values) can be converted into a transformation matrix 222 using known techniques so that the transformation matrix 222, when applied to the reconstructed reference frame 225, provides a projective transformation from the current scene pose 221 to the reference scene pose 223. As shown, the transformation matrix 222 can be applied to the reconstructed reference frame 225 by the projective transformation module 213 to generate a reprojected reconstructed reference frame 226. The reprojected reconstructed reference frame 226 is provided to the motion estimation module 202 and the motion compensation module 209. In the example shown, the difference finder 212 generates the scene pose difference data. In other examples, the encoder 200 can receive such scene pose difference data as a transformation matrix 222 or in any other suitable format.
[0044] As discussed, an input frame 224 is provided for encoding by the encoder 200. In the context of the system 100, the input frame 224 may be the rendered frame 124. However, as discussed herein, the input frame 224 may be any suitable frame for encoding (e.g., an input image or a frame captured by an image capture device, a rendered frame, an augmented reality frame, etc.). As shown, the input frame 224 may be encoded based in part on a reprojected reconstructed reference frame 226 to generate a bitstream 235. For example, in the context of the encoder 100, the bitstream 235 may correspond to the bitstream 125. The bitstream 235 may have any suitable format (e.g., a format compliant with a standard (e.g., AVC, HEVC, etc.)).
[0045] For example, the encoder 200 may divide the input frame 224 into blocks of different sizes, which may be either temporally (inter-frame) predicted via the motion estimation module 202 and the motion compensation module 209, or spatially (intra-frame) predicted via the intra prediction module 201. Such encoding decisions may be implemented via a selection switch 215 under the control of an encoding controller (not shown). As shown, the motion estimation module 202 may use the reprojected reconstructed reference frame 226 as a motion compensation reference frame. That is, the motion estimation module 202 may use a block of the input frame 224 to search for the best matching block for the reprojected reconstructed reference frame 226 (and other motion compensation reference frames, if used), and may use a reference index to the reprojected reconstructed reference frame 226 and a motion vector to reference the best matching block. When more than one motion compensation reference frame is used for motion search, a reference index may be used to indicate the motion compensation reference frame used for the block. When only one motion compensation reference frame (e.g., the reprojected reconstructed reference frame 226) is used, the reference index may be omitted. The motion vectors and reference indices (if necessary) for these blocks are provided as motion vectors and reference indices 227 from the motion estimation module 202 for encoding into the bitstream 235 via the entropy encoder 214 .
[0046] As discussed, the reprojected reconstructed reference frame 226 is used as a motion compensated reference frame via the motion estimation module 202 and the motion compensation module 209 (and via the motion compensation module 309 of the decoder 300 discussed herein below). In an embodiment, only the reprojected reconstructed reference frame 226 is used as a motion compensated reference frame. In other embodiments, the reprojected reconstructed reference frame 226 and other frames are used as motion compensated reference frames. In an embodiment, the reprojected reconstructed reference frame 226 may be used instead of a standard-based reconstructed reference frame, so that the reprojected reconstructed reference frame 226 replaces the reconstructed reference frame and all other encodings may be compliant with the standard. In another embodiment, the reprojected reconstructed reference frame 226 may be added to the available frames, and the standard may need to be extended so that an indicator of the reprojected reconstructed reference frame 226 may be provided via the bitstream 235, etc.
[0047] In an embodiment, both the reprojected reconstructed reference frame 226 and the reconstructed reference frame 225 are used as motion compensation reference frames. For example, the motion estimation module 202 can use both the reconstructed reference frame 225 and the reprojected reconstructed reference frame 226 as motion estimation reference frames to perform motion estimation on the input frame 224 in a block-by-block manner, such that the first block of the input frame 224 is referenced to the reconstructed reference frame 225 (e.g., via a reference index and a motion vector) for motion compensation, and the second block of the input frame 224 is referenced to the reprojected reconstructed reference frame 226 (e.g., via a different reference index and another motion vector) for motion compensation.
[0048] In addition, although the generation of one reprojected reconstructed reference frame 226 is discussed, one or more additional reprojected reconstructed reference frames may be generated based on applying different transformation matrices to the reconstructed reference frame 225. For example, multiple projection transformations (assuming each has different scene pose difference data) may be applied to generate multiple reprojected reconstructed reference frames, which may all be provided to the motion estimation module 202 and the motion compensation module 209 (and the motion compensation module 309) for motion compensation. When a block references a specific reprojected reconstructed reference frame in the reprojected reconstructed reference frames, the reference may be indicated by the reference index of the motion vector and the reference index 227. For example, a first reprojected reconstructed reference frame generated by applying a first projection transformation to the reconstructed reference frame 225 (e.g., using the scene pose difference data between the reconstructed reference frame 225 and the input frame 224) and a second reprojected reconstructed reference frame generated by applying a second projection transformation to the reconstructed reference frame 225 (e.g., using the scene pose difference data between the reconstructed reference frame 225 and the frame before the input frame 224) may both be used as motion compensation reference frames. Alternatively or additionally, one or more projective transformations may be applied to other reconstructed reference frames (e.g., other past reconstructed reference frames) to generate one or more reprojected reconstructed reference frames, as described herein with respect to Figure 6 as discussed further.
[0049] Continue to refer to Figure 2 Based on the use of intra-frame or inter-frame coding, the difference between the source pixels of each block of the input frame 224 and the predicted pixels for each block (e.g., the difference between the pixels of the input frame 224 and the reprojected reconstructed reference frame 226 when the reprojected reconstructed reference frame 226 is used as the motion compensation reference frame shown or other motion compensation reference frames are being used) can be derived via the difference finder 203 to generate a prediction residual for the block. The difference or prediction residual is converted to the frequency domain via the transform module 204 (e.g., based on discrete cosine transform, etc.) and converted to quantization coefficients via the quantization module 205. These quantization coefficients, motion vectors and reference indexes 227, and various control signals can be entropy encoded via the entropy encoder 214 to generate an encoded bitstream 235, which can be sent or transmitted (etc.) to a decoder.
[0050] In addition, as part of the local decoding loop, the quantized prediction residual coefficients can be inverse quantized via an inverse quantization module 206 and inverse transformed via an inverse transform module 207 to generate a reconstructed difference or residual. The reconstructed difference or residual can be combined via an adder 208 with a reference block from a motion compensation module 209 (which can use pixels from the reprojected reconstructed reference frame 226 when the reprojected reconstructed reference frame 226 is used as a motion compensation reference frame as shown or is using pixels of other motion compensation reference frames) or an intra-frame decoding module 210 to generate a reconstructed block, which can be provided to a deblocking filter module 211 for deblocking filtering as shown to provide a reconstructed reference frame for use by another input frame. For example, the reconstructed reference frame (e.g., the reconstructed reference frame 225) can be stored in a frame buffer.
[0051] Thus, encoder 200 can use the reprojected reconstructed reference frame 226 to more efficiently encode input frame 224 relative to using only reconstructed reference frame 225. Example results of these encoding coefficients are further discussed herein with respect to Table 1. Bitstream 235 can then be stored, sent to a remote device, etc. for subsequent decoding to generate a reconstructed frame corresponding to input frame 224 for presentation to a user.
[0052] Figure 3 A block diagram of an example decoder 300 for video decoding using a reprojected reconstructed reference frame arranged in accordance with at least some implementations of the present disclosure is shown. For example, the decoder 300 may be implemented as the decoder 113 in the system 110. As shown, the decoder 300 may include a projective transformation module 313, a difference finder 312, an entropy decoder 305, an inverse quantization module 306, an inverse transformation module 307, an adder 308, a motion compensation module 309, an intra-frame decoding module 310, a switch 314, and a deblocking filter module 311. The decoder 300 may include additional modules and / or interconnections that are not shown for clarity of presentation.
[0053] As shown, decoder 300 may receive current scene pose 221, reference scene pose 223, and input bitstream 235 (e.g., an input bitstream corresponding to or representing a video frame encoded using one or more reprojected reconstructed reference frames), and decoder 300 may generate frame 230 for presentation. For example, decoder 300 may receive input bitstream 235, which may have any suitable format (e.g., a format compliant with a standard (e.g., AVC, HEVC, etc.)). As discussed with respect to encoder 200, a difference between current scene pose 221 and reference scene pose 223 may be determined by differencer 312 to generate scene pose difference data, which in the context of encoder 200 and decoder 300 is provided by transformation matrix 222. As discussed, 6-DOF scene pose difference data may be converted to transformation matrix 222 using known techniques, such that when applied to reconstructed reference frame 225, transformation matrix 222 provides a projective transformation from current scene pose 221 to reference scene pose 223. Transformation matrix 222 can be applied to reconstructed reference frame 225 by projection transformation module 313 to generate reprojected reconstructed reference frame 226. As shown, reprojected reconstructed reference frame 226 is provided to motion compensation module 309. In the illustrated example, differencer 312 generates scene gesture difference data. In other examples, decoder 300 can receive such scene gesture difference data as transformation matrix 222 or in any other suitable format. In an embodiment, decoder 300 receives such scene gesture difference data by decoding a portion of bitstream 235.
[0054] For example, the decoder 300 may receive the bitstream 235 via an entropy decoder 305, which may decode the motion vectors and reference indices 227 and block-based quantized prediction residual coefficients from the bitstream 235. As shown, the motion vectors and reference indices 227 are provided to a motion compensation module 309. The quantized prediction residual coefficients are inverse quantized via an inverse quantization module 306 and inverse transformed via an inverse transform module 307 to generate a reconstructed block-based difference or residual (e.g., a prediction residual block). The reconstructed difference or residual is combined with a reference block from the motion compensation module 309 (which may use pixels from the reprojected reconstructed reference frame 226, or pixels from other motion compensated reference frames) or an intra decoding module 310 via an adder 308 to generate a reconstructed block. For example, for each block, one of the motion compensation module 309 or the intra-frame decoding module 310 can provide a reference block for addition to the corresponding reconstructed difference or residual for the block under the control of a switch 314, which is controlled by a control signal decoded from the bitstream 235. As shown, the reconstructed block is provided to a deblocking filter module 311 for deblocking filtering to provide a reconstructed reference frame for use by another input frame and presentation to a user (if desired). For example, the reconstructed reference frame (e.g., the reconstructed reference frame 225) can be stored in a frame buffer for use in decoding other frames and for final presentation to a user. For example, the frame 230 for presentation can be sent directly to a display, or it can be sent for additional reprojection, as discussed with respect to the rendering module 114.
[0055] As discussed, the reprojected reconstructed reference frame 226 is used as a motion compensated reference frame via the motion compensation module 309. In an embodiment, only the reprojected reconstructed reference frame 226 is used as a motion compensated reference frame. In other embodiments, the reprojected reconstructed reference frame 226 and other frames are used as motion compensated reference frames. As discussed, the reprojected reconstructed reference frame 226 can be used instead of a standard-based reconstructed reference frame, so that the reprojected reconstructed reference frame 226 replaces the reconstructed reference frame and all other encodings can be compliant with the standard. In another embodiment, the reprojected reconstructed reference frame 226 can be added to the available frames, and the standard may need to be extended so that an indicator of the reprojected reconstructed reference frame 226 can be provided via the bitstream 235, etc.
[0056] In an embodiment, the reprojected reconstructed reference frame 226 and the reconstructed reference frame 225 are used as motion compensation reference frames. For example, the motion compensation module 309 can perform motion compensation by acquiring pixel data from the reprojected reconstructed reference frame 226 and / or the reconstructed reference frame 225 under the control of the motion vector and the reference index 227. In addition, although the discussion is about generating one reprojected reconstructed reference frame 226, one or more additional reprojected reconstructed reference frames can be generated based on applying different transformation matrices to the reconstructed reference frame 225. For example, multiple projection transformations can be applied (assuming that each has different scene pose difference data) to generate multiple reprojected reconstructed reference frames, which can all be provided to the motion compensation module 309 for motion compensation. When a block references a specific reprojected reconstructed reference frame in the reprojected reconstructed reference frames, the motion compensation module 309 can perform motion compensation by acquiring pixel data from any available reprojected reconstructed reference frame. For example, a first reprojected reconstructed reference frame is generated by applying a first projection transformation to the reconstructed reference frame 225 (e.g., using scene pose difference data between the reconstructed reference frame 225 and the input frame 224), a second reprojected reconstructed reference frame is generated by applying a second projection transformation to the reconstructed reference frame 225 (e.g., using scene pose difference data between the reconstructed reference frame 225 and a frame before the input frame 224), and both are used as motion compensation reference frames. Alternatively or additionally, one or more projection transformations may be applied to other reconstructed reference frames (e.g., other past reconstructed reference frames) to generate one or more reprojected reconstructed reference frames.
[0057] Figure 4 4 is a flow chart illustrating an example process 400 for encoding a video using a reprojected reconstructed reference frame, arranged in accordance with at least some implementations of the present disclosure. The process 400 may include: Figure 4 One or more operations 401-409 are shown. Process 400 may form at least a portion of a video encoding process. As non-limiting examples, process 400 may form at least a portion of a video encoding process or a video decoding process.
[0058] Process 400 begins at operation 401, where a reconstructed reference frame corresponding to a first scene pose is generated. The reconstructed reference frame may be reconstructed using any one or more suitable techniques. For example, reference blocks for a frame may be determined using intra-frame decoding and / or motion compensation techniques, and each reference block may be combined with a prediction residual (if any) to form a reconstructed reference block. The reconstructed reference blocks may be combined or fused into a frame, and the frame may be deblock filtered to generate a reconstructed reference frame. For example, the reconstructed reference frame may correspond to the reconstructed reference frame 225 discussed with respect to the encoder 200 and the decoder 300.
[0059] Processing can continue at operation 402, where scene pose difference data for a scene pose change from a first scene pose (corresponding to a reconstructed reference frame) to a second scene pose (corresponding to a closer evaluation of the scene) following the first scene pose is received or generated. As discussed, the scene pose difference data indicates a scene change pose over time. The scene pose difference data can be in any suitable format and can be applied to a frame, as discussed with respect to operation 404. In an embodiment, the scene pose difference data is a transformation matrix. For example, each pixel coordinate or some pixel coordinates of a frame (e.g., a reconstructed reference frame) can be matrix multiplied with a transformation matrix to provide new or reprojected pixel coordinates for a pixel, so that a reprojected frame (e.g., a reprojected reconstructed reference frame) is generated. In an embodiment, the scene pose difference data is 6 degrees of freedom difference data (e.g., Δx, Δy, Δz, Δα, Δβ, and Δγ values), which can be converted into a transformation matrix and / or applied to a frame (e.g., a reconstructed reference frame) to generate a reprojected frame (e.g., a reprojected reconstructed reference frame). In an embodiment, the scene pose difference data is a motion vector field, which can be applied to a frame (eg, a reconstructed reference frame) to generate a reprojected frame (eg, a reprojected reconstructed reference frame).
[0060] Processing can continue at operation 403, wherein the scene pose difference data can be optionally evaluated so that the application of the scene pose difference data to the reconstructed reference frame is conditional on the evaluation. The scene pose difference data can be evaluated, for example, to determine whether the difference of the scene pose is large enough to guarantee the cost of performing reprojection. For example, if the difference of the scene pose or one or more or all amplitude values corresponding to the scene pose difference data are less than a threshold value, reprojection can be skipped. In some embodiments, while skip reprojection at a decoder (e.g., decoder 300) can be performed in response to a skip reprojection indicator in a bitstream (e.g., bitstream 235), operation 403 can be performed at an encoder (e.g., encoder 200).
[0061] Figure 5 is a flow chart illustrating an example process 500 for conditionally applying frame reprojection based on evaluating scene pose difference data, arranged in accordance with at least some implementations of the present disclosure. The process 500 may include: Figure 5 One or more operations 501-504 are shown.
[0062] Process 500 begins at operation 501, where one or more scene change difference magnitude values (SCDMVs) are generated. The scene change difference magnitude values may include any one or more values that indicate the magnitude of scene pose difference data (e.g., the magnitude of a change in scene pose). For example, in the context of 6-DOF difference data or any DOF difference data, the scene change difference magnitude value may include the sum of the squares of each DOF difference or increment (e.g., Δx 2 +Δy 2 +Δz 2 +Δα 2 +Δβ 2 +Δγ 2 ), the sum of the squares of the translation components (e.g., Δx 2 +Δy 2 +Δz 2 ), etc. In the context of a translation matrix, the scene change difference magnitude value may include the sum of squares of matrix coefficients, etc. In the context of a motion vector field, the scene change difference magnitude value may include the average absolute motion vector value for the motion vector field, the average of the sum of squares of the x-component and y-component of the motion vector in the motion vector field, etc.
[0063] Processing can continue at operation 502, where the scene change difference magnitude value is compared to a threshold value. As shown, if the scene change difference magnitude value corresponding to the scene pose difference data exceeds the threshold value, processing continues at operation 503, where a projection transformation is applied to the corresponding reconstructed reference frame. If not, processing continues at operation 504, where the projection transformation is skipped and the scene pose difference data is discarded.
[0064] In the illustrated embodiment, a single scene change difference amplitude value is compared to a single threshold value, and when the scene change difference amplitude value exceeds the threshold value, the projective transformation is applied. In another embodiment, when the scene change difference amplitude value meets or exceeds the threshold value, the projective transformation is applied. In an embodiment, multiple scene change difference amplitude values must all meet or exceed their respective threshold values. In an embodiment, for a projective transformation to be applied, each degree of freedom employed is required to exceed a threshold value. In an embodiment, for a projective transformation to be applied, the scene change difference amplitude value (e.g., one or more scene change difference amplitude values) must meet or exceed a first threshold value but not exceed a second threshold value, wherein the first threshold value is less than the second threshold value.
[0065] return Figure 4, processing can continue at operation 404, where a projection transformation is applied. For example, when the evaluation provided in operation 403 is adopted, or in all instances where the evaluation is not used, a projection transformation can be conditionally applied based on the evaluation. The projection transformation can be applied using any one or more suitable techniques. For example, the application of the projection transformation can depend on the format of the scene pose difference data. In the context of scene pose difference data conversion or conversion to a transformation matrix, each pixel coordinate or some pixel coordinates of the reconstructed reference frame can be matrix multiplied with the transformation matrix to provide new or reprojected pixel coordinates for the pixel, so that a reprojected frame (e.g., a reprojected reconstructed reference frame) is generated. When the scene pose difference data is 6-degree-of-freedom differential data (e.g., Δx, Δy, Δz, Δα, Δβ, and Δγ values) or differential data for fewer degrees of freedom, etc., the 6-degree-of-freedom differential data can be converted into a transformation matrix and / or applied to the reconstructed reference frame to generate a reprojected frame. In embodiments where the scene pose difference data is a motion vector field, the motion vector field may be applied to the reconstructed reference frame (e.g., on a block-by-block basis) to reposition pixels corresponding to each block to new locations based on a corresponding motion vector for the block. As discussed, the projective transformation applied at operation 404 may be based on the scene pose difference, where the scene pose difference varies over time.
[0066] Figure 6 An example of a plurality of reprojected reconstructed reference frames for use in video encoding arranged in accordance with at least some implementations of the present disclosure is shown. Figure 6 As shown, the scene pose change context 600 includes a reference scene pose 223 (P ref ) of the reconstructed reference frame 225, as discussed herein. The reference scene pose 223 may be a scene pose at the time when a frame corresponding to the reconstructed reference frame 225 is presented to the user, the time when a frame corresponding to the reconstructed reference frame 225 is rendered, etc. Figure 6 As shown, the reference scene posture 223 and the current scene posture 221 (P curr ) provides scene pose difference data 601 (ΔP = P curr -P ref), which can be in any format discussed herein. Scene pose difference data 601 is applied to reconstructed reference frame 225 to generate a reprojected reconstructed reference frame 226. As discussed, current scene pose 221 can be based on a closer scene pose measurement, or current scene pose 221 can be based on the projected scene pose (using extrapolation or similar techniques). In addition, scene pose difference data 601 (and / or current scene pose 221) can be used to render input frame 224, as discussed herein. As shown, reprojected reconstructed reference frame 226 is then used for motion estimation and motion compensation 602 (performed by encoder 200), or only for motion compensation (performed by decoder 300) to encode input frame 224. That is, reprojected reconstructed reference frame 226 is used as a motion compensated reference frame for encoding input frame 224, as discussed herein.
[0067] In addition, as shown in the scene pose change context 600, one or more additional reprojected reconstructed reference frames can be generated and used for motion estimation and motion compensation 602. For example, the motion estimation and motion compensation 602 can search a set of motion compensated reference frames 607, including one or more reprojected reconstructed reference frames and one or more reconstructed reference frames without reprojection (e.g., reconstructed reference frame 225). During the motion estimation search (e.g., at the encoder 200), for a block of the input frame 224, the best matching block is found from any motion compensated reference frame 607, and the best matching block is referenced using the frame reference and motion vector. During motion compensation (e.g., at the encoder 200 or decoder 300), the frame reference and motion vector are used to access the best matching block (e.g., reference block) among the motion compensated reference frames 607, and the best matching block is added with the reconstructed prediction residual to form a reconstructed block, which is combined with other blocks to reconstruct the frame, as discussed herein.
[0068] In an embodiment, the reconstructed reference frame 605 has a corresponding reference scene pose 604 (P ref2 ), wherein the reference scene pose 604 is before the reference scene pose 223. The reference scene pose 604 may be a scene pose at the time when the frame corresponding to the reconstructed reference frame 605 is presented to the user, the time when the frame corresponding to the reconstructed reference frame 605 is rendered, etc. The reference scene pose 604 and the current scene pose 221 (P curr ) provides scene pose difference data 610 (ΔP2 = P curr -P ref2), which may be in any format discussed herein. The scene pose difference data 610 is applied to the reconstructed reference frame 605 to generate a reprojected reconstructed reference frame 606. As shown, the reprojected reconstructed reference frame 606 is then used for motion estimation and motion compensation 602 (performed by the encoder 200) or only for motion compensation (performed by the decoder 300) as part of the motion compensated reference frame 607 to encode the input frame 224. For example, using multiple reprojected reconstructed reference frames can improve the encoding efficiency of the input frame 224.
[0069] The projective transformation discussed herein may reproject or distort the reconstructed reference frame (or a portion thereof) in any suitable manner (e.g., providing translation of objects in the frame, zoom-in or zoom-out effects for the frame, rotation of the frame, distortion of the frame, etc.). The reconstructed reference frame may be characterized as a reference frame, a reconstructed frame, etc., and the reprojected reconstructed reference frame may be characterized as a distorted reconstructed reference frame, a distorted reference frame, a reprojected reference frame, etc.
[0070] In some embodiments, after the projective transformation, the reprojected or warped reference frame may be further processed and then provided as a motion estimation / compensation reference frame. For example, zoom-in, zoom-out, and rotation operations may provide pixels that are moved outside the footprint of the reconstructed reference frame (e.g., the original size and shape of the reconstructed reference frame). In these contexts, pixels of the resulting frame after the projective transformation may be altered, eliminated, or additional pixel values may be added to fill the gaps so that the reprojected reconstructed reference frame used for motion estimation / compensation reference has the same size and shape as the reconstructed reference frame (and the same size and shape as the frame to be encoded using the reprojected reconstructed reference frame as a reference frame).
[0071] Figure 7 700 illustrates example post-processing of a reprojected reconstructed reference frame following a zoom up operation, arranged in accordance with at least some implementations of the present disclosure. Figure 7 As shown, after applying the projective transformation, the resulting reprojected reconstructed reference frame 701 has a larger size (e.g., h2x w2) than the original size (e.g., h1x w1) of the corresponding reconstructed reference frame 702 used to generate (via the projective transformation in question) the resulting reprojected reconstructed reference frame 701. The resulting reprojected reconstructed reference frame 701 may be characterized as a warped reconstructed reference frame, a resulting reconstructed reference frame, etc.
[0072] As shown, in an embodiment where the resulting reprojected reconstructed reference frame 701 has a size larger than the original size of the reconstructed reference frame or a portion of the reprojected reconstructed reference frame 701 is outside the original size of the reconstructed reference frame, a bounding box 703 having the same size and shape as the original size of the reconstructed reference frame (and the size and shape of the input frame to be encoded) is applied to the resulting reprojected reconstructed reference frame 701, and scaling (704) is applied to the pixel values of the resulting reprojected reconstructed reference frame 701 within the bounding box 703 to generate a reprojected reconstructed reference frame 706 having the same size, shape, and pixel density as the reconstructed reference frame (and the size and shape of the input frame to be encoded). In the illustrated embodiment, the bounding box 703 has the same size and shape as the original size of the reconstructed reference frame (and the size and shape of the input frame to be encoded). In other embodiments, a larger reprojected reconstructed reference frame 706 may be generated if supported by the implemented encoding / decoding architecture. For example, if a larger size is supported, the bounding box 703 may be larger than the original size of the reconstructed reference frame. In these examples, bounding box 703 has a size that is larger than the original size of the reconstructed reference frame, up to the maximum supported reference frame size.
[0073] For example, the zoom-up (e.g., moving closer to the user's viewpoint) projection transformation causes the resulting reprojected reconstructed reference frame 701 to be scaled to a greater resolution than the original reconstructed reference frame 702. In this context, the encoder 200 and decoder 300 may still require a full resolution reference frame. However, the zoom-up operation allocates a larger surface as discussed. Using the pitch and initial x, y coordinates of the reconstructed reference frame, a bounding box 703 is applied to the resulting reprojected reconstructed reference frame 701, and via scaling (704), a reprojected reconstructed reference frame 706 with full resolution is provided to the encoder 200 and decoder 300 (e.g., in a frame buffer, etc.), so that the reprojected reconstructed reference frame 706 (which may correspond to the reprojected reconstructed reference frame 227) will thereby correspond to the reference frame native resolution. These techniques allow the rest of the encoder 200 and decoder 300 to operate normally with respect to motion estimation / compensation, etc. It should be understood that using these techniques, pixel information about boundary pixels 705 will be lost. However, since similar scene poses will be used to generate frames to be encoded using the reprojected reconstructed reference frame 706 as a reference frame (eg, input frame 224 / frame for rendering 230 ), it is expected that this pixel information will not be needed during motion estimation / compensation.
[0074] Although described with respect to a zoom-in operation for generating the resulting reprojected reconstructed reference frame 701, any transformation or distortion of the frame that provides a larger original resolution or pixels outside the original size of the reconstructed reference frame can be subjected to the discussed bounding box and scaling techniques to generate a reprojected reconstructed reference frame having the same resolution as the original reconstructed reference frame. For example, a frame rotation transformation may provide pixels outside the original reconstructed reference frame that can be eliminated prior to the encoding / decoding process. In other embodiments, after the discussed projective transformation produces a zoom-in or similar effect, the boundary pixels 705 or portions thereof can be used for motion estimation / compensation if the encoding / decoding architecture supports it.
[0075] Figure 8 8 shows example post-processing 800 of a reprojected reconstructed reference frame following a zoom-out operation, arranged in accordance with at least some implementations of the present disclosure. Figure 8 As shown, after applying the projective transformation, the resulting reprojected reconstructed reference frame 801 has a larger size (e.g., h2x w2) than the original size (e.g., h1x w1) of the corresponding reconstructed reference frame 802 used to generate (via the projective transformation in question) the resulting reprojected reconstructed reference frame 801. The resulting reprojected reconstructed reference frame 801 can be characterized as a warped reconstructed reference frame, a resulting reconstructed reference frame, etc.
[0076] As shown, in embodiments where the resulting reprojected reconstructed reference frame 801 has a size smaller than the original size of the reconstructed reference frame 802 or a portion of the reprojected reconstructed reference frame 801 is within an edge of the reconstructed reference frame 802 and does not extend to the edge, an edge pixel generation operation 805 is applied to the resulting reprojected reconstructed reference frame 801 to generate a reprojected reconstructed reference frame 804 having the same size, shape, and pixel density as the reconstructed reference frame 802 (as well as the size and shape of the input frame to be encoded). In the illustrated embodiment, gaps 803 between outer edges (e.g., one or more edges) of the reprojected reconstructed reference frame 801 and corresponding edges of the reconstructed reference frame 802 are filled with corresponding constructed pixel values of the reprojected reconstructed reference frame 804. The constructed pixel values may be generated using any one or more suitable techniques (e.g., pixel replication techniques, etc.).
[0077] For example, the zoom down projection transformation causes the resulting reprojected reconstructed reference frame 801 to be scaled to a resolution smaller than the original reconstructed reference frame 802. As discussed above with respect to the zoom up operation, the encoder 200 and decoder 300 may still require a full resolution reference frame. Figure 8The zoom reduction shown, edge pixels (e.g., pixels used to fill gap 803) can be copied as discussed to fill in the missing pixels. Such pixel copying can be performed by pixel copying, pixel value extrapolation, etc. As shown, the reprojected reconstructed reference frame 804 with full resolution is provided to the encoder 200 and decoder 300 (e.g., in a frame buffer, etc.), so that the reprojected reconstructed reference frame 804 (which can correspond to the reprojected reconstructed reference frame 226) will thereby correspond to the reference frame native resolution. These techniques allow the rest of the encoder 200 and decoder 300 to operate normally with respect to motion estimation / compensation, etc.
[0078] Although described with respect to a zoom-out operation for generating the resulting reprojected reconstructed reference frame 601, any transformation or distortion that provides a frame with a resolution smaller than the original resolution may be subjected to the discussed pixel construction techniques to generate a reprojected reconstructed reference frame having the same resolution as the original reconstructed reference frame. For example, a frame rotation transformation may provide a pixel gap relative to the original reconstructed reference frame (which may have been constructed prior to the encoding / decoding process).
[0079] In addition, in some embodiments, reference Figure 1 , if the system 100 generates a rendered frame with barrel distortion, a reconstructed reference frame can be generated by: removing the barrel distortion; applying reprojection (e.g., a projective transformation); and reapplying the barrel distortion. These techniques generate a reprojected reconstructed reference frame from the same perspective as the current view. In addition, the distortion from the barrel distortion may change the size and shape of the object by a significant amount, which can be alleviated by removing the barrel distortion for reprojection. As discussed herein, reprojection produces a more similar view between the input frame and the reprojected reconstructed reference frame than using a reconstructed reference frame without reprojection.
[0080] return Figure 4 As discussed with reference to operation 404 of , in some embodiments, the projective transform is applied to the entire reconstructed reference frame to generate a resulting reprojected reconstructed reference frame. Such full frame projective transform application may provide simplicity of implementation. In other embodiments, the projective transform is applied only to one or more portions of the reconstructed reference frame to generate a resulting reprojected reconstructed reference frame. For example, one or more objects or regions of interest may be determined within the reprojected reconstructed reference frame such that the one or more regions of interest do not include a background of the reprojected reconstructed reference frame, and the projective transform may be applied only to the one or more regions of interest or background.
[0081] Fig. 9 2 shows an example projection transformation applied only to a region of interest, arranged in accordance with at least some implementations of the present disclosure. Fig. 9 As shown, a region of interest 902 may be provided within a reconstructed reference frame 901 such that the reconstructed reference frame 901 includes the region of interest 902 and a background 903 excluding the region of interest 902. The region of interest 902 may be determined or provided within the reconstructed reference frame 901 using any one or more suitable techniques. In an embodiment, the region of interest 902 (e.g., the coordinates of the region of interest 902) is provided by the application module 101 to the encoder 103 and the decoder 113 via a communication link 131 (e.g., within the bitstream 125, or next to the bitstream 125). In some embodiments, the application module 101 may determine the region of interest 902 such that the region of interest 902 is a rendered entity (e.g., a part of a game, etc.). In other embodiments, the region of interest 902 may be determined using object detection, object tracking, etc.
[0082] As shown, in an embodiment, the projective transformation 904 is applied only to the region of interest 902 to generate a distorted or reprojected region of interest 906 of the reprojected reconstructed reference frame 905, and not to the background 903. In other embodiments, the projective transformation 904 is applied only to the background 903 to generate a distorted or reprojected background of the reprojected reconstructed reference frame 905, and not to the region of interest 902. These techniques may not provide, for example, for distortion or reprojection of an object that is known to be fixed relative to the background that is changing. For example, if the object moves with the viewer (e.g., a ball in front of the viewer) while the background around the object moves in 6-DOF, etc., as discussed herein, it may be advantageous to not apply the projective transformation to the ball (e.g., without motion within the region of interest 902) while applying the projective transformation to the background 903. Similarly, when only the region of interest 902 is changing relative to the viewer (e.g., such that the background 903 is not changing or is only panning), it may be advantageous to apply the projective transformation only to the object of interest 902 while keeping the background 903 unchanged. In the illustrated embodiment, a single rectangular region of interest is provided. However, any number and shape of regions of interest may be implemented.
[0083] return Figure 4 As discussed with respect to full-frame projective transformations, operation 405 may be applied when the projective transformation is applied only to the region of interest 902 or the background 903. For example, when the region of interest is expanded due to the projective transformation, it may be scaled to within the original size of the region of interest 902. When the region of interest is smaller than the size of the region of interest 902 due to the projective transformation, pixels from the background 903 may be used or pixel reconstruction (e.g., replication) may be used to fill the gap.
[0084] Processing may continue along either the encoding path or the decoding path from optional operation 405, as shown in process 400. For example, encoder 200 and decoder 300 perform operations 401-405 in the same manner so that both have the same reprojected reconstructed reference frame for motion compensation (e.g., performed by motion compensation module 209 and motion compensation module 309, respectively). It should be understood that any mismatch between the motion compensation frames used by encoder 200 and decoder 300 will cause disruptions in the encoding process.
[0085] For the encoding processing path, processing may continue at operation 406, where motion estimation and motion compensation are performed using the reprojected reconstructed reference frame generated at operation 404 and / or operation 405. For example, as discussed with respect to encoder 200, a motion estimation search is performed on a block-by-block basis (e.g., by motion estimation module 202) for blocks of input frame 224 using the reprojected reconstructed reference frame as a motion compensated frame (e.g., by searching portions of some or all of the reprojected reconstructed reference frame). The best matching block is indicated by a reference index (e.g., indicating a reference frame (if more than one is used)) and a motion vector. Additionally, motion compensation is performed (e.g., by motion compensation module 209) to reconstruct the block by taking the best matching block and adding a corresponding reconstructed prediction residual, as discussed herein.
[0086] Processing may continue at operation 407, where the reference index and motion vector and the transformed and quantized prediction residual (e.g., the difference between the block of the input frame 224 and the corresponding best matching block after the difference is transformed and quantized) are encoded into a bitstream. The bitstream may be compliant with a standard (e.g., AVC, HEVC, etc.) or non-standard compliant, as discussed herein.
[0087] For the decoding processing path, processing may continue at operation 408, where motion compensation is performed (e.g., by motion compensation module 209). For example, a bitstream (e.g., the bitstream generated at operation 407) may be decoded to provide reference indices and motion vectors for motion compensation and a reconstructed prediction residual (e.g., the decoded residual after inverse quantization and inverse transformation). Motion compensation is performed to reconstruct the block by obtaining the best matching block indicated by the reference index and motion vector (for a reference frame (including a reconstructed reference frame after reprojection as discussed herein if more than one is used)) and adding the corresponding reconstructed prediction residual to the obtained best matching block.
[0088] Processing may continue at operation 409, where the frame is reconstructed using the reconstructed blocks generated at operation 408 and any intra-decoded reconstructed blocks to generate a reconstructed frame, thereby generating a frame for presentation. The reconstructed frame may optionally be deblock filtered to generate a reconstructed frame for presentation (and for reference to subsequently decoded frames). The reconstructed frame may be stored in a frame buffer, for example, for use as a reference frame and for display via a display device.
[0089] The discussed techniques can improve compression efficiency, particularly in contexts with complex scene pose changes. For example, for a use case of a video sequence generated based on a user playing a game where the user moves close to an object in the game (where head motion is inevitable), the following improvements have been observed. The video sequence is encoded with constant quality (e.g., the PSNR results are very similar, as shown in Table 1 below). The first row in Table 1 (labeled "Normal") corresponds to the encoding of the video sequence without using the discussed reprojection techniques. The second row labeled "Reference Frame Reprojection" corresponds to encoding the same video sequence with the reference frame reprojected based on scene pose difference data or information (e.g., based on the movement of the HMD), as discussed in this article. As shown in Table 1, compression for the test sequence is improved by more than 50%. The encoding improvement is that more motion vectors find better matches (e.g., 93% for inter-block or inter-coding unit (CU) compared to 79%), and fewer bits are spent on motion vectors, indicating that blocks find closer matches due to reprojection.
[0090]
[0091] Fig.10 1 is a flow chart illustrating an example process 1000 for video encoding using a reprojected reconstructed reference frame, arranged in accordance with at least some implementations of the present disclosure. The process 1000 may include: Fig.10 One or more operations 1001-1004 are shown. Process 1000 may form at least a portion of a video encoding process. As a non-limiting example, process 1000 may form at least a portion of a video encoding process, a video decoding process, a video pre-processing, or a video post-processing for a video undertaken by system 100 as discussed herein. Fig.11 System 1100 describes process 1000.
[0092] Fig.11 is an illustrative diagram of an example system 1100 for video encoding using a reprojected reconstructed reference frame, arranged in accordance with at least some implementations of the present disclosure. Fig.11As shown, the system 1100 may include a graphics processor 1101, a central processing unit 1102, and a memory 1103. The system 1100 may also include a scene posture tracking module 111 and / or a display 115. In addition, as shown in the figure, the graphics processor 1101 may include or implement a rendering module 102 and / or a rendering module 114. In addition, the central processing unit 1102 may include or implement an application module 101, an encoder 103, 200, and / or a decoder 113, 300. For example, as a system (e.g., a host system, etc.) implemented to generate a compressed bitstream from a rendered or captured frame, the system 1100 may include a rendering module 102 and encoders 103, 200 (e.g., encoder 103 and / or encoder 200 or components of one or both thereof). As a system (e.g., a sink, a display system, etc.) implemented to decompress a bitstream to generate a frame for presentation, the system 1100 may include a rendering module 114, an application module 101, a decoder 113, 300 (e.g., a decoder 113 and / or a decoder 300 or a component of one or both), a scene pose tracking module 111, and / or a display 115. For example, the system 1100 may implement the system 100 and / or the system 110. In the example of the system 1100, the memory 1103 may store video content (e.g., video frames, reconstructed reference frames after reprojection, bitstream data, scene pose data, scene pose difference data, or any other data or parameters discussed herein).
[0093] The graphics processor 1101 may include any number and type of graphics processors or processing units that can provide the operations discussed herein. These operations can be implemented via software or hardware or a combination thereof. In an embodiment, the modules shown in the graphics processor 1101 can be implemented via circuits, etc. For example, the graphics processor 1101 may include circuits dedicated to rendering frames, manipulating video data to generate compressed bitstreams, and / or circuits dedicated to manipulating compressed bitstreams to generate video data to provide the operations discussed herein. For example, the graphics processor 1101 may include electronic circuits for manipulating and changing memory to accelerate the creation of video frames in a frame buffer and / or for manipulating and changing memory to accelerate the creation of bitstreams based on images or frames of video.
[0094] The central processor 1102 may include any number and type of processing units or modules that may provide control and other high-level functions for the system 1100 and / or provide the operations discussed herein. For example, the central processor 1102 may include electronic circuits for executing instructions by performing basic arithmetic, logic, control, input / output operations, etc., as specified by the instructions of a computer program.
[0095] The memory 1103 may be any type of memory, such as a volatile memory (e.g., a static random access memory (SRAM), a dynamic random access memory (DRAM), etc.) or a non-volatile memory (e.g., a flash memory, etc.). In an embodiment, the memory 1103 may be configured to store video data (e.g., pixel values, control parameters, bitstream data, or any other video data, frame data, or any other data discussed herein). In a non-limiting example, the memory 1103 may be implemented by a cache memory. In an embodiment, one or more parts of the rendering module 102 and / or the rendering module 114 may be implemented via an execution unit (EU) of the graphics processor 1101. The execution unit may include, for example, programmable logic or circuitry (e.g., one or more logic cores that may provide a large number of programmable logic functions). In an embodiment, the rendering module 102 and / or the rendering module 114 may be implemented via dedicated hardware (e.g., a fixed function circuit, etc.). The fixed function circuit may include dedicated logic or circuitry, and may provide a set of fixed function entry points that may be mapped to the dedicated logic for a fixed purpose or function.
[0096] In the illustrated embodiment, the rendering module 102 and / or the rendering module 114 are implemented by the graphics processor 1101. In other embodiments, one or both of the rendering module 102 and / or the rendering module 114 or components thereof are implemented by the central processor 1102. Similarly, in the illustrated embodiment, the application module 101, the encoder 103, 200, and the decoder 113, 300 are implemented by the central processor 1102. In other embodiments, one, some, all, or components thereof of the application module 101, the encoder 103, 200, and the decoder 113, 300 are implemented by the graphics processor 1101. In some embodiments, one, some, all, or components thereof of the application module 101, the encoder 103, 200, and the decoder 113, 300 are implemented by a dedicated image or video processor.
[0097] return Fig.10 As discussed above, process 1000 may begin at operation 1001, where a reconstructed reference frame corresponding to a first scene pose is generated. The reconstructed reference frame may be generated using any one or more suitable techniques. For example, a reconstructed reference frame may be generated by determining reference blocks for a frame using intra-frame decoding and / or motion compensation techniques (at an encoder or decoder), and each reference block may be combined with a prediction residual (if any) to form a reconstructed reference block. The reconstructed reference blocks may be combined or fused into a frame, and the frame may be deblocking filtered to generate a reconstructed reference frame. For example, the reconstructed reference frame may correspond to the reconstructed reference frame 225 discussed with respect to the encoder 200 and / or decoder 300.
[0098] Processing can continue at operation 1002, wherein scene pose difference data indicating a scene pose change from a first scene pose to a second scene pose subsequent to the first scene pose is received or generated. The scene pose difference data may include any suitable data format and may be received or generated using any one or more suitable techniques. In an embodiment, the scene pose difference data includes a transformation matrix, 6-DOF differential data, a motion vector field, etc., as discussed herein.
[0099] In an embodiment, scene pose difference data is generated based on a difference between a first scene pose and a measured second scene pose measured at a time subsequent to the time corresponding to the first scene pose. In addition, the second scene pose can be used to render a frame, as discussed herein. In an embodiment, scene pose difference data is predicted using an extrapolation technique, etc. In an embodiment, scene pose difference data is predicted by extrapolating second scene pose difference data indicating a change from a third scene pose to a second scene pose of the first scene pose, wherein the first scene pose is subsequent to the third scene pose.
[0100] Processing may continue at operation 1003, where a projection transformation is applied to at least a portion of the reconstructed reference frame based on the scene pose difference data to generate a reprojected reconstructed reference frame. The projection transformation may be applied using any one or more suitable techniques. In an embodiment, the projection transformation includes both an affine projection (e.g., an affine projection component) and a non-affine projection (e.g., a non-affine projection component), the non-affine projection including at least one of a zoom projection, a barrel distortion projection, and a spherical rotation projection.
[0101] As discussed herein, the projective transform may be applied to the entire reconstructed reference frame or only a portion of the reconstructed reference frame. In an embodiment, the projective transform is applied to the entire reconstructed reference frame. In an embodiment, the process 1000 further comprises: determining a region of interest of the reconstructed reference frame and a background region of the reconstructed reference frame excluding the region of interest, and applying the projective transform comprises: applying the projective transform only to the region of interest, or only to the background of the reconstructed reference frame.
[0102] In addition, post-processing may be provided (after applying the projective transformation) to generate a reprojected reconstructed reference frame (e.g., a final frame of a format to be used as a motion compensated reference frame). In an embodiment, applying the projective transformation includes: applying a zoom-in transformation to the reconstructed reference frame to generate a first reprojected reconstructed reference frame having a size larger than the size of the reconstructed reference frame, and process 1000 also includes: applying a bounding box having the same size as the reconstructed reference frame to the first reprojected reconstructed reference frame, and scaling a portion of the first reprojected reconstructed reference frame within the bounding box to the size and resolution of the reconstructed reference frame to generate the reprojected reconstructed reference frame. In an embodiment, applying the projective transformation includes: applying a zoom-out transformation to the reconstructed reference frame to generate a first reprojected reconstructed reference frame having a size smaller than the size of the reconstructed reference frame, and process 1000 also includes: generating edge pixels adjacent to at least one edge of the first reprojected reconstructed reference frame to provide a reprojected reconstructed reference frame having the same size and resolution as the reconstructed reference frame. In an embodiment, applying the projection transformation includes: applying a spherical rotation to the reconstructed reference frame to generate a first reprojected reconstructed reference frame, and processing 1000 also includes: generating edge pixels adjacent to at least one edge of the first reprojected reconstructed reference frame to provide a reprojected reconstructed reference frame having the same size and resolution as the reconstructed reference frame.
[0103] In some embodiments, a projective transformation may be applied conditional on the estimated scene pose difference data. In an embodiment, at least one scene change difference amplitude value corresponding to the scene pose difference data is compared with a threshold value, and the application of the projective transformation to at least a portion of the reconstructed reference frame is conditional on the scene change difference amplitude value meeting or exceeding the threshold value. In an embodiment, at least one scene change difference amplitude value corresponding to the scene pose difference data is compared with a first threshold value and a second threshold value greater than the first threshold value, and the application of the projective transformation to at least a portion of the reconstructed reference frame is conditional on the scene change difference amplitude value meeting or exceeding the first threshold value but not exceeding the second threshold value.
[0104] In some embodiments, the discussed application of the projective transform to the reconstructed reference frame may be performed concurrently with other operations to reduce lag time or delay in the process. In an embodiment, the process 1000 further comprises at least one of: rendering the second frame at least partially concurrently with said applying the projective transform; and receiving the bitstream at least partially concurrently with said applying the projective transform.
[0105] Processing may continue at operation 1004, where motion compensation is performed to generate a current reconstructed frame using the reprojected reconstructed reference frame as a motion compensated reference frame. The motion compensation may be performed at the encoder (e.g., as part of a local loop) or at the decoder. For example, the motion vector and frame reference index information may be used to obtain a block from the reprojected reconstructed reference frame for use in reconstructing the current reconstructed frame.
[0106] In some embodiments, only the reprojected reconstructed reference frame is used as a motion compensation reference frame. In other embodiments, an additional motion compensation reference frame is used. In an embodiment, performing motion compensation further comprises: using both the reconstructed reference frame (e.g., without applying a projection transformation) and the reprojected reconstructed reference frame as motion compensation reference frames, performing motion compensation in a block-by-block manner, such that the first block of the current reconstructed frame is motion compensated with reference to the reconstructed reference frame, and the second block of the current reconstructed frame is motion compensated with reference to the reprojected reconstructed reference frame. In an embodiment, processing 1000 further comprises: generating a second reconstructed reference frame corresponding to a third scene pose, wherein the third scene pose is before the first scene pose; receiving second scene pose difference data indicating a scene pose change from the third scene pose to the second scene pose; based on the second scene pose difference data, applying a second projection transformation to at least a portion of the second reconstructed reference frame to generate a second reprojected reconstructed reference frame, so that performing motion compensation on the current frame uses both the reprojected reconstructed reference frame and the second reprojected reconstructed reference frame as motion compensation reference frames.
[0107] The various components of the systems described herein may be implemented in software, firmware, and / or hardware, and / or any combination thereof. For example, the various components of the systems 100, 110, 1100 may be provided, at least in part, by hardware such as a computing system-on-chip (SoC) such as may be found in a computing system (e.g., a smart phone). It will be appreciated by those skilled in the art that the systems described herein may include additional components not yet described in the corresponding figures. For example, the systems discussed herein may include additional components not yet described for clarity (e.g., a bitstream multiplexer or demultiplexer module, etc.).
[0108] Although implementations of the example processing discussed herein may include undertaking all operations shown in the order shown, the present disclosure is not limited thereto, and in various examples, implementations of the example processing herein may include only a subset of the operations shown, operations performed in a different order than shown, or additional operations.
[0109] In addition, one or more operations discussed herein may be undertaken in response to instructions provided by one or more computer program products. These program products may include signal-bearing media that provide instructions that, when executed by, for example, a processor, may provide the functionality described herein. A computer program product may be provided by one or more machine-readable media in any form. Thus, for example, a processor including one or more graphics processing units or processor cores may undertake one or more blocks of the example processing herein in response to program codes and / or instructions or instruction sets transmitted to the processor by one or more machine-readable media. Typically, a machine-readable medium may transmit software in the form of program codes and / or instructions or instruction sets that may enable any device and / or system described herein to implement the techniques, modules, components, etc. discussed herein.
[0110] As used in any implementation described herein, the term "module" refers to any combination of software logic, firmware logic, hardware logic, and / or circuitry configured to provide the functionality described herein. Software may be embodied as a software package, code, and / or instruction set or instructions, and "hardware" used in any implementation described herein may include, for example, hard-wired circuits, programmable circuits, state machine circuits, fixed function circuits, execution unit circuits, and / or firmware that stores instructions executed by programmable circuits, either alone or in any combination. Modules may be collectively or individually embodied as circuits that form part of a larger system (e.g., an integrated circuit (IC), a system on a chip (SoC), etc.).
[0111] Fig.12 is an illustrative diagram of an example system 1200 arranged in accordance with at least some implementations of the present disclosure. In various implementations, the system 1200 may be a mobile system, but the system 1200 is not limited to this context. For example, the system 1200 may be incorporated into a personal computer (PC), a laptop, an ultra-laptop, a tablet, a touchpad, a portable computer, a handheld computer, a palmtop, a personal digital assistant (PDA), a cellular phone, a combination cellular phone / PDA, a television, a smart device (e.g., a smart phone, a smart tablet, or a smart TV), a mobile Internet device (MID), a messaging device, a data communication device, a camera (e.g., a point-and-shoot camera, a super zoom camera, a digital single-lens reflex (DSLR) camera), a virtual reality device, an augmented reality device, and the like.
[0112] In various implementations, system 1200 includes a platform 1202 coupled to a display 1220. Platform 1202 may receive content from a content device such as content services device(s) 1230 or content delivery device(s) 1240 or other similar content sources. A navigation controller 1250 including one or more navigation features may be used to interact with, for example, platform 1202 and / or display 1220. Each of these components is described in greater detail below.
[0113] In various implementations, the platform 1202 may include any combination of a chipset 1205, a processor 1210, memory 1212, an antenna 1213, storage 1214, a graphics subsystem 1215, applications 1216, and / or a radio 1218. The chipset 1205 may provide intercommunication between the processor 1210, the memory 1212, the storage 1214, the graphics subsystem 1215, the applications 1216, and / or the radio 1218. For example, the chipset 1205 may include a storage adapter (not depicted) capable of providing intercommunication with the storage 1214.
[0114] Processor 1210 may be implemented as a complex instruction set computer (CISC) or reduced instruction set computer (RISC) processor, an x86 instruction set compatible processor, a multi-core or any other microprocessor or central processing unit (CPU). In various implementations, processor 1210 may be a dual-core processor, a dual-core mobile processor, etc.
[0115] The memory 1212 may be implemented as a volatile memory device such as, but not limited to, a random access memory (RAM), a dynamic random access memory (DRAM), or a static RAM (SRAM).
[0116] Storage 1214 may be implemented as a non-volatile storage device (e.g., but not limited to, a magnetic disk drive, an optical disk drive, a tape drive, an internal storage device, an attached storage device, flash memory, battery-backed SDRAM (synchronous DRAM), and / or a network accessible storage device). In various implementations, storage 1214 may include technology for adding storage performance-enhancing protection to valuable digital media, such as when multiple hard drives are included.
[0117] The graphics subsystem 1215 may perform processing of images (e.g., still images or video) for display. The graphics subsystem 1215 may be, for example, a graphics processing unit (GPU) or a visual processing unit (VPU). An analog or digital interface may be used to communicatively couple the graphics subsystem 1215 and the display 1220. For example, the interface may be any of a High Definition Multimedia Interface, a DisplayPort, wireless HDMI, and / or wireless HD compliant technologies. The graphics subsystem 1215 may be integrated into the processor 1210 or the chipset 1205. In some implementations, the graphics subsystem 1215 may be a standalone device communicatively coupled to the chipset 1205.
[0118] The graphics and / or video processing techniques described herein can be implemented in various hardware architectures. For example, graphics and / or video functions can be integrated into a chipset. Alternatively, a discrete graphics and / or video processor can be used. As another implementation, graphics and / or video functions can be provided by a general-purpose processor including a multi-core processor. In other embodiments, these functions can be implemented in consumer electronic devices.
[0119] Radio 1218 may include one or more radios capable of sending and receiving signals using a variety of suitable wireless communication technologies. These technologies may involve communications across one or more wireless networks. Example wireless networks include, but are not limited to, wireless local area networks (WLANs), wireless personal area networks (WPANs), wireless metropolitan area network (WMANs), cellular networks, and satellite networks. When communicating across these networks, radio 1218 may operate in accordance with one or more applicable standards in any version.
[0120] In various implementations, the display 1220 may include any television type monitor or display. The display 1220 may include, for example, a computer display screen, a touch screen display, a video monitor, a television-like device, and / or a television. The display 1220 may be digital and / or analog. In various implementations, the display 1220 may be a holographic display. In addition, the display 1220 may be a transparent surface that can receive visual projections. These projections may convey various forms of information, images, and / or objects. For example, these projections may be visual overlays for mobile augmented reality (MAR) applications. Under the control of one or more software applications 1216, the platform 1202 may display a user interface 1222 on the display 1220.
[0121] In various implementations, content services device(s) 1230 may be hosted by any national, international, and / or independent service and thus may access platform 1202 via the Internet, for example. Content services device(s) 1230 may be coupled to platform 1202 and / or to display 1220. Platform 1202 and / or content services device(s) 1230 may be coupled to network 1260 to communicate (e.g., send and / or receive) media information to and from network 1260. Content delivery device(s) 1240 may also be coupled to platform 1202 and / or to display 1220.
[0122] In various implementations, content services device(s) 1230 may include a cable box, a personal computer, a network, a telephone, an Internet-enabled device or appliance capable of transmitting digital information and / or content, and any other similar device capable of delivering content unidirectionally or bidirectionally via network 1260 or in a direct manner between content providers and platform 1202 and / or display 1220. It should be appreciated that content may be delivered unidirectionally and / or bidirectionally to and from any one of the components of system 1200 and content providers via network 1260. Examples of content may include any media information, including, for example, video, music, medical and gaming information, and the like.
[0123] Content services device 1230 can receive content (e.g., cable television program delivery including media information, digital information, and / or other content). Examples of content providers can include any cable or satellite television or radio or Internet content providers. The examples provided are not intended to limit implementations according to the present disclosure in any way.
[0124] In various implementations, platform 1202 may receive control signals from navigation controller 1250 having one or more navigation features. For example, the navigation features may be used to interact with user interface 1222. In various embodiments, the navigation may be a pointing device that may be a computer hardware component (specifically, a human interface device) that allows a user to input spatial (e.g., continuous and multi-dimensional) data into a computer. Many systems (e.g., graphical user interfaces (GUIs) and televisions and monitors) allow a user to control data using physical gestures and provide it to a computer or television.
[0125] The movement of the navigation features may be replicated on a display (e.g., display 1220) by movement of a pointer, cursor, focus ring, or other visual indicator displayed on the display. For example, under the control of software application 1216, the navigation features located on the navigation may be mapped to virtual navigation features displayed on user interface 1222. In various embodiments, rather than being a separate component, the navigation features may be integrated into platform 1202 and / or display 1220. However, the present disclosure is not limited to the elements or in the context shown or described herein.
[0126] In various implementations, drivers (not shown) may include, for example, when enabled, technology for enabling a user to instantly turn on and off platform 1202 (such as a television) with the touch of a button after initial booting. Program logic may also allow platform 1202 to stream content to a media adapter or other content service device 1230 or content delivery device 1240 even when the platform is "off." In addition, for example, chipset 1205 may include hardware and / or software support for 5.1 surround sound audio and / or high definition 7.1 surround sound audio. Drivers may include a graphics driver for an integrated graphics platform. In various embodiments, the graphics driver may include a peripheral component interconnect (PCI) express graphics card.
[0127] In various implementations, any one or more of the components shown in system 1200 may be integrated. For example, platform 1202 and content service device 1230 may be integrated, or platform 1202 and content delivery device 1240 may be integrated, or platform 1202, content service device 1230, and content delivery device 1240 may be integrated. In various embodiments, platform 1202 and display 1220 may be an integrated unit. For example, display 1220 and content service device 1230 may be integrated, or display 1220 and content delivery device 1240 may be integrated. These examples are not meant to limit the present disclosure.
[0128] In various embodiments, system 1200 may be implemented as a wireless system, a wired system, or a combination of the two. When implemented as a wireless system, system 1200 may include components and interfaces suitable for communicating via a wireless shared medium (e.g., one or more antennas, transmitters, receivers, transceivers, amplifiers, filters, control logic, etc.). Examples of wireless shared media may include portions of a wireless spectrum (e.g., an RF spectrum, etc.). When implemented as a wired system, system 1200 may include components and interfaces suitable for communicating via a wired communication medium (e.g., an input / output (I / O) adapter, a physical connector for connecting an I / O adapter to a corresponding wired communication medium, a network interface card (NIC), a disk controller, a video controller, an audio controller, etc.). Examples of wired communication media may include wires, cables, metal leads, printed circuit boards (PCBs), backplanes, switch structures, semiconductor materials, twisted pair wires, coaxial cables, optical fibers, etc.
[0129] Platform 1202 can establish one or more logical channels or physical channels to transfer information. Information can include media information and control information. Media information can refer to any data representing the content of the schematic for the user. Examples of content can include, for example, data from voice conversations, video conferences, streaming video, electronic mail ("email") messages, voice mail messages, alphanumeric symbols, graphics, images, videos, text, etc. Data from voice conversations can be, for example, voice information, silent periods, background noise, comfort noise, tones, etc. Control information can refer to any data representing commands, instructions, or control words for an automated system. For example, control information can be used to route media information through a system, or to command a node to process media information in a predetermined manner. However, embodiments are not limited to Fig.12 The elements or context shown or described in the.
[0130] As described above, the system 1200 may be embodied in varying physical styles or shapes. Fig.13 An example digital device 1300 is shown, arranged in accordance with at least some implementations of the present disclosure. In some examples, system 1200 may be implemented via device 1300. In other examples, system 1100 or portions thereof may be implemented via device 1300. In various embodiments, for example, device 1300 may be implemented as a mobile computing device with wireless capabilities. A mobile computing device may refer to any device with a processing system and a mobile power source or power supply (e.g., one or more batteries).
[0131] As mentioned above, examples of mobile computing devices may include a personal computer (PC), a laptop, an ultra-laptop, a tablet, a touchpad, a portable computer, a handheld computer, a palmtop, a personal digital assistant (PDA), a cellular phone, a combination cellular phone / PDA, a smart device (e.g., a smart phone, a smart tablet, or a smart mobile television), a mobile Internet device (MID), a messaging device, a data communications device, a camera, etc.
[0132] Examples of mobile computing devices may also include computers that are arranged to be worn by a person (e.g., a wrist computer, a finger computer, an earring computer, an eyeglass computer, a belt-clip computer, an arm-band computer, a shoe computer, a clothing computer, and other wearable computers). In various embodiments, for example, the mobile computing device may be implemented as a smart phone capable of executing computer applications as well as voice communications and / or data communications. Although some embodiments may be described with the mobile computing device implemented as a smart phone by way of example, it will be appreciated that other embodiments may also be implemented using other wireless mobile computing devices. The embodiments are not limited in this context.
[0133] like Fig.13As shown, device 1300 may include a housing with a front 1301 and a rear 1302. Device 1300 includes a display 1304, an input / output (I / O) device 1306, and an integrated antenna 1308. Device 1300 may also include a navigation feature 1312. I / O device 1306 may include any suitable I / O device for inputting information into a mobile computing device. Examples for I / O device 1306 may include an alphanumeric keyboard, a numeric keypad, a touch pad, an input key, a button, a switch, a microphone, a speaker, a voice recognition device, and software, etc. Information may also be input into device 1300 by means of a microphone (not shown), or may be digitized by a voice recognition device. As shown, device 1300 may include a camera 1305 (e.g., including a lens, an aperture, and an imaging sensor) and a flash 1310 integrated into the rear 1302 (or other places) of device 1300. In other examples, the camera 1305 and flash 1310 can be integrated into the front 1301 of the device 1300, or both a front camera and a rear camera can be provided. For example, the camera 1305 and flash 1310 can be components of a camera module for generating image data that is processed into streaming video that is output to the display 1304 and / or transmitted remotely from the device 1300 via the antenna 1308.
[0134] Various embodiments can be realized using hardware elements, software elements or a combination of the two.The example of hardware elements can include processors, microprocessors, circuits, circuit elements (e.g., transistors, resistors, capacitors, inductors, etc.), integrated circuits, application specific integrated circuits (ASICs), programmable logic devices (PLDs), digital signal processors (DSPs), field programmable gate arrays (FPGAs), logic gates, registers, semiconductor devices, chips, microchips, chipsets, etc.The example of software can include software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, processes, software interfaces, application program interfaces (APIs), instruction sets, computing codes, computer codes, code segments, computer code segments, words, values, symbols or any combination thereof.Determining whether to use hardware elements and / or software elements to realize an embodiment can vary according to any number of factors, such as desired computing rate, power level, thermal tolerance, processing cycle budget, input data rate, output data rate, memory resources, data bus speed and other design or performance constraints.
[0135] One or more aspects of at least one embodiment may be implemented by representative instructions stored on a machine-readable medium representing various logic within a processor, which when read by a machine causes the machine to fabricate logic to perform the techniques described herein. Such representations, known as "IP cores," may be stored on a tangible machine-readable medium and provided to various customers or manufacturing sites to load into a manufacturing machine that actually builds the logic or processor.
[0136] Although the specific features described herein have been described with reference to various implementations, this description is not intended to be understood in a limiting sense. Therefore, various modifications of the implementations described herein and other implementations that are obvious to those skilled in the art to which the present disclosure belongs are considered to be within the spirit and scope of the present disclosure.
[0137] The following examples pertain to further embodiments.
[0138] In one or more first embodiments, a computer-implemented method for video encoding includes: generating a reconstructed reference frame corresponding to a first scene pose; receiving scene pose difference data indicating a scene pose change from the first scene pose to a second scene pose subsequent to the first scene pose; based on the scene pose difference data, applying a projection transformation to at least a portion of the reconstructed reference frame to generate a reprojected reconstructed reference frame; and performing motion compensation using the reprojected reconstructed reference frame as a motion compensation reference frame to generate a current reconstructed frame.
[0139] In one or more second embodiments, for any one of the first embodiments, the projection transformation includes both affine projection and non-affine projection, the non-affine projection includes at least one of zoom projection, barrel distortion projection and spherical rotation projection, and the scene pose difference data includes one of a transformation matrix, 6-DOF difference data and a motion vector field.
[0140] In one or more third embodiments, for any one of the first and second embodiments, the projective transformation is applied to the entire reconstructed reference frame, and the method further includes at least one of: rendering a second frame at least partially simultaneously with applying the projective transformation; and receiving a bitstream at least partially simultaneously with applying the projective transformation.
[0141] In one or more fourth embodiments, for any one of the first to third embodiments, the performing motion compensation includes: using both the reconstructed reference frame and the reprojected reconstructed reference frame as motion compensation reference frames, performing motion compensation in a block-by-block manner, so that the first block of the current reconstructed frame is motion compensated with reference to the reconstructed reference frame, and the second block of the current reconstructed frame is motion compensated with reference to the reprojected reconstructed reference frame.
[0142] In one or more fifth embodiments, for any one of the first to fourth embodiments, the method further includes: determining a region of interest of the reconstructed reference frame and a background region of the reconstructed reference frame excluding the region of interest, wherein applying the projection transformation includes: applying the projection transformation only to one of the region of interest and the background of the reconstructed reference frame.
[0143] In one or more sixth embodiments, for any one of the first to fifth embodiments, applying the projection transformation includes: applying a zoom magnification transformation to the reconstructed reference frame to generate a first reprojected reconstructed reference frame having a size larger than that of the reconstructed reference frame, and the method also includes: applying a bounding box having the same size as the reconstructed reference frame to the first reprojected reconstructed reference frame, and scaling the portion of the first reprojected reconstructed reference frame within the bounding box to the size and resolution of the reconstructed reference frame to generate a reprojected reconstructed reference frame.
[0144] In one or more seventh embodiments, for any one of the first to sixth embodiments, applying the projection transformation includes: applying a zoom reduction transformation to the reconstructed reference frame to generate a first reprojected reconstructed reference frame having a size smaller than that of the reconstructed reference frame, and the method also includes: generating edge pixels adjacent to at least one edge of the first reprojected reconstructed reference frame to provide a reprojected reconstructed reference frame having the same size and resolution as the reconstructed reference frame.
[0145] In one or more eighth embodiments, for any one of the first to seventh embodiments, applying the projection transformation includes: applying a spherical rotation to the reconstructed reference frame to generate a first reprojected reconstructed reference frame, and the method also includes: generating edge pixels adjacent to at least one edge of the first reprojected reconstructed reference frame to provide a reprojected reconstructed reference frame having the same size and resolution as the reconstructed reference frame.
[0146] In one or more ninth embodiments, for any one of the first to eighth embodiments, scene pose difference data is predicted by extrapolating second scene pose difference data indicating a second scene pose change from a third scene pose to the first scene pose, wherein the first scene pose is after the third scene pose.
[0147] In one or more tenth embodiments, for any one of the first to ninth embodiments, the method further includes: comparing at least one scene change difference amplitude value corresponding to the scene posture difference data with a threshold, wherein applying the projection transformation to at least a portion of the reconstructed reference frame is conditional on the scene change difference amplitude value satisfying or exceeding the threshold.
[0148] In one or more eleventh embodiments, for any one of the first to tenth embodiments, the method also includes: generating a second reconstructed reference frame corresponding to a third scene posture, wherein the third scene posture is before the first scene posture; receiving second scene posture difference data indicating a scene posture change from the third scene posture to the second scene posture; and based on the second scene posture difference data, applying a second projection transformation to at least a portion of the second reconstructed reference frame to generate a second reprojected reconstructed reference frame, wherein motion compensation is performed on the current frame using both the reprojected reconstructed reference frame and the second reprojected reconstructed reference frame as motion compensation reference frames.
[0149] In one or more twelfth embodiments, a system for video encoding comprises: a memory for storing a reconstructed reference frame corresponding to a first scene pose; and a processor coupled to the memory, the processor being configured to: apply a projection transformation to at least a portion of the reconstructed reference frame based on scene pose difference data to generate a reprojected reconstructed reference frame, wherein the scene pose difference data indicates a scene pose change from the first scene pose to a second scene pose subsequent to the first scene pose; and perform motion compensation using the reprojected reconstructed reference frame as a motion compensation reference frame to generate a current reconstructed frame.
[0150] In one or more thirteenth embodiments, for any one of the twelfth embodiments, the projection transformation includes both affine projection and non-affine projection, the non-affine projection includes at least one of zoom projection, barrel distortion projection and spherical rotation projection, and the scene pose difference data includes one of a transformation matrix, 6-degree-of-freedom difference data and a motion vector field.
[0151] In one or more fourteenth embodiments, for any one of the twelfth embodiment and the thirteenth embodiment, the processor performs motion compensation including: the processor uses both the reconstructed reference frame and the reprojected reconstructed reference frame as motion compensation reference frames, and performs motion compensation in a block-by-block manner, so that the first block of the current reconstructed frame is motion compensated with reference to the reconstructed reference frame, and the second block of the current reconstructed frame is motion compensated with reference to the reprojected reconstructed reference frame.
[0152] In one or more fifteenth embodiments, for any one of the twelfth to fourteenth embodiments, the processor is further used to: determine a region of interest of the reconstructed reference frame and a background region of the reconstructed reference frame excluding the region of interest, wherein the processor applies the projection transformation including: the processor applies the projection transformation only to one of the region of interest and the background of the reconstructed reference frame.
[0153] In one or more sixteenth embodiments, for any one of the twelfth to fifteenth embodiments, the processor is further used to: predict scene gesture difference data based on second scene gesture difference data extrapolated to indicate a second scene gesture change from a third scene gesture to the first scene gesture, wherein the first scene gesture is after the third scene gesture.
[0154] In one or more seventeenth embodiments, for any one of the twelfth to sixteenth embodiments, the processor is further used to: compare at least one scene change difference amplitude value corresponding to the scene posture difference data with a threshold, wherein the processor applies the projection transformation to at least a portion of the reconstructed reference frame on the condition that the scene change difference amplitude value meets or exceeds the threshold.
[0155] In one or more eighteenth embodiments, for any one of the twelfth to seventeenth embodiments, the processor is further used to: generate a second reconstructed reference frame corresponding to the third scene posture, wherein the third scene posture is before the first scene posture; receive second scene posture difference data indicating a scene posture change from the third scene posture to the second scene posture; and based on the second scene posture difference data, apply a second projection transformation to at least a portion of the second reconstructed reference frame to generate a second reprojected reconstructed reference frame, wherein the processor performs motion compensation on the current frame, including: the processor uses both the reprojected reconstructed reference frame and the second reprojected reconstructed reference frame as motion compensation reference frames.
[0156] In one or more nineteenth embodiments, a system for video encoding comprises: a module for generating a reconstructed reference frame corresponding to a first scene pose; a module for receiving scene pose difference data indicating a scene pose change from the first scene pose to a second scene pose subsequent to the first scene pose; a module for applying a projection transformation to at least a portion of the reconstructed reference frame based on the scene pose difference data to generate a reprojected reconstructed reference frame; and a module for performing motion compensation using the reprojected reconstructed reference frame as a motion compensation reference frame to generate a current reconstructed frame.
[0157] In one or more twentieth embodiments, for any one of the nineteenth embodiments, the projection transformation includes both affine projection and non-affine projection, the non-affine projection includes at least one of zoom projection, barrel distortion projection and spherical rotation projection, and the scene pose difference data includes one of a transformation matrix, 6-degree-of-freedom difference data and a motion vector field.
[0158] In one or more twenty-first embodiments, for any one of the nineteenth and twentieth embodiments, the projection transformation is applied to the entire reconstructed reference frame, and the system also includes at least one of: a module for rendering a second frame at least partially simultaneously with the application of the projection transformation; and a module for receiving a bitstream at least partially simultaneously with the application of the projection transformation.
[0159] In one or more twenty-second embodiments, for any one of the nineteenth to twenty-first embodiments, a module for performing motion compensation includes: a module for using both the reconstructed reference frame and the reprojected reconstructed reference frame as motion compensation reference frames, and performing motion compensation in a block-by-block manner, so that the first block of the current reconstructed frame is motion compensated with reference to the reconstructed reference frame, and the second block of the current reconstructed frame is motion compensated with reference to the reprojected reconstructed reference frame.
[0160] In one or more twenty-third embodiments, at least one machine-readable medium comprises a plurality of instructions, which in response to being executed on a computing device cause the computing device to perform video encoding by the following steps: generating a reconstructed reference frame corresponding to a first scene pose; receiving scene pose difference data indicating a scene pose change from the first scene pose to a second scene pose subsequent to the first scene pose; based on the scene pose difference data, applying a projection transformation to at least a portion of the reconstructed reference frame to generate a reprojected reconstructed reference frame; and performing motion compensation using the reprojected reconstructed reference frame as a motion compensation reference frame to generate a current reconstructed frame.
[0161] In one or more twenty-fourth embodiments, for any one of the twenty-third embodiments, the projection transformation includes both affine projection and non-affine projection, the non-affine projection includes at least one of zoom projection, barrel distortion projection and spherical rotation projection, and the scene pose difference data includes one of a transformation matrix, 6-degree-of-freedom difference data and a motion vector field.
[0162] In one or more twenty-fifth embodiments, for any one of the twenty-third and twenty-fourth embodiments, the performing motion compensation includes: using both the reconstructed reference frame and the reprojected reconstructed reference frame as motion compensation reference frames, performing motion compensation in a block-by-block manner, so that the first block of the current reconstructed frame is motion compensated with reference to the reconstructed reference frame, and the second block of the current reconstructed frame is motion compensated with reference to the reprojected reconstructed reference frame.
[0163] In one or more twenty-sixth embodiments, for any one of the twenty-third to twenty-fifth embodiments, the machine-readable medium also includes multiple instructions, which, in response to being executed on the computing device, cause the computing device to perform video encoding through the following steps: determining a region of interest of the reconstructed reference frame and a background region of the reconstructed reference frame excluding the region of interest, wherein applying the projection transformation includes applying the projection transformation only to one of the region of interest and the background of the reconstructed reference frame.
[0164] In one or more twenty-seventh embodiments, for any one of the twenty-third to twenty-sixth embodiments, the machine-readable medium also includes multiple instructions, which, in response to being executed on the computing device, cause the computing device to perform video encoding through the following steps: predicting scene pose difference data by extrapolating second scene pose difference data indicating a change from a third scene pose to a second scene pose of the first scene pose, wherein the first scene pose is after the third scene pose.
[0165] In one or more twenty-eighth embodiments, for any one of the twenty-third to twenty-seventh embodiments, the machine-readable medium also includes multiple instructions, which, in response to being executed on the computing device, cause the computing device to perform video encoding through the following steps: comparing at least one scene change difference amplitude value corresponding to the scene posture difference data with a threshold, wherein applying the projection transformation to at least a portion of the reconstructed reference frame is conditional on the scene change difference amplitude value satisfying or exceeding the threshold.
[0166] In one or more twenty-ninth embodiments, for any one of the twenty-third to twenty-eighth embodiments, the machine-readable medium also includes multiple instructions, which, in response to being executed on the computing device, cause the computing device to perform video encoding through the following steps: generating a second reconstructed reference frame corresponding to a third scene pose, wherein the third scene pose is before the first scene pose; receiving second scene pose difference data indicating a scene pose change from the third scene pose to the second scene pose; and based on the second scene pose difference data, applying a second projection transformation to at least a portion of the second reconstructed reference frame to generate a second reprojected reconstructed reference frame, wherein motion compensation is performed on the current frame using both the reprojected reconstructed reference frame and the second reprojected reconstructed reference frame as motion compensation reference frames.
[0167] In one or more thirtieth embodiments, at least one machine-readable medium may include a plurality of instructions, which in response to being executed on a computing device cause the computing device to perform a method as described in any one of the above embodiments.
[0168] In one or more thirty-first embodiments, an apparatus or system may include a module for executing a method as described in any one of the above embodiments.
[0169] It should be understood that the embodiments are not limited to the embodiments described above, but can be practiced through modification and alteration without departing from the scope of the appended claims. For example, the above examples may include a specific combination of features. However, the above embodiments are not limited to this, and in various implementations, the above embodiments may include only a subset of these features, a different order of these features, a different combination of these features, and / or additional features in addition to these features explicitly listed. Therefore, the scope of the embodiments should be determined with reference to the drawings together with the full scope of equivalents to which these claims belong.
Claims
1. A computer-implemented method for video encoding, comprising: generating a reconstructed reference frame corresponding to a first scene pose, wherein the first scene pose indicates a pose of a scene with respect to a first viewpoint, and wherein the first viewpoint corresponds to a first position and orientation of a system for video encoding relative to the scene in three-dimensional (3D) space; generating scene pose difference data indicating a change in scene pose from the first scene pose to a second scene pose subsequent to the first scene pose, wherein the second scene pose is associated with a second viewpoint corresponding to a second position and orientation of the system relative to the scene in 3D space, and wherein the scene pose difference data is generated by tracking movement of the system from the first position and orientation to the second position and orientation; Based on the scene pose difference data, applying a projective transformation to at least a portion of the reconstructed reference frame to generate a reprojected reconstructed reference frame; and The reprojected reconstructed reference frame is used as a motion compensation reference frame to perform motion compensation to generate a current reconstructed frame.
2. The method of claim 1, wherein: The projection transformation includes both affine projection and non-affine projection, the non-affine projection includes at least one of zoom projection, barrel distortion projection and spherical rotation projection, and wherein the scene pose difference data includes one of a transformation matrix, 6-DOF difference data and a motion vector field.
3. The method according to claim 1 or 2, wherein: The projective transformation is applied to the entire reconstructed reference frame, and the method further comprises at least one of the following: rendering a second frame at least partially concurrently with said applying the projective transformation; and A bitstream is received at least partially concurrently with said applying the projective transform.
4. The method according to claim 1 or 2, wherein: The performing motion compensation comprises: Using both the reconstructed reference frame and the reprojected reconstructed reference frame as motion compensation reference frames, motion compensation is performed in a block-by-block manner, so that the first block of the current reconstructed frame is motion compensated with reference to the reconstructed reference frame, and the second block of the current reconstructed frame is motion compensated with reference to the reprojected reconstructed reference frame.
5. The method according to claim 1 or 2, further comprising: A region of interest of the reconstructed reference frame and a background region of the reconstructed reference frame excluding the region of interest are determined, wherein applying the projective transformation comprises: applying the projective transformation to only one of the region of interest and the background region of the reconstructed reference frame.
6. The method of claim 1, wherein: Applying the projective transformation includes applying a zoom magnification transformation to the reconstructed reference frame to generate a first reprojected reconstructed reference frame having a size larger than the size of the reconstructed reference frame, and the method further includes: applying a bounding box having the same size as the reconstructed reference frame to the first reprojected reconstructed reference frame; and The portion of the first reprojected reconstructed reference frame within the bounding box is scaled to a size and a resolution of the reconstructed reference frame to generate the reprojected reconstructed reference frame.
7. The method of claim 1, wherein: Applying the projective transformation includes applying a zoom-down transformation to the reconstructed reference frame to generate a first reprojected reconstructed reference frame having a size smaller than that of the reconstructed reference frame, and the method further includes: Edge pixels adjacent to at least one edge of the first reprojected reconstructed reference frame are generated to provide a reprojected reconstructed reference frame having the same size and resolution as the reconstructed reference frame.
8. The method of claim 1, wherein: Applying the projective transformation includes applying a spherical rotation to the reconstructed reference frame to generate a first reprojected reconstructed reference frame, and the method further includes: Edge pixels adjacent to at least one edge of the first reprojected reconstructed reference frame are generated to provide a reprojected reconstructed reference frame having the same size and resolution as the reconstructed reference frame.
9. The method according to claim 1 or 2, further comprising: The scene pose difference data is predicted by extrapolating second scene pose difference data indicating a change from a third scene pose to a second scene pose of the first scene pose, wherein the first scene pose is subsequent to the third scene pose.
10. The method according to claim 1 or 2, further comprising: Compare at least one scene change difference amplitude value corresponding to the scene posture difference data with a threshold, wherein: Applying the projective transformation to at least a portion of the reconstructed reference frame is conditional on the scene change difference magnitude value meeting or exceeding the threshold.
11. The method according to claim 1 or 2, further comprising: generating a second reconstructed reference frame corresponding to a third scene pose, wherein the third scene pose is before the first scene pose; receiving second scene pose difference data indicating a scene pose change from the third scene pose to the second scene pose; and Based on the second scene pose difference data, a second projection transformation is applied to at least a portion of the second reconstructed reference frame to generate a second reprojected reconstructed reference frame, wherein: Motion compensation is performed on the current frame using both the reprojected reconstructed reference frame and the second reprojected reconstructed reference frame as motion compensation reference frames.
12. A system for video encoding, comprising: a memory for storing a reconstructed reference frame corresponding to a first scene pose, wherein the first scene pose indicates a pose of the scene relative to a first viewpoint, and wherein the first viewpoint corresponds to a first position and orientation of the system relative to the scene in three-dimensional (3D) space; and A processor, coupled to the memory, configured to: generating scene pose difference data indicating a change in scene pose from the first scene pose to a second scene pose subsequent to the first scene pose, wherein the second scene pose is associated with a second viewpoint corresponding to a second position and orientation of the system relative to the scene in 3D space, and wherein the scene pose difference data is generated by tracking movement of the system from the first position and orientation to the second position and orientation; Based on the scene pose difference data, applying a projective transformation to at least a portion of the reconstructed reference frame to generate a reprojected reconstructed reference frame; and The reprojected reconstructed reference frame is used as a motion compensation reference frame to perform motion compensation to generate a current reconstructed frame.
13. The system of claim 12, wherein: The projection transformation includes both affine projection and non-affine projection, the non-affine projection includes at least one of zoom projection, barrel distortion projection and spherical rotation projection, and wherein the scene pose difference data includes one of a transformation matrix, 6-DOF difference data and a motion vector field.
14. The system of claim 12 or 13, wherein: The processor performs motion compensation including: The processor uses both the reconstructed reference frame and the reprojected reconstructed reference frame as motion compensation reference frames, and performs motion compensation in a block-by-block manner, so that the first block of the current reconstructed frame is motion compensated with reference to the reconstructed reference frame, and the second block of the current reconstructed frame is motion compensated with reference to the reprojected reconstructed reference frame.
15. The system of claim 12 or 13, wherein: The processor is further used to: determine a region of interest of the reconstructed reference frame and a background region of the reconstructed reference frame excluding the region of interest, wherein the processor applies the projection transformation including: the processor applies the projection transformation only to one of the region of interest and the background region of the reconstructed reference frame.
16. The system of claim 12 or 13, wherein: The processor is further configured to: The scene pose difference data is predicted based on extrapolating second scene pose difference data indicating a change from a third scene pose to the first scene pose, wherein the first scene pose is subsequent to the third scene pose.
17. The system of claim 12 or 13, wherein: The processor is further used to compare at least one scene change difference amplitude value corresponding to the scene posture difference data with a threshold, wherein the processor applies the projection transformation to at least a portion of the reconstructed reference frame on the condition that the scene change difference amplitude value meets or exceeds the threshold.
18. The system of claim 12 or 13, wherein: The processor is further configured to: generating a second reconstructed reference frame corresponding to a third scene pose, wherein the third scene pose is before the first scene pose; receiving second scene pose difference data indicating a scene pose change from the third scene pose to the second scene pose; and applying a second projection transformation to at least a portion of the second reconstructed reference frame based on the second scene pose difference data to generate a second reprojected reconstructed reference frame, The processor performing motion compensation on the current frame includes: the processor using both the reprojected reconstructed reference frame and the second reprojected reconstructed reference frame as motion compensation reference frames.
19. A machine-readable medium comprising a plurality of instructions which, in response to being executed on a computing device, cause the computing device to perform video encoding by: Generate a reconstructed reference frame corresponding to the first scene pose, where The first scene pose indicates a pose of a scene relative to a first viewpoint, and wherein the first viewpoint corresponds to a first position and orientation of a system for video encoding relative to the scene in three-dimensional (3D) space; generating scene pose difference data indicating a change in scene pose from the first scene pose to a second scene pose subsequent to the first scene pose, wherein the second scene pose is associated with a second viewpoint corresponding to a second position and orientation of the system relative to the scene in 3D space, and wherein the scene pose difference data is generated by tracking movement of the system from the first position and orientation to the second position and orientation; Based on the scene pose difference data, applying a projective transformation to at least a portion of the reconstructed reference frame to generate a reprojected reconstructed reference frame; and The reprojected reconstructed reference frame is used as a motion compensation reference frame to perform motion compensation to generate a current reconstructed frame.
20. The machine-readable medium of claim 19, wherein: The projection transformation includes both affine projection and non-affine projection, the non-affine projection includes at least one of zoom projection, barrel distortion projection and spherical rotation projection, and wherein the scene pose difference data includes one of a transformation matrix, 6-DOF difference data and a motion vector field.
21. The machine-readable medium of claim 19 or 20, wherein: The performing of motion compensation includes: using both the reconstructed reference frame and the reprojected reconstructed reference frame as motion compensation reference frames, performing motion compensation in a block-by-block manner, so that the first block of the current reconstructed frame is motion compensated with reference to the reconstructed reference frame, and the second block of the current reconstructed frame is motion compensated with reference to the reprojected reconstructed reference frame.
22. The machine-readable medium of claim 19 or 20, further comprising a plurality of instructions that, in response to being executed on the computing device, cause the computing device to perform video encoding by: Determine a region of interest of the reconstructed reference frame and a background region of the reconstructed reference frame excluding the region of interest, wherein: Applying the projective transformation includes applying the projective transformation only to one of a region of interest and a background region of the reconstructed reference frame.
23. The machine-readable medium of claim 19 or 20, further comprising a plurality of instructions that, in response to being executed on the computing device, cause the computing device to perform video encoding by: The scene pose difference data is predicted by extrapolating second scene pose difference data indicating a change from a third scene pose to a second scene pose of the first scene pose, wherein The first scene gesture is subsequent to the third scene gesture.
24. The machine-readable medium of claim 19 or 20, further comprising a plurality of instructions that, in response to being executed on the computing device, cause the computing device to perform video encoding by: Compare at least one scene change difference amplitude value corresponding to the scene posture difference data with a threshold, wherein: Applying the projective transformation to at least a portion of the reconstructed reference frame is conditional on the scene change difference magnitude value meeting or exceeding the threshold.
25. The machine-readable medium of claim 19 or 20, further comprising a plurality of instructions that, in response to being executed on the computing device, cause the computing device to perform video encoding by: Generate a second reconstructed reference frame corresponding to the third scene pose, wherein The third scene posture is before the first scene posture; receiving second scene pose difference data indicating a scene pose change from the third scene pose to the second scene pose; as well as Based on the second scene pose difference data, a second projection transformation is applied to at least a portion of the second reconstructed reference frame to generate a second reprojected reconstructed reference frame, wherein: Motion compensation is performed on the current frame using both the reprojected reconstructed reference frame and the second reprojected reconstructed reference frame as motion compensation reference frames.
Citation Information
Patent Citations
Moving image estimating system
US20020114392A1