Systems and Methods for Motion Vectors in Game Generation

By injecting the motion vector data generated by the graphics engine into the video codec engine and skipping the motion estimation step, the problems of long and poor video encoding time in the prior art are solved, and more efficient video encoding and better video quality are achieved.

CN110959169BActive Publication Date: 2025-06-10ZENIMAX MEDIA INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN201880040171.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-12-08
Filing Date
2018-04-20
Publication Date
2025-06-10
Estimated Expiration
2039-05-11

AI Technical Summary

Technical Problem

When existing video encoding methods deal with delay-sensitive game video streams, they have large calculations and long encoding time, resulting in poor video quality and affecting player experience.

Method used

By injecting the motion vector data generated by the graphics engine into the video codec engine, the complex motion estimation steps are skipped, thereby reducing encoding time and calculation amount and improving video quality.

Benefits of technology

Significantly reduces encoding time and computational volume, improves the quality of encoded videos, and preserves the original bitstream data format to maintain interoperability of hardware devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN110959169B_ABST
    Figure CN110959169B_ABST
Patent Text Reader

Abstract

Systems and methods for integrated graphics rendering are disclosed. In some embodiments, the systems and methods utilize a graphics engine, a video encoding engine, and a remote client encoding engine to render graphics over a network. The systems and methods involve generating per-pixel motion vectors that are converted to per-tile motion vectors at the graphics engine. The graphics engine injects these per-tile motion vectors into the video encoding engine such that the video encoding engine can convert these vectors into encoded video data for transmission to the remote client encoding engine.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of priority of U.S. Provisional Application No. 62 / 488,526, filed on April 21, 2017, and U.S. Provisional Application No. 62 / 596,325, filed on December 8, 2017. Background Art

[0003] In remote gaming applications where server - side games are controlled by client players, the application has attempted to encode the video output from a three - dimensional (3D) graphics engine in real time using existing or custom encoders. However, the interactive nature of video games, especially the player feedback loop between video output and player input, makes game video streams more sensitive to latency than traditional video streams. Existing video encoding methods can trade computational power for a reduction in encoding time, with few other alternatives. New methods of integrating the encoding process into the video rendering process can significantly reduce encoding time, while also reducing computational power, improving the quality of the encoded video, and preserving the original bit - stream data format to maintain interoperability with existing hardware devices.

[0004] Existing video encoding standards only contain color and temporal information in an image sequence to improve video encoding time, size, or quality. Some encoding standards, such as those in the MPEG standard family, use computationally intensive block - based motion estimation methods based on the color data contained in the video to approximate object motion. These block - based motion estimation methods have historically significantly reduced the size of the encoded video, but are a significant source of latency in real - time video stream environments.

[0005] Integrating the encoding process into the video rendering process provides access to other data sources that can be used for encoding improvements. For example, some 3D graphics engines, such as those included in game engines, may already have generated motion vectors that perfectly describe the motion of each pixel on each video frame. By providing the final rendered frame and injecting appropriately formatted motion vector data into the encoder, the most computationally complex and time - consuming step in the video encoder - motion estimation - can be skipped for each frame. Additionally, the motion vectors provided by the graphics engine are more accurate than those estimated by block - based motion estimation algorithms, which improves the quality of the encoded video.

[0006] The two fields of video encoding and real - time graphics rendering have traditionally run separately and independently. By integrating the graphics engine and the encoder to leverage their respective advantages, the encoding time can be reduced to a level sufficient to support latency - highly - sensitive streaming applications.

[0007] These and other attendant advantages of the present invention will become apparent from the deficiencies of the techniques described below.

[0008] For example, U.S. Patent Application Publication No. 2015 / 0228106A1 (“the ‘106 Publication”) discloses techniques for decoding video data to generate a sequence of decoded blocks of video images. This technique allows each decoded block of a video image to be used as a separate texture for a corresponding polygon of a geometric surface when a codec engine generates decoded blocks. The ‘106 Publication technique describes the integration between a codec engine and a 3D graphics engine, where the codec engine decodes encoded video data to generate a video image to be mapped, and the 3D graphics engine partially renders a display picture by performing texture mapping of the video image to a geometric surface. However, compared with the present invention, this technique is insufficient, at least because: it does not disclose or use a graphics engine that provides properly formatted motion vector data and a final rendered frame for injection into a video codec engine, such that the video codec engine does not need to perform any motion estimation before transmitting the encoded video data to a remote client encoding engine. In contrast, the improvement of the present invention to computer technology reduces the encoding time and computational amount, improves the quality of the encoded video, and preserves the original bitstream data format to maintain interoperability.

[0009] U.S. Patent Application Publication No. 2011 / 0261885A1 (“the ‘885 Publication”) discloses systems and methods for reducing bandwidth through the integration of motion estimation and macroblock coding. In this system, acquired video data can be used to perform motion estimation to generate motion estimation-related information, including motion vectors. Using the corresponding video data cached in a buffer, these motion vectors can be mapped to a current macroblock. Similarly, compared with the present invention, the ‘885 Publication technique is insufficient, at least because: it does not disclose or use a graphics engine that provides properly formatted motion vector data and a final rendered frame for injection into a video codec engine, such that the video codec engine does not need to perform any motion estimation before transmitting the encoded video data to a remote client encoding engine. Therefore, the technique disclosed in the ‘885 Publication does not provide the reduction in encoding time and computational amount and the improvement in the quality of the encoded video provided by the present invention.

[0010] It is evident from the above discussion of the prior art in this technical field that there is a need in the art to improve the current computer technology related to video encoding in a gaming environment. SUMMARY OF THE INVENTION

[0011] Accordingly, an object of the exemplary embodiments disclosed herein is to address the drawbacks in the art and provide a system and method for graphics generation that uses a networked server architecture running a graphics engine, a video codec engine, and a remote client encoding engine to transmit encoded video data, wherein the graphics engine provides appropriately formatted motion vector data and final rendered frames for injection into the video codec engine.

[0012] Another object of the present invention is to provide a system and method for graphics generation, wherein the video codec engine does not need to perform any motion estimation before transmitting the encoded video data to the remote client encoding engine.

[0013] Another object of the present invention is to provide a system and method for graphics generation, wherein the graphics engine converts per-pixel motion vectors to per-block motion vectors.

[0014] Yet another object of the present invention is to provide a system and method for graphics generation, wherein per-pixel motion vectors are generated by adding per-pixel motion vectors to camera velocity using a compute shader to obtain a per-pixel result, and wherein the per-pixel result is stored in a motion vector buffer.

[0015] Still another object of the present invention is to provide a system and method for graphics generation, wherein per-block motion vector data is injected into the video encoding engine in real time by the graphics engine along with chroma subsampled video frames. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The present invention will be better understood and its numerous attendant advantages will be readily obtained by reference to the following detailed description when considered in conjunction with the accompanying drawings, wherein:

[0017] Figure 1 is a block diagram showing a 3D graphics engine rendering video for encoding and transmission to a client.

[0018] Figure 2 is a flowchart showing the steps required to reduce latency by injecting motion vectors generated by a 3D graphics engine into Figure 4 a modified encoding process;

[0019] Figure 3 is a diagram showing the transformation of per-pixel motion vectors generated in a graphics engine to per-macroblock motion vectors for injection into an encoding engine;

[0020] Figure 4 is a flowchart showing the changes required for the video encoding process used in Figure 1 ; DETAILED DESCRIPTION

[0021] In describing the preferred embodiments of the present invention shown in the accompanying drawings, specific terms will be employed for the sake of clarity. However, the present invention is not intended to be limited to the specific terms so chosen, and it should be understood that each specific term includes all technical equivalents that operate in a similar manner to achieve a similar purpose. For purposes of illustration, several preferred embodiments of the present invention are described, and it is understood that the present invention may be embodied in other forms not specifically shown in the drawings.

[0022] In applications where a 3D graphics engine is rendering video for real-time encoding and transmission, the graphics engine and the encoder can be more tightly coupled to reduce the total computation time and computational overhead. For each video frame, the per-pixel motion vector data that has already been generated by the graphics engine can be converted to per-block motion vector data and injected into the codec engine to bypass the motion estimation step, which is the only most complex and computationally intensive step in the encoding process. In a graphics engine that uses a reconstruction filter in a plausible motion blur method, the per-pixel motion vector may already have been computed for each video frame. The conversion from per-pixel motion vector to per-block motion vector can be performed by finding the average vector for each 16×16 pixel macroblock. This conversion is performed in the 3D graphics engine, so only a small portion of the original motion vector data needs to be passed from the 3D graphics engine to the encoding engine. In cases where the graphics engine and the encoding engine do not share memory, this also helps to reduce memory bandwidth consumption. Injecting the per-block motion vector into the codec engine completely skips the motion estimation step without significantly altering the rest of the encoding process.

[0023] Figures 1 to 4 An example technique for improving video encoding in a video streaming application is shown, where a 3D graphics engine generates accompanying motion vector data during the process of rendering video frames.

[0024] Figure 1An example system is shown in which video is rendered and encoded for transmission to a remote client 116. A 3D graphics engine 100 running in a memory 106 on a server architecture 120 passes video and supplementary motion vector information about the rendered video frames to a codec engine (herein referred to as a codec or encoder) 102, which generates an encoded bitstream 108 for transmission to a client computer system 116. The server architecture 120 is any combination of hardware or software capable of supporting the functions of the graphics engine and the codec engine. In the given example, the graphics engine 100 can be implemented as, for example, a GPU that executes video game software 104 loaded into a computer-readable memory 106, while the codec engine 102 can be implemented as a CPU running video encoding software. The encoding engine 102 generates encoded video data 108 for transmission to a remote client computer system 116, which includes a remote encoding engine (codec) 110 that decodes the bitstream for playback on a display 114 driven by a display controller 112. The remote client computer system 116 is any combination of hardware, device, or software capable of decoding and displaying the encoded bitstream 108.

[0025] Figure 2 Steps required to achieve faster encoding times by reusing existing supplementary data from the rendering process during video encoding are shown. In step 202, as a normal operating function of the graphics engine 100 located at the server 120, supplementary data must first be generated. As GPUs become more powerful and prevalent, real-time per-pixel motion vector generation has become a common feature of modern video game engines. During the rendering of 2D video frames from a 3D scene, the 3D graphics engine generates auxiliary outputs during the color generation process for use as inputs to subsequent post-processing processes. The auxiliary outputs can include information written to three memory locations of an accumulation, color, or velocity buffer, which are respectively allocated for temporarily storing information about pixel depth, pixel color, and pixel motion.

[0026] In a common implementation of motion blur (referred to as a reconstruction filter for plausible motion blur), the per-pixel velocity from the velocity buffer is first downsampled into a smaller number of slices, where each slice is assumed to be the maximum velocity from a group of pixels. Then, the slices are masked using the per-pixel depth in the accumulation buffer, and the result is applied to the per-pixel color in the color buffer to generate motion blur. There are various variants of the reconstruction filter method that improve fidelity, performance, or both, but the concept remains similar, and the velocity buffer contains the per-pixel motion between two adjacent frames. Although "velocity" is the term used in the graphics engine terminology and "motion vector" is the term used in the video coding terminology, these terms are functionally equivalent, and per-pixel velocity and per-pixel motion vector are the same thing. The velocity buffer contains supplementary data in the form of pixel motion vectors, which will be reused during the video coding process.

[0027] In step 204, the graphics engine 100 located at the server 120 converts the per-pixel motion vectors into per-block motion vectors based on the macroblock size to be used for encoding. The H.264 codec defaults to using 16x16 pixel macroblocks and can optionally be further subdivided. It is possible to average together 256 per-pixel motion vectors to provide a single average vector as the per-block motion vector. Combining Figure 3 This process is described in further detail.

[0028] In step 206, the per-macroblock motion vector information is injected into the encoding engine / encoder 102 located at the server 120, thus bypassing the motion estimation step. In a software implementation of the encoder, it is possible to completely disable the motion estimation step, thus significantly saving CPU computation time. The time saved in the CPU should be sufficient to offset the additional time required to calculate the average vector (in step 204) in the GPU and transfer it to the CPU.

[0029] In step 208, since the per-block motion vectors provided by the graphics engine 100 are interchangeable with the motion vectors calculated in a typical motion estimation step, the encoding starts with the motion compensation step (step 208). As described in further detail in Figure 4 The remainder of the video coding process is not significantly different from the typical motion compensation, residual calculation, and encoding steps performed by coding standards that use motion estimation techniques.

[0030] Figure 3Further details are shown of the transformation from per-pixel motion vectors to per-macroblock motion vectors that occurs in the graphics engine 100. During the color generation phase, the 3D graphics engine 100 located at the server 120 generates per-pixel motion vectors and stores this data in the velocity buffer 300 also located at the server 120. The velocity buffer 300 can contain only data for dynamic objects and does not include the motion information imparted by the player camera motion. To obtain the motion vector information for each pixel in the image space, the compute shader 302 combines the vectors in the velocity buffer 300 with the camera velocity of all static objects that are not yet included in the velocity buffer and stores the per-pixel result in the motion vector buffer 304. The camera velocity is the 2D projection of the rotational and translational camera motion during a frame. A particular graphics engine may use a slightly different method to calculate these per-pixel motion vectors for the entire screen space, but the concept remains the same.

[0031] The default macroblock size used by the H.264 encoder is 16×16, but it is capable of being subdivided into smaller sizes down to a minimum of 4×4. In Figure 3 the example, a 4×4 macroblock 306 is used as a simplified case, but the method should be extrapolated to match the macroblock sizes used in the encoder. For the 4×4 macroblock 306, 16 per-pixel motion vectors 308 are stored in the motion vector buffer 304. These per-pixel motion vectors 308 need to be transformed 312 into a single per-macroblock motion vector 310 that can be injected into the encoder for motion compensation as shown in Figure 4 . An arithmetic mean of a set of per-pixel vectors 308 is a transformation 312 method with low computational complexity and short computation time.

[0032] Optional modifications can be made to the arithmetic mean transformation 312 to improve quality at the cost of additional computational complexity or amount of computation. For example, prior to the arithmetic mean calculation, a vector median filtering technique can be used to remove discontinuities in the macroblock vector field to ensure that the per-macroblock motion vector 310 represents the majority of the pixels in the macroblock 306. Since the final per-macroblock motion vector is derived from the pixel-perfect motion vectors initially calculated based on known object motion data, the per-macroblock motion vector is always a more accurate representation than the motion vectors calculated by existing block-based motion estimation algorithms, where these motion estimation algorithms can only derive motion based on pixel color data.

[0033] Figure 4 shows the injection of the motion vectors generated in the graphics engine 100 of the server 120 in Figure 1 into Figure 1A method for skipping the computationally complex motion estimation process in the encoding engine 102 of server 120. As described in detail below, the resulting bitstream of the encoded video data 108 is transmitted to the remote client computer system 116. Figure 4 The method shown illustrates the inter-frame encoding process for a single frame, specifically the encoding process for P-frames defined by the MPEG video codec standard family. Since motion compensation 406 is not performed in I-frame generation, intra-frame (I-frame) generation is not changed. Once the chroma subsampled video frame 402 and per-block motion vector data 404 are available, they are transmitted from the graphics engine 100. The game-generated motion vectors 404 are used to circumvent the motion vector generation that would otherwise occur in the typical motion estimation 426 step, as described in the H.264 / MPEG-4 AVC standard. The motion estimation 426 step will be skipped and can be disabled in the software implementation of the encoding engine. Skipping the block-based motion estimation 426 step will significantly reduce the encoding time, which will greatly offset the time for converting the velocity buffer data to the appropriate format as described in Figure 3 connection.

[0034] The motion vectors 404 that have been converted to the appropriate macroblock size can be used directly without any changes to the motion compensation 406. The result of the motion compensation 406 is combined with the input chroma subsampled video frame 402 to form a residual image 430, which is typically processed by the residual transform & scaling 408, quantization 410, and scan 412 steps, which typically occur within existing hardware or software video encoders.

[0035] If the decoding standard selected by the implementation requires a deblocking step, the deblocking step must be performed. The deblocking settings 420 and the deblocked image 428 are calculated by applying the inverse quantization 414, inverse transform & scaling 416, and then deblocking 418 algorithms of the encoding standard. The scanned coefficients 412 are combined with the deblocking settings 420 and encoded in the entropy encoder 422, and then transmitted as the bitstream 108 to the remote client computer system 116 for decoding at the codec 110 of the remote client computer system. The deblocked image 428 becomes the input for the motion compensation 406 of the next frame. The bitstream (including the encoded video data) 108 retains the same format as defined by the encoding standard used in the implementation, such as H.264 / MPEG-4 AVC. This example is specific to the H.264 / MPEG-4 AVC standard and can generally be used for similar encoding standards that use the motion estimation 426 and motion compensation 406 techniques.

[0036] Example 1: Benchmark tests show a reduction in encoding time

[0037] The motion estimation step in traditional H.264 compatible encoding is usually the most computationally complex and time-consuming step. As discussed in this paper, reusing the motion vectors generated by the game can significantly reduce the encoding time.

[0038] In the test environment, the graphics engine produces output at a resolution of 1280×720 at a rate of 60 frames per second. The encoding time was captured from running the x264 encoder in a single-threaded manner. Running the single-threaded encoder will produce a longer encoding time than actual use, but the measurements are normalized to one core so that they can be directly compared to each other. The encoding time was first measured using the unmodified motion estimation within the encoder and then re-measured in the same environment with the game-generated motion estimation feature enabled.

[0039] A low-motion area was selected, including the player's hand, weapon, and the first-person view of a fixed wall. The player's hand and weapon were animated with a slight "swing", generating a small amount of pixel motion in a relatively small screen space. The results of this test are shown in Table 1 below, which shows the latency results when using and not using the game-generated motion estimation technique described in this paper. At low intensity, where the game-generated motion estimation was disabled, the unmodified encoding time was 12 ms. When the game-generated motion estimation was enabled, the encoding time was reduced by 3 ms to 9 ms. Similar latency reductions were shown for the average motion intensity scenario and the high motion intensity scenario, with a 17.6% reduction for the average motion intensity scenario and a 15% to 30% reduction in the high latency scenario. These results indicate that the latency is significantly reduced when the game-generated motion estimation is enabled.

[0040]

[0041] Table 1: Latency Results at Different Motion Intensities

[0042] The test environment also showed that there is additional overhead when converting the game-generated per-pixel motion vectors to per-macroblock motion vectors for the encoder. However, this overhead is significantly lower than the encoding time reduction described in the previous section. Using the graphics engine to generate video at a resolution of 1280×720, the conversion of per-pixel to per-macroblock motion vectors takes 0.02 ms. The measured encoder time savings are three orders of magnitude higher than the increased time cost of encoding using the game-generated motion vectors.

[0043] The foregoing description and drawings are to be understood as illustrative only of the principles of the invention. The invention is not limited to the preferred embodiment and can be implemented in various ways that will be apparent to those of ordinary skill in the art. Many applications of the invention will readily occur to those skilled in the art. Accordingly, it is not desired to limit the invention to the specific examples disclosed or to the exact architecture and operation shown and described. All suitable modifications and equivalents may be employed within the scope of the invention.

Claims

1. A computer-implemented method for generating graphics, comprising the steps of: generating one or more per-pixel motion vectors of a dynamic object and storing the one or more per-pixel motion vectors in a velocity buffer, wherein the velocity buffer contains only data for the dynamic object and does not include motion information imparted by camera movement; using a compute shader to combine the one or more per-pixel motion vectors in the velocity buffer with the camera velocity of all static objects not yet included in the velocity buffer to obtain a per-pixel result, wherein the camera velocity is a 2D projection of rotational and translational camera movement during a frame; storing the per-pixel result in a motion vector buffer; converting the one or more per-pixel motion vectors to one or more per-block motion vectors; injecting the per-block motion vectors into a video encoding engine; the video encoding engine converting the one or more per-block motion vectors to encoded video data; and transmitting the encoded video data to a remote client computer system.

2. The method according to claim 1, wherein, a graphics engine injects the per-block motion vector data and one or more chroma subsampled video frames into the video encoding engine in real time.

3. The method according to claim 1, wherein, the encoded video data is decoded for playback on the remote client computer system.

4. The method according to claim 1, wherein, the video encoding engine performs motion compensation and residual transformation to convert the one or more per-block motion vectors to encoded video data.

5. The method according to claim 1, wherein, the encoded video data is prepared for transmission to the remote client encoding engine by applying one or more inverse quantization algorithms, inverse transformation and scaling and / or deblocking.

6. The method according to claim 1, wherein, a transform method applying arithmetic mean is used to convert the one or more per-pixel vectors to one or more per-block motion vectors.

7. A computer-implemented graphics generation system, comprising: a graphics engine configured to: generate one or more per-pixel motion vectors of a dynamic object and store the one or more per-pixel motion vectors in a velocity buffer, wherein the velocity buffer contains only data for the dynamic object and does not include motion information imparted by camera movement; using a compute shader to combine the one or more per-pixel motion vectors in the velocity buffer with the camera velocity of all static objects not yet included in the velocity buffer to obtain a per-pixel result, wherein the camera velocity is a 2D projection of rotational and translational camera movement during a frame; storing the per-pixel result in a motion vector buffer; converting the per-pixel motion vectors to one or more per-block motion vectors; and directly injecting the per-block motion vectors into a video codec engine; and The video codec engine, which is configured to: convert each block motion vector into encoded video data and transmit the encoded video data to a remote client computer system.

8. The system according to claim 7, wherein, the graphics engine simultaneously injects each block motion vector data and one or more chroma subsampled video frames into the video codec engine in real time.

9. The system according to claim 7, wherein, the video codec engine performs motion compensation and residual transformation to convert the one or more per-block motion vectors into encoded video data.

10. The system according to claim 7, wherein, the encoded video data is prepared for transmission to the remote client computer system by applying one or more inverse quantization algorithms, inverse transformation and scaling and / or deblocking.

11. The system according to claim 7, wherein, the encoded video data is configured to be decoded and played on a display driven by a display controller.

12. The system according to claim 7, wherein, the graphics engine converts the one or more per-pixel vectors into one or more per-block motion vectors using a transformation method that applies arithmetic mean.

Citation Information

Patent Citations

  • Method and system for bandwidth reduction through integration of motion estimation and macroblock encoding

    US20110261885A1

  • Low latency video texture mapping via tight integration of codec engine with 3D graphics engine

    US20150228106A1

  • Moving image coding device, moving image decoding device, moving image coding / decoding system, moving image coding method and moving image decoding method

    US20120213278A1

  • Motion based adaptive rendering

    US20150379727A1

  • Layer-based video decoding

    US20160150231A1