Image rendering method, apparatus, and medium

By acquiring extrapolated frame data and residual data, and using a differential reference coding structure to correct image frames, the visual artifact problem of decoding devices at object edges and occluded areas is solved, achieving low-latency and high-efficiency video stream rendering.

CN121603657BActive Publication Date: 2026-05-01VASTAI TECH (SHANGHAI) INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
VASTAI TECH (SHANGHAI) INC
Filing Date
2026-01-31
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In existing technologies, decoding devices lack depth information and motion vectors, resulting in visual artifacts such as ghosting or holes at object edges and occluded areas. Furthermore, the non-causal generation mechanism of traditional decoding devices leads to increased transmission latency, failing to meet low-latency requirements.

Method used

An image rendering method is adopted, which constructs image frames by acquiring extrapolated frame data and residual data, and uses a differential reference coding structure to correct the extrapolated frame data, thereby improving the construction efficiency and generation quality of image frames.

Benefits of technology

It improves the efficiency and quality of image frame construction, reduces network transmission latency, and ensures the smoothness and image quality of video streams, especially maintaining image continuity in high-frequency interactive scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121603657B_ABST
    Figure CN121603657B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to image rendering methods, apparatuses and media. The method proposed herein includes: obtaining a code stream from an encoding device; at a first time, obtaining extrapolation frame data corresponding to a target frame sequence number by decoding the code stream, the extrapolation frame data being generated by performing an extrapolation frame operation on a first reference frame; in response to obtaining first residual data corresponding to the target frame sequence number from the code stream before a second time, constructing an image frame corresponding to the target frame sequence number based on the extrapolation frame data and the first residual data, the second time being determined based on a display time of the target frame sequence number, the first residual data being generated by the encoding device based on a difference between interpolation frame data and the extrapolation frame data, the interpolation frame data being generated by performing an interpolation frame operation on the first reference frame and a second reference frame; and displaying the image frame at a display time corresponding to the target frame sequence number.
Need to check novelty before this filing date? Find Prior Art

Description

Image rendering methods, apparatus and media Technical Field

[0001] The exemplary embodiments disclosed herein relate generally to the field of computers, and more particularly to image rendering methods, apparatus and media. Background Technology

[0002] In the field of video processing, video smoothness is a crucial factor in ensuring user experience and platform competitiveness. In some scenarios, video encoding and decoding technologies are used to guarantee video smoothness. Therefore, the efficiency of image frame construction has become a key focus. Summary of the Invention

[0003] In a first aspect of this disclosure, an image rendering method is provided. The method includes: acquiring a bitstream from an encoding device; at a first moment, acquiring extrapolated frame data corresponding to a target frame sequence number by decoding the bitstream, the extrapolated frame data being generated by performing an extrapolation operation on a first reference frame; in response to obtaining first residual data corresponding to the target frame sequence number from the bitstream before a second moment, constructing an image frame corresponding to the target frame sequence number based on the extrapolated frame data and the first residual data, the second moment being determined based on a display moment of the target frame sequence number, the first residual data being generated by the encoding device based on the difference between interpolated frame data and extrapolated frame data corresponding to the target frame sequence number, the interpolated frame data being generated by performing an interpolation operation on a first reference frame and a second reference frame; and displaying the image frame at the display moment corresponding to the target frame sequence number.

[0004] In a second aspect, an apparatus for image rendering is proposed. The apparatus includes a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform the method according to the first aspect of this disclosure.

[0005] In a third aspect, a non-transitory computer-readable storage medium is proposed. This non-transitory computer-readable storage medium stores instructions that cause a processor to perform the method according to the first aspect of this disclosure.

[0006] In this way, embodiments of the present disclosure construct a special differential reference coding structure. In response to the time when the target frame number is displayed, the difference between the interpolated frame data and the extrapolated frame data corresponding to the target frame number (e.g., the first residual data) is obtained. The extrapolated frame data can be corrected based on the difference to construct the corresponding image frame, thereby improving the construction efficiency and generation quality of the image frame.

[0007] The present invention is provided to present, in a simplified form, the selection of concepts further described below in the detailed description. The present invention is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description

[0008] The above and other objects, features, and advantages of exemplary embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.

[0009] Figure 1 shows a block diagram illustrating an example video codec system according to some embodiments of the present disclosure.

[0010] Figure 2 shows a block diagram of an example video encoder according to some embodiments of the present disclosure.

[0011] Figure 3 shows a block diagram of an example video decoder according to some embodiments of the present disclosure.

[0012] Figure 4 shows a flowchart of an example image rendering process according to some embodiments of the present disclosure.

[0013] Figure 5 shows a block diagram of a computing device in which various embodiments of the present disclosure may be implemented. Detailed Implementation

[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0015] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.

[0016] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0017] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.

[0018] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information beyond what is necessary for basic functions will not affect the user's use of basic functions.

[0019] Traditional decoding methods, lacking depth information and motion vectors, are prone to producing visual artifacts such as ghosting or holes at object edges and in occluded areas when extrapolating predictions based solely on two-dimensional images. Furthermore, the non-causal generation mechanism of traditional decoding devices (requiring future frames as references) can lead to increased transmission latency, failing to meet the low-latency requirements of various scenarios.

[0020] This disclosure provides an image rendering scheme. The scheme includes: acquiring a bitstream from an encoding device; at a first moment, acquiring extrapolated frame data corresponding to a target frame number by decoding the bitstream, the extrapolated frame data being generated by performing an extrapolation operation on a first reference frame; in response to obtaining first residual data corresponding to the target frame number from the bitstream before a second moment, constructing an image frame corresponding to the target frame number based on the extrapolated frame data and the first residual data, the second moment being determined based on the display time of the target frame number, the first residual data being generated by the encoding device based on the difference between interpolated frame data and extrapolated frame data corresponding to the target frame number, the interpolated frame data being generated by performing an interpolation operation on a first reference frame and a second reference frame; and displaying the image frame at the display time corresponding to the target frame number.

[0021] The embodiments of this disclosure construct a special differential reference coding structure. In response to the time when the target frame number is displayed, the difference between the interpolated frame data and the extrapolated frame data corresponding to the target frame number (e.g., the first residual data) is obtained. The extrapolated frame data can be corrected based on the difference to construct the corresponding image frame, thereby improving the construction efficiency and generation quality of the image frame.

[0022] The following section provides a detailed description of various example implementations of this scheme, with reference to the accompanying drawings.

[0023] Example environment:

[0024] Figure 1 is a block diagram illustrating an example video encoding / decoding system 100 from which the techniques of this disclosure may be utilized. As shown, the video encoding / decoding system 100 may include encoding devices (e.g., source device 110) and decoding devices (e.g., destination device 120). The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and a first I / O interface 116.

[0025] Video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, interfaces for receiving video data from video content providers, computer graphics systems for generating video data, and / or combinations thereof.

[0026] Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded representation of the video data. The bitstream may include encoded images and associated data. An encoded image is an encoded representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. First I / O interface 116 may include a modulator / demodulator and / or a transmitter. Encoded video data can be directly transmitted to destination device 120 via network 130A through first I / O interface 116. Encoded video data may also be stored on storage medium / server 130B for access by destination device 120.

[0027] The destination device 120 may include a second I / O interface 126, a video decoder 124, and a display device 122. The second I / O interface 126 may include a receiver and / or a modem. The second I / O interface 126 may acquire encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or it may be external to the destination device 120, which is configured to interface with an external display device.

[0028] The video encoder 114 and the video decoder 124 can operate according to video compression standards such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other existing and / or future standards.

[0029] Figure 2 is an example block diagram illustrating a video encoder 114 according to some embodiments of the present disclosure. The video encoder 114 can be configured to implement any or all of the technologies disclosed herein.

[0030] In the example of Figure 2, the video encoder 114 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 114. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure. As an example, the processor can include, but is not limited to, an implementable architecture capable of running on heterogeneous platforms such as CPUs / GPUs / NPUs / AI accelerators / hardware decoders.

[0031] In some embodiments, the video encoder 114 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, a first inverse quantization unit 210, a first inverse transform unit 211, a first reconstruction unit 212, a first buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a first motion compensation unit 205, and a first intra-frame prediction unit 206.

[0032] In other examples, the video encoder 114 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, in which at least one reference picture is the picture in which the current video block is located.

[0033] Furthermore, although some components (such as motion estimation unit 204 and first motion compensation unit 205) can be integrated, for illustrative purposes, these components are shown separately in the example of Figure 2.

[0034] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 114 and the video decoder 124 can support various video block sizes.

[0035] The mode selection unit 203 can, for example, select one of several coding modes (intra-coding or inter-coding) based on the error result, and provide the resulting intra-coded or inter-coded block to the residual generation unit 207 to generate residual block data, and provide it to the first reconstruction unit 212 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 203 can select an intra-inter-prediction joint prediction (CIIP) mode, in which prediction is based on inter-prediction signals and intra-prediction signals. In the case of inter-prediction, the mode selection unit 203 can also select a resolution for the block based on the motion vector (e.g., sub-pixel precision or integer pixel precision).

[0036] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from the first buffer 213 with the current video block. First motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from the first buffer 213 other than the image associated with the current video block.

[0037] The motion estimation unit 204 and the first motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-strip, P-strip, or B-strip. As used herein, an "I-strip" can refer to a portion of an image composed of macroblocks, all of which are based on macroblocks within the same image. Furthermore, as used herein, in some aspects, "P-strip" and "B-strip" can refer to portions of an image composed of macroblocks independent of the same image.

[0038] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search reference images in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. First motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0039] Alternatively, in other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search reference images in list 0 to find a reference video block for the current video block, and can also search reference images in list 1 to find another reference video block for the current video block. Motion estimation unit 204 can then generate multiple reference indices and multiple motion vectors, the multiple reference indices indicating multiple reference images containing multiple reference video blocks in lists 0 and 1, and the multiple motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. Motion estimation unit 204 can output the multiple reference indices and multiple motion vectors of the current video block as motion information for the current video block. First motion compensation unit 205 can generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.

[0040] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process. Alternatively, in some embodiments, the motion estimation unit 204 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0041] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 124 that the current video block has the same motion information as another video block.

[0042] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 124 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0043] As discussed above, the video encoder 114 can transmit motion vectors via signaling in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 114 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.

[0044] The first intra-prediction unit 206 can perform intra-prediction on the current video block. When the first intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.

[0045] The residual generation unit 207 can generate residual data for the current video block by subtracting (or more) predicted video blocks from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components in the current video block.

[0046] In other examples, such as in skip mode, residual data for the current video block may not exist, and residual generation unit 207 may not perform a subtraction operation.

[0047] Transform unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0048] After the transform unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0049] The first inverse quantization unit 210 and the first inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The first reconstruction unit 212 can add the reconstructed residual video block to the corresponding sample points of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the first buffer 213.

[0050] After the video block is reconstructed in the first reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.

[0051] Entropy encoding unit 214 can receive data from other functional components of video encoder 114. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.

[0052] Figure 3 is a block diagram illustrating an example of a video decoder 124 according to some embodiments of the present disclosure. The video decoder 124 may be configured to perform any or all of the techniques of the present disclosure.

[0053] In the example of Figure 3, the video decoder 124 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 124. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0054] In the example of Figure 3, the video decoder 124 includes an entropy decoding unit 301, a second motion compensation unit 302, a second intra-frame prediction unit 303, a second inverse quantization unit 304, a second inverse transform unit 305, a second reconstruction unit 306, and a second buffer 307. In some examples, the video decoder 124 may perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 114.

[0055] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy-encoded video data, and second motion compensation unit 302 can determine motion information from the entropy-decoded video data, including motion vectors, motion vector precision, reference picture list indices, and other motion information. Second motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, which involves deriving several most likely candidates based on data from neighboring PBs and reference pictures. Motion information typically includes horizontal motion vector displacement values ​​and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B-strip, an identifier of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially or temporally neighboring blocks.

[0056] The second motion compensation unit 302 can generate motion compensation blocks, possibly performing interpolation based on an interpolation filter. Identifiers for interpolation filters used with sub-pixel precision can be included in the syntax elements.

[0057] The second motion compensation unit 302 can use the interpolation filter used by the video encoder 114 during the encoding of the video block to calculate the interpolated values ​​of the sub-integer pixels for the reference block. The second motion compensation unit 302 can determine the interpolation filter used by the video encoder 114 based on the received syntax information, and the second motion compensation unit 302 can use the interpolation filter to generate the prediction block.

[0058] The second motion compensation unit 302 may use at least some of the syntax information to determine the size of the blocks used to encode the encoded video sequence (multiple frames) and / or (multiple stripes), segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a “strip” can refer to a data structure that can be decoded independently of other stripes of the same image in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A strip can be an entire image or a region of an image.

[0059] The second intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. The second dequantization unit 304 dequantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The second inverse transform unit 305 applies an inverse transform.

[0060] The second reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the second motion compensation unit 302 or the second intra-frame prediction unit 303. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a second buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.

[0061] As an example, such a display device can be any terminal device with display capabilities, such as AR (Augmented Reality) devices, VR (Virtual Reality) devices, MR (Mixed Reality) devices, and other XR devices.

[0062] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.

[0063] Example process:

[0064] Figure 4 shows a flowchart of an example image rendering process 400 according to some embodiments of the present disclosure. Process 400 can be implemented at the destination device 120. Process 400 will now be described with reference to Figure 1.

[0065] As shown in Figure 4, in step 410, the destination device 120 obtains the bitstream from the encoding device.

[0066] In some embodiments, the application scenarios of this stream may include, but are not limited to: end-to-end streaming media, video conferencing, monitoring, virtual reality (VR), augmented reality (AR), cloud gaming scenarios, etc., which will not be elaborated here.

[0067] In step 420, the target device 120 obtains the extrapolated frame data corresponding to the target frame sequence number by decoding the bitstream at the first moment.

[0068] As an example, after acquiring the bitstream from the encoding device (e.g., at the first moment), the destination device 120 can decode the acquired bitstream to obtain interpolated frame data corresponding to the target frame sequence number. As an example, the destination device 120 can decode the acquired bitstream based on any suitable algorithm or standard, and there is no limitation here.

[0069] In some embodiments, such extrapolated frame data may be generated by performing an extrapolation operation on a first reference frame. As an example, such extrapolated frame data may be generated by performing an extrapolation operation on a first reference frame that has already been rendered. This first reference frame may, for example, be represented as rendering frame i-2, corresponding to frame number i-2.

[0070] As an example, an encoding device (e.g., source device 110) can determine the state of the next image frame of a first reference frame (e.g., rendered frame i-2) based on its geometric and motion information to generate corresponding extrapolated frame data. In some embodiments, such motion information may include, but is not limited to, optical flow information and motion trends of the first reference frame. This extrapolated frame data may, for example, be represented as extrapolated frame i-1, corresponding to frame number i-1.

[0071] In some embodiments, in response to acquiring the extrapolated frame data, the destination device 120 may store the extrapolated frame data in a buffer. As an example, such a buffer may be a memory region temporarily storing the extrapolated frame data, and such a memory region may be presented in the form of a circular queue or a priority queue. It is understood that the above is merely illustrative, and the memory region may be presented in any suitable form.

[0072] In some embodiments, the encoding device may also generate interpolated frame data corresponding to the target frame sequence number. In some embodiments, the interpolated frame data may be generated by performing an interpolation operation on a first reference frame and a second reference frame.

[0073] As an example, the encoding device may generate interpolated frame data by performing an interpolation frame operation on a first reference frame (e.g., rendering frame i-2, with frame number i-2) and a second reference frame (e.g., the next rendering frame i after rendering frame i-2, with frame number i). Such interpolated frame data may be represented, for example, as interpolated frame i-1, with frame number i-1.

[0074] In some embodiments, the generation time of the extrapolated frame data corresponding to the target frame sequence number may be earlier than the generation time of the interpolated frame data corresponding to the target frame sequence number. As an example, the generation time of the extrapolated frame i-1 may be earlier than the generation time of the interpolated frame i-1. For example, the interpolated frame i-1 may be generated before the display time of the extrapolated frame i-1.

[0075] In some embodiments, after generating the interpolated frame data corresponding to the target frame number, the encoding device can generate first residual data based on the difference between the interpolated frame data and the extrapolated frame data corresponding to the target frame number. For example, the encoding device can determine the difference between the interpolated frame i-1 and the extrapolated frame i-1 corresponding to the target frame number i-1, thereby generating the corresponding first residual data.

[0076] In some embodiments, the difference between the interpolated frame data and the extrapolated frame data can be, for example, the pixel difference between the two. Further, the encoding device can generate first residual data based on this pixel difference.

[0077] In some embodiments, the encoding device can encode the first residual data generated based on the interpolated frame data and the extrapolated frame data corresponding to the target frame number into the bitstream. Further, the target device 120 can determine the first residual data by decoding the bitstream.

[0078] In some embodiments, in response to acquiring interpolated frame data or storing interpolated frame data in a buffer, the target device 120 may initiate a waiting window to determine whether the first residual data corresponding to the target frame number can be obtained within a preset time period.

[0079] In step 430, the target device 120 decodes the first residual data corresponding to the target frame number from the bitstream before the second time, and constructs an image frame corresponding to the target frame number based on the extrapolated frame data and the first residual data.

[0080] In some embodiments, the second moment here can be determined based on the display moment of the target frame number (e.g., frame number i-1). Such a display moment can be understood, for example, as the moment when the image frame corresponding to the target frame number is presented. As an example, assuming the display moment of the target frame number is j, then the second moment here can also be represented as j, that is, the moment when the image frame corresponding to the target frame number is presented can be j.

[0081] As an example, in some scenarios (e.g., low latency, good network), the target device 120 can construct an image frame corresponding to the target frame number based on the extrapolated frame data and the first residual data after decoding the bitstream before the second time step, in response to the first residual data.

[0082] As an example, after the first residual data is determined, the destination device 120 can construct an image frame corresponding to the target frame number based on the first residual data and the extrapolated frame data corresponding to the target frame number. For example, after obtaining the first residual data, the destination device 120 can determine the residual correction amount based on this first residual data. Further, the destination device 120 can perform pixel-by-pixel addition and fusion of the extrapolated frame data (e.g., extrapolated frame i-1) and the residual correction amount to construct an image frame corresponding to the target frame number.

[0083] In this way, the embodiments of this disclosure can greatly optimize the utilization of network transmission bandwidth while ensuring image quality restoration. The embodiments of this disclosure set the encoding reference frame of the interpolated frame to the interpolated frame at the same display time. Since the two are highly similar in image content, the encoding device only needs to encode the minor differences between them (e.g., the first residual data). Based on this, the data packet size used for image quality restoration can be reduced, and the image quality can be restored from the lossy predicted state to the true state generated by the rendering engine without significantly increasing the network load, achieving a balance between image quality and bandwidth.

[0084] In step 440, the target device 120 displays the image frame at the display time corresponding to the target frame sequence number.

[0085] Taking the target frame number (for example, frame number i-1) as an example, if the display time is 00:30, the target device 120 can display the image frame generated above at 00:30.

[0086] In some embodiments, in response to the failure to decode the first residual data corresponding to the target frame number from the bitstream before the second time, the target device 120 can obtain interpolated frame data from the buffer. Furthermore, the target device 120 can display the interpolated image frame corresponding to the interpolated frame data.

[0087] As an example, in some scenarios (e.g., high latency or network jitter), the target device 120, in response to not decoding the first residual data corresponding to the target frame number from the bitstream before the display time of the target frame number (e.g., frame number i-1), can obtain the extrapolated frame data corresponding to the target frame number from the buffer and can display the extrapolated image frame corresponding to the extrapolated frame data.

[0088] In this way, the embodiments of this disclosure can effectively build a forward-looking buffer on the client side by generating and sending future interpolated frames in advance on the server side. When network jitter causes the data packets of the actual rendered frame to fail to arrive on time, the client can immediately and seamlessly switch to the pre-received interpolated frame for display.

[0089] Based on this, the mechanism effectively avoids the physical delay caused by network transmission, ensuring that the screen remains smooth and continuous when users are engaging in high-frequency interactions (such as fast movement or shooting in competitive games), eliminating the screen freezing or input lag caused by waiting for data in traditional cloud gaming.

[0090] In traditional low-latency transmission schemes, if interpolated frames with poor prediction quality are used directly as references for subsequent frames, the prediction error will accumulate in subsequent P-frames, eventually causing blurry images or artifact divergence.

[0091] To address the error propagation problem in video encoding and improve video stream quality, in some embodiments, in response to obtaining first residual data corresponding to the target frame number (e.g., frame number i-1) from the bitstream after a second time step, the target device 120 can generate reference data corresponding to the target frame number based on the interpolated frame data (e.g., interpolated frame i-1) and the first residual data. Furthermore, the target device 120 can store the reference data in a buffer.

[0092] In some embodiments, after displaying the interpolated image frame corresponding to the interpolated frame data, the target device 120, in response to decoding the first residual data corresponding to the target frame number from the bitstream, can generate reference data corresponding to the target frame number based on the interpolated frame data and the first residual data, and store it in a buffer.

[0093] As an example, after obtaining the first residual data, the destination device 120 can determine the residual correction amount based on this first residual data. Furthermore, the destination device 120 can perform pixel-by-pixel addition and fusion of the extrapolated frame data and the residual correction amount to generate reference data corresponding to the target frame number, and store it in a buffer.

[0094] In this way, the embodiments of this disclosure can satisfy the decoding dependencies of subsequent frames, thereby repairing the reference chain of subsequent frames while maintaining smooth playback.

[0095] In some embodiments, the target device 120 may use reference data in the buffer to construct at least one subsequent image frame. As an example, such at least one subsequent image frame may include a rendered frame, an interpolated frame, or an extrapolated frame corresponding to the subsequent frame number of the target frame number, etc.

[0096] Specifically, the target device 120 can obtain second residual data corresponding to the subsequent frame number of the target frame number (e.g., interpolated frame i-1) by decoding the bitstream. Further, the target device 120 can construct the subsequent image frame corresponding to the subsequent frame number based on the reference data and the second residual data in the buffer. In some embodiments, the second residual data is generated by the encoding device based on the difference between the rendered frame data and the interpolated frame data corresponding to the subsequent frame number.

[0097] In this way, embodiments of this disclosure can retroactively correct the reference frame buffer of the decoding device using high-quality interpolated frames. Even if an interpolated image frame corresponding to the interpolated frame data is displayed, subsequent arriving correction data will silently update the reference frame in the background. This mechanism cuts off the error propagation link, ensuring that the decoding of subsequent video frames can always be based on accurate image data, maintaining the long-term stability of the video stream.

[0098] In some embodiments, this bitstream can be encoded frame by frame:

[0099] For interpolated frame data, the encoding device may use a first reference frame as a reference frame to generate the interpolated frame data. As an example, the encoding device may use a first reference frame (e.g., render frame i-2) as a reference frame to generate a subsequent interpolated frame or interpolated frame data (e.g., interpolated frame i-1) of the first reference frame.

[0100] For interpolated frame data, the encoding device can use the extrapolated frame data as a reference frame. The encoding device can determine the difference information between the interpolated frame data and this extrapolated frame data, and encode the difference information to obtain the first residual data. As an example, the encoding device can use the extrapolated frame i-1 as a reference frame. The encoding device can determine the difference information between the interpolated frame i-1 and the extrapolated frame i-1, and encode the difference information. Here, the difference information can be represented as pixel difference.

[0101] For the second reference frame, the encoding device can encode with reference to interpolated frame data (e.g., interpolated frame i-1, corresponding to frame number i-1). As an example, the encoding device can encode with reference to interpolated frame data (e.g., interpolated frame i-1, corresponding to frame number i-1) to obtain the second reference frame (e.g., rendered frame i, corresponding to frame number i).

[0102] For future backup purposes, in some embodiments, the encoding device may also obtain the next extrapolated frame corresponding to the second reference frame. For example, the encoding device may encode with the second reference frame (e.g., rendering frame i, with corresponding frame number i) as a reference to obtain the extrapolated frame corresponding to the second reference frame (e.g., extrapolated frame i+1, with corresponding frame number i+1).

[0103] Furthermore, the encoding device can obtain a bitstream that includes the first residual data, the second reference frame, and the next extrapolated frame of the second reference frame after encoding.

[0104] In this way, the embodiments of this disclosure can improve the efficiency of video encoding under the same network bandwidth, and can transmit higher resolution or lower distortion images.

[0105] In other embodiments, the bitstream may also be encoded as a group of images:

[0106] In response to completing rendering for the first image frame, the encoding device (e.g., source device 110) can generate data units, such as groups of pictures (GOPs), based on the first image frame, the second image frame, and the third image frame. In some embodiments, the encoding device can be any suitable device compatible with a predetermined encoding standard, which can be configured to encode image frames or image blocks, and will not be elaborated here.

[0107] In some embodiments, the first image frame here can be a rendered frame. As an example, this first image frame can be an image frame that has been rendered (e.g., rendered frame i, corresponding to frame number i).

[0108] In some embodiments, the second image frame may be a previous interpolated frame generated based on the first image frame. As an example, this second image frame may be an interpolated frame generated based on the first image frame and at least one historical rendering frame.

[0109] Taking a historical rendering frame as an example, the encoding device can perform an interpolation operation based on a first image frame (e.g., rendering frame i) and this historical rendering frame (e.g., rendering frame i-2, corresponding to frame number i-2) to generate a second image frame (e.g., interpolated frame i-1, corresponding to frame number i-1). Specifically, the encoding device can, for example, generate the interpolated frame i-1 (the second image frame) between these two rendering frames based on rendering frame i-2, the camera pose change matrix of rendering frame i, and the motion vectors of objects in the scene.

[0110] In some embodiments, the third image frame can be an extrapolated frame generated based on the first image frame. As an example, to reduce transmission latency, the encoding device can perform an extrapolation operation based on rendered frame i after rendering it, to generate a third image frame (e.g., extrapolated frame i+1, corresponding to frame number i+1). Specifically, the encoding device can, for example, determine the state of the next image frame of rendered frame i based on its geometric and motion information (e.g., optical flow information, motion trend) to generate the extrapolated frame i+1.

[0111] Furthermore, the encoding device (e.g., source device 110) can generate a bitstream using the encoded data units. As an example, the encoding device can generate a bitstream using a first image frame (e.g., rendered frame i), a second image frame (e.g., interpolated frame i-1), and a third image frame (e.g., extrapolated frame i+1) from the encoded data units. In some embodiments, the bitstream can also be referred to as a compressed video stream, such as a one-dimensional binary sequence converted from encoded image frames.

[0112] In some embodiments, the first image frame in the data unit can be encoded as a first B-frame in the bitstream, the second image frame can be encoded as a second B-frame in the bitstream, and the third image frame can be encoded as a P-frame in the bitstream. In some embodiments, a P-frame can be represented as a predicted frame, which can be generated based on preceding I-frames or P-frames. A B-frame can be represented as a bidirectional predicted frame, which can be generated by referencing preceding and following I-frames or P-frames. An I-frame can be represented as an intra picture, a complete image frame that can exist independently of other frames.

[0113] In some embodiments, the first B-frame in the bitstream may be encoded with reference to the second B-frame and the P-frame. As an example, the first B-frame may be encoded with reference to the second image frame and the third image frame; for instance, the first B-frame may be encoded with reference to the interpolated frame i-1 and the extrapolated frame i+1.

[0114] In some embodiments, the second B-frame in the bitstream may be encoded with reference to the P-frame and the preceding extrapolation frame of the first image frame. As an example, the second B-frame may be encoded with reference to the third image frame and the preceding extrapolation frame of the first image frame. For instance, the second B-frame may be encoded with reference to the following extrapolation frame i+1 and the preceding extrapolation frame i-1 of rendering frame i.

[0115] In some embodiments, P-frames in the bitstream may be encoded with reference to the preceding extrapolation frame of the first image frame. As an example, a P-frame may be encoded with reference to the preceding extrapolation frame i-1 of the rendered frame i.

[0116] Since the preceding and following extrapolated frames of the first image frame depict the general outline of the scene's motion, the second image frame and the first image frame, located in the middle, have a high degree of geometric similarity to these two reference frames. In B-frame mode, the encoding device can efficiently compress data using bidirectional motion vectors, and can reconstruct high-quality image frames by transmitting very little residual information.

[0117] In some embodiments, after the destination device 120 receives the encoded data unit, it can store the third reference frame (e.g., interpolated frame i+1) in a buffer. The destination device 120 can also decode the residual information corresponding to the frame number of the second frame image (e.g., interpolated frame i-1) (such residual information is determined, for example, based on the difference between interpolated frame i-1 and interpolated frame i-1). Further, the destination device 120 can correct the interpolated frame (e.g., interpolated frame i-1) corresponding to the frame number of the second image frame based on the residual information to display the corrected interpolated frame i-1.

[0118] To make it easier to understand, the following explanation will use a racing game scenario running in the cloud as an example.

[0119] Taking a scenario where the player is cornering at high speed as an example, this scenario places high demands on the smoothness of the graphics and the responsiveness of the controls.

[0120] On the server side, when rendering the real-world image of frame 100 (e.g., frame number 100), the system can use the vehicle speed and track geometry information provided by the game engine to predict and generate an interpolated frame for frame 101 in advance. In some embodiments, this interpolated frame can be a future image calculated based on motion patterns. This interpolated frame can be encoded as a reference frame and sent to the client along with the data packet of frame 100, and stored in a buffer.

[0121] When the server actually renders the real image of the 101st frame in the next moment, the server can compare this real image with the previously predicted interpolated frame, obtain the difference between the two (e.g., pixel difference), and generate a residual correction package with a very small amount of data (e.g., the first residual data).

[0122] In some scenarios (e.g., when network conditions are good), the client receives subsequent data containing residual correction packets in response to the rendering of the 101st frame. Furthermore, the client can merge the predicted extrapolated frames stored in the buffer with the corrected data from the newly received residual correction packets to generate and display a high-quality 101st frame consistent with the rendering. In this mode, the user sees a perfectly corrected high-definition image, and because the correction data is small, it does not significantly increase bandwidth pressure or decoding burden.

[0123] In some scenarios (e.g., poor network conditions), data packets containing real-world footage or corrected data may be delayed. For example, if the client's monitoring mechanism detects a missing frame (frame 101) that is about to be displayed, it can trigger a low-latency mode. The system directly retrieves the predicted frame 101, which was received and cached in advance at the previous moment, and sends it to the display, thus accurately continuing the race car's trajectory and ensuring a continuous gaming experience.

[0124] Finally, when the data packet finally arrives at the client, although the system no longer displays the already past 101st frame, it silently updates the video decoding device's buffer in the background using the correction data within it. In this step, the correction data in the data packet corrects for errors that may be caused by extrapolation prediction, ensuring that subsequent frames 102 and 103 have a correct reference base during decoding. This avoids screen tearing or artifact spread caused by the accumulation of prediction errors, achieving a seamless transition from network fluctuations to normal operation.

[0125] Example devices and equipment:

[0126] Figure 5 shows a block diagram of a computing device 500 in which various embodiments of the present disclosure may be implemented. The computing device 500 may be implemented as a source device 110 (or video encoder 114) or a destination device 120 (or video decoder 124), or may be included in a source device 110 (or video encoder 114) or a destination device 120 (or video decoder 124).

[0127] It should be understood that the computing device 500 shown in FIG5 is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.

[0128] As shown in Figure 5, the computing device 500 includes a general-purpose computing device 500. The computing device 500 may include at least one or more processors or processing units 510, memory 520, storage units 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560.

[0129] In some embodiments, the computing device 500 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, a large computing device, etc., provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 500 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).

[0130] Processing unit 510 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 520. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of computing device 500. Processing unit 510 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.

[0131] Computing device 500 typically includes various computer storage media. Such media can be any media accessible by computing device 500, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 530 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 500.

[0132] The computing device 500 may also include additional removable / non-removable storage media, and volatile / non-volatile storage media. Although not shown in Figure 5, a disk drive for reading from and / or writing to a removable non-volatile disk, and an optical disk drive for reading from and / or writing to a removable non-volatile optical disk, may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.

[0133] Communication unit 540 communicates with another computing device via a communication medium. Furthermore, the functionality of the components in computing device 500 can be implemented by a single computing cluster or multiple computing machines that can communicate via communication connections. Therefore, computing device 500 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.

[0134] Input device 550 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 560 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 540, computing device 500 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 500 can also communicate with one or more devices that enable a user to interact with computing device 500, or, if needed, with any device that enables computing device 500 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via an input / output (I / O) interface (not shown).

[0135] In some embodiments, some or all of the components of computing device 500 may be deployed in a cloud computing architecture, rather than being integrated into a single device. In a cloud computing architecture, components may be remotely provided and work together to achieve the functionality described herein. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (WAN), such as the Internet, using suitable protocols. For example, a cloud computing provider provides applications via a WAN that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding data, may be stored on servers at remote locations. Computing resources in a cloud computing environment may be consolidated or distributed across remote data center locations. Cloud computing infrastructure may provide services through shared data centers, although to users they appear as a single access point. Therefore, a cloud computing architecture can be used to provide the components and functionality described herein from service providers at remote locations. Alternatively, the components and functionality described herein may be provided by conventional servers or installed directly or otherwise on client devices.

[0136] In embodiments of this disclosure, computing device 500 may be used to implement video encoding / decoding. Memory 520 may include one or more video codec modules 525 having one or more program instructions. These modules are accessible and executable by processing unit 510 to perform the functions of the various embodiments described herein.

[0137] In an example embodiment of performing video encoding, input device 550 may receive video data as input 570 to be encoded. The video data may be processed, for example, by video codec module 525 to generate an encoded bitstream. The encoded bitstream may be provided as output 580 via output device 560.

[0138] In an example embodiment of performing video decoding, input device 550 may receive an encoded bitstream as input 570. The encoded bitstream may be processed, for example, by video codec module 525 to generate decoded video data. The decoded video data may be provided as output 580 via output device 560.

[0139] While this disclosure has been specifically shown and described with reference to preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. These variations are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.

Claims

1. An image rendering method, characterized in that, The method includes: acquiring a bitstream from an encoding device; at a first moment, acquiring extrapolated frame data corresponding to a target frame number by decoding the bitstream, the extrapolated frame data being generated by performing an extrapolation operation on a first reference frame; in response to obtaining first residual data corresponding to the target frame number from the bitstream before a second moment, constructing an image frame corresponding to the target frame number based on the extrapolated frame data and the first residual data, the second moment being determined based on the display time of the target frame number, the first residual data being generated by the encoding device based on the difference between interpolated frame data corresponding to the target frame number and the extrapolated frame data, the interpolated frame data being generated by performing an interpolation operation on the first reference frame and a second reference frame; and displaying the image frame at the display time corresponding to the target frame number.

2. The method according to claim 1, characterized in that, The method further includes: in response to acquiring the extrapolated frame data, storing the extrapolated frame data in a buffer.

3. The method according to claim 2, characterized in that, The method further includes: in response to not obtaining first residual data corresponding to the target frame number from the bitstream before the second time step, obtaining the interpolated frame data from the buffer; and displaying the interpolated image frame corresponding to the interpolated frame data.

4. The method according to claim 3, characterized in that, The method further includes: in response to decoding the first residual data corresponding to the target frame number from the bitstream after the second time step, generating reference data corresponding to the target frame number based on the extrapolated frame data and the first residual data; and storing the reference data in the buffer.

5. The method according to claim 4, characterized in that, The method further includes: constructing at least one subsequent image frame using the reference data in the buffer.

6. The method according to claim 5, characterized in that, Constructing at least one subsequent image frame using the reference data in the buffer includes: obtaining second residual data corresponding to the subsequent frame number of the target frame number by decoding the bitstream, the second residual data being generated by the encoding device based on the difference between the rendered frame data and the interpolated frame data corresponding to the subsequent frame number; and constructing the subsequent image frame corresponding to the subsequent frame number based on the reference data in the buffer and the second residual data.

7. The method according to claim 1, characterized in that, The bitstream is generated by the encoding device based on the following process: in response to completing the rendering of a first image frame, a data unit is generated based on the first image frame, the second image frame, and the third image frame, wherein the first image frame is a rendered frame, the second image frame is a previous interpolated frame generated based on the first image frame, and the third image frame is a subsequent interpolated frame generated based on the first image frame. And by encoding the data units, the bitstream is generated.

8. The method according to claim 7, characterized in that, The first image frame in the data unit is encoded as a first B-frame in the bitstream, the second image frame is encoded as a second B-frame in the bitstream, and the third image frame is encoded as a P-frame in the bitstream.

9. The method according to claim 8, characterized in that, The first B-frame is encoded with reference to the second B-frame and the P-frame.

10. The method according to claim 8, characterized in that, The second B frame is encoded with reference to the P frame and the preceding extrapolated frame of the first image frame.

11. The method according to claim 8, characterized in that, The P-frame is encoded with reference to the preceding extrapolation frame of the first image frame.

12. An apparatus for image rendering, characterized in that, The apparatus includes a processor and a non-transitory memory having instructions thereon, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 11.

13. A non-transitory computer-readable storage medium for storing instructions, characterized in that, The instructions cause the processor to execute the instructions of the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Reference selection for video interpolation or extrapolation

    CN101919255A

  • Image processing method and device and computer equipment

    CN113469930A