Video encoding and decoding method and apparatus
By using a filter determined from rendering information to create a virtual reference frame, the method addresses low video compression efficiency by reducing residual values and data amounts in the bitstream, enhancing encoding and decoding performance.
Patent Information
- Application Number
- JP2025529935
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-23
- Filing Date
- 2023-05-17
- Publication Date
- 2025-11-28
AI Technical Summary
Existing video compression technologies suffer from low efficiency due to large video bitstream data and poor similarity between reference frames and current images, leading to increased data amounts and reduced performance.
The method involves determining a first type filter from a set of specified filters based on rendering information to create a virtual reference frame that improves the similarity with the second image, reducing residual values and data amounts in the bitstream.
This approach enhances video encoding efficiency by reducing residual values and data amounts in the bitstream, improving compression performance and decoding efficiency.
Smart Images

Figure 2025538566000001_ABST
Abstract
Description
[Technical Field]
[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims priority to Chinese Patent Application No. 202211478458.9, entitled "VIDEO ENCODING AND DECODING METHOD AND APPARATUS," filed with the State Intellectual Property Office of the People's Republic of China on November 23, 2022, the entire contents of which are incorporated herein by reference.
[0002] [Technical field] This application relates to the field of computers, and more particularly to video encoding and decoding methods and apparatus. [Background technology]
[0003] In video encoding and decoding technologies, video compression technology is particularly important. In a video compression system, redundant information inherent in a video sequence may be reduced or eliminated by performing spatial (intra) prediction and / or temporal (inter) prediction. In the video encoding process of video compression, an encoder randomly selects one or more reference frames from the encoded frames of a video image, obtains a predicted block corresponding to a current image block from the reference frame, calculates a residual value between the predicted block and the current image block, and performs quantization encoding on the residual value. However, the predicted block is selected by the encoder from the reference frame based on motion compensation and motion estimation. The reference frame is an image obtained through processing by a loop filter, and since the reference frame is an image corresponding to a neighboring frame of the current image block, the similarity between the reference frame and the current image block is low, and the difference between the predicted block and the current image block is large. As a result, the residual value between the predicted block and the current image block becomes large, the amount of data of the video bitstream obtained through encoding becomes large, the performance of the video compression process becomes low, and the video encoding efficiency is affected. Therefore, how to provide a more effective video encoding and decoding method is currently an urgent problem to be solved. Summary of the Invention
[0004] This application provides a video encoding and decoding method and apparatus to solve the problem that the amount of data of the bitstream obtained by compressing the video is large and the video compression rate is low.
[0005] According to a first aspect, this application provides a video encoding method. The video encoding method may be applied to an encoding and decoding system, or an encoder side that supports the encoding and decoding system in realizing the video encoding method. For example, the encoder side includes a video encoder. Here, an example in which the encoder side executes the video encoding method provided in this embodiment is used for description. The video encoding method includes: first, the encoder side obtains a source video and rendering information corresponding to the source video; second, the encoder side determines a first type filter from a plurality of specified type filters based on the rendering information; third, the encoder side performs interpolation on a first decoded image by using the first type filter to obtain a virtual reference frame of a second image; and fourth, the encoder side encodes a second image based on the virtual reference frame of the second image to obtain a bitstream corresponding to the second image.
[0006] The source video includes a plurality of source images, the rendering information indicates processing parameters used in a process of generating a bitstream by an encoder based on the plurality of source images, the processing parameters indicating that a first image and a second image of the plurality of source images have at least a partial overlapping region, and the first decoded image is a reconstructed image obtained by encoding and then decoding the first image of the plurality of images.
[0007] In this embodiment, the encoder side determines a first type filter from a plurality of specified type filters, and the first type filter is compatible with a filter used in the image rendering process. This helps improve the similarity between the virtual reference frame determined by the encoder side and the second image by using the first type filter, compared to the problem of low compression performance due to low similarity between the reference frame and the second image when a fixed filter is used to process the reconstructed image to obtain the reference frame. Therefore, when the encoder side encodes the second image based on the virtual reference frame, the residual value corresponding to the second image is reduced, and the amount of data corresponding to the second image in the bitstream is reduced. This improves compression performance and video coding efficiency.
[0008] In a possible implementation, the first type of filter determined by the encoder side based on the rendering information includes: the encoder side consults a mapping table specified based on the rendering information to obtain a filter corresponding to each processing parameter in the rendering information; and the encoder side determines the first type of filter from all filters corresponding to the rendering information.
[0009] The mapping table indicates at least one type of filter, among a plurality of types of filters, that corresponds to each processing parameter.
[0010] The mapping table indicates a filter corresponding to each processing parameter in the rendering information, and the correspondence between the processing parameter and the filter is determined based on the filter used in each processing process of the image rendering engine. The encoder side consults the mapping table based on the acquired processing parameters to determine a first type of filter. The first type of filter matches the filter used by the image rendering engine, so that using the first type of filter improves the temporal-spatial correlation between the virtual reference frame acquired by the encoder side and the second image, thereby reducing the residual value of the second image determined by the encoder side based on the virtual reference frame. This improves the encoding effect.
[0011] In a possible implementation, the mapping table further indicates a priority of each type of filter among the multiple types of filters. The encoder side's determining the first type of filter from the filters corresponding to the rendering information includes: the encoder side obtains priorities of all filters corresponding to the rendering information, and the encoder side uses a filter with the highest filter priority among all filters as the first type of filter.
[0012] The encoder selects a filter from the mapping table that corresponds to the processing parameters and has the highest priority as a first type filter based on the processing parameters included in the rendering information. The filter with the highest priority also has the highest degree of compatibility with the filter used by the image rendering engine. The encoder performs interpolation on the reference frame using the filter with the highest degree of compatibility to obtain a virtual reference frame. The temporal-spatial correlation between the virtual reference frame and the second image is improved, resulting in a reduction in the residual value corresponding to the second image and obtained by the encoder based on the virtual reference frame. This improves the encoding effect, reduces the amount of data obtained through encoding, and improves compression performance.
[0013] In a possible implementation, the encoder side's determining a first type filter from all filters corresponding to the rendering information includes: the encoder side performs prediction on all filters corresponding to the rendering information to obtain at least one prediction result; and the encoder side selects a target prediction result that satisfies a condition specified by the encoding information from the at least one prediction result, and uses the filter corresponding to the target prediction result as the first type filter.
[0014] One prediction result corresponds to one of all the filters, and the prediction result indicates encoding information obtained by precoding the second image by the encoder side by using the filter, and the encoding information includes at least one of a predicted bit rate and a distortion rate of the second image.
[0015] For example, the above-mentioned precoding is a process of performing inter prediction on a second image to obtain a predicted block of the second image, obtaining a residual value of a corresponding image block in the second image based on the predicted block, and then performing a transformation process.
[0016] The encoder performs precoding by using all filters corresponding to the processing parameters in the rendering information. The encoder filters the prediction results obtained through precoding based on specified conditions and uses the filter corresponding to the target prediction result that satisfies the specified conditions as a first-type filter. As a result, the filter that actually satisfies the preset conditions among all filters is determined as the first-type filter. By using the first-type filter, the similarity between the virtual reference frame obtained by the encoder through interpolation and the second image is maximized. During the process of encoding the second image by the encoder, the similarity between the virtual reference frame and the second image is maximized, resulting in the smallest residual value corresponding to the second image determined by the encoder by encoding the second image based on the virtual reference frame. This improves the encoding compression effect.
[0017] In a possible example, the rendering information includes one or a combination of a depth map, an albedo map, and post-processing parameters.
[0018] For the above example in which the rendering information includes processing parameters, this application provides a step of determining a first type filter or first type filters corresponding to each processing parameter. If the rendering information includes a depth map, the first type filter is determined based on a relationship between pixels in the depth map and edge information of objects in the depth map. For example, the depth map may correspond to a Catmull-Rom filter or a bilinear filter.
[0019] Alternatively, if the rendering information includes an albedo map, the first type of filter is determined based on a relationship between pixels in the albedo map and edge information of objects in the albedo map. For example, the albedo map may correspond to a Catmull-Rom filter or a B-spline filter.
[0020] Alternatively, if the post-processing parameters included in the rendering information include anti-aliasing parameters, the first type of filter is the filter indicated by the anti-aliasing parameters.
[0021] Alternatively, if the post-processing parameters included in the rendering information include instructions for the execution of a motion blur module, the first type of filter is a bilinear filter.
[0022] Alternatively, the first type of filter is a Catmull-Rom filter.
[0023] For example, the edge information of the object in the depth map may be obtained in the following manner: For example, the encoder side performs Sobel filtering on the depth map to obtain the filtering result, and then performs Otsu binarization on the filtering result to obtain the edge information of the object in the depth map.
[0024] The encoder determines a first type of filter corresponding to the rendering information based on the correspondence between each processing parameter and the filter, and the first type of filter has a certain degree of compatibility with the filter used by the image rendering engine. The encoder performs interpolation on the reference frame by using the first type of filter to obtain a virtual reference frame, which improves the degree of compatibility between the virtual reference frame and the second image. When the encoder encodes the second image based on the virtual reference frame, the obtained residual value corresponding to the second image is reduced, which improves the encoding effect.
[0025] In a possible implementation, the encoder side performs interpolation on the first decoded image by using a first type filter determined based on the rendering information to obtain a virtual reference frame of the second image, which includes: first, generating a blank image whose size matches that of the reference block; then, performing interpolation on the reference block based on the first correspondence and the partitioned graphic motion vector map to determine pixel values of all pixels in the blank image; and finally, obtaining a virtual reference frame based on the pixel values of all pixels in the blank image.
[0026] The first decoded image includes one or more reference blocks corresponding to the first image, and the first correspondence relationship indicates a positional mapping relationship between pixels in the first image and pixels in the second image in an overlapping region between the first image and the second image.
[0027] The encoder performs interpolation on the reference blocks in the first decoded image based on the first correspondence to obtain a virtual reference frame for the second image. The similarity between the virtual reference frame and the second image is higher than the similarity between the first decoded image and the second image. Therefore, the residual value between the virtual reference frame and the second image is reduced, ensuring image quality consistency and reducing the amount of image data in the bitstream. This improves the coding effect.
[0028] In a possible implementation, the first correspondence is obtained by using the following method: the encoder side first generates a blank image whose size matches that of the source image; then, based on the size of the first graphics motion vector map and the size of the blank image, the encoder side determines a second correspondence between pixel positions in the first graphics motion vector map and pixel positions in the blank image; finally, the encoder side performs interpolation on the graphics motion vector map by using an interpolation filter to determine the pixel value and offset of each pixel in the blank image.
[0029] When the size of the graphics motion vector does not match the size of the source image, the encoder performs interpolation on the graphics motion vector map to obtain a graphics motion vector map obtained through interpolation. The size of the graphics motion vector map obtained through interpolation matches the size of the first decoded image, and the pixels of the graphics motion vector map obtained through interpolation match the pixels of the first decoded image. The encoder performs interpolation on the first decoded image based on the graphics motion vector map obtained through interpolation to obtain a virtual reference frame. This improves the fit between the virtual reference frame and the second image. When the encoder encodes the second image based on the virtual reference frame, the residual value corresponding to the image is reduced, improving compression performance.
[0030] In a possible implementation, the video encoding method includes: first, an encoder side determines target discrete parameters of pixels in a first decoded image based on a graphic motion vector map indicated in the rendering information; second, a first type filter on the encoder side performs interpolation on a reference block by using a first filter kernel, and calculates first discrete parameters of pixels in a virtual reference block obtained through interpolation for each channel; third, the encoder side determines a first discrete parameter that has a minimum difference from the target discrete parameter from at least one first discrete parameter corresponding to all filter kernels, and uses the filter kernel corresponding to the first discrete parameter as a parameter of the first type filter.
[0031] The target discrete parameter includes a target variance or a target covariance. The first discrete parameter includes a first variance or a first covariance. The first filter kernel is one of a plurality of specified filter kernels.
[0032] The encoder determines first discrete parameters that have the smallest difference from the target discrete parameters, and uses a first filter kernel corresponding to the first discrete parameters as parameters of a first type of filter. The encoder performs interpolation on the first decoded image by using the first type of filter based on the parameters of the first type of filter to obtain a virtual reference frame. The fit between the virtual reference frame and the second image is improved. When the encoder encodes the second image based on the virtual reference frame, the residual value corresponding to the second image is reduced, and the amount of image data in the bitstream is reduced. This improves compression performance.
[0033] In a possible implementation, the video encoding method further includes: the encoder side writes type information of the first type filter into a bitstream.
[0034] The encoder transmits type information of the first type filter to the decoder, so that the decoder does not need to perform a process of determining the first type filter based on the rendering information, which improves video decoding efficiency.
[0035] According to a second aspect, this application provides a video decoding method. The video decoding method may be applied to an encoding and decoding system, or a decoder side that supports the encoding and decoding system in realizing the video decoding method. For example, the decoder side includes a video decoder. Here, an example in which the decoder side executes the video decoding method provided in this embodiment is used for description. The video decoding method includes: first, the decoder side obtains a bitstream and rendering information corresponding to the bitstream, where the bitstream includes multiple image frames; second, the decoder side determines a first type filter from multiple specified type filters based on the rendering information; third, the decoder side performs interpolation on a first decoded image corresponding to the first image frame by using the first type filter to obtain a virtual reference frame of a second image frame; and fourth, the decoder side decodes the second image frame based on the virtual reference frame to obtain a second decoded image of the second image frame.
[0036] The rendering information indicates processing parameters used in a process of generating a bitstream based on a plurality of source images, the processing parameters indicating that a source image corresponding to a first image frame and a source image corresponding to a second image frame of the plurality of image frames have an area that at least partially overlaps.
[0037] For example, the bitstream may be transmitted by the encoder side to the decoder side.
[0038] The decoder side selects a first type of filter from a plurality of specified type filters based on the processing parameters indicated by the rendering information, and obtains a virtual reference frame for the second image frame by using the first type of filter, thereby avoiding the problem of poor decoding efficiency due to the reference frame being obtained by processing the first decoded image using a fixed filter. The degree of match between the virtual reference frame and the source image corresponding to the second image frame is improved. When the decoder side decodes the second image frame based on the virtual reference frame and the processing capability of the decoder side is consistent, the amount of data processed by the decoder side is reduced, and decoding efficiency is improved.
[0039] In a possible implementation, the first type filter determined by the decoder side based on the rendering information includes: the decoder side consults a mapping table specified based on the rendering information to obtain a filter corresponding to each processing parameter in the rendering information; and the decoder side determines the first type filter from all filters corresponding to the rendering information.
[0040] The mapping table indicates at least one type of filter, among a plurality of types of filters, that corresponds to each processing parameter.
[0041] The decoder side consults the specified mapping table based on the rendering information to determine a first type of filter, where the first type of filter has a certain degree of compatibility with the filter used by the image rendering engine. The decoder side performs interpolation on the reference frame by using the first type of filter to obtain a virtual reference frame. The temporal-spatial correlation between the virtual reference frame and the second image frame is improved. In other words, the degree of compatibility between the virtual reference frame and the source image corresponding to the second image frame is improved. When the decoder side decodes the second image frame based on the virtual reference frame, the amount of data processed is reduced, and decoding efficiency is improved.
[0042] In a possible implementation, the mapping table further indicates a priority of each type of filter among the multiple types of filters. The decoder side's determining the first type of filter from the filters corresponding to the rendering information includes: the decoder side obtains priorities of all filters corresponding to the rendering information, and the decoder side uses the filter with the highest filter priority among all filters as the first type of filter.
[0043] The decoder side selects a filter from the mapping table that corresponds to the processing parameters and has the highest priority as a first type filter based on the processing parameters included in the rendering information. The filter with the highest priority also has the highest degree of match with the filter used by the image rendering engine. The decoder side performs interpolation on the reference frame by using the filter with the highest degree of match to obtain a virtual reference frame. The degree of match between the virtual reference frame and the source image corresponding to the second image frame is higher. Therefore, when the decoder side performs decoding based on the virtual reference frame, the amount of data processed is reduced and decoding efficiency is improved.
[0044] In a possible example, the rendering information includes one or a combination of a depth map, an albedo map, and post-processing parameters.
[0045] For the above example where the rendering information includes processing parameters, this application provides a step of determining a first type filter or first type filters corresponding to each processing parameter. If the rendering information includes a depth map, the first type filters are determined based on a relationship between pixels in the depth map and edge information of objects in the depth map.
[0046] Alternatively, if the rendering information includes an albedo map, the first type of filter is determined based on a relationship between pixels in the albedo map and edge information of objects in the albedo map.
[0047] Alternatively, if the post-processing parameters included in the rendering information include anti-aliasing parameters, the first type of filter is the filter indicated by the anti-aliasing parameters.
[0048] Alternatively, if the post-processing parameters included in the rendering information include instructions for the execution of a motion blur module, the first type of filter is a bilinear filter.
[0049] Alternatively, the first type of filter is a Catmull-Rom filter.
[0050] For example, the edge information of the object in the depth map may be obtained in the following manner: For example, the decoder side performs Sobel filtering on the depth map to obtain the filtering result, and then performs Otsu binarization on the filtering result to obtain the edge information of the object in the depth map.
[0051] The decoder side determines a first type of filter corresponding to the rendering information based on the correspondence between each processing parameter and the filter, and the first type of filter has a certain degree of compatibility with the filter used by the image rendering engine. The decoder side performs interpolation on the reference frame by using the first type of filter to obtain a virtual reference frame, which improves the degree of compatibility between the virtual reference frame and the source image corresponding to the second image frame. When the decoder side decodes the second image frame based on the virtual reference frame, the amount of data processed is reduced and decoding efficiency is improved.
[0052] In a possible implementation manner, before the decoder side performs interpolation on the first decoded image corresponding to the first image frame by using the first type filter determined based on the rendering information, the video decoding method further includes: the decoder side obtains filter information, and the decoder side determines a first type filter and filtering parameters corresponding to the first type filter based on the filter type and parameters indicated in the filter information.
[0053] The filtering parameters indicate the processing parameters used in the process of the decoder side obtaining the virtual reference frame by using the first type filter.
[0054] For example, after the encoder side determines a first type filter and filtering parameters corresponding to the first type filter, the encoder side writes the first type filter and filtering parameters corresponding to the first type filter into a bitstream and sends the bitstream to the decoder side.
[0055] The decoder side directly obtains the first type filter and the filtering parameters corresponding to the first type filter, so that the decoder side does not need to perform a processing process for determining the first type filter based on the rendering information, which improves video decoding efficiency.
[0056] In a possible implementation, the decoder side performs interpolation on a first decoded image corresponding to a first image frame by using a first type filter determined based on rendering information to obtain a virtual reference frame of a second image frame, the method comprising: first generating a blank image whose size is consistent with that of the reference block; then performing interpolation on the reference block based on the first correspondence and the partitioned graphic motion vector map to determine pixel values of all pixels in the blank image; and finally obtaining a virtual reference frame based on the pixel values of all pixels in the blank image.
[0057] The first decoded image includes one or more reference blocks corresponding to the first image, and the first correspondence relationship indicates a positional mapping relationship between pixels in the source image corresponding to the first image frame and pixels in the source image corresponding to the second image frame in an overlapping region between the source image corresponding to the first image frame and the source image corresponding to the second image frame.
[0058] The decoder side performs interpolation on the reference block in the first decoded image based on the first correspondence relationship to obtain a virtual reference frame of the second image. The similarity between the virtual reference frame and the second image frame is higher than the similarity between the first decoded image and the second image frame. When the decoder side decodes the second image frame based on the virtual reference frame, the amount of data processed is reduced, and decoding efficiency is improved.
[0059] In a possible implementation, the first correspondence is obtained by using the following method: the decoder first generates a blank image whose size matches that of the first decoded image; then, the decoder determines a second correspondence between pixel positions in the first graphics motion vector map and pixel positions in the blank image based on the size ratio of the first graphics motion vector map to the blank image; finally, the decoder performs interpolation on the graphics motion vector map by using an interpolation filter to determine the pixel value and offset of each pixel in the blank image.
[0060] The size of the first decoded image matches the size of the source image.
[0061] When the size of the graphics motion vector does not match the size of the first decoded image, the decoder side performs interpolation on the graphics motion vector map to obtain a graphics motion vector map obtained through interpolation. The size of the graphics motion vector map obtained through interpolation matches the size of the first decoded image, and the graphics motion vector map obtained through interpolation matches the first decoded image. The decoder side performs interpolation on the first decoded image based on the graphics motion vector map obtained through interpolation to obtain a virtual reference frame. This improves the degree of match between the virtual reference frame and the second image frame. When the decoder side decodes the second image frame based on the virtual reference frame, the amount of data processed is reduced and decoding efficiency is improved.
[0062] In a possible implementation, the video decoding method further includes: first, the decoder side determines target discrete parameters of pixels in the first decoded image based on the graphic motion vector map indicated in the rendering information; second, a first type filter on the decoder side performs interpolation on the reference block by using a first filter kernel, and calculates first discrete parameters of pixels in the virtual reference block obtained through interpolation for each channel; third, the decoder side determines a first discrete parameter having a minimum difference from the target discrete parameter from at least one first discrete parameter corresponding to all filter kernels, and the decoder side uses the filter kernel corresponding to the first discrete parameter as a parameter of the first type filter.
[0063] The target discrete parameter includes a target variance or a target covariance. The first discrete parameter includes a first variance or a first covariance. The first filter kernel is one of a plurality of specified filter kernels.
[0064] The decoder determines a first discrete parameter that has the smallest difference from the target discrete parameter, and uses a first filter kernel corresponding to the first discrete parameter as a parameter of a first type filter. The decoder performs interpolation on the first decoded image by using the first type filter based on the parameter of the first type filter to obtain a virtual reference frame. The degree of fit between the virtual reference frame and the second image frame is improved. When the decoder decodes the second image frame based on the virtual reference frame, the amount of data processed is reduced, and decoding efficiency is improved.
[0065] According to a third aspect, this application provides a video encoding device. The video encoding device is used on an encoder side and in a video encoding and decoding system including the encoder side. The video encoding device includes modules configured to perform the video encoding method according to the first aspect or any one of optional implementation manners of the first aspect. For example, the video encoding device includes a first acquisition module, a first interpolation module, and an encoding module. The first acquisition module is configured to acquire a source video and rendering information corresponding to the source video. The first interpolation module is configured to determine a first type of filter based on the rendering information and perform interpolation on a first decoded image by using the first type of filter to obtain a virtual reference frame of a second image. The encoding module is configured to encode the second image based on the virtual reference frame of the second image.
[0066] The source video includes a plurality of source images, the rendering information indicates processing parameters used in a process of generating a bitstream based on the plurality of source images, the processing parameters indicating that a first image and a second image of the plurality of source images have at least a partial overlapping region, the first type of filter is one of a plurality of specified types of filters, and the first decoded image is a reconstructed image obtained by encoding and then decoding a first image of the plurality of images, the first image being encoded before the second image.
[0067] For more detailed implementation details of the video encoding device, please refer to the description of any implementation details of the second aspect and the following specific implementation details, and the details will not be described again here.
[0068] According to a fourth aspect, this application provides a video decoding device. The video decoding device is used at a decoder side and in a video encoding and decoding system including the decoder side. The video decoding device includes modules configured to perform the video decoding method according to the second aspect or any one of optional implementation manners of the second aspect. For example, the video decoding device includes a second acquisition module, a second interpolation module, and a decoding module. The second acquisition module is configured to acquire a bitstream and rendering information corresponding to the bitstream. The second interpolation module is configured to determine a first type of filter based on the rendering information and perform interpolation on a first decoded image corresponding to the first image frame by using the first type of filter to obtain a virtual reference frame for the second image frame. The decoding module is configured to decode the second image frame based on the virtual reference frame to obtain a second decoded image corresponding to the second image frame.
[0069] The bitstream includes a plurality of image frames, the rendering information indicates processing parameters used in a process of generating the bitstream based on a plurality of source images, the processing parameters indicating that a source image corresponding to a first image frame and a source image corresponding to a second image frame of the plurality of image frames have at least a partial overlapping region, and the first type of filter is one of a plurality of specified types of filters.
[0070] For more detailed implementation details of the video decoding device, please refer to the description of any implementation details of the first aspect and the following specific implementation details, and the details will not be described again here.
[0071] According to a fifth aspect, the present application provides a chip including a processor and a power supply circuit, the power supply circuit configured to supply power to the processor, the processor configured to execute a method according to the first aspect and any one of the possible implementations of the first aspect, and / or the processor configured to execute a method according to the second aspect and any one of the possible implementations of the second aspect. The method is configured to perform any one of the methods.
[0072] According to a sixth aspect, the present application provides a codec including a memory and a processor. The memory is configured to store computer instructions that, when executed, cause the processor to implement a method according to the first aspect and any one of possible implementations of the first aspect, and / or, when executed, cause the processor to implement a method according to the second aspect and any one of possible implementations of the second aspect.
[0073] According to a seventh aspect, the application provides an encoding and decoding system, including an encoder side and a decoder side, wherein the encoder side is configured to encode a plurality of images based on rendering information corresponding to the plurality of source images to obtain a bitstream corresponding to the plurality of images, thereby realizing the method according to the first aspect and any one of the possible realization manners of the first aspect.
[0074] The decoder side is configured to decode the bitstream based on rendering information corresponding to the plurality of source images to obtain the plurality of decoded images, and to realize the method according to the second aspect and any one of the possible realization methods of the second aspect.
[0075] According to an eighth aspect, the present application provides a computer-readable storage medium, the storage medium storing a computer program or instructions, the computer program or instructions, when executed by a processing device, realizing a method according to the first aspect or any one of the optional implementations of the first aspect, and / or the computer program or instructions, when executed by a processing device, realizing a method according to the second aspect or any one of the optional implementations of the second aspect.
[0076] According to a ninth aspect, the application provides a computer program product, the computer program product comprising computer programs or instructions which, when executed by a processing device, result in the method according to the first aspect or any one of the optional implementations of the first aspect, and / or which, when executed by a processing device, result in the method according to the second aspect or any one of the optional implementations of the second aspect.
[0077] For the beneficial effects of the third to ninth aspects, please refer to the description of the first aspect or any one of the implementation methods of the first aspect, or the second aspect or any one of the implementation methods of the second aspect. Details will not be described again here. In this application, based on the implementation methods according to the above aspects, the implementation methods may be further combined to provide more implementation methods. [Brief explanation of the drawings]
[0078] [Figure 1] 1 is an exemplary block diagram of a video encoding and decoding system according to the present application; [Figure 2] 1 is a diagram of the structure of a video encoder according to the present application; [Figure 3] 1 is a diagram of the structure of a video decoder according to this application; [Figure 4] 1 is a schematic flow chart of a temporal anti-aliasing method according to the present application; [Figure 5]1 is a schematic flow chart of a video encoding method according to the present application; [Figure 6] 1 is a diagram of a reference block interpolation method according to the present application; [Figure 7] 1 is a diagram of a graphics motion vector interpolation method according to the present application. [Figure 8] 1 is a schematic flow chart of video decoding according to the present application; [Figure 9] 1 is a diagram of the structure of a video encoding device according to this application; [Figure 10] 1 is a diagram of the structure of a video decoding device according to the present application; [Figure 11] 1 is a diagram of the architecture of a computing device according to the present application. DETAILED DESCRIPTION OF THE INVENTION
[0079] In this application, the encoder side determines a first type filter from a plurality of specified type filters, and the first type filter is matched with a filter used in the image rendering process. This helps improve the similarity between the virtual reference frame determined by the encoder side and the second image by using the first type filter. Therefore, when the encoder side encodes the second image based on the virtual reference frame, the residual value corresponding to the second image is reduced, and the amount of data corresponding to the second image in the bitstream is reduced. This improves compression performance and video coding efficiency.
[0080] The decoder side selects a first type filter from a plurality of specified type filters based on the processing parameters indicated by the rendering information, and obtains a virtual reference frame for the second image frame by using the first type filter, thereby avoiding the problem of low decoding efficiency due to the reference frame being obtained by processing the decoded image using a fixed filter. The degree of match between the virtual reference frame and the source image corresponding to the image to be decoded is improved. When the decoder side decodes the second image frame based on the virtual reference frame and the processing capability of the decoder side is consistent, the amount of data processed by the decoder side is reduced, and video decoding efficiency is improved.
[0081] The following describes the solutions provided in this application with reference to embodiments: For a clear and concise description of the following embodiments, a brief description of the related art is provided first.
[0082] Video encoding is the process of compressing multiple frames of images contained in a video into a bitstream.
[0083] Video decoding is the process of recovering a bitstream into a reconstructed image of multiple frames according to specific syntax rules and processing methods.
[0084] To realize video encoding and decoding, this application provides a video encoding and decoding system. FIG. 1 is an exemplary block diagram of a video encoding and decoding system according to this application. The video encoding and decoding system includes an encoder side 100 and a decoder side 200. The encoder side 100 generates coded video data (alternatively called a bitstream). Therefore, the encoder side 100 may be referred to as a video encoding device. The decoder side 200 may decode the bitstream (e.g., a video including one or more image frames) generated by the encoder side 100. Therefore, the decoder side 200 may be referred to as a video decoding device. Various implementations of the encoder side 100, the decoder side 200, or both may include one or more processors and memories coupled to the one or more processors. The memory may include, but is not limited to, random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer.
[0085] 1, encoder 100 includes source video 110, video encoder 120, and output interface 130. In some examples, output interface 130 may include a modulator / demodulator (modem) and / or a transmitter. Source video 110 may include a video capture device (e.g., a video camera), a video archive containing previously captured video data, a video feed-in interface for receiving video data from a video content provider, and / or a computer graphics system, e.g., an image rendering engine, for generating video data, or a combination of the above video sources.
[0086] The video encoder 120 may encode video data from the source video 110. In some examples, the encoder side 100 transmits the bitstream directly to the decoder side 200 through the output interface 130 via the link 300. In other examples, the bitstream may be further stored in a storage device 400 for later access by the decoder side 200 for decoding and / or playback.
[0087] In the example of Figure 1, decoder side 200 includes input interface 230, video decoder 220, and display device 210. In some examples, input interface 230 includes a receiver and / or a modem. Input interface 230 may receive encoded video data via link 300 and / or from storage device 400. Display device 210 may be integrated with decoder side 200 or may be external to decoder side 200. Generally, display device 210 displays decoded video data. Display device 210 may include multiple types of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices.
[0088] This application provides a possible video encoder based on the video encoding and decoding system shown in Figure 1. Figure 2 is a diagram of the structure of a video encoder according to this application.
[0089] Video encoder 120 includes an inter predictor 121, an intra predictor 122, a transformer 123, a quantizer 124, an inverse quantizer 126, an inverse transformer 127, a filter unit 128, and a memory 129. Inverse quantizer 126 and inverse transformer 127 are configured to reconstruct image blocks. Filter unit 128 is configured to implement one or more loop filters, such as a deblocking filter, an adaptive loop filter, and a sample adaptive offset filter.
[0090] Memory 129 may store video data encoded by components of video encoder 120. The video data stored in memory 129 may be obtained from source video 110. Memory 129 may be a reference picture memory that stores reference video data used by video encoder 120 to encode video data in intra- or inter-coding mode. Memory 129 may be dynamic random access memory (DRAM), magnetic RAM (MRAM), resistive RAM (RRAM), or other types of memory devices.
[0091] Regarding the video encoding procedure, the operation procedure of the video encoder will be described below with reference to the contents of FIG.
[0092] After inter predictor 121 and intra predictor 122 generate predictive blocks for a current image block (alternatively referred to as a second image) based on source data (source video), video encoder 120 subtracts the predictive blocks from the current image block to be encoded to form residual image blocks. The residual video data in the residual blocks may be included in one or more transform units (TUs) and used by transformer 123. Transformer 123 converts the residual video data into residual transform coefficients through a transform such as a discrete cosine transform or a conceptually similar transform. Transformer 123 may convert the residual video data from the pixel value domain to a transform domain, such as the frequency domain.
[0093] The transformer 123 may send the resulting transform coefficients to the quantizer 124, which quantizes the transform coefficients to further reduce the bit rate. In some examples, the quantizer 124 may then perform a scan of a matrix containing the quantized residual transform coefficients. Alternatively, the entropy encoder 125 may perform the scan.
[0094] After quantization, entropy encoder 125 performs entropy coding on the quantized transform coefficients. For example, entropy encoder 125 may perform context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioned entropy (PIPE) coding, or other entropy coding methods or techniques. After performing entropy coding, entropy encoder 125 may transmit the coded bitstream to a video decoder or archive the coded bitstream for subsequent transmission or retrieval by a video decoder. Entropy encoder 125 may also perform entropy coding on syntax elements of the current image block being coded.
[0095] The inverse quantizer 126 and the inverse transformer 127 perform inverse quantization and inverse transformation, respectively, to reconstruct the residual block in the pixel domain, e.g., for later use as a reference block in a reference image. The video encoder 120 adds the reconstructed residual block to the prediction block generated by the inter predictor 121 or the intra predictor 122 to generate a reconstructed image or a reconstructed image block. The filter unit 128 can be applied to the reconstructed image block to reduce distortion, e.g., block artifacts. The reconstructed image or the reconstructed image block is then stored in memory 129 as a reference block (alternatively referred to as a first decoded image) and may be used by the inter predictor 121 as a reference block for performing inter prediction on blocks in a subsequent video frame or image.
[0096] As shown in FIG. 2, the video encoder 120 is further externally connected to an interpolation filter unit 140. The interpolation filter unit 140 is configured to perform interpolation on the first decoded image to obtain a virtual reference frame and store the virtual reference frame in the memory 129. The video encoder 120 uses the virtual reference frame to assist in encoding the second image. The filters in the interpolation filter unit are determined by the encoder 100 based on acquired prior information, which may be processing parameters used in the process of performing rendering by using a computer image rendering technique to obtain the source video. The similarity between the virtual reference frame and the second image is higher than the similarity between the first decoded image and the second image.
[0097] If possible, interpolation filter unit 140 may be located inside video encoder 120.
[0098] It should be understood that other structural variations of video encoder 120 can be used to encode the video stream. For example, for some image blocks or image frames, video encoder 120 may directly quantize the residual signal, and no processing by transformer 123 is required, and correspondingly, no processing by inverse transformer 127 is required. Alternatively, for some image blocks or image frames, video encoder 120 does not generate residual data, and correspondingly, no processing by transformer 123, quantizer 124, inverse quantizer 126, and inverse transformer 127 is required. Alternatively, video encoder 120 may directly store reconstructed image blocks as reference blocks, and no processing by filter unit 128 is required. Alternatively, quantizer 124 and inverse quantizer 126 in video encoder 120 may be combined.
[0099] 3 is a diagram of the structure of a video decoder according to an embodiment of this application. The video decoder 220 includes an entropy decoder 221, an inverse quantizer 222, an inverse transformer 223, a filter unit 224, a memory 225, an inter predictor 226, and an intra predictor 227. The video decoder 220 may perform a decoding process that is substantially the reverse of the encoding process described with respect to the video encoder 120 in FIG. 2. First, a residual block or residual value is obtained through the entropy decoder 221, the inverse quantizer 222, and the inverse transformer 223, and the bitstream is decoded to determine whether intra prediction or inter prediction is used for the current video block. If intra prediction is performed, the intra predictor 227 constructs prediction information based on pixel values of pixels in a surrounding reconstructed region according to the intra prediction method used. If inter prediction is performed, the inter predictor 226 needs to obtain motion information through analysis, determine a reference block in a reconstructed image based on the motion information obtained through analysis, and use pixel values of pixels in the block as prediction information. By performing a filtering operation on the prediction information and the residual information, the reconstruction information can be obtained.
[0100] In addition, the video decoder 220 is further externally connected to an interpolation filter unit 240. For the function of the interpolation filter unit 240, please refer to the above content of the interpolation filter unit 140 externally connected to the video decoder 120. The details will not be described again here.
[0101] If possible, the interpolation filter unit 240 may be located inside the video encoder 220.
[0102] Optionally, the source video is obtained by performing rendering on input data by a rendering side, where the input data is data obtained by a display device in communication with the rendering side, or the input data is data received by the rendering side.
[0103] For example, the rendering side includes an image rendering engine (for example, V-Ray, Unreal, or Unity). In the process in which the image rendering engine performs rendering on input data, process map data and parameter information are generated.
[0104] The process map data includes a rendered image generated when the rendering side processes input data, an intermediate map between the input data and the rendered image, etc. For example, the intermediate map may include, but is not limited to, a graphics motion vector map, a low-quality rendering result, a depth map, a position map, a normal map, an albedo map, a specular intensity map, etc.
[0105] The parameter information in the above rendering process is post-processing parameters used by the rendering side to perform post-processing on the rendered image, such as mesh ID, material ID, motion blur parameters and anti-aliasing parameters.
[0106] It should be noted that the above process map data and parameter information may also be collectively referred to as intermediate data corresponding to the source video, rendering information used when the encoder encodes the source video, etc.
[0107] In a possible example, the image rendering engine of the decoder side 200 may perform image rendering based on input data on the display side to obtain one or more of process map data and parameter information. Since the decoder side 200 obtains all or part of the intermediate data, the encoder side 100 is prevented from sending all of the intermediate data to the decoder side 200, thereby reducing the amount of intermediate data transmitted by the encoder side 100 to the decoder side 200 and saving bandwidth.
[0108] For example, the display-side input data may be one or more of the camera (angle of view) position, light source position / intensity / color, etc. The rendering-side input data may be one or more of the object geometry, object initial position, light source information, etc. The graphic motion vector map is an image that indicates the positional correspondence between pixels in two frames of a source video image. The motion blur parameters are instructions for the image rendering engine to perform motion blur. The anti-aliasing parameters are filtering methods used when the image rendering engine performs TAA.
[0109] Optionally, the image rendering engine may further perform post-processing on the rendered image. For the purposes of explanation, the following uses an example in which the post-processing refers to anti-aliasing processing.
[0110] The rendering side performs temporal anti-aliasing (TAA) processing on the rendered image to obtain multiple images contained in the source video and solve the aliasing problem of the rendered image. FIG. 4 is a schematic flowchart of a temporal anti-aliasing method according to this application. The difference between multiple frames of the rendered image is small. For example, only a few pixels (e.g., one or two pixels) are offset between the previous frame and the current frame. A first pixel in the current frame is used as an example. The rendering side determines a second pixel in the previous frame of the current frame that corresponds to the first pixel based on the graphic motion vector map, and the rendering side performs a weighted average on the pixel value of the first pixel and the pixel value of the second pixel to obtain a final pixel value. The rendering side fills the first pixel in the current frame with the final pixel value and determines pixel values for each pixel in the first image to obtain the first image. The first pixel is any one of all pixels in the current frame.
[0111] The rendering side may be an encoder side 100, which may be on a cloud server. The display side may be a decoder side 200, which may be on a client.
[0112] The rendering side determining the pixel value of the second pixel may include: the encoder side 100 uses a filter to filter the second pixel or pixels surrounding the second pixel (e.g., a 3x3 pixel range) to obtain the pixel value of the second pixel. The filtering method of the filter may be one of average filtering, median filtering, Gaussian filtering, etc.
[0113] To reduce the load on the terminal side (e.g., the display side or the decoder side), the image rendering engine may be located on the cloud side (e.g., the cloud-side server or the cloud-side encoder side). The encoder side 100 performs image rendering on multiple images to obtain a source video, and encodes the source video to obtain a bitstream. Furthermore, the encoder side 100 transmits the bitstream to the decoder side 200, and the decoder side 200 decodes the bitstream for playback.
[0114] In a possible scenario, in a cloud gaming scenario, the encoder side 100 processes multiple images by using an image rendering engine to obtain game pictures, encodes the game pictures to obtain a bitstream, and reduces the amount of data of the game pictures to be transmitted. The decoder side 200 decodes the bitstream for playback. A user operates the game on the decoder side 200 (e.g., a mobile device).
[0115] A specific implementation of the image encoding method provided in this embodiment will be described in detail below with reference to the accompanying drawings.
[0116] Since the encoder side 100 and the decoder side 200 are located on a cloud server and a client, respectively, after the encoder side 100 compresses the source video to obtain a bitstream, the encoder side 100 transmits the bitstream to the decoder side 200 to improve data transmission efficiency. Therefore, this application provides a video encoding method. FIG. 5 is a schematic flowchart of the video encoding method according to this application. The video encoding method may be applied to the video encoding and decoding system shown in FIG. 1 or the video encoder shown in FIG. 2. For example, the video encoding method may be performed by the encoder side 100 or the video encoder 120. Here, an example in which the encoder side 100 performs the image encoding method provided in this embodiment is used for explanation. As shown in FIG. 5, the image encoding method provided in this embodiment includes the following steps S510 to S540.
[0117] S510: The encoder side 100 acquires a source video and rendering information corresponding to the source video.
[0118] The source video includes multiple source images. For further details of the source video, please refer to the description in the above embodiment. The details will not be described again here.
[0119] The rendering information indicates processing parameters used by the encoder 100 in generating a bitstream based on multiple source images in the source video. The processing parameters may include at least one of the process map data and parameter information described above. Furthermore, the graphics motion vector map in the processing parameters indicates that a first image and a second image among the multiple source images have at least a partially overlapping region, and that pixels in the overlapping region have a positional correspondence.
[0120] S520: The encoder side 100 determines a first type of filter based on the rendering information.
[0121] The first type of filter is one of a plurality of specified types of filters.
[0122] The encoder side 100 may determine the first type of filter from the specified filters based on the correspondence between the processing parameters in the rendering information and the specified filters.
[0123] For example, the rendering information acquired by the encoder side 100 may include one or a combination of multiple processing parameters shown in the following Table 1. Table 1 shows at least one of multiple types of filters corresponding to each processing parameter. [Table 1]
[0124] Anti-aliasing parameters and motion blur parameters are included in the post-processing parameters. A filter corresponding to a priority of 5 is an alternative filter. When the rendering information does not include a depth map, an albedo map, anti-aliasing parameters, or post-processing parameters, a Catmull-Rom filter corresponding to a priority of 5 is directly used as the first type of filter. For example, a higher priority corresponding to a filter indicates a higher degree of compatibility between the filter and the filter used by the image rendering engine. The encoder 100 performs interpolation on the reference frame by using a filter with a higher degree of compatibility to obtain a virtual reference frame. The temporal-spatial correlation between the virtual reference frame and the second image is improved, and the encoder 100 determines a prediction block based on the virtual reference frame. The similarity between the prediction block and the image block corresponding to the second image is improved. This improves the encoding effect, reduces the amount of data obtained through encoding, and improves compression performance.
[0125] Table 1 is merely an example provided in this application and should not be understood as a limitation to this application. In some cases, Table 1 further includes more processing parameters, and the processing parameters may correspond to one or more filters, and the priority corresponding to the filters may be before or after the priority corresponding to the filters shown in Table 1. For example, the processing parameters in Table 1 may further include a normal map, and the filter corresponding to the normal map is a Mitchell-Netravali filter, and the priority of the Mitchell-Netravali filter is before the priority corresponding to the filter used for TAA.
[0126] If possible, the encoder side 100 consults the mapping table in Table 1 based on the rendering information to obtain a filter corresponding to each processing parameter in the rendering information. The encoder side 100 determines a first type filter from all filters corresponding to the rendering information.
[0127] The mapping table indicates the filter corresponding to each processing parameter in the rendering information, and the correspondence between the processing parameter and the filter is determined based on the filter used in each processing process of the image rendering engine. The encoder side 100 consults the mapping table based on the acquired processing parameters to determine a first type of filter. The first type of filter matches the filter used by the image rendering engine, so that using the first type of filter improves the temporal-spatial correlation between the virtual reference frame acquired by the encoder side 100 and the second image, thereby reducing the residual value of the second image determined by the encoder side 100 based on the virtual reference frame. This improves the encoding effect.
[0128] The following provides three possible examples for the encoder side 100 to determine the first type of filter based on the processing parameters included in the rendering information.
[0129] In a first possible example, the rendering information obtained by the encoder side 100 includes only anti-aliasing parameters.
[0130] The encoder side 100 consults Table 1 based on the anti-aliasing parameters to obtain the filter used for TAA corresponding to the anti-aliasing parameters. Because the rendering information only includes the anti-aliasing parameters, the filter used for TAA is directly used as the first type filter.
[0131] In a second possible example, the rendering information obtained by the encoder side 100 includes a depth map and anti-aliasing parameters.
[0132] The encoder side 100 consults Table 1 based on the depth map and the anti-aliasing parameters, and finds that the filter priority corresponding to the depth map and the filter priority corresponding to the anti-aliasing parameters are 1 and 3, respectively, based on the priority of each type of filter in the multiple types of filters shown in Table 1. The encoder side 100 selects the filter with the highest priority, i.e., the "Catmull-Rom filter / Bilinear filter" with a priority of 1. The encoder side 100 determines one filter from the "Catmull-Rom filter / Bilinear filter" as the first type of filter.
[0133] Regarding the process in which the encoder side 100 determines one filter from "Catmull-Rom filter / Bilinear filter" as the first type filter, the following provides a possible implementation method, and the details are not described here.
[0134] The encoder side 100 selects a filter from the mapping table that corresponds to the processing parameters and has the highest priority as a first type filter based on the processing parameters included in the rendering information. The filter with the highest priority also has the highest degree of compatibility with the filter used by the image rendering engine. The encoder side 100 performs interpolation on the reference frame to obtain a virtual reference frame by using the filter with the highest degree of compatibility. The temporal-spatial correlation between the virtual reference frame and the second image is improved, and the residual value corresponding to the second image and obtained by the encoder side 100 based on the virtual reference frame is reduced. This improves the encoding effect, reduces the amount of data obtained through encoding, and improves compression performance.
[0135] In a third possible example, the encoder side 100 obtains the depth map and anti-aliasing parameters contained in the rendering information.
[0136] The encoder side 100 obtains all filters corresponding to the depth map and the anti-aliasing parameters, and separately precodes the second image by using all the filters. Each filter in all the filters obtains a corresponding prediction result. The encoder side 100 selects a target prediction result that satisfies a specified condition from all the prediction results, and uses the filter corresponding to the target prediction result as a first type filter.
[0137] The prediction result indicates coding information obtained by precoding the second image by using the filter, and the coding information includes at least one of a predicted bit rate and a distortion rate of the second image.
[0138] For example, the above precoding is a process of performing inter-prediction on a second image to obtain a predicted block of the second image, obtaining a residual value or a residual block of the second image based on the predicted block, and then performing a conversion process. After performing precoding by using all of the above filters, the encoder side 100 obtains the corresponding predicted bit rate and distortion rate. The above specified condition may be a prediction result in which the distortion rate is minimized when the predicted bit rate reaches a specified value, or a prediction result in which the predicted bit rate is minimized when the distortion rate is less than a specified value.
[0139] The encoder performs precoding by using all filters corresponding to the processing parameters in the rendering information. The encoder filters the prediction results obtained through precoding based on specified conditions and uses the filter corresponding to the target prediction result that satisfies the specified conditions as the first-type filter. As a result, the filter that actually satisfies the preset conditions among all filters is determined as the first-type filter. By using the first-type filter, the similarity between the virtual reference frame obtained by the encoder 100 through interpolation and the second image is maximized. During the process of encoding the second image by the encoder 100, the similarity between the virtual reference frame and the second image is maximized, thereby achieving an optimal balance between the bit rate and distortion rate determined when the encoder 100 encodes the second image based on the virtual reference frame. This improves the encoding compression effect.
[0140] If possible, a first type of filter is pre-configured on the encoder side 100 when the TAA algorithm of the image rendering engine of the encoder side 100 is determined.
[0141] Still referring to Figure 5, the video encoding method provided in this embodiment further includes step S530.
[0142] S530: The encoder side 100 performs interpolation on the first decoded image by using a first type filter to obtain a virtual reference frame of the second image.
[0143] The first decoded image is a reconstructed image obtained by encoding and then decoding a first image of a plurality of images, and the reconstructed image is an image obtained through processing by the filter unit 128 in FIG. 2. The first image is encoded before the second image. The first decoded image includes one or more reference blocks corresponding to the first image. For example, the first decoded image includes M×N reference blocks, where M and N are both integers greater than or equal to 1. For the contents of the obtained reconstructed image, please refer to the description shown in FIG. 2. Details will not be described again here. The obtained virtual reference frame may be stored in the memory 129 in FIG. 2.
[0144] The following provides a possible concrete example. An example in which the encoder side performs interpolation on any reference block included in the first decoded image is used for explanation. Figure 6 is a diagram of a reference block interpolation method according to this application. The method includes the following steps 1 to 3.
[0145] Step 1: The encoder side 100 generates a blank image whose size matches the size of the reference block.
[0146] When the size of the graphic motion vector map is consistent with the size of the first decoded image, the encoder side 100 directly generates a blank image whose size is consistent with the size of the reference block, and divides the graphic motion vector map into M×N images, so that the size of the images obtained through division is consistent with the size of the reference block. When both M and N are 1, the encoder side 100 does not divide the graphic motion vector map. The sizes of the first decoded image, the first image, and the second image are the same.
[0147] When the size of the graphics motion vector map does not match the size of the first decoded image, the encoder side 100 performs interpolation on the graphics motion vector map, so that the size of the graphics motion vector map obtained through the interpolation matches the size of the first decoded image.
[0148] For the implementation manner in which the encoder side 100 performs interpolation on the graphic motion vector map, Fig. 7 provides a possible implementation manner, the details of which will not be described here.
[0149] Step 2: The encoder side 100 performs interpolation on the reference block based on the first correspondence and the divided graphic motion vector map to determine the pixel values of all pixels in the blank image.
[0150] The first correspondence relationship indicates that the first image and the second image indicated by the graphics motion vector map in the rendering information have at least a partially overlapping area, and indicates a positional mapping relationship between pixels in the first image and pixels in the second image in the partially overlapping area.
[0151] Since the first decoded image is obtained by encoding and then decoding the first image, there is a positional correspondence between the pixels in the first decoded image and the pixels in the first image. The size of the blank image obtained in the first step is the same as the size of the reference block. Therefore, it can be understood that the positions of pixels in the blank image correspond to the positions of pixels in the corresponding image block in the second image. Based on the positional mapping relationship between the pixels in the first image and the pixels in the second image, there is a positional correspondence between the pixels in the reference block and the pixels in the blank image.
[0152] By using a first type of filter, the encoder side 100 performs interpolation filtering on each pixel in the reference block based on the correspondence between the pixel positions in the reference block and the pixel positions in the blank image to determine the pixel value corresponding to each pixel position in the blank image, and further obtains the pixel values of all pixels in the virtual reference block.
[0153] For example, the position mapping relationship between pixel A in the reference block and pixel A' in the blank image in the first correspondence relationship is used as an example for explanation. As shown in FIG. 6, the encoder side 100 obtains the correspondence relationship between the position of a pixel in the reference block and the position of a pixel in the blank image based on the position mapping relationship between the pixel in the first image and the pixel in the second image. For example, there is a position correspondence relationship between pixel A in the reference block and pixel A' in the blank image. The encoder side 100 performs interpolation filtering on pixel A, or pixel A and pixels surrounding pixel A, by using a first type of filter to obtain a corresponding pixel value. Then, the encoder side 100 fills in the pixel value into pixel A' in the blank image to obtain the color of pixel A'.
[0154] The encoder side 100 calculates the pixel value of each pixel in the blank image according to the above processing steps of pixel A' to obtain the pixel values of all pixels in the blank image.
[0155] Step 3: The encoder side 100 obtains a virtual reference frame based on the pixel values of all pixels in the blank image.
[0156] The encoder side 100 fills in the pixel values of all pixels in the blank image with corresponding pixels in the blank image to obtain a virtual reference block. The encoder side 100 performs the processes in steps 1 and 2 described above for each reference block included in the first decoded image to obtain pixel values of all pixels in the blank image corresponding to each reference block. In this way, a corresponding virtual reference block is obtained. The encoder side 100 integrates the virtual reference blocks to obtain a virtual reference frame. When the encoder side processes each reference block included in the first decoded image in steps 1 and 2 described above, different first type filters may be used. For example, different first type filters are used for different reference blocks.
[0157] The encoder performs interpolation on the reference blocks in the first decoded image based on the first correspondence to obtain a virtual reference frame for the second image. The similarity between the virtual reference frame and the second image is higher than the similarity between the first decoded image and the second image. Therefore, the residual value between the virtual reference frame and the second image is reduced, ensuring image quality consistency and reducing the amount of image data in the bitstream. This improves the coding effect.
[0158] Still referring to Figure 5, the video encoding method provided in this embodiment further includes step S540.
[0159] S540: The encoder side 100 encodes the second image based on the virtual reference frame of the second image to obtain a bitstream corresponding to the second image.
[0160] The encoder 100 performs inter prediction based on a virtual reference frame of the second image to determine a predicted frame corresponding to the second image. The predicted frame includes a plurality of predicted blocks. The encoder 100 obtains corresponding residual values (also called residual blocks) based on the image blocks in the second image and the corresponding predicted blocks, and determines motion vectors (MVs) for the image blocks in the second image and the corresponding predicted blocks. The encoder 100 encodes the residual values and motion vectors to encode the second image and obtain a corresponding bitstream.
[0161] For the process by which the encoder side 100 determines multiple prediction blocks corresponding to the second image, this application provides the following possible examples.
[0162] In a possible example, the encoder side 100 determines, from the memory 129, a reference block having the highest degree of match with an image block in the second image as a predicted block of the second image by using the inter predictor 121 in the video encoder 120 shown in FIG.
[0163] The reference block with the best fit comprises the hypothetical reference block.
[0164] For example, the encoder side 100 may perform an image block search according to a full search (FS) method or a fast search method, and then calculate the matching degree between the searched reference block and the image block in the second image according to a method such as a mean-square error (MSE) method and a mean absolute deviation (MAD) method. In this way, the reference block with the best matching degree is obtained.
[0165] For the encoder side 100 to determine the motion vectors of the image blocks in the second image and the corresponding prediction blocks, this application provides the following possible examples.
[0166] In a possible example, the encoder side 100 determines the relative displacement between the image block and the prediction block based on the position of the image block in the second image and the position of the prediction block in the corresponding reference frame, and the relative displacement is the motion vector of the image block.For example, if the motion vector is (1,1), this indicates that the position of the image block in the second image is moved by (1,1) relative to the position of the prediction block in the corresponding reference frame.
[0167] If possible, the encoder side 100 may transmit type information of the first type of filter to the decoder side 200. The type information may indicate, for example, a Catmull-Rom filter, a bilinear filter, or a B-spline filter.
[0168] Example 1: The encoder side 100 may write type information of a first type of filter into a bitstream and send the bitstream to the decoder side 200.
[0169] Example 2: The encoder side 100 may transmit type information of the first type filter as rendering information to the decoder side 200. For example, the encoder side 100 compresses the rendering information including the type information of the first type filter, and then transmits the compressed rendering information to the decoder side 200.
[0170] In this embodiment, the encoder side determines a first type filter from a plurality of specified type filters, and the first type filter is compatible with a filter used in the image rendering process. This helps improve the similarity between the virtual reference frame determined by the encoder side and the second image by using the first type filter, compared to the problem of low compression performance due to low similarity between the reference frame and the second image when a fixed filter is used to process the reconstructed image to obtain the reference frame. Therefore, when the encoder side encodes the second image based on the virtual reference frame, the residual value corresponding to the second image is reduced, and the amount of data corresponding to the second image in the bitstream is reduced. This improves compression performance and video coding efficiency.
[0171] For the above case where the encoder side 100 optionally determines one filter from "Catmull-Rom filter / Bilinear filter" as the first type filter, the following provides a possible implementation scheme.
[0172] When the rendering parameters include a depth map, the encoder side 100 determines a first type of filter from "Catmull-Rom filter / Bilinear filter" based on the correspondence between the edge information of the object in the depth map and the reference block in the first decoded image corresponding to the depth map.
[0173] For example, the encoder 100 first performs Sobel filtering on the depth map to obtain the filtering result. Then, the encoder 100 performs Otsu binarization on the filtering result to obtain edge information of the object in the depth map. Finally, the encoder 100 determines a first type filter corresponding to each reference block from "Catmull-Rom filter / Bilinear filter" based on the correspondence between the edge information of the object and each reference block in the first decoded image or an image block in the depth map that corresponds to the reference block. For example, when the position of a pixel in the reference block in the first decoded image is the position indicated by the edge information, a Catmull-Rom filter is used as the first type filter for the reference block. When the position of a pixel in the reference block in the first decoded image is not the position indicated by the edge information, a bilinear filter is used as the first type filter for the reference block.
[0174] Optionally, when the processing parameter with the highest priority among the determined rendering parameters is an albedo map and the filter is a "Catmull-Rom filter / B-spline filter", the filter is determined as a first type filter from the "Catmull-Rom filter / B-spline filter". Possible examples are provided below.
[0175] The encoder 100 first performs Sobel filtering on the albedo map to obtain the filtering result. Then, the encoder 100 performs Otsu binarization on the filtering result to obtain edge information of the object in the albedo map. Finally, the encoder 100 determines a first type filter corresponding to each reference block from the "Catmull-Rom filter / B-spline filter" selection based on the correspondence between the edge information of the object and each reference block in the first decoded image or an image block in the depth map that corresponds to the reference block. For example, when the position of a pixel in a reference block in the first decoded image is at the position indicated by the edge information, a Catmull-Rom filter is used as the first type filter for the reference block. When the position of a pixel in a reference block in the first decoded image is not at the position indicated by the edge information, a B-spline filter is used as the first type filter for the reference block.
[0176] The encoder side performs interpolation on different reference blocks by using different first-type filters based on different positional relationships between the edge information of the object and the reference blocks, thereby improving the similarity between the virtual reference block obtained by the encoder side through interpolation and the corresponding image block in the second image. When the encoder side encodes the image block based on the virtual reference block, the residual value corresponding to the image block is reduced, which improves compression performance and video coding efficiency.
[0177] For the above case where the encoder side 100 performs interpolation on the graphics motion vector map when the size of the graphics motion vector map does not match the size of the first decoded image, the following provides a possible implementation scheme.
[0178] The encoder 100 first generates a blank image whose size matches that of the source image or the graphics motion vector map. Then, the encoder 100 determines a second correspondence between pixel positions in the first graphics motion vector map and pixel positions in the blank image based on the size ratio of the first graphics motion vector map to the blank image. Finally, the encoder 100 performs interpolation on the graphics motion vector map or the first decoded image using an interpolation filter to determine pixel values and offsets for each pixel in the blank image. When the size of the graphics motion vector map obtained through interpolation matches that of the first decoded image, a first correspondence indicated by the graphics motion vector map obtained through interpolation is obtained. For an explanation of the first correspondence, refer to step 2 in FIG. 6. The size of the source image matches that of the first decoded image.
[0179] For example, Figure 7 is a diagram of a graphics motion vector interpolation method according to this application. An example is used for explanation, in which the first graphics motion vector map is the original graphics motion vector map output by the image rendering engine, and the size of the first graphics motion vector map is less than the size of the first decoded image.
[0180] The size of the first graphics motion vector map is 4x4, and the size of the first decoded image is 8x8. The encoder side 100 generates a blank image whose size matches the size of the first decoded image, for example, an 8x8 blank image. The encoder side 100 determines that the size ratio of the first graphics motion vector map to the blank image is 1:4, that is, one pixel in the first graphics motion vector map corresponds to four pixels in the blank image. As shown in Figure 6, when the coordinates of pixel C in the first graphics motion vector map are (1,1) and pixel C has an offset (1,1), the four corresponding pixels in the blank image are C(1,1), C1(1,2), C2(2,1), and C3(2,2). The encoder side 100 uses an interpolation filter to perform interpolation on the first graphics motion vector map based on a second correspondence between the position of pixel C in the first graphics motion vector map and the positions of pixels C, C1, C2, and C3 in the blank image to obtain pixel values and offsets of pixels C, C1, C2, and C3 in the second graphics motion vector map. For example, all offsets of pixels in the blank image are (2,2). The encoder side 100 may obtain the first correspondence based on the offsets.
[0181] For the above interpolation filters, the application provides one of several possible examples, such as a bilinear filter, a nearest filter, and a bicubic filter.
[0182] When the graphics motion vectors do not match the source image, the encoder performs interpolation on the graphics motion vector map to obtain a graphics motion vector map obtained through interpolation. The size of the graphics motion vector map obtained through interpolation matches the size of the first decoded image, and the pixels of the graphics motion vector map obtained through interpolation match the pixels of the first decoded image. The encoder performs interpolation on the first decoded image based on the graphics motion vector map obtained through interpolation to obtain a virtual reference frame. This improves the fit between the virtual reference frame and the second image. When the encoder encodes the second image based on the virtual reference frame, the residual value corresponding to the image is reduced, improving compression performance.
[0183] Optionally, when the first type filter determined in step S520 is a filter for which a filter kernel can be specified, the encoder side 100 further determines the size of the filter kernel used by the filter. The following provides a possible implementation scheme.
[0184] The encoder side 100 first determines target discrete parameters for pixels in the first decoded image based on the graphics motion vector map indicated in the rendering information.
[0185] In a possible example, when the size of the graphics motion vector map in the rendering information does not match the size of the first decoded image, the encoder side 100 extracts image blocks in the overlapping area in the first decoded image based on an indication that the first image and the second image have at least a partially overlapping area indicated by the graphics motion vector map obtained through interpolation, and calculates the discrete parameters of the pixels in the image blocks on each channel as target discrete parameters.
[0186] In another possible example, when the size of the graphics motion vector map in the rendering information matches the size of the first decoded image, the encoder side 100 calculates the discrete parameters of the pixels in the current reference block in the first decoded image on each channel as the target discrete parameters. The current reference block may indicate the reference block for which interpolation is being performed.
[0187] Each channel of a pixel includes three color channels: a luminance channel (Y), a blue density offset channel (Cb), and a red density offset channel (Cr). The target discrete parameters include a target variance and a target covariance.
[0188] Second, the first type filter on the encoder side 100 performs interpolation on the reference block by using a first filter kernel and calculates first discrete parameters of pixels in the virtual reference block obtained through interpolation for each channel.
[0189] The first filter kernel is one of a plurality of specified filter kernels, the plurality of filter kernels having different sizes.
[0190] For example, the first filter kernel is a 3x3 filter kernel, and weights are specified for the 3x3 filter kernel. The first type filter performs interpolation filtering on the reference block by using the 3x3 filter kernel to obtain pixel values of each pixel in the virtual reference block. The encoder side 100 calculates a first variance or a first covariance of the pixels in the virtual reference block on each channel based on the pixel values of each pixel in the virtual reference block to obtain first discrete parameters of the virtual reference block.
[0191] The weights in the filter kernel may be obtained based on a graphics motion vector map. For example, the graphics motion vector map indicates a positional mapping relationship between pixels in a first image and pixels in a second image, and the mapping relationship is not necessarily an integer-pixel mapping relationship. After an object in the first image moves, a second image is acquired. A pixel located at a first position in the first image is offset to a second position in the second image, and the second position may be an integer-pixel position or may be between multiple pixel positions in the second image. If the distance value corresponding to the second position is 0.5, this indicates that the second position is located at the center of two pixel positions in the second image. The encoder side 100 determines the weights in the filter kernel based on the pixel position mapping relationship.
[0192] Third, the encoder side 100 determines a first discrete parameter having a minimum difference from a target discrete parameter from at least one first discrete parameter corresponding to all the filter kernels, and uses the filter kernel corresponding to the first discrete parameter as a parameter of a first type filter.
[0193] The encoder 100 determines first discrete parameters that have the smallest difference from the target discrete parameters, and uses a first filter kernel corresponding to the first discrete parameters as the parameters of a first type of filter. The encoder performs interpolation on the first decoded image by using the first type of filter based on the parameters of the first type of filter to obtain a virtual reference frame. The fit between the virtual reference frame and the second image is improved. When the encoder encodes the second image based on the virtual reference frame, the residual value corresponding to the second image is reduced, and the amount of image data in the bitstream is reduced. This improves compression performance.
[0194] After the encoder side 100 encodes the source video or one or more source images in the source video according to the above video encoding method, the encoder side 100 obtains a bitstream corresponding to the source video or the source image. The encoder side 100 transmits the bitstream to the decoder side 200, and the decoder side 200 decodes the bitstream for playback. FIG. 8 is a schematic flowchart of a video decoding method according to this application. The video decoding method may be applied to the video encoding and decoding system shown in FIG. 1 or the video decoder shown in FIG. 3. For example, the video decoding method may be performed by the decoder side 200 or the video decoder 220. Here, an example in which the decoder side 200 performs the image decoding method provided in this embodiment is used for explanation. As shown in FIG. 8, the image encoding method provided in this embodiment includes the following steps S810 to S840.
[0195] S810: The decoder side 200 obtains a bitstream and rendering information corresponding to the bitstream.
[0196] The rendering information indicates processing parameters used by the decoder side 200 in the process of generating a bitstream based on multiple source images. For details of the processing parameters, please refer to the description of the processing parameters at S510 in Figure 5. Details will not be described again here. The graphics motion vector map in the processing parameters indicates that a source image corresponding to a first image frame and a source image corresponding to a second image frame among the multiple image frames have at least a partial overlapping area.
[0197] For example, the rendering information may be sent by the encoder side 100 to the decoder side 200 .
[0198] In a possible example, the rendering information obtained by the decoder side 200 includes one or a combination of the processing parameters in the mapping table shown in Table 1.
[0199] S820: The decoder side 200 determines a first type of filter based on the rendering information.
[0200] The first type of filter is one of a plurality of specified types of filters.
[0201] For example, the decoder side 200 consults Table 1 based on the processing parameters indicated by the rendering information to obtain a first type of filter corresponding to the processing parameters. Table 1 indicates at least one type of filter among multiple types of filters corresponding to the processing parameters, and each type of filter has a corresponding priority.
[0202] Regarding the content of the decoder side 200 determining the first type of filter based on the rendering information, refer to the example in which the encoder side 100 determines the first type of filter based on the processing parameters included in the rendering information, and the details will not be described again here.
[0203] If possible, the decoder side 200 may consult a pre-defined mapping table based on the rendering information to obtain a filter corresponding to each processing parameter in the rendering information. The decoder side 200 determines a first type filter from all filters corresponding to the rendering information.
[0204] For the relevant contents of the mapping table, please refer to the description of determining the first type filter in S520 in Fig. 5 and Table 1. The details will not be described again here.
[0205] If possible, the decoder side 200 obtains filter information, and determines a first type of filter and filtering parameters corresponding to the first type of filter based on the filter type and parameters indicated in the filter information.
[0206] The filtering parameters indicate the processing parameters used in the process of the decoder side 200 obtaining the virtual reference frame by using a first type of filter, such as the size of the filter kernel and the corresponding weight of the filter.
[0207] For example, the rendering information or bitstream may include filter information.
[0208] In a possible example, a first type of filter is pre-configured on the decoder side 200 when the TAA algorithm of the image rendering engine on the encoder side 100 is determined.
[0209] S830: The decoder side 200 performs interpolation on a first decoded image corresponding to a first image frame by using a first type filter to obtain a virtual reference frame of a second image frame.
[0210] For example, the decoder side 200 first generates a blank image whose size matches that of the reference block. Then, the decoder side 200 performs interpolation on the reference block based on the first correspondence and the divided graphic motion vector map to determine pixel values of all pixels in the blank image. Finally, the decoder side 200 obtains a virtual reference frame based on the pixel values of all pixels in the blank image. The first correspondence is determined based on the graphic motion vector map in the rendering information. The reference block is a first decoded image, and the first decoded image is any one of multiple reference blocks corresponding to the first image.
[0211] For a detailed description of the above example, please refer to the contents shown in Figure 6. The details will not be described again here.
[0212] The encoder side performs interpolation on the reference blocks in the first decoded image based on the first correspondence relationship to obtain a virtual reference frame of the second image. The similarity between the virtual reference frame and the second image frame is higher than the similarity between the first decoded image and the second image frame. When the decoder side 200 decodes the second image frame based on the virtual reference frame, the amount of data processed is reduced and decoding efficiency is improved.
[0213] S840: The decoder side 200 decodes the second image frame based on the virtual reference frame to obtain a second decoded image corresponding to the second image frame.
[0214] The decoder side 200 performs processing such as inverse quantization and inverse transform on the second image frame to obtain residual blocks, motion vectors, etc. corresponding to the second image frame. The decoder side 200 performs image reconstruction based on data such as residual blocks and motion vectors and the virtual reference frame determined in S830 to obtain a second decoded image corresponding to the second image frame.
[0215] For example, the decoder side 200 processes a second image frame in the bitstream by using an entropy decoder 221, an inverse quantizer 222, and an inverse transformer 223 in a video decoder 220 shown in Fig. 3 to obtain a residual block between the second image frame and a corresponding virtual reference frame. The video decoder 220 performs image block reconstruction based on the residual block, the motion vector, and the virtual reference frame. The video decoder 220 processes a plurality of reconstructed image blocks by using a filter unit 224 to obtain a second decoded image corresponding to the second image frame.
[0216] The decoder side 200 selects a first type of filter from a plurality of specified type filters based on the processing parameters indicated by the rendering information, and obtains a virtual reference frame for the second image frame by using the first type of filter, thereby avoiding the problem of poor decoding efficiency due to the reference frame being obtained by processing the first decoded image using a fixed filter. The degree of match between the virtual reference frame and the source image corresponding to the second image frame is improved. When the decoder side 200 decodes the second image frame based on the virtual reference frame and the processing capability of the decoder side is consistent, the amount of data processed by the decoder side is reduced, and decoding efficiency is improved.
[0217] When the size of the graphics motion vector map obtained by the decoder side 200 does not match the size of the first decoded image, the decoder side 200 performs interpolation on the graphics motion vector map. For example, the decoder side 200 first generates a blank image whose size matches the size of the first decoded image. Then, the decoder side 200 determines a second correspondence between the positions of pixels in the first graphics motion vector map and the positions of pixels in the blank image based on the size ratio of the first graphics motion vector map to the blank image. Finally, the decoder side 200 performs interpolation on the graphics motion vector map by using an interpolation filter to determine the pixel value and offset of each pixel in the blank image.
[0218] The size of the first decoded image matches the size of the source image.
[0219] For a detailed description of the above example, please refer to the contents shown in Figure 7. The details will not be described again here.
[0220] When the graphics motion vector does not match the first decoded image, the decoder side performs interpolation on the graphics motion vector map to obtain a graphics motion vector map obtained through interpolation. The size of the graphics motion vector map obtained through interpolation matches the size of the first decoded image, and the graphics motion vector map obtained through interpolation matches the first decoded image. The decoder side performs interpolation on the first decoded image based on the graphics motion vector map obtained through interpolation to obtain a virtual reference frame. This improves the degree of match between the virtual reference frame and the second image frame. When the decoder side decodes the second image frame based on the virtual reference frame, the amount of data processed is reduced, which improves decoding efficiency.
[0221] Optionally, when the first type of filter determined in step S820 is a filter that can specify a filter kernel, the decoder side 200 further determines the size of the filter kernel used by the filter. The following provides a possible implementation scheme.
[0222] The decoder side 200 first determines the target discrete parameters of the pixels in the first decoded image based on the graphics motion vector map indicated in the rendering information.
[0223] Second, the first type filter 200 on the decoder side performs interpolation on the reference block by using a first filter kernel, and calculates first discrete parameters of pixels in the virtual reference block obtained through interpolation for each channel.
[0224] Third, the decoder side 200 determines a first discrete parameter having the smallest difference from the target discrete parameter from at least one first discrete parameter corresponding to all the filter kernels, and the decoder side 200 uses the filter kernel corresponding to the first discrete parameter as a parameter of a first type filter.
[0225] The process of determining the filter kernel at the decoder side 200 refers to the process of determining the size of the filter kernel at the encoder side 100, and the details will not be described again here.
[0226] The decoder side 200 determines the first discrete parameters that have the smallest difference from the target discrete parameters, and uses the first filter kernel corresponding to the first discrete parameters as the parameters of the first type filter. The decoder side 200 performs interpolation on the first decoded image by using the first type filter based on the parameters of the first type filter to obtain a virtual reference frame. The fit between the virtual reference frame and the second image frame is improved. When the decoder side 200 decodes the second image frame based on the virtual reference frame, the amount of data processed is reduced. This improves decoding efficiency.
[0227] Above, the video encoding method provided in this application is described in detail with reference to Figures 4 to 7. Below, the video encoding device provided in this application is described with reference to Figure 9. Figure 9 is a structural diagram of the video encoding device according to this application. The video encoding device 900 may be configured to implement the encoder-side functions in the above method embodiments, and thus can also realize the beneficial effects of the above method embodiments.
[0228] As shown in Figure 9, the video encoding device 900 includes a first acquisition module 910, a first interpolation module 920, and an encoding module 930. The video encoding device 900 is configured to implement the encoder side functions in the method embodiments corresponding to Figures 4 to 7. In a possible example, the specific process by which the video encoding device 900 is configured to implement the above video encoding method includes the following processes:
[0229] The first acquisition module 910 is configured to acquire a source video and rendering information corresponding to the source video, the source video including a plurality of source images, the rendering information indicating processing parameters to be used in a process of generating a bitstream based on the plurality of source images, the processing parameters indicating that a first image and a second image of the plurality of source images have at least a partial overlapping region.
[0230] The first interpolation module 920 is configured to determine a first type filter based on the rendering information, and to perform interpolation on the first decoded image by using the first type filter to obtain a virtual reference frame for the second image, where the first type filter is one of a plurality of specified types of filters, and the first decoded image is a reconstructed image obtained by encoding and then decoding a first image of the plurality of images, where the first image is encoded before the second image.
[0231] The encoding module 930 is configured to encode the second image based on the virtual reference frame of the second image.
[0232] To further realize the functions in the method embodiments shown in Figures 4 to 7, this application further provides a video encoding device 900. The video encoding device 900 further includes a first filter kernel determination module 940.
[0233] The first filter kernel determination module 940 is configured to determine target discrete parameters for pixels in the first decoded image based on the rendering information, the target discrete parameters including a target variance or a target covariance, and to use a first filter kernel corresponding to a first discrete parameter having a smallest difference from the target discrete parameters as a parameter of a first type filter, the first discrete parameters indicating discrete parameters obtained by the first type filter by performing interpolation on the pixels based on the first filter kernel and performing calculations based on the pixels obtained through the interpolation, the first filter kernel being one of a plurality of specified filter kernels, and the first discrete parameters including a first variance or a first covariance.
[0234] It should be understood that the encoder side in the above embodiment may correspond to the video encoding device 900, and may correspond to the corresponding entities for performing the methods in Figures 4 to 7 according to the embodiments of this application. Furthermore, the operations and / or functions of the modules in the video encoding device 900 are separately used to realize the corresponding procedures of the methods in the corresponding embodiments of Figures 4 to 7. For the sake of brevity, the details will not be described again here.
[0235] Above, the video decoding method provided in this application has been described in detail with reference to Figures 3 to 8. Below, the video decoding device provided in this application will be described with reference to Figure 10. Figure 10 is a structural diagram of the video decoding device according to this application. The video decoding device 1000 may be configured to implement the decoder-side functions in the above method embodiments, and thus can also realize the beneficial effects of the above method embodiments.
[0236] As shown in Figure 10, the video decoding device 1000 includes a second acquisition module 1010, a second interpolation module 1020, and a decoding module 1030. The video decoding device 1000 is configured to implement the decoder-side functions in the method embodiments corresponding to Figures 3 and 8. In a possible example, the specific process by which the video decoding device 1000 is configured to implement the above video decoding method includes the following processes:
[0237] The second acquisition module 1010 is configured to acquire a bitstream and rendering information corresponding to the bitstream, the bitstream including a plurality of image frames, the rendering information indicating processing parameters used in a process of generating the bitstream based on a plurality of source images, the processing parameters indicating that a source image corresponding to a first image frame and a source image corresponding to a second image frame of the plurality of image frames have at least a partial overlapping region.
[0238] The second interpolation module 1020 is configured to determine a first type filter based on the rendering information, and perform interpolation on a first decoded image corresponding to the first image frame by using the first type filter to obtain a virtual reference frame of the second image frame. The first type filter is one of a plurality of specified types of filters.
[0239] The decoding module 1030 is configured to decode the second image frame based on the virtual reference frame to obtain a second decoded image corresponding to the second image frame.
[0240] To further realize the functions in the method embodiments shown in Figures 3 and 8, this application further provides a video decoding device. The video decoding device 1000 further includes a filter information acquisition module 1040 and a second filter kernel determination module 1050.
[0241] The filter information acquisition module 1040 is configured to acquire filter information, and determine a first type filter and filtering parameters corresponding to the first type filter based on a filter type and parameters indicated in the filter information, where the filtering parameters indicate processing parameters used in the process of acquiring a virtual reference frame by using the first type filter.
[0242] The second filter kernel determination module 1050 is configured to determine target discrete parameters for pixels in the first decoded image based on the rendering information, the target discrete parameters including a target variance or a target covariance, and to use a first filter kernel corresponding to a first discrete parameter having a smallest difference from the target discrete parameters as a parameter of a first type filter, the first discrete parameters indicating discrete parameters obtained by the first type filter by performing interpolation on the pixels based on the first filter kernel and performing calculations based on the pixels obtained through the interpolation, the first filter kernel being one of a plurality of specified filter kernels, and the first discrete parameters including a first variance or a first covariance.
[0243] It should be understood that the decoder side according to the embodiments of this application may correspond to the video decoding device 1000 in the embodiments of this application, and may correspond to the corresponding entities for performing the methods according to the embodiments of this application in Figures 3 and 8. Furthermore, the operations and / or functions of the modules in the video decoding device 1000 are separately used to realize the corresponding procedures of the methods in the corresponding embodiments in Figures 3 and 8. For the sake of brevity, the details will not be described again here.
[0244] An embodiment of this application provides a computer device. Figure 11 is a diagram of the structure of the computer device according to this application. The computer device may be used in the encoding and decoding system shown in Figure 1, and the computer device may be either the encoder side 100 or the decoder side 200.
[0245] The computer device 1100 may specifically be a computing device having computing capabilities, such as a mobile phone, a tablet computer, a television (which may also be called a smart television, smart screen, or large screen device), a notebook computer, an ultra-mobile personal computer (UMPC), a handheld computer, a netbook, a personal digital assistant (PDA), a wearable electronic device (e.g., a smart watch, smart band, or smart glasses), an in-vehicle device, a virtual reality device, or a server.
[0246] 11, the processing device 1100 may include a processor 1110, a memory 1120, a communication interface 1130, a bus 1140, etc. The processor 1110, the memory 1120, and the communication interface 1130 are connected via the bus 1140.
[0247] It can be understood that the structures shown in the embodiments of the present invention do not constitute specific limitations on the processing device. In some other embodiments of this application, the processing device may include more or fewer components than those shown in the drawings, or some components may be combined, or some components may be divided, or a different component arrangement may be used. The components shown in the drawings may be realized by hardware, software, or a combination of software and hardware.
[0248] The processor 1110 may include one or more processing units. For example, the processor 1110 may include an application processor (AP), a modem processor, a central processing unit (CPU), a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, a neural-network processing unit (NPU), etc. The different processing units may be separate devices or may be integrated into one or more processors.
[0249] An internal memory may also be located in the processor 1110 and configured to store instructions and data. In some embodiments, the internal memory in the processor 1110 is a cache memory. The internal memory may store instructions or data that have just been used or that are periodically used by the processor 1110. When the processor 1110 needs to use the instructions or data again, the processor may retrieve the instructions or data directly from the internal memory. This avoids repeated accesses, reduces the latency of the processor 1110, and improves system efficiency.
[0250] The memory 1120 may be configured to store computer-executable program code. Executable program code includes instructions. The processor 1110 executes the instructions stored in the internal memory 1120 to perform various functional applications and data processing of the processing device 1100. The internal memory 1120 may include a program storage area and a data storage area. The program storage area may store an operating system, applications required by at least one function (e.g., a coding function or a transmission function), etc. The data storage area may store data (e.g., bitstreams and reference frames) created in processes using the processing device 1100, etc. Furthermore, the internal memory 1120 may include high-speed random access memory or non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or universal flash storage (UFS).
[0251] The communication interface 1130 is configured to facilitate communication between the processing device 1100 and external devices or components. In this embodiment, the communication interface 1130 is configured to exchange data with other processing devices.
[0252] The bus 1140 may include paths configured to transmit information between the above components (e.g., the processor 1110, the memory 1120, and the communication interface 1130). In addition to a data bus, the bus 1140 may further include a power bus, a control bus, a status signal bus, etc. However, for clarity of explanation, various buses are illustrated in the figures as the bus 1140. The bus 1140 may be a peripheral component interconnect express (PCIe) high-speed bus, an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a compute express link (CXL), a cache coherent interconnect for accelerators (CCIX), etc.
[0253] 11, it should be noted that the processing device 1100 includes only one processor 1110 and one memory 1120 as an example. Here, the processor 1110 and the memory 1120 respectively represent types of components or devices. In a particular embodiment, the number of components or devices of each type may be determined based on service requirements.
[0254] When the video encoding (or video decoding) device is realized by hardware, the hardware may be realized by using a processor or a chip. In the following, an example in which the hardware is a chip is used for explanation. The chip includes a processor configured to realize the functions of the encoder side and / or the decoder side in the above method. In a possible design, the chip further includes a power supply circuit configured to provide power to the processor. The chip may directly include a chip, or may include a chip and other discrete components.
[0255] All or part of the above embodiments may be realized by using software, hardware, firmware, or any combination thereof. When software is used to realize the above embodiments, all or part of the embodiments may be realized in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer programs or instructions are loaded and executed on a computer, all or part of the procedures or functions in the embodiments of this application are executed. The computer may be a general-purpose computer, a special-purpose computer, a computer network, a network device, user equipment, or other programmable device. The computer program or instructions may be stored in a computer-readable storage medium or transmitted from a computer-readable storage medium to another computer-readable storage medium. For example, the computer program or instructions may be transmitted from a website, computer, server, or data center to another website, computer, server, or data center via a wired or wireless method. The computer-readable storage medium may be any available medium accessible by a computer, or a data storage device such as a server or data center that integrates one or more available media. The media available may be magnetic media such as floppy disks, hard disks or magnetic tapes, optical media such as digital video discs (DVDs), or semiconductor media such as solid-state drives (SSDs).
[0256] The above description is merely a specific embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications or replacements that can be easily conceived by those skilled in the art within the technical scope disclosed in this application shall fall within the scope of protection of this application. Therefore, the scope of protection of this application shall be subject to the scope of protection of the claims.
Claims
1. A video decoding method performed by a decoder side, comprising: obtaining a bitstream and rendering information corresponding to the bitstream, the bitstream including a plurality of image frames, the rendering information indicating processing parameters used in a process of generating the bitstream based on a plurality of source images, the processing parameters indicating that a source image corresponding to a first image frame and a source image corresponding to a second image frame of the plurality of image frames have at least a partial overlapping region; performing interpolation on a first decoded image corresponding to the first image frame to obtain a virtual reference frame of the second image frame by using a first type filter determined based on the rendering information, wherein the first type filter is one of a plurality of specified type filters; decoding the second image frame based on the virtual reference frame to obtain a second decoded image corresponding to the second image frame; A method comprising:
2. The first type filter determined based on the rendering information includes: Querying a mapping table specified based on the rendering information to obtain a filter corresponding to each processing parameter in the rendering information, the mapping table indicating at least one type of filter corresponding to each processing parameter among the plurality of types of filters; determining the first type of filter from all filters corresponding to the rendering information; The method of claim 1 , comprising:
3. the mapping table further indicates a priority of each type of filter among the plurality of types of filters; Determining the first type of filter from all filters corresponding to the rendering information includes: obtaining priorities of all the filters corresponding to the rendering information; using the filter having the highest filter priority among all of said filters as said first type filter; The method of claim 2 , comprising:
4. The method of claim 1 , wherein the rendering information comprises one or a combination of a depth map, an albedo map, and post-processing parameters.
5. If the rendering information includes the depth map, the first type of filter is determined based on a relationship between pixels in the depth map and edge information of objects in the depth map, or If the rendering information includes the albedo map, the first type of filter is determined based on a relationship between pixels in the albedo map and edge information of objects in the albedo map; or If the post-processing parameters included in the rendering information include an anti-aliasing parameter, the first type filter is a filter indicated by the anti-aliasing parameter, or if the post-processing parameters included in the rendering information include instructions for the execution of a motion blur module, the first type of filter is a bilinear filter, or The method of claim 4 , wherein the first type of filter is a Catmull-Rom filter.
6. before performing interpolation on a first decoded image corresponding to the first image frame by using a first type filter determined based on the rendering information; obtaining filter information; determining the first type filter and filtering parameters corresponding to the first type filter based on the filter type and parameters indicated in the filter information, the filtering parameters indicating processing parameters used in the process of obtaining the virtual reference frame by using the first type filter; 6. The method of claim 1, comprising:
7. performing interpolation on a first decoded image corresponding to the first image frame by using a first type filter determined based on the rendering information to obtain a virtual reference frame for the second image frame, performing interpolation on the first decoded image based on a first correspondence relationship to obtain pixel values of all pixels in the virtual reference frame, the first correspondence relationship indicating a positional mapping relationship between pixels in the source image corresponding to the first image frame and pixels in the source image corresponding to the second image frame in the overlapping region between the source image corresponding to the first image frame and the source image corresponding to the second image frame; obtaining the virtual reference frame based on the pixel values of all of the pixels; 7. The method of claim 1, comprising:
8. The first correspondence relationship is determined in the following manner: When the size of the graphics motion vector map in the rendering information does not match the size of the source image, creating a blank image whose size matches the size of the source image; determining a second correspondence between pixel coordinates in the graphics motion vector map and pixel coordinates in the blank image; performing interpolation on the graphics motion vector map based on the second correspondence by using a specified filter to obtain the first correspondence; The method of claim 7, wherein the signal is obtained by using
9. determining target discrete parameters for pixels in the first decoded image based on the rendering information, the target discrete parameters including a target variance or a target covariance; a step of using a first filter kernel corresponding to a first discrete parameter having a minimum difference from the target discrete parameter as a parameter of the first type filter, the first discrete parameter indicating a discrete parameter obtained by the first type filter by performing interpolation on the pixel based on the first filter kernel and performing calculation based on the pixel obtained through interpolation, the first filter kernel being one of a plurality of specified filter kernels, and the first discrete parameter including a first variance or a first covariance; The method of claim 7 or 8, further comprising:
10. A video encoding method performed by an encoder, comprising: obtaining a plurality of source images and rendering information corresponding to the plurality of source images, the rendering information indicating processing parameters to be used in a process of generating a bitstream based on the plurality of source images, the processing parameters indicating that a first image and a second image of the plurality of source images have at least a partially overlapping region; performing interpolation on a first decoded image to obtain a virtual reference frame for the second image by using a first type filter determined based on the rendering information, the first type filter being one of a plurality of specified types of filters, and the first decoded image being a reconstructed image obtained by encoding and then decoding the first image of the plurality of images; encoding the second image based on the virtual reference frame to obtain a bitstream corresponding to the second image; A method comprising:
11. The first type filter determined based on the rendering information includes: Querying a mapping table specified based on the rendering information to obtain a filter corresponding to each processing parameter in the rendering information, the mapping table indicating at least one type of filter corresponding to each processing parameter among the plurality of types of filters; determining the first type of filter from all filters corresponding to the rendering information; The method of claim 10, comprising:
12. the mapping table further indicates a priority of each type of filter among the plurality of types of filters; Determining the first type of filter from all filters corresponding to the rendering information includes: obtaining priorities of all the filters corresponding to the rendering information; using the filter having the highest filter priority among all of said filters as said first type filter; The method of claim 11 , comprising:
13. Determining the first type of filter from all filters corresponding to the rendering information includes: performing prediction for all the filters corresponding to the rendering information to obtain at least one prediction result, where one prediction result corresponds to one of all the filters, and the prediction result indicates coding information obtained by precoding the second image by using the filter, and the coding information includes at least one of a predicted bit rate and a distortion rate of the second image; selecting a target prediction result whose coding information satisfies a specified condition from the at least one prediction result, and using a filter corresponding to the target prediction result as the first type filter; The method of claim 11 , comprising:
14. The method of any one of claims 10 to 13, wherein the rendering information comprises one or a combination of a depth map, an albedo map, and post-processing parameters.
15. If the rendering information includes the depth map, the first type of filter is determined based on a relationship between pixels in the depth map and edge information of objects in the depth map, or If the rendering information includes the albedo map, the first type of filter is determined based on a relationship between pixels in the albedo map and edge information of objects in the albedo map; or If the post-processing parameters included in the rendering information include an anti-aliasing parameter, the first type filter is a filter indicated by the anti-aliasing parameter, or if the post-processing parameters included in the rendering information include instructions for the execution of a motion blur module, the first type of filter is a bilinear filter, or The method of claim 14 , wherein the first type of filter is a Catmull-Rom filter.
16. performing interpolation on a first decoded image to obtain a virtual reference frame for the second image by using a first type filter determined based on the rendering information; performing interpolation on the first decoded image based on a first correspondence relationship to obtain pixel values of all pixels in the virtual reference frame, the first correspondence relationship indicating a positional mapping relationship between pixels in the first image and pixels in the second image in the overlapping region between the first image and the second image; obtaining the virtual reference frame based on the pixel values of all of the pixels; 16. The method of any one of claims 10 to 15, comprising:
17. The first correspondence relationship is determined in the following manner: When the size of the graphics motion vector map in the rendering information does not match the size of the source image, creating a blank image whose size matches the size of the source image; determining a second correspondence between pixel coordinates in the graphics motion vector map and pixel coordinates in the blank image; performing interpolation on the graphics motion vector map based on the second correspondence by using a specified filter to obtain the first correspondence; The method of claim 16, wherein the signal is obtained by using
18. determining target discrete parameters for pixels in the first decoded image based on the rendering information, the target discrete parameters including a target variance or a target covariance; a step of using a first filter kernel corresponding to a first discrete parameter having a minimum difference from the target discrete parameter as a parameter of the first type filter, the first discrete parameter indicating a discrete parameter obtained by the first type filter by performing interpolation on the pixel based on the first filter kernel and performing calculation based on the pixel obtained through interpolation, the first filter kernel being one of a plurality of specified filter kernels, and the first discrete parameter including a first variance or a first covariance; 18. The method of claim 16 or 17, further comprising:
19. 19. The method of claim 10, further comprising the step of writing type information of the first type of filter into the bitstream.
20. 1. A video decoding device, comprising: a second acquisition module configured to acquire a bitstream and rendering information corresponding to the bitstream, the bitstream including a plurality of image frames, the rendering information indicating processing parameters used in a process of generating the bitstream based on a plurality of source images, the processing parameters indicating that a source image corresponding to a first image frame and a source image corresponding to a second image frame of the plurality of image frames have at least a partial overlapping region; a second interpolation module configured to perform interpolation on a first decoded image corresponding to the first image frame to obtain a virtual reference frame for the second image frame by using a first type filter determined based on the rendering information, wherein the first type filter is one of a plurality of specified type filters; a decoding module configured to decode the second image frame based on the virtual reference frame to obtain a second decoded image corresponding to the second image frame; An apparatus comprising:
21. The second interpolation module: The method is further configured to query a mapping table specified based on the rendering information to obtain a filter corresponding to each processing parameter in the rendering information, wherein the mapping table indicates at least one type of filter corresponding to each processing parameter among the plurality of types of filters; The apparatus of claim 20 , further configured to determine the first type of filter from all filters corresponding to the rendering information.
22. the mapping table further indicates a priority of each type of filter among the plurality of types of filters; 22. The apparatus of claim 21, wherein the second interpolation module is further configured to obtain priorities of all the filters corresponding to the rendering information, and use a filter having a highest filter priority among all the filters as the first type filter.
23. 23. The apparatus of any one of claims 20 to 22, wherein the rendering information comprises one or a combination of a depth map, an albedo map, and post-processing parameters.
24. If the rendering information includes the depth map, the first type of filter is determined based on a relationship between pixels in the depth map and edge information of objects in the depth map, or If the rendering information includes the albedo map, the first type of filter is determined based on a relationship between pixels in the albedo map and edge information of objects in the albedo map; or If the post-processing parameters included in the rendering information include an anti-aliasing parameter, the first type filter is a filter indicated by the anti-aliasing parameter, or if the post-processing parameters included in the rendering information include instructions for the execution of a motion blur module, the first type of filter is a bilinear filter, or 24. The apparatus of claim 23, wherein the first type of filter is a Catmull-Rom filter.
25. 25. The apparatus of claim 20, further comprising: a filter information acquisition module configured to acquire filter information and determine the first type filter and filtering parameters corresponding to the first type filter based on a filter type and parameters indicated in the filter information, the filtering parameters indicating processing parameters used in a process of acquiring the virtual reference frame by using the first type filter.
26. The second interpolation module: further configured to perform interpolation on the first decoded image based on a first correspondence relationship to obtain pixel values of all pixels in the virtual reference frame, the first correspondence relationship indicating a positional mapping relationship between pixels in the source image corresponding to the first image frame and pixels in the source image corresponding to the second image frame in the overlapping region between the source image corresponding to the first image frame and the source image corresponding to the second image frame; 26. The apparatus of claim 20, further configured to obtain the virtual reference frame based on the pixel values of all the pixels.
27. The correspondence relationship is determined in the following manner: When the size of the graphics motion vector map in the rendering information does not match the size of the source image, creating a blank image whose size matches the size of the source image; determining a second correspondence between pixel coordinates in the graphics motion vector map and pixel coordinates in the blank image; performing interpolation on the graphics motion vector map based on the second correspondence by using a specified filter to obtain the first correspondence; 27. The apparatus of claim 26, wherein the signal is obtained by using
28. 28. The apparatus of claim 26 or 27, further comprising: a second filter kernel determination module configured to determine target discrete parameters for pixels in the first decoded image based on the rendering information, the target discrete parameters including a target variance or a target covariance, and to use a first filter kernel corresponding to a first discrete parameter having a minimum difference from the target discrete parameter as a parameter of the first type filter, the first discrete parameter indicating a discrete parameter obtained by the first type filter by performing interpolation on the pixel based on the first filter kernel and performing calculations based on the pixel obtained through the interpolation, the first filter kernel being one of a plurality of specified filter kernels, and the first discrete parameter including a first variance or a first covariance.
29. 1. A video encoding device, comprising: a first acquisition module configured to acquire a plurality of source images and rendering information corresponding to the plurality of source images, the rendering information indicating processing parameters to be used in a process of generating a bitstream based on the plurality of source images, the processing parameters indicating that a first image and a second image of the plurality of source images have at least a partially overlapping region; a first interpolation module configured to perform interpolation on a first decoded image to obtain a virtual reference frame for the second image by using a first type filter determined based on the rendering information, the first type filter being one of a plurality of specified types of filters, and the first decoded image being a reconstructed image obtained by encoding and then decoding the first image of the plurality of images; an encoding module configured to encode the second image based on the virtual reference frame; An apparatus comprising:
30. The first interpolation module: The method is further configured to query a mapping table specified based on the rendering information to obtain a filter corresponding to each processing parameter in the rendering information, wherein the mapping table indicates at least one type of filter corresponding to each processing parameter among the plurality of types of filters; 30. The apparatus of claim 29, further configured to determine the first type of filter from all filters corresponding to the rendering information.
31. the mapping table further indicates a priority of each type of filter among the plurality of types of filters; 31. The apparatus of claim 30, wherein the first interpolation module is further configured to obtain priorities of all the filters corresponding to the rendering information, and use a filter having a highest filter priority among all the filters as the first type filter.
32. The first interpolation module: and further configured to perform prediction for all the filters corresponding to the rendering information to obtain at least one prediction result, one prediction result corresponding to one of the filters, the prediction result indicating coding information obtained by precoding the second image by using the filter, the coding information including at least one of a predicted bit rate and a distortion rate of the second image; 31. The apparatus of claim 30, further configured to: select, from the at least one prediction result, a target prediction result whose coding information satisfies a specified condition; and use a filter corresponding to the target prediction result as the first type filter.
33. 33. The apparatus of any one of claims 29 to 32, wherein the rendering information comprises one or a combination of a depth map, an albedo map, and post-processing parameters.
34. If the rendering information includes the depth map, the first type of filter is determined based on a relationship between pixels in the depth map and edge information of objects in the depth map, or If the rendering information includes the albedo map, the first type of filter is determined based on a relationship between pixels in the albedo map and edge information of objects in the albedo map; or If the post-processing parameters included in the rendering information include an anti-aliasing parameter, the first type filter is a filter indicated by the anti-aliasing parameter, or if the post-processing parameters included in the rendering information include instructions for the execution of a motion blur module, the first type of filter is a bilinear filter, or 34. The apparatus of claim 33, wherein the first type of filter is a Catmull-Rom filter.
35. The first interpolation module: further configured to perform interpolation on the first decoded image based on a first correspondence relationship to obtain pixel values of all pixels in the virtual reference frame, the first correspondence relationship indicating a positional mapping relationship between pixels in the first image and pixels in the second image in the overlapping region between the first image and the second image; 35. The apparatus of any one of claims 29 to 34, further configured to obtain the virtual reference frame based on the pixel values of all the pixels.
36. The first correspondence relationship is determined in the following manner: When the size of the graphics motion vector map in the rendering information does not match the size of the source image, creating a blank image whose size matches the size of the source image; determining a second correspondence between pixel coordinates in the graphics motion vector map and pixel coordinates in the blank image; performing interpolation on the graphics motion vector map based on the second correspondence by using a specified filter to obtain the first correspondence; 36. The apparatus of claim 35, wherein the signal is obtained by using
37. 37. The apparatus of claim 35 or 36, further comprising: a first filter kernel determination module configured to determine target discrete parameters for pixels in the first decoded image based on the rendering information, the target discrete parameters including a target variance or a target covariance, and to use a first filter kernel corresponding to a first discrete parameter having a minimum difference from the target discrete parameter as a parameter of the first type filter, the first discrete parameter indicating a discrete parameter obtained by the first type filter by performing interpolation on the pixel based on the first filter kernel and performing calculations based on the pixel obtained through the interpolation, the first filter kernel being one of a plurality of specified filter kernels, and the first discrete parameter including a first variance or a first covariance.
38. 38. The apparatus of claim 29, wherein the encoding module is further configured to write type information of the first type of filter into the bitstream.
39. A chip including a processor and a power supply circuit, the power supply circuit is configured to provide power to the processor; A chip, wherein the processor is configured to perform the method of any one of claims 1 to 9 and / or the processor is configured to perform the method of any one of claims 10 to 19.
40. 1. A codec including a memory and a processor, The memory is configured to store computer instructions that, when executed, cause the processor to implement a method according to any one of claims 1 to 9 and / or that, when executed, cause the processor to implement a method according to any one of claims 10 to 19.
41. An encoding and decoding system including an encoder side and a decoder side, the encoder side is configured to encode a plurality of images based on rendering information corresponding to the plurality of source images to obtain a bitstream corresponding to the plurality of images, thereby implementing the method according to any one of claims 10 to 19; 10. The encoding and decoding system, wherein the decoder side is configured to decode the bitstream based on rendering information corresponding to the plurality of source images to obtain a plurality of decoded images, thereby realizing the method according to any one of claims 1 to 9.
42. 1. A computer-readable storage medium, comprising: A computer-readable storage medium storing a computer program or instructions, the computer program or instructions being executed by a processing device to implement the method of any one of claims 1 to 9, and / or the computer program or instructions being executed by a processing device to implement the method of any one of claims 10 to 19.
43. A computer program product containing computer programs or instructions, A computer program product, the computer program or the instructions of which, when executed by a processing device, cause the method of any one of claims 1 to 9 to be realized, and / or the computer program or the instructions of which, when executed by a processing device, cause the method of any one of claims 10 to 19 to be realized.
Citation Information
Patent Citations
Image encoding method, image decoding method, image encoding device, image decoding device and program therefor
JP2012080162A
Apparatus and method for video motion compensation using a selectable interpolation filter
JP2018530244A
A coding scheme for immersive video using asymmetric downsampling and machine learning
JP2022548374A
METHOD AND SYSTEM FOR VIDEO CODING USING REFERENCE REGIONS - Patent application
JP2023522845A
Video encoding method and apparatus, video decoding method and apparatus, and programs therefor
US20150189276A1