Encoding method, decoding method, code stream, encoder, decoder, and storage medium

By predicting the virtual reference image in different domain orders of the reconstruction reference image on the encoder and decoder side, selecting the optimal target virtual reference image for inter-frame prediction, the problem of data mismatch in the neural network virtual reference image generation technology is solved, and the encoding and decoding performance is improved.

WO2025138262A1PCT designated stage expired Publication Date: 2025-07-03GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/143636
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

The existing virtual reference image generation technology based on neural networks has data mismatch problems in the training stage and the actual application stage, resulting in low similarity between the generated virtual reference image and the original image, affecting inter-frame prediction and video encoding and codec performance.

Method used

By performing virtual reference image prediction on at least two reconstructed reference images in a first time domain order consistent with the original time domain order and a second time domain order opposite to the original time domain order, the first virtual reference image and the second virtual reference image are determined, and the target virtual reference image is selected according to the cost of distortion for inter prediction, the time domain order is characterized by image-level syntax identification information.

Benefits of technology

Improves the correlation of virtual reference images and the accuracy of inter-frame prediction, and improves the encoding and decoding performance without increasing the decoding complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023143636_03072025_PF_FP_ABST
    Figure CN2023143636_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The present application discloses an encoding method, a decoding method, a code stream, an encoder, a decoder, and a storage medium. The decoding method comprises: parsing a code stream, to determine image-level syntax identification information corresponding to a current image; determining at least two reconstruction reference images corresponding to the current image; and on the basis of the image-level syntax identification information, determining to perform virtual reference image prediction on the at least two reconstruction reference images in a first temporal order or a second temporal order, to determine target virtual reference images, wherein the first temporal order represents the original temporal order of the at least two reconstruction reference images, the second temporal order is opposite to the first temporal order, and the target virtual reference images are used for performing inter-frame prediction on the current image. In this way, the quality of generated target virtual reference images can be improved, thereby improving the encoding and decoding performance.
Need to check novelty before this filing date? Find Prior Art

Description

Coding and decoding method, code stream, encoder, decoder and storage medium Technical Field

[0001] The present application relates to the field of video coding and decoding technology, and in particular to a coding and decoding method, a bit stream, an encoder, a decoder, and a storage medium. Background Art

[0002] For the inter-frame prediction module, the decoded picture buffer (DPB) stores several reconstructed images, which can be used as reference frames for the current frame to perform operations such as inter-frame motion estimation and motion compensation.

[0003] At present, in the research work of Neural Network based Video Coding (NNVC), a Neural Network based Virtual Reference Frame Generation (NNVRF) technology has been proposed. It can predict the reconstructed image selected from the DPB and generate a virtual reference image for inter-frame prediction coding of the current frame.

[0004] However, in the current NNVRF method, the training data used in the network training stage may not match the data in actual applications, which reduces the similarity between the virtual reference image generated by NNVRF and the original image, and further reduces the performance of inter-frame prediction and video encoding and decoding based on the virtual reference image.

[0005] Summary of the Invention

[0006] The embodiments of the present application provide a coding and decoding method, a bit stream, an encoder, a decoder, and a storage medium, which can improve the performance of inter-frame prediction and video coding and decoding based on virtual reference images.

[0007] The technical solution of the embodiment of the present application can be implemented as follows:

[0008] In a first aspect, an embodiment of the present application provides a decoding method, applied to a decoder, the method comprising:

[0009] Parse the code stream to determine the image-level syntax identification information corresponding to the current image;

[0010] Determining at least two reconstructed reference images corresponding to the current image;

[0011] According to the image-level syntax identification information, it is determined to perform virtual reference image prediction on the at least two reconstructed reference images in a first time domain order or a second time domain order, and a target virtual reference image is determined; the first time domain order represents the original time domain order of the at least two reconstructed reference images; the second time domain order is opposite to the first time domain order; the target virtual reference image is used to perform inter-frame prediction on the current image.

[0012] In a second aspect, an embodiment of the present application provides an encoding method, applied to an encoder, the method comprising:

[0013] determining at least two reconstructed reference images corresponding to the current image;

[0014] Based on the at least two reconstructed reference images, performing virtual reference image prediction according to a first temporal order and a second temporal order, respectively, to determine a first virtual reference image and a second virtual reference image; the first temporal order represents an original temporal order of the at least two reconstructed reference images; and the second temporal order is opposite to the first temporal order;

[0015] According to the distortion cost, a target virtual reference image is determined from the first virtual reference image and the second virtual reference image, and image-level syntax identification information is determined according to the time domain order corresponding to the target virtual reference image; the target virtual reference image is used to perform inter-frame prediction on the current image.

[0016] In a third aspect, an embodiment of the present application provides a code stream, wherein the code stream is generated by bit encoding based on information to be encoded; wherein the information to be encoded includes at least one of the following:

[0017] Image-level syntax identification information;

[0018] The picture-level syntax identification information is used to indicate that virtual reference image prediction is performed on at least two reconstructed reference images corresponding to the current image in the first time domain order or the second time domain order to determine a target virtual reference image.

[0019] In a fourth aspect, an embodiment of the present application provides a decoder, comprising:

[0020] The parsing part is configured to parse the code stream and determine the image-level syntax identification information corresponding to the current image;

[0021] A first determining part is configured to determine at least two reconstructed reference images corresponding to the current image;

[0022] The first prediction part is configured to determine, based on the image-level syntax identification information, to perform virtual reference image prediction on the at least two reconstructed reference images in a first time domain order or a second time domain order, and to determine a target virtual reference image; the first time domain order represents the original time domain order of the at least two reconstructed reference images; the second time domain order is opposite to the first time domain order; the target virtual reference image is used to perform inter-frame prediction on the current image.

[0023] In a fifth aspect, an embodiment of the present application provides a decoder, the decoder comprising a first memory and a first processor; wherein,

[0024] a first memory for storing a computer program capable of running on the first processor;

[0025] The first processor is configured to execute the decoding method as described in the first aspect when running a computer program.

[0026] In a sixth aspect, an embodiment of the present application provides an encoder, comprising:

[0027] A second determining part is configured to determine at least two reconstructed reference images corresponding to the current image;

[0028] a second prediction section configured to perform virtual reference image prediction based on the at least two reconstructed reference images according to a first temporal order and a second temporal order, respectively, to determine a first virtual reference image and a second virtual reference image; the first temporal order representing an original temporal order of the at least two reconstructed reference images; and the second temporal order being opposite to the first temporal order;

[0029] The third determination part is configured to determine the target virtual reference image from the first virtual reference image and the second virtual reference image according to the distortion cost, and determine the image-level syntax identification information according to the time domain order corresponding to the target virtual reference image; the target virtual reference image is used to perform inter-frame prediction on the current image.

[0030] In a seventh aspect, an embodiment of the present application provides an encoder, comprising a second memory and a second processor; wherein,

[0031] a second memory for storing a computer program capable of running on the second processor;

[0032] The second processor is used to execute the encoding method as described in the second aspect when running the computer program.

[0033] In an eighth aspect, an embodiment of the present application provides a storage medium storing a computer program, which, when executed, implements the decoding method as described in the first aspect, or implements the encoding method as described in the second aspect.

[0034] The present application provides a coding and decoding method, a bitstream, an encoder, a decoder, and a storage medium. At the encoder end, virtual reference image prediction can be performed on at least two reconstructed reference images corresponding to a current image according to a first temporal order consistent with the original temporal order and a second temporal order opposite to the original temporal order, respectively, to determine a first virtual reference image and a second virtual reference image. A target virtual reference image for inter-frame prediction of the current image is determined from the first and second virtual reference images based on a distortion cost, and picture-level syntax identification information is determined based on the temporal order corresponding to the target virtual reference image. Thus, upon parsing picture-level syntax identification information from the bitstream, the decoder end can perform virtual reference image prediction on at least two reconstructed reference images corresponding to the current image according to the temporal order represented by the picture-level syntax identification information, and determine a target virtual reference image for inter-frame prediction of the current image. The present application can determine a target virtual reference image with a higher correlation and lower distortion with the current image for inter-frame prediction by comparing virtual reference images predicted based on reconstructed reference images with different temporal orders, without increasing decoding complexity, thereby improving coding and decoding performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] FIG1 is a schematic diagram of a reference image for inter-frame bidirectional prediction;

[0036] FIG2 is a schematic diagram of an encoding process based on a virtual reference image generation technology;

[0037] FIG3 is a schematic diagram of the structure of a virtual reference image generation network;

[0038] FIG4 is a schematic diagram of optical flow information estimation for virtual reference image prediction;

[0039] FIG5 is a second schematic diagram of optical flow information estimation for virtual reference image prediction;

[0040] FIG6 is a schematic block diagram of a composition of an encoder provided in an embodiment of the present application;

[0041] FIG7 is a schematic block diagram of a decoder provided in an embodiment of the present application;

[0042] FIG8 is a schematic diagram of a network architecture of a coding and decoding system provided in an embodiment of the present application;

[0043] FIG9 is a flowchart diagram 1 of an encoding method provided in an embodiment of the present application;

[0044] FIG10 is a first schematic diagram of a sequence of inputting a reconstructed reference image into an NNVRF network according to an embodiment of the present application;

[0045] FIG11 is a second schematic diagram of a sequence of inputting a reconstructed reference image into an NNVRF network according to an embodiment of the present application;

[0046] FIG12 is a schematic diagram of a candidate image block provided in an embodiment of the present application;

[0047] FIG13 is a second flow chart of an encoding method provided in an embodiment of the present application;

[0048] FIG14 is a second schematic diagram of the structure of an encoder provided in an embodiment of the present application;

[0049] FIG15 is a schematic diagram of a flowchart of a decoding method provided in an embodiment of the present application;

[0050] FIG16 is a schematic diagram of the structure of an encoder provided in an embodiment of the present application;

[0051] FIG17 is a schematic diagram of a specific hardware structure of an encoder provided in an embodiment of the present application;

[0052] FIG18 is a schematic diagram of the structure of a decoder provided in an embodiment of the present application;

[0053] FIG19 is a schematic diagram of a specific hardware structure of a decoder provided in an embodiment of the present application;

[0054] FIG20 is a schematic diagram of the composition structure of a coding and decoding system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0055] In order to enable a more detailed understanding of the features and technical contents of the embodiments of the present application, the implementation of the embodiments of the present application is described in detail below with reference to the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present application.

[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0057] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0058] It should also be pointed out that the terms "first\second\third" involved in the embodiments of the present application are only used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.

[0059] Digital video compression technology primarily compresses massive amounts of digital video data for easier transmission and storage. With the surge in Internet video usage and increasing demand for higher-quality video, while existing digital video compression standards can save significant amounts of video data, there is still a need for better digital video compression technologies to reduce bandwidth and traffic pressures associated with digital video transmission.

[0060] For the inter-frame prediction module, the decoded picture buffer (DPB) stores several reconstructed images, which can be used as reference frames (reference images) of the current frame (current image) to perform operations such as inter-frame motion estimation and motion compensation.

[0061] In the most advanced Versatile Video Coding (VVC) standard, three types of frames (images) are defined: I, P, and B. I frames represent intra-frame coded frames, while P and B frames represent inter-frame coded frames. P frames can only refer to reference frames that are before the current coded frame in time, that is, only forward prediction can be used. B frames can not only refer to reference frames that are before the current coded frame in time, but also refer to reference frames that are after the current coded frame in time, that is, forward prediction, backward prediction, and bidirectional prediction can be used. As shown in Figure 1, in bidirectional prediction, the information of two reference images can be used simultaneously, and the weighted average of the two reference macroblocks can be used as the reference information of the current block.

[0062] VVC establishes two reference picture lists (RPLs) for inter-frame prediction, including list0 (L0) for forward prediction and list1 (L1) for backward prediction. A single image can appear in different positions in both lists at the same time, providing a high degree of flexibility in the selection of multiple reference images. For each block in a P frame, only the reference image in L0 can be used, while each block in a B frame can use the reference images in both L0 and L1 (in bidirectional prediction, one image is selected from L0 and one from L1 as a reference).

[0063] Specifically, taking the random access coding configuration (RA) in VVC as an example, the default setting of IntraPeriod = 32, that is, an I frame is inserted every 32 frames. Under this configuration, the corresponding L0 and L1 settings are shown in Table 1. Table 1 shows the VVC RA configuration when IntraPeriod = 32.

[0064] Table 1

[0065] According to Table 1, taking POC = 8 as an example, its reference L0 is frame 0 and frame 16, and L1 is frame 16 and frame 32. As can be seen from Table 1, for the current frame, traditional reference frames are all encoded reconstructed frames, which are at a certain temporal distance from the current frame. Therefore, there may be a certain difference in image content or image quality between traditional reference frames and the current frame. The performance of inter-frame prediction is highly dependent on the content and quality of the reference frames. Reference frames with less compression distortion or more relevant content can reduce the prediction residual. Therefore, if some methods can be used to synthesize better reference images, the encoding performance of the current frame can be improved.

[0066] In recent years, with the development of deep learning technology, neural network-based virtual reference image generation technology has gradually developed. Deep neural networks play an important role in image and video generation tasks. Many neural network-based interpolation techniques have been proposed, which achieve good generation results by predicting the changes in temporal optical flow information through the network.

[0067] Among them, within the current research work of the Audio Video Coding Standard (AVS) Working Group, the High Performance-Modular Artificial Intelligence Model (HPM-ModAI) employs a convolutional neural network-based virtual reference image generation method. This method, when used for inter-frame prediction, can generate a reference image closer to the current image for inter-frame prediction reference. The virtual reference image tool adopted by HPM-ModAI 12.0 is based on the classic interpolation network algorithm IFRNet-L. Its network input is the two most adjacent reconstructed images in the time domain, and its output is a predicted virtual reference frame.

[0068] Current NNVC research has proposed a virtual reference image generation technique (NNVRF) based on frame interpolation. NNVRF is a reference image generation method based on a deep neural network and designed for the VVC layered coding structure. As shown in Figure 2, the encoder in the hybrid frame coding mode reads unequal pixels from the original video sequences of different color formats, including luminance and chrominance components. In other words, the encoder reads a black and white or color image. It then divides the image into blocks and encodes the image in units of blocks. The encoder includes an intra-frame prediction unit, an inter-frame prediction unit, a transform unit, a quantization unit, a scaling unit, an inverse transform and inverse quantization unit, a loop filter unit, and an entropy coding unit. Among them, the intra-frame prediction unit only refers to the information of the same image to predict the pixel information within the current block to eliminate spatial redundancy; the inter-frame prediction unit can refer to the image information of different frames and use motion estimation to search for the motion vector information that best matches the current block to eliminate temporal redundancy; the transform unit converts the predicted image block to the frequency domain and redistributes the energy. Combined with the quantization unit, it can remove information that the human eye is not sensitive to to eliminate visual redundancy; the entropy coding unit can eliminate character redundancy based on the current context model and the probability information of the binary code stream; the loop filter unit mainly processes the pixels after inverse transformation and inverse quantization to compensate for the distortion information and provide a better reference for subsequent encoded pixels.

[0069] As shown in Figure 2, the images in the reference image list are obtained by acquiring reconstructed images from the decoded image cache unit. The images in the reference image list can include reference image list L0 and reference image list L1. The encoder selects one reconstructed image from each of reference image list L0 and reference image list L1 as two reconstructed reference images, which are input into the virtual reference image generation network. After processing by the virtual reference image generation network, a virtual reference image is generated. The virtual reference image is then inserted into reference image list L0 and reference image list L1 for inter-frame prediction coding of the current image.

[0070] The above-mentioned virtual reference image generation network is based on the virtual reference image generation technology NNVRF, and its network structure can be shown in Figure 3. The network input is a reconstructed reference image selected from the reference image list L0 and the reference image list L1, which is input into a three-level hierarchical neural network to extract feature information of different granularities. Among them, the interpolation algorithm module 30 (such as Small IFRNet) is used to predict or estimate the optical flow information of the two reconstructed reference images to obtain two optical flow information, which respectively represent the optical flow information between each reconstructed reference image and the current image. The two optical flow information enters channel a' of the dual-feature processing model 32, and the first-scale optical flow information feature extraction is performed to obtain two first-scale optical flow features; the two first-scale optical flow features are downsampled once and enter channel b', and the second-scale optical flow information feature extraction is performed to obtain two second-scale optical flow features; the two second-scale optical flow features are downsampled once and enter channel c', and the third-scale optical flow information feature extraction is performed to obtain two third-scale optical flow features; thus, two optical flow features at three scales are obtained. Furthermore, the convolution (Convolution, Conv) module 31 extracts image features from the two input reconstructed reference images to obtain two initial feature information. The two initial feature information corresponding to the two reconstructed reference images are input into channel a of the dual-feature processing model 32 for first-scale image feature extraction to obtain two first-scale image features. The two first-scale image features are downsampled once and input into channel b for second-scale image feature extraction to obtain two second-scale image features. The two second-scale image features are further downsampled once and input into channel c for third-scale image feature extraction to obtain two third-scale image features. Thus, two image features at three scales are obtained. The optical flow features and the image features of the first scale are further processed by the first feature processing module 33 based on optical flow and the first scale feature processing module 36 to extract deeper feature information and obtain first scale features. The optical flow features and the image features of the second scale are further processed by the second feature processing module 34 based on optical flow and the second scale feature processing module 37 to obtain second scale features. The optical flow features and the image features of the third scale are further processed by the third feature processing module 35 based on optical flow and the third scale feature processing module 38 to obtain third scale features. The first scale features, the second scale features, and the third scale features are aggregated by the feature aggregation prediction module 39 (such as GridNet), and prediction is performed based on the aggregated features to obtain a virtual reference image.

[0071] Based on the above-mentioned NNVRF scheme and virtual reference image generation network, the applicant analyzed the training process of the NNVRF scheme and found that there is a training method that interchanges the order of the two reconstructed images used as training data input during the training phase to achieve data augmentation. The following pseudo code is shown: if random.randint(0,1): gt_path_list.reverse() input_path_list.reverse()

[0072] That is, during the training process of the NNVRF scheme, there is a certain probability that the playback order of the two reconstructed reference images used as training sample input is opposite. Accordingly, the NNVRF neural network will perform machine learning and network training based on the two reconstructed reference images input in the opposite playback order to predict the virtual reference image. However, in actual coding tests / applications, the order of the reconstructed reference images input to the NNVRF neural network is not interchanged. It can be seen that there is a certain mismatch between the training process of the neural network in the current NNVRF scheme and the coding test / application process, which may reduce the quality of the virtual reference frames predicted and generated by the NNVRF scheme, thereby affecting the accuracy of inter-frame prediction and encoding and decoding performance.

[0073] This application analyzes the effect of the playback order of interchanged input reconstructed reference images on optical flow estimation. As shown in Figure 4, the two reconstructed reference images are input into the virtual reference image generation network shown in Figure 3 in the original playback order (reconstructed reference image 1, reconstructed reference image 2). The object in the image moves from the lower left corner to the upper right corner. In the predicted virtual reference image, the object is in the middle position. As shown in Figure 5, the same two reconstructed images are input into the virtual reference image generation network after the playback order is swapped (reconstructed reference image 2, reconstructed reference image 1). The object in the image moves from the upper right corner to the lower left corner. Although the estimated optical flow information is in the opposite direction, the object is still in the middle position in the predicted virtual reference image. This shows that for the prediction of virtual reference images, interchanging the playback order of the input images has certain physical significance.

[0074] This application studies the order interchange of reconstructed reference images input into the NNVRF neural network during the virtual image generation process of inter-frame prediction, and proposes an input information interchange method to further optimize the encoding performance of the virtual reference image generation tool by interchanging the order of the input images of the virtual reference image generation network.

[0075] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0076] Referring to Figure 6, which shows a schematic block diagram of the composition of an encoder provided in an embodiment of the present application. As shown in Figure 6, the encoder 100 may include a transform and quantization unit 101, an intra-frame estimation unit 102, an intra-frame prediction unit 103, a motion compensation unit 104, a motion estimation unit 105, an inverse transform and inverse quantization unit 106, a filter control analysis unit 107, a filtering unit 108, an encoding unit 109 and a decoded image cache unit 110, etc., wherein the filtering unit 108 can implement deblocking filtering and sample adaptive offset (SAO) filtering, and the encoding unit 109 can implement header information encoding and context-based adaptive binary arithmetic coding (CABAC).For the input original video signal, a video coding block can be obtained by dividing the coding tree unit (CTU). Then, the residual pixel information obtained after intra-frame or inter-frame prediction is transformed by the transformation and quantization unit 101, including transforming the residual information from the pixel domain to the transform domain and quantizing the obtained transform coefficients to further reduce the bit rate; the intra-frame estimation unit 102 and the intra-frame prediction unit 103 are used to perform intra-frame prediction on the video coding block. Specifically, the intra-frame estimation unit 102 and the intra-frame prediction unit 103 are used to determine the intra-frame prediction mode to be used to encode the video coding block; the motion compensation unit 104 and the motion estimation unit 105 are used to perform inter-frame prediction coding on the received video coding block relative to one or more blocks in one or more reference frames to provide temporal prediction information; the motion estimation performed by the motion estimation unit 105 is the process of generating a motion vector, which can estimate the motion of the video coding block. The motion compensation unit 104 then calculates the motion vector based on the motion vector determined by the motion estimation unit 105. After determining the intra-frame prediction mode, the intra-frame prediction unit 103 is further configured to provide the selected intra-frame prediction data to the encoding unit 109, and the motion estimation unit 105 also sends the calculated motion vector data to the encoding unit 109. In addition, the inverse transform and inverse quantization unit 106 is configured to reconstruct the video coding block and reconstruct a residual block in the pixel domain. The reconstructed residual block is subjected to the filter control analysis unit 107 and the filtering unit 108 to remove the block effect artifacts. The reconstructed residual block is then added to a predictive block in the frame of the decoded image buffer unit 110 to generate a reconstructed video coding block. The encoding unit 109 is configured to encode various coding parameters and quantized transform coefficients. In the CABAC-based coding algorithm, the context content can be based on adjacent coding blocks and can be used to encode information indicating the determined intra-frame prediction mode, and output the code stream of the video signal. The decoded image buffer unit 110 is configured to store the reconstructed video coding block for prediction reference. As the video image encoding proceeds, new reconstructed video encoding blocks are continuously generated, and these reconstructed video encoding blocks are stored in the decoded image buffer unit 110 .

[0077] Refer to Figure 7, which shows a block diagram of a decoder provided by an embodiment of the present application. As shown in Figure 7, the decoder 200 includes a decoding unit 201, an inverse transform and inverse quantization unit 202, an intra-frame prediction unit 203, a motion compensation unit 204, a filtering unit 205 and a decoded image cache unit 206, etc., wherein the decoding unit 201 can implement header information decoding and CABAC decoding, and the filtering unit 205 can implement deblocking filtering and SAO filtering. After the input video signal is encoded in Figure 6, the code stream of the video signal is output; the code stream is input to the decoder 200, and first passes through the decoding unit 201 to obtain the decoded transform coefficient; the transform coefficient is processed by the inverse transform and inverse quantization unit 202 to generate a residual block in the pixel domain; the intra-frame prediction unit 203 can be used to generate prediction data for the current video decoding block based on the determined intra-frame prediction mode and the data of the previously decoded block from the current frame or picture; the motion compensation unit 204 is to determine the prediction information for the video decoding block by analyzing the motion vector and other associated syntax elements, and use The prediction information is used to generate a predictive block for the video decoding block being decoded; a decoded video block is formed by summing the residual block from the inverse transform and inverse quantization unit 202 with the corresponding predictive block generated by the intra-frame prediction unit 203 or the motion compensation unit 204; the decoded video signal passes through the filtering unit 205 to remove blocking artifacts, thereby improving video quality; the decoded video block is then stored in the decoded image buffer unit 206, which stores reference images used for subsequent intra-frame prediction or motion compensation, and is also used for outputting the video signal, thereby obtaining the restored original video signal.

[0078] Furthermore, the embodiment of the present application also provides a network architecture of a coding and decoding system including an encoder and a decoder, wherein FIG8 shows a schematic diagram of a network architecture of a coding and decoding system provided by the embodiment of the present application. As shown in FIG8 , the network architecture includes one or more electronic devices 13 to 1N and a communication network 01, wherein the electronic devices 13 to 1N can perform video interaction through the communication network 01. During implementation, the electronic device can be various types of devices with video coding and decoding functions. For example, the electronic device can include a smart phone, a tablet computer, a personal computer, a personal digital assistant, a navigator, a digital phone, a video phone, a television, a sensing device, a server, etc., which is not specifically limited in the embodiment of the present application. Here, the decoder or encoder described in the embodiment of the present application can be the above-mentioned electronic device.

[0079] It should be noted that the method of the embodiment of the present application is primarily applied to the prediction portion shown in Figure 6 and the prediction portion shown in Figure 7. That is, the embodiment of the present application can be applied to both the encoder and the decoder, or even simultaneously, but this embodiment of the present application is not specifically limited thereto. Furthermore, the prediction portion here may include the inter-frame prediction portion. Specifically, the method of the embodiment of the present application is applied to the virtual reference image construction portion of the inter-frame prediction portion.

[0080] It should also be noted that, on the encoding side, the "current image" specifically refers to the image currently to be inter-frame prediction encoded; on the decoding side, the "current image" specifically refers to the image currently to be inter-frame prediction decoded. In the embodiments of the present application, an "image block" can be a coding unit (CU), a coding tree unit (CTU), or even a prediction unit (PU) or a transform unit (TU), without specific limitation here.

[0081] In one embodiment of the present application, referring to FIG9 , a schematic flow chart of an encoding method provided by an embodiment of the present application is shown. As shown in FIG9 , the method may include:

[0082] S101: Determine at least two reconstructed reference images corresponding to a current image.

[0083] It should be noted that the encoding method of the embodiments of the present application is applied to an encoder. Furthermore, the encoding method may specifically refer to an inter-frame prediction method. The method primarily improves the process of generating a virtual reference image based on the NNVRF scheme during inter-frame prediction, thereby improving the encoding performance of inter-frame prediction.

[0084] In an embodiment of the present application, when the encoder performs inter-frame prediction on the current image and executes the process of constructing a virtual reference image in the inter-frame prediction, it determines at least two reconstructed images from the reconstructed images that have been encoded as at least two reconstructed reference images corresponding to the current image.

[0085] Exemplarily, when the encoder performs inter-frame prediction on the current image, it extracts two reconstructed images from the decoded image buffer DPB as reconstructed reference images according to the input image information required by the NNVRF technology, so as to be used as input images of the prediction network.

[0086] S102 : Based on at least two reconstructed reference images, perform virtual reference image prediction according to a first temporal order and a second temporal order respectively to determine a first virtual reference image and a second virtual reference image.

[0087] In this embodiment of the present application, a virtual reference image is predicted based on at least two reconstructed reference images in a first temporal order to generate a first virtual reference image. A virtual reference image is predicted based on at least two reconstructed reference images in a second temporal order to generate a first virtual reference image. The first temporal order represents the original temporal order of the at least two reconstructed reference images, and the second temporal order is the opposite of the first temporal order.

[0088] That is, the original temporal order of the at least two reconstructed reference images, i.e., the first temporal order, is used as the input order of the prediction network. The prediction network performs virtual reference image prediction on the at least two reconstructed reference images input in the first temporal order to determine the first virtual reference image. Furthermore, the original temporal order of the at least two reconstructed reference images is adjusted and input into the prediction network in a second temporal order. The prediction network performs virtual reference image prediction on the at least two reconstructed reference images input in the second temporal order to determine the second virtual reference image.

[0089] In some embodiments, for a case where two reconstructed reference images are determined, as shown in FIG10 , taking the two reconstructed reference images as reconstructed reference image 1 and reconstructed reference image 2, and the prediction network as the NNVRF network shown in FIG3 as an example, reconstructed reference image 1 and reconstructed reference image 2 are input into the NNVRF network in a first temporal order: reconstructed reference image 1, then reconstructed reference image 2, and the NNVRF network outputs a first virtual reference image. As shown in FIG11 , the input order of reconstructed reference image 1 and reconstructed reference image 2 is swapped, i.e., the information input into the NNVRF network remains reconstructed reference image 1 and reconstructed reference image 2, but is input into the NNVRF network in a second temporal order: reconstructed reference image 2, then reconstructed reference image 1, and the NNVRF network outputs a second virtual reference image.

[0090] In some embodiments, when N reconstructed reference images are determined (N is greater than 2), the N reconstructed reference images are input into the prediction network in a first temporal order, such as F0, F1, ..., FN-1, FN, and the prediction network outputs a first virtual reference image. The N reconstructed reference images are input into the prediction network in a second temporal order, such as FN, FN-1, ..., F1, F0, and the prediction network outputs a second virtual reference image.

[0091] In some embodiments, the encoder determines candidate image blocks at the same position in at least two reconstructed reference images; in a first time domain order, successively traverses the candidate image blocks in each reconstructed reference image of the at least two reconstructed reference images to perform virtual reference image prediction and determine a first virtual reference image; in a second time domain order, successively traverses the candidate image blocks in each reconstructed reference image of the at least two reconstructed reference images to perform virtual reference image prediction and determine a second virtual reference image.

[0092] As an example, each image block in the current image can be determined as a candidate image block. That is, virtual reference image prediction can be performed by sequentially traversing each image block in each of the at least two reconstructed reference images in a first temporal order to determine the first virtual reference image. Furthermore, virtual reference image prediction can be performed by sequentially traversing each image block in each of the at least two reconstructed reference images in a second temporal order to determine the second virtual reference image.

[0093] As another example, some image blocks in the current image can be used as candidate image blocks. Exemplarily, at least one image block can be determined as a candidate image block at every preset interval in the current image. As shown in FIG12 , one image block can be determined as a candidate image block every four CTUs (the image block filled with oblique lines in FIG12 ). It should be noted that FIG12 preliminarily divides the CTU blocks in the current image, takes each 2*2 CTU as an area, and determines the upper left corner image block in each area as a candidate image block, thereby determining one candidate image block every four CTUs. Candidate image blocks can also be determined based on other interval methods, interval sizes, or the number of candidate image blocks. The specific selection is made according to the actual situation and is not limited in the embodiments of the present application.

[0094] It can be seen that traversing some blocks in the current image to predict the virtual reference image can save the prediction workload of the prediction network. For example, for the candidate image blocks filled with diagonal lines in Figure 12, two predictions are required, one in the first temporal order and the other in the second temporal order, to obtain the first and second virtual reference images. However, for the blank image blocks in Figure 12, only one prediction is required, based on the temporal order corresponding to the target virtual reference image.

[0095] In some embodiments, the prediction network may include: an interpolation algorithm module, a convolution module, a feature processing module and a feature aggregation module. The interpolation algorithm module is used to estimate the optical flow information of at least two reconstructed reference images input into the prediction network in a first time domain order, and determine the first optical flow feature information between each reconstructed reference image and the current image; or, to estimate the optical flow information of at least two reconstructed reference images input into the prediction network in a second time domain order, and determine the second optical flow feature information between each reconstructed reference image and the current image. The convolution module is used to extract image feature information from each reconstructed reference image as the initial feature information corresponding to each reconstructed reference image. The feature processing module is used to perform deeper feature processing on the optical flow feature information and the image feature information at each scale, extract deeper feature information, and obtain intermediate feature information of different scales. The feature aggregation module is used to aggregate the intermediate feature information of different scales for prediction and determine the virtual reference image.

[0096] For example, the network structure of the prediction network can refer to the network structure of the virtual reference image generation network shown in FIG3 . The feature processing module is equivalent to including the dual-feature processing model 32, the first feature processing module 33 based on optical flow, the second feature processing module 34 based on optical flow, the third feature processing module 35 based on optical flow, the first scale feature processing module 36, the second scale feature processing module 37, and the third scale feature processing module 38 in FIG3 .

[0097] In some embodiments, the interpolation algorithm module in the prediction network can be used to predict the optical flow information of the candidate image blocks in the at least two reconstructed reference images input in the first time domain order to determine at least two first optical flow feature information; the convolution module in the prediction network can be used to extract the feature information of the candidate image blocks in the at least two reconstructed reference images respectively to determine at least two initial feature information; the at least two first optical flow feature information and the at least two initial feature information are downsampled at least once, and feature information extraction is performed on the at least two first optical flow feature information and the at least two initial feature information each time downsampled to determine first intermediate feature information of at least one scale; the feature aggregation prediction module in the prediction network can be used to aggregate the first intermediate feature information of at least one scale and perform prediction to determine the first virtual reference image.

[0098] In some embodiments, the interpolation algorithm module in the prediction network can be used to predict the optical flow information of the candidate image blocks in the at least two reconstructed reference images input in the second time domain order to determine at least two second optical flow feature information; the convolution module in the prediction network can be used to extract the initial feature information of the candidate image blocks in the at least two reconstructed reference images respectively to determine at least two initial feature information; the at least two second optical flow feature information and the at least two initial feature information are downsampled at least once, and feature information extraction is performed on the at least two second optical flow feature information and the at least two initial feature information each time downsampled to determine second intermediate feature information of at least one scale; the feature aggregation prediction module in the prediction network can be used to aggregate the second intermediate feature information of at least one scale and perform prediction to determine the second virtual reference image.

[0099] It should be noted that, according to the actual coding and decoding needs, the network structure or weight parameters of the prediction network can be adjusted based on FIG3 . The specific selection is based on the actual situation and is not limited in the embodiment of the present application.

[0100] In some embodiments, the present invention can determine whether to enable the process of performing virtual reference image prediction on at least two reconstructed reference images in two input orders to generate two virtual reference images in S102 based on preset sequence-level syntax identification information. As shown in FIG13 ,

[0101] S1021: Determine preset sequence-level syntax identification information.

[0102] In an embodiment of the present application, the preset sequence-level syntax identification information is used to indicate whether to enable at least two reconstructed reference images to be input into the prediction network in a first time domain order and a second time domain order to obtain two virtual reference images.

[0103] The preset sequence-level syntax identification information may be a parameter written in a profile, a value of a flag, or a parameter in an encoder configuration file, which is not specifically limited here.

[0104] S1022. When the preset sequence-level syntax identification information is a first value, based on at least two reconstructed reference images, virtual reference image prediction is performed according to a first time domain order and a second time domain order respectively to determine a first virtual reference image and a second virtual reference image; the first value represents the virtual reference image prediction image that enables the first time domain order and the second time domain order.

[0105] In this embodiment of the present application, when the value of the preset sequence-level syntax identification information is a first value, it indicates that virtual reference image prediction in the first temporal order and the second temporal order is enabled. This means that virtual reference image prediction can be performed on at least two reconstructed reference images in the first temporal order or the second temporal order to determine the target virtual reference image. Exemplarily, the preset sequence-level syntax identification information can be the lsf_enable_flag flag; the first value can be set to 1 or to true. The first value can be in parameter form or in numeric form. The specific selection depends on the actual situation and is not limited in this embodiment of the present application.

[0106] In some embodiments, when the preset sequence-level syntax identification information is a second value, a prediction network is used to perform virtual reference image prediction on at least two reconstructed reference images input in a first time domain order to determine a target virtual reference image.

[0107] Here, when the value of the preset sequence-level syntax identification information is a second value, it indicates that virtual reference picture prediction in the second temporal order is not enabled. That is, at least two reconstructed reference pictures are input into the prediction network only according to the first temporal order, i.e., the original temporal order. The prediction network performs virtual reference picture prediction based on the at least two reconstructed reference pictures input in the first temporal order. The resulting predicted picture is used as the target virtual reference picture, so that the target virtual reference picture is used for inter-frame prediction coding.

[0108] In the embodiments of the present application, the first value and the second value are different. In some embodiments, the second value can be set to 0 or set to false. The second value can be in the form of a parameter or a number. The specific selection is based on the actual situation and is not limited by the embodiments of the present application.

[0109] It should be noted that, when the value of the sequence-level syntax identification information is preset to be the second value, the picture-level syntax identification information may not be determined.

[0110] In some embodiments, the encoder may encode preset sequence-level syntax identification information and write the resulting coded bits into the bitstream to indicate to the decoder whether the current sequence enables virtual reference image prediction in both the first time domain order and the second time domain order as input information for the prediction network.

[0111] Exemplarily, the sequence header definition of the preset sequence-level syntax identification information lsf_enable_flag may be as follows:

[0112] S103 : Determine a target virtual reference image from the first virtual reference image and the second virtual reference image according to the distortion cost, and determine picture-level syntax identification information according to a temporal order corresponding to the target virtual reference image.

[0113] In this embodiment of the present application, based on a first virtual reference image and a second virtual reference image generated by a prediction network, the encoder determines a first distortion cost between the first virtual reference image and the current image, and a second distortion cost between the second virtual reference image and the current image. Based on the first distortion cost and the second distortion cost, a target virtual reference image is determined between the first and second virtual reference images. The target virtual reference image is used to perform inter-frame prediction on the current image.

[0114] For example, the target virtual reference image can be determined based on the smaller distortion cost between the first distortion cost and the second distortion cost. For example, if the first distortion cost is smaller than the second distortion cost, the first virtual reference image is determined as the target virtual reference image. If the second distortion cost is smaller than the first distortion cost, the second virtual reference image is determined as the target virtual reference image.

[0115] In some embodiments, the distortion cost can be determined through image comparison. Specifically, a first distortion cost is determined by comparing a first virtual reference image with the current image, and a second distortion cost is determined by comparing a second virtual reference image with the current image. Thus, based on the first and second distortion costs, the virtual reference image with the lower distortion cost, i.e., the first or second virtual reference image that is closer to the current image, is determined as the target reference image.

[0116] Exemplarily, the first distortion cost is determined based on the distortion cost between the candidate image block in the first virtual reference image and the image block at the same position in the current image; the second distortion cost is determined based on the distortion cost between the candidate image block in the second virtual reference image and the image block at the same position in the current image.

[0117] In some embodiments, the distortion cost can also be determined by comparing the coding costs. For example, inter-frame prediction coding is performed based on the first virtual reference image to determine the first candidate coding information, and the first distortion cost is determined by determining the coding cost of the first candidate coding information; inter-frame prediction coding is performed based on the second virtual reference image to determine the second candidate coding information, and the second distortion cost is determined by determining the coding cost of the second candidate coding information. In some embodiments, the coding cost can be a rate-distortion cost, or other cost algorithms can be used. The specific selection is based on the actual situation and is not limited in the embodiments of the present application.

[0118] In some embodiments, when the target virtual reference image is determined, the image-level syntax identification information is determined according to the time domain order corresponding to the target virtual reference image, so as to instruct the decoder through the image-level syntax identification information to input the at least two reconstructed reference images corresponding to the current image into the prediction network in an order.

[0119] In some embodiments, when the time domain order corresponding to the target virtual reference image is the first time domain order, the value of the image-level syntax identification information is set to the third value; when the time domain order corresponding to the target virtual reference image is the second time domain order, the value of the image-level syntax identification information is set to the fourth value.

[0120] In an embodiment of the present application, the third value and the fourth value are different. Here, the third value and the fourth value can be in parameter form or in digital form. Specifically, the image-level syntax identification information here can be a parameter written in the profile or the value of a flag. Exemplarily, the third value can be set to 0 and the fourth value can be set to 1; or, the third value can also be set to false and the fourth value can also be set to true. The specific selection is made according to the actual situation and is not limited in the embodiment of the present application.

[0121] In some embodiments, when the picture-level syntax identification information is determined, the preset picture-level syntax identification information is encoded, and the obtained encoded bits are written into the bitstream.

[0122] For example, the picture level syntax identification information can be represented by picture_ifs_enable_flag. The picture header definition of the preset picture level syntax identification information can be as follows:

[0123] In some embodiments, when determining the target virtual reference image, the encoder inserts the target virtual reference image into the forward prediction reference image list and the backward prediction reference image list corresponding to the current image, performs inter-frame prediction on the current image based on the forward prediction reference image list and the backward prediction reference image list, determines the predicted image corresponding to the current image, and then encodes the current image based on the predicted image, generates encoding information of the current image and writes it into the bitstream.

[0124] Exemplarily, the process of generating the target virtual reference image in the above S102-S103 can be implemented by an input information sequence adjustment (Input Frame Switch, IFS) unit and a prediction network unit. Exemplarily, the prediction network unit can be implemented by an NNVRF network. Based on Figure 6, the positions of the IFS unit 111 and the prediction network unit 112 in the encoder can be as shown in Figure 14. The order of at least two reconstructed reference images input to the prediction network unit 112 is adjusted by the IFS unit 111, so that the prediction network unit 112 generates a first virtual reference image and a second virtual reference image. And the first virtual reference image and the second virtual reference image are compared with the rate-distortion cost of the current image by the IFS unit 111 to determine the target virtual reference image.

[0125] It is understood that at the encoder, virtual reference image prediction can be performed on at least two reconstructed reference images corresponding to the current image according to a first temporal order consistent with the original temporal order and a second temporal order opposite to the original temporal order, respectively, to determine a first virtual reference image and a second virtual reference image. A target virtual reference image for inter-frame prediction of the current image is determined from the first and second virtual reference images based on the distortion cost, and picture-level syntax identification information is determined based on the temporal order corresponding to the target virtual reference images. Thus, upon parsing the picture-level syntax identification information from the bitstream, the decoder can perform virtual reference image prediction on the at least two reconstructed reference images corresponding to the current image according to the temporal order represented by the picture-level syntax identification information, and determine a target virtual reference image for inter-frame prediction of the current image. In this way, by comparing virtual reference images predicted based on reconstructed reference images of different temporal orders, a target virtual reference image with a higher correlation with the current image and lower distortion can be determined for inter-frame prediction without increasing decoding complexity, thereby improving coding performance.

[0126] In one embodiment of the present application, referring to FIG15 , a flowchart of a decoding method provided by an embodiment of the present application is shown. As shown in FIG15 , the method may include:

[0127] S201: Parse the code stream to determine the picture-level syntax identification information corresponding to the current picture.

[0128] S202: Determine at least two reconstructed reference images corresponding to the current image.

[0129] S203 : Determine, according to the picture-level syntax identification information, to perform virtual reference image prediction on at least two reconstructed reference images in the first temporal order or the second temporal order, and determine a target virtual reference image.

[0130] It should be noted that the decoding method of the embodiments of the present application is applied to a decoder. Furthermore, the decoding method may specifically refer to an inter-frame prediction method. The method primarily improves the process of generating virtual reference frames based on the NNVRF scheme during inter-frame prediction, thereby improving the encoding performance based on inter-frame prediction.

[0131] It should also be noted that, in this embodiment of the present application, the picture-level syntax identification information is used to indicate whether to use a first temporal order or a second temporal order for virtual reference image prediction of at least two reconstructed reference images corresponding to the current image to determine a target virtual reference image. The first temporal order represents the original temporal order of the at least two reconstructed reference images; the second temporal order is the opposite of the first temporal order; and the target virtual reference image is used for inter-frame prediction of the current image.

[0132] In some embodiments, if the value of the image-level syntax identification information is the third value, then the image-level syntax identification information indicates that virtual reference image prediction is performed on at least two reconstructed reference images input in the first time domain order; if the value of the image-level syntax identification information is the fourth value, then the image-level syntax identification information indicates that virtual reference image prediction is performed on at least two reconstructed reference images input in the second time domain order. In an embodiment of the present application, the third value and the fourth value are different. Here, the third value and the fourth value can be in parameter form or in digital form. Specifically, the image-level syntax identification information here can be a parameter written in the profile or the value of a flag, which is not specifically limited here.

[0133] For example, the picture-level syntax identification information can be represented by picture_ifs_enable_flag. The third value can be set to 0, and the fourth value can be set to 1; alternatively, the third value can be set to false, and the fourth value can be set to true. If picture_ifs_enable_flag is 0, it indicates that at least two reconstructed reference images are input into the prediction network in the first temporal order for virtual reference image prediction. If picture_ifs_enable_flag is 1, it indicates that at least two reconstructed reference images are input into the prediction network in the second temporal order for virtual reference image prediction.

[0134] That is, in the embodiment of the present application, when the picture-level syntax identification information is the third value, the prediction network performs virtual reference image prediction on at least two reconstructed reference images input in the first temporal order to determine the target virtual reference image; when the picture-level syntax identification information is the fourth value, the prediction network performs virtual reference image prediction on at least two reconstructed reference images input in the second temporal order to determine the target virtual reference image. In the embodiment of the present application, the structure and processing of the prediction network on the decoder side are consistent with those described on the encoder side and will not be repeated here.

[0135] In some embodiments, the decoder predicts optical flow information of at least two reconstructed reference images input in a second time domain order through an interpolation algorithm module in the prediction network to determine at least two second optical flow feature information; extracts feature information of at least two reconstructed reference images respectively through a convolution module in the prediction network to determine at least two initial feature information; downsamples at least two second optical flow feature information and at least two initial feature information at least once through a feature processing module in the prediction network, and extracts feature information of at least two second optical flow feature information and at least two initial feature information each time downsampled to determine second intermediate feature information of at least one scale; aggregates the second intermediate feature information of at least one scale and performs prediction through a feature aggregation module in the prediction network to determine a target virtual reference image.

[0136] In some embodiments, before parsing to obtain the picture-level syntax identification information, the decoder parses the code stream to determine the preset sequence-level syntax identification information corresponding to the current sequence where the current image is located; when the preset sequence-level syntax identification information is the first value, the decoder continues to parse the code stream to determine the picture-level syntax identification information.

[0137] In this embodiment of the present application, when the value of the preset sequence-level syntax identification information is the first value, it indicates that virtual reference image prediction of the first temporal order and the second temporal order is enabled for the current sequence in which the current image resides. In other words, virtual reference image prediction of the first temporal order and the second temporal order is enabled for each image in the current sequence. The description of the preset sequence-level syntax identification information is consistent with that on the encoder side and is not repeated here.

[0138] In some embodiments, when the value of the preset sequence-level syntax identification information is a second value, it indicates that virtual reference image prediction of the second temporal order is not enabled for the current sequence. When the value of the preset sequence-level syntax identification information is the second value, the decoder determines to perform virtual reference image prediction on at least two reconstructed reference images in the first temporal order to determine a target virtual reference image. That is, when the value of the preset sequence-level syntax identification information is the second value, the decoder no longer continues to parse the picture-level syntax identification information. For each picture in the current sequence, virtual reference image prediction is performed according to the original temporal order of the at least two reconstructed reference pictures corresponding to the picture to determine the target virtual reference picture.

[0139] In some embodiments, after determining the target virtual reference image, the decoder inserts the target virtual reference image into the forward prediction reference image list and the backward prediction reference image list corresponding to the current image. Inter-frame prediction decoding is then performed on the current image based on the forward prediction reference image list and the backward prediction reference image list to determine the predicted image corresponding to the current image. Furthermore, inter-frame prediction decoding can be performed based on the predicted image corresponding to the current image to obtain a decoded reconstructed image corresponding to the current image.

[0140] It is understood that, when the decoder parses the picture-level syntax identifier information from the bitstream, it can perform virtual reference image prediction on at least two reconstructed reference images corresponding to the current image based on the temporal order represented by the picture-level syntax identifier information, and determine a target virtual reference image for use in inter-frame prediction of the current image. In this way, a target virtual reference image with higher correlation and lower distortion with the current image can be determined for inter-frame prediction without increasing decoding complexity, thereby improving decoding performance.

[0141] Below, taking application in actual scenarios as an example, an encoding method applied to an encoder in an embodiment of the present application is introduced as follows:

[0142] Based on the encoder structure in Figure 14, when constructing a virtual reference image, the encoder first extracts two reconstructed images from the decoded image buffer (DPB) as two reconstructed reference images according to the input image information required by the NNVRF technology. These images are then used as input information for the NNVRF network. The following steps are then performed:

[0143] Step a) Determine whether the IFS module is enabled for the current sequence based on the sequence-level flag ifs_enable_flag (equivalent to the preset sequence-level syntax identification information). If ifs_enable_flag is 1, the IFS module is enabled for the current sequence, and the process jumps to step b). If ifs_enable_flag is 0, the IFS module is not enabled for the current sequence, and the process jumps to step c).

[0144] Step b) For the current image in the current sequence, swap the order of the two reconstructed reference images input to the NNVRF network, then input them back into the NNVRF network, and predict and output a virtual image vrf_switch (equivalent to the second virtual reference image). Jump to step c).

[0145] Step c) For the current image in the current sequence, input the two reconstructed reference images into the NNVRF network in their original temporal order, and predict the output to obtain the virtual image vrf_normal (equivalent to the first virtual reference image). If ifs_enable_flag is 1, jump to step d); if ifs_enable_flag is 0, jump to step f).

[0146] Step d) Compare the virtual image vrf_switch and the virtual image vrf_normal with the original image of the current image and calculate the distortion cost. The distortion cost determined by comparing the virtual image vrf_normal with the current image is calculated as D normal The distortion cost determined by comparing the virtual image vrf_switch with the current image is calculated as D IFS Comparing the two distortion costs, if D IFS <D normal , then the input information interchange permission flag picture_ifs_enable_flag (equivalent to the picture-level syntax identification information) corresponding to the current image is assigned to 1, and the virtual image vrf_switch is used as the target virtual reference image corresponding to the current image; if D IFS ≥D normal , the input information interchange enable flag picture_ifs_enable_flag corresponding to the current image is set to 0, and the virtual image vrf_normal is used as the target virtual reference image. Jump to step e).

[0147] Step e) encodes the input information interchange enabling flag picture_ifs_enable_flag of the current image into the bitstream, and then jumps to step f).

[0148] Step f) inserting the target virtual reference image into the reference image lists L0 and L1, and performing inter-frame coding on the current image. If the current image has been processed, the next image is loaded for processing and the process jumps to step b).

[0149] For example, the code of the above encoding process may be as follows:

[0150] Based on the above encoding process, a decoding method applied to a decoder according to an embodiment of the present application is introduced as follows:

[0151] When the decoder constructs a virtual reference image, it first extracts two reconstructed images from the decoded image buffer (DPB) as two reconstructed reference images according to the input image information required by the NNVRF technology. It then uses these images as input information for the NNVRF network. The following steps are then performed:

[0152] In step a'), the IFS module is determined to be enabled for the current sequence based on the sequence-level flag ifs_enable_flag. If ifs_enable_flag is 1, the IFS module is enabled for the current sequence, and the process jumps to step b'). If ifs_enable_flag is 0, the IFS module is not enabled for the current sequence, and the process jumps to step d').

[0153] In step b'), for the current picture in the current sequence, the input information interchange enable flag picture_ifs_enable_flag of the current picture is analyzed. If picture_ifs_enable_flag is 1, the process jumps to step c'); if picture_ifs_enable_flag is 0, the process jumps to step d').

[0154] In step c'), for the current image in the current sequence, the order of the two reconstructed reference images input to the network is swapped, and then the images are input to the NNVRF network. The predicted output is the virtual image vrf_switch, which serves as the target virtual reference image. The process then jumps to step e').

[0155] In step d'), for the current image in the current sequence, input the two reconstructed reference images into the NNVRF network in their original temporal order, and predict the output virtual image vrf_normal as the target virtual reference image. The process then proceeds to step e').

[0156] Step e') inserts the target virtual reference image vrf_switch or vrf_normal into the reference image lists L0 and L1, and performs inter-frame decoding on the current image. If the current image has been processed, the next image is loaded for processing and the process jumps to step b').

[0157] Based on the existing NNVRF tool, the applicant implemented the encoding and decoding method of this application and tested its performance. Under the general test conditions of Random Access configuration, the JVET-specified general sequence was tested. The performance of some sequences was compared with the current NNVRF network that does not adjust the input order. The results are shown in Table 2.

[0158] Table 2

[0159] As can be seen from Table 2, the encoding and decoding method of the present embodiment can improve encoding and decoding performance based on the neural network virtual reference image generation technology. Furthermore, due to the coarse granularity of frame-level decision making, performance is currently better at small resolutions without increasing decoding complexity, further improving encoding and decoding performance.

[0160] In yet another embodiment of the present application, based on the same inventive concept as the aforementioned embodiment, FIG16 is a schematic diagram showing the composition structure of an encoder provided by an embodiment of the present application. As shown in FIG16 , the encoder 130 may include: a second determination part 1301, a second prediction part 1302, and a third determination part 1303; wherein:

[0161] The second determining part 1301 is configured to determine at least two reconstructed reference images corresponding to the current image;

[0162] The second prediction section 1302 is configured to perform virtual reference image prediction based on the at least two reconstructed reference images according to a first temporal order and a second temporal order, respectively, to determine a first virtual reference image and a second virtual reference image; the first temporal order represents an original temporal order of the at least two reconstructed reference images; and the second temporal order is opposite to the first temporal order.

[0163] The third determination part 1302 is configured to determine the target virtual reference image from the first virtual reference image and the second virtual reference image based on the distortion cost, and determine the image-level syntax identification information based on the time domain order corresponding to the target virtual reference image; the target virtual reference image is used to perform inter-frame prediction on the current image.

[0164] In some embodiments, the third determination part 1303 is further configured to determine a first distortion cost between the first virtual reference image and the current image, and a second distortion cost between the second virtual reference image and the current image; and determine the target virtual reference image based on the first distortion cost and the second distortion cost.

[0165] In some embodiments, the second prediction part 1302 is further configured to determine preset sequence-level syntax identification information; when the preset sequence-level syntax identification information is a first value, based on the at least two reconstructed reference images, virtual reference image prediction is performed according to the first time domain order and the second time domain order, respectively, to determine the first virtual reference image and the second virtual reference image; the first value representation enables virtual reference image prediction of the first time domain order and the second time domain order.

[0166] In some embodiments, the second prediction part 1302 is further configured to determine candidate image blocks at the same position in the at least two reconstructed reference images; to traverse the candidate image blocks in each of the at least two reconstructed reference images in the first time domain order to perform virtual reference image prediction and determine the first virtual reference image; and to traverse the candidate image blocks in each of the at least two reconstructed reference images in the second time domain order to perform virtual reference image prediction and determine the second virtual reference image.

[0167] In some embodiments, the second determining part 1301 is further configured to determine each image block in the current image as the candidate image block.

[0168] In some embodiments, the second determining part 1301 is further configured to use some image blocks in the current image as the candidate image blocks.

[0169] In some embodiments, the second determining part 1301 is further configured to determine at least one image block in the current image at every preset interval as the candidate image block.

[0170] In some embodiments, the third determination part 1303 is further configured to determine the first distortion cost based on the distortion cost between the candidate image block in the first virtual reference image and the image block at the same position in the current image; and determine the second distortion cost based on the distortion cost between the candidate image block in the second virtual reference image and the image block at the same position in the current image.

[0171] In some embodiments, the third determination part 1303 is further configured to perform inter-frame prediction coding based on the first virtual reference image, determine first candidate coding information, and determine the first distortion cost by determining the coding cost of the first candidate coding information; perform inter-frame prediction coding based on the second virtual reference image, determine second candidate coding information, and determine the second distortion cost by determining the coding cost of the second candidate coding information.

[0172] In some embodiments, the second prediction part 1302 is further configured to perform virtual reference image prediction on the at least two reconstructed reference images input according to the first time domain order through a prediction network to determine the target virtual reference image when the preset sequence-level syntax identification information is a second value; the second value indicates that the virtual reference image prediction of the second time domain order is not enabled.

[0173] In some embodiments, the encoder 130 further includes an encoding part, which is configured to encode the preset sequence-level syntax identification information and write the obtained encoded bits into a bitstream.

[0174] In some embodiments, the third determination part 1303 is further configured to set the value of the image-level syntax identification information to a third value when the time domain order corresponding to the target virtual reference image is the first time domain order; and to set the value of the image-level syntax identification information to a fourth value when the time domain order corresponding to the target virtual reference image is the second time domain order.

[0175] In some embodiments, the encoding part is further configured to encode the preset picture-level syntax identification information and write the obtained encoded bits into the bitstream.

[0176] In some embodiments, the second prediction part 1302 is further configured to perform optical flow information prediction on the candidate image blocks in the at least two reconstructed reference images input according to the second time domain sequence through a frame interpolation algorithm module in the prediction network to determine at least two second optical flow feature information;

[0177] The convolution module in the prediction network is used to extract feature information of the candidate image blocks in the at least two reconstructed reference images respectively to determine at least two initial feature information; the feature processing module in the prediction network is used to downsample the at least two second optical flow feature information and the at least two initial feature information at least once, and feature information extraction is performed on the at least two second optical flow feature information and the at least two initial feature information each time downsampled to determine second intermediate feature information of at least one scale; the feature aggregation prediction module in the prediction network is used to aggregate the second intermediate feature information of at least one scale and perform prediction to determine the second virtual reference image.

[0178] In some embodiments, the encoding part is further configured to insert the target virtual reference image into the forward prediction reference image list and the backward prediction reference image list corresponding to the current image, perform inter-frame prediction on the current image according to the forward prediction reference image list and the backward prediction reference image list, and determine the predicted image corresponding to the current image.

[0179] It should be noted that the description of the above device embodiment is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. For technical details not disclosed in the device embodiment of the present invention, please refer to the description of the method embodiment of the present invention for understanding.

[0180] It is understandable that in the embodiments of the present application, a "part" can be a part of a circuit, a part of a processor, a part of a program or software, etc., and of course it can also be a module, or it can be non-modular. Moreover, the various components in this embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional modules.

[0181] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the portion that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0182] Therefore, an embodiment of the present application provides a storage medium (i.e., a computer-readable storage medium), which is applied to the encoder 130. The computer-readable storage medium stores a computer program, and when the computer program is executed by the first processor, it implements the encoding method described in any one of the aforementioned embodiments.

[0183] Based on the composition of the above-mentioned encoder 130 and the computer-readable storage medium, refer to Figure 17, which shows a specific hardware structure diagram of the encoder 130 provided in an embodiment of the present application. As shown in Figure 17, the encoder 130 may include: a second communication interface 1401, a second memory 1402 and a second processor 1403; each component is coupled together through a second bus system 1404. It can be understood that the second bus system 1404 is used to realize the connection and communication between these components. In addition to the data bus, the second bus system 1404 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are labeled as the second bus system 1404 in Figure 17. Among them:

[0184] The second communication interface 1401 is used to receive and send signals when sending and receiving information with other external network elements;

[0185] The second memory 1402 is used to store computer programs that can be run on the second processor 1403;

[0186] The second processor 1403 is configured to, when running the computer program, execute:

[0187] determining at least two reconstructed reference images corresponding to the current image;

[0188] Based on the at least two reconstructed reference images, performing virtual reference image prediction according to a first temporal order and a second temporal order, respectively, to determine a first virtual reference image and a second virtual reference image; the first temporal order represents an original temporal order of the at least two reconstructed reference images; and the second temporal order is opposite to the first temporal order;

[0189] According to the distortion cost, a target virtual reference image is determined from the first virtual reference image and the second virtual reference image, and image-level syntax identification information is determined according to the time domain order corresponding to the target virtual reference image; the target virtual reference image is used to perform inter-frame prediction on the current image.

[0190] It is understood that the second memory 1402 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The second memory 1402 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0191] The second processor 1403 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in the second processor 1403. The above-mentioned second processor 1403 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the second memory 1402 , and the second processor 1403 reads the information in the second memory 1402 and completes the steps of the above method in combination with its hardware.

[0192] It is to be understood that these embodiments described in the present application can be implemented with hardware, software, firmware, middleware, microcode or its combination.For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (Application Specific Integrated Circuits, ASIC), digital signal processor (Digital Signal Processing, DSP), digital signal processing equipment (DSP Device, DSPD), programmable logic device (Programmable Logic Device, PLD), field programmable gate array (Field-Programmable Gate Array, FPGA), general-purpose processor, controller, microcontroller, microprocessor, other electronic units for performing functions described in the present application or its combination.For software implementation, the technology described in the present application can be realized by the module (such as process, function etc.) that performs functions described in the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.

[0193] Optionally, as another embodiment, the second processor 1403 is further configured to execute the encoding method described in any one of the aforementioned embodiments when running the computer program.

[0194] In yet another embodiment of the present application, based on the same inventive concept as the aforementioned embodiment, FIG18 is a schematic diagram showing the structure of a decoder provided by the embodiment of the present application. As shown in FIG18 , the decoder 150 may include: a parsing part 1501, a first determination part 1502, and a first prediction part 1503; wherein:

[0195] Parsing section 1501 is configured to parse the code stream and determine the picture-level syntax identification information corresponding to the current picture;

[0196] The first determining part 1502 is configured to determine at least two reconstructed reference images corresponding to the current image.

[0197] The first prediction part 1503 is configured to determine the virtual reference image prediction of the at least two reconstructed reference images in the first time domain order or the second time domain order according to the image-level syntax identification information, and determine the target virtual reference image; the first time domain order represents the original time domain order of the at least two reconstructed reference images; the second time domain order is opposite to the first time domain order; the target virtual reference image is used to perform inter-frame prediction on the current image.

[0198] In some embodiments, the first prediction part 1503 is further configured to perform virtual reference image prediction on the at least two reconstructed reference images input in the first time domain order through a prediction network when the image level syntax identification information is a third value, and determine the target virtual reference image.

[0199] In some embodiments, the first prediction part 1503 is further configured to perform virtual reference image prediction on the at least two reconstructed reference images input in the second time domain order through a prediction network when the image level syntax identification information is a fourth value, and determine the target virtual reference image.

[0200] In some embodiments, the parsing part 1501 is further configured to parse the code stream to determine the preset sequence-level syntax identification information corresponding to the current sequence where the current image is located; when the preset sequence-level syntax identification information is a first value, continue to parse the code stream to determine the image-level syntax identification information; the first value represents that the virtual reference image prediction of the first time domain order and the second time domain order is enabled for the current sequence where the current image is located.

[0201] In some embodiments, the first prediction part 1503 is further configured to determine, when the preset sequence-level syntax identification information is a second value, to perform virtual reference image prediction on the at least two reconstructed reference images in the first time domain order to determine the target virtual reference image; the second value indicates that the virtual reference image prediction of the second time domain order is not enabled for the current sequence where the current image is located.

[0202] In some embodiments, the first prediction part 1503 is further configured to perform optical flow information prediction on the at least two reconstructed reference images input in the second time domain order through the interpolation algorithm module in the prediction network, and determine at least two second optical flow feature information; extract the feature information of the at least two reconstructed reference images respectively through the convolution module in the prediction network, and determine at least two initial feature information; downsample the at least two second optical flow feature information and the at least two initial feature information at least once through the feature processing module in the prediction network, and extract feature information of the at least two second optical flow feature information and the at least two initial feature information each time downsampled, and determine second intermediate feature information of at least one scale; aggregate the second intermediate feature information of at least one scale and perform prediction through the feature aggregation module in the prediction network to determine the target virtual reference image.

[0203] In some embodiments, the decoder 150 also includes a decoding part, which is configured to insert the target virtual reference image into the forward prediction reference image list and the backward prediction reference image list corresponding to the current image, perform inter-frame prediction decoding on the current image according to the forward prediction reference image list and the backward prediction reference image list, and determine the predicted image corresponding to the current image.

[0204] It should be noted that the description of the above device embodiment is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. For technical details not disclosed in the device embodiment of the present invention, please refer to the description of the method embodiment of the present invention for understanding.

[0205] It is understood that in this embodiment, a "portion" may be a circuit portion, a processor portion, a program portion, or software portion, and may also be a module or non-modular. Furthermore, the various components in this embodiment may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional modules.

[0206] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, this embodiment provides a storage medium (i.e., a computer-readable storage medium) for use in decoder 150. The computer-readable storage medium stores a computer program that, when executed by a second processor, implements any of the decoding methods described in the aforementioned embodiments.

[0207] Based on the composition of the above-mentioned decoder 150 and the computer-readable storage medium, refer to Figure 19, which shows a specific hardware structure diagram of the decoder 150 provided in an embodiment of the present application. As shown in Figure 19, the decoder 150 may include: a first communication interface 1601, a first memory 1602 and a first processor 1603; each component is coupled together through a first bus system 1604. It can be understood that the first bus system 1604 is used to realize the connection and communication between these components. In addition to the data bus, the first bus system 1604 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are labeled as the first bus system 1604 in Figure 19. Among them:

[0208] The first communication interface 1601 is used to receive and send signals when sending and receiving information with other external network elements;

[0209] A first memory 1602 is used to store computer programs that can be run on the first processor 1603;

[0210] The first processor 1603 is configured to, when running the computer program, execute:

[0211] Parse the code stream to determine the image-level syntax identification information corresponding to the current image;

[0212] Determining at least two reconstructed reference images corresponding to the current image;

[0213] According to the image-level syntax identification information, it is determined to perform virtual reference image prediction on the at least two reconstructed reference images in a first time domain order or a second time domain order, and a target virtual reference image is determined; the first time domain order represents the original time domain order of the at least two reconstructed reference images; the second time domain order is opposite to the first time domain order; the target virtual reference image is used to perform inter-frame prediction on the current image.

[0214] Optionally, as another embodiment, the first processor 1603 is further configured to execute the decoding method described in any one of the aforementioned embodiments when running the computer program.

[0215] It can be understood that the hardware functions of the first memory 1602 and the second memory 1402 are similar, and the hardware functions of the first processor 1603 and the second processor 1403 are similar; they are not described in detail here.

[0216] In yet another embodiment of the present application, referring to FIG20 , a schematic diagram of the structure of a coding and decoding system provided by an embodiment of the present application is shown. As shown in FIG20 , the coding and decoding system 170 may include an encoder 1701 and a decoder 1702 .

[0217] In the embodiment of the present application, the encoder 1701 may be the encoder described in any one of the aforementioned embodiments, and the decoder 1702 may be the decoder described in any one of the aforementioned embodiments.

[0218] It should be noted that, in this application, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0219] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0220] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0221] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0222] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0223] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims. Industrial Applicability

[0224] At the encoder end, virtual reference image prediction can be performed on at least two reconstructed reference images corresponding to the current image according to a first time domain order that is consistent with the original time domain order, and a second time domain order that is opposite to the original time domain, respectively, to determine the first virtual reference image and the second virtual reference image. The target virtual reference image for inter-frame prediction of the current image is determined from the first virtual reference image and the second virtual reference image based on the distortion cost, and the image-level syntax identification information is determined based on the time domain order corresponding to the target virtual reference image. In this way, when the decoder parses the image-level syntax identification information from the bitstream, it can perform virtual reference image prediction on at least two reconstructed reference images corresponding to the current image according to the time domain order represented by the image-level syntax identification information, and determine the target virtual reference image to be applied to inter-frame prediction of the current image. The present application can determine the target virtual reference image with higher correlation and lower distortion with the current image for inter-frame prediction by comparing the virtual reference images predicted based on reconstructed reference images with different time domain orders, without increasing the decoding complexity, thereby improving the encoding and decoding performance.

Claims

1. A decoding method, comprising: Analyzing a bitstream to determine image-level syntax identification information corresponding to a current image; Determining at least two reconstructed reference images corresponding to the current image; According to the image-level syntax identification information, determining to perform virtual reference image prediction on the at least two reconstructed reference images in a first temporal order or a second temporal order to determine a target virtual reference image; The first temporal order represents the original temporal order of the at least two reconstructed reference images; the second temporal order is opposite to the first temporal order; The target virtual reference image is used for inter-frame prediction of the current image.

2. The method according to claim 1, wherein The step of, according to the image-level syntax identification information, determining to perform virtual reference image prediction on the at least two reconstructed reference images in a first temporal order or a second temporal order to determine a target virtual reference image, includes: When the image-level syntax identification information is a third value, performing virtual reference image prediction on the at least two reconstructed reference images input in the first temporal order through a prediction network to determine the target virtual reference image.

3. The method according to claim 1, wherein The step of, according to the image-level syntax identification information, determining to perform virtual reference image prediction on the at least two reconstructed reference images in a first temporal order or a second temporal order to determine a target virtual reference image, includes: When the image-level syntax identification information is a fourth value, performing virtual reference image prediction on the at least two reconstructed reference images input in the second temporal order through a prediction network to determine the target virtual reference image.

4. The method according to claim 1, wherein, The step of analyzing a bitstream to determine image-level syntax identification information corresponding to a current image includes: Analyzing the bitstream to determine preset sequence-level syntax identification information corresponding to a current sequence where the current image is located; When the preset sequence-level syntax identification information is a first value, continuing to analyze the bitstream to determine the image-level syntax identification information; the first value represents enabling virtual reference image prediction in the first temporal order and the second temporal order for the current sequence where the current image is located.

5. The method according to claim 4, wherein The method further includes: When the preset sequence-level syntax identification information is a second value, determining to perform virtual reference image prediction on the at least two reconstructed reference images in the first temporal order to determine the target virtual reference image; the second value represents not enabling virtual reference image prediction in the second temporal order for the current sequence where the current image is located.

6. The method according to claim 3, wherein The step of, through a prediction network, performing virtual reference image prediction on the at least two reconstructed reference images input in the second temporal order to determine the target virtual reference image, includes: Through an interpolation algorithm module in the prediction network, performing optical flow information prediction on the at least two reconstructed reference images input in the second temporal order to determine at least two second optical flow feature information; Through a convolution module in the prediction network, respectively extracting feature information of the at least two reconstructed reference images to determine at least two initial feature information; Through the feature processing module in the prediction network, perform at least one downsampling on the at least two second optical flow feature information and the at least two initial feature information, and extract feature information from the at least two second optical flow feature information and the at least two initial feature information for each downsampling to determine second intermediate feature information at at least one scale; Through the feature aggregation module in the prediction network, aggregate the second intermediate feature information at at least one scale and perform prediction to determine the target virtual reference image.

7. The method according to any one of claims 1-6, wherein, The method further includes: Insert the target virtual reference image into the forward prediction reference image list and the backward prediction reference image list corresponding to the current image, and perform inter-frame prediction decoding on the current image according to the forward prediction reference image list and the backward prediction reference image list to determine the prediction image corresponding to the current image.

8. An encoding method, including: Determine at least two reconstructed reference images corresponding to the current image; Based on the at least two reconstructed reference images, perform virtual reference image prediction respectively in the first time domain order and the second time domain order to determine a first virtual reference image and a second virtual reference image; The first time domain order represents the original time domain order of the at least two reconstructed reference images; the second time domain order is opposite to the first time domain order; According to the distortion cost, determine the target virtual reference image from the first virtual reference image and the second virtual reference image, and determine the picture-level syntax identification information according to the time domain order corresponding to the target virtual reference image; the target virtual reference image is used for inter-frame prediction of the current image.

9. The method according to claim 8, wherein The determining the target virtual reference image from the first virtual reference image and the second virtual reference image according to the distortion cost includes: Determine a first distortion cost between the first virtual reference image and the current image, and a second distortion cost between the second virtual reference image and the current image; Determine the target virtual reference image according to the first distortion cost and the second distortion cost.

10. The method according to claim 8 or 9, wherein, The performing virtual reference image prediction respectively in the first time domain order and the second time domain order based on the at least two reconstructed reference images to determine the first virtual reference image and the second virtual reference image includes: Determine preset sequence-level syntax identification information; When the preset sequence-level syntax identification information is a first value, perform virtual reference image prediction respectively in the first time domain order and the second time domain order based on the at least two reconstructed reference images to determine the first virtual reference image and the second virtual reference image; the first value represents enabling virtual reference image prediction in the first time domain order and the second time domain order.

11. The method according to claim 10, wherein, The performing virtual reference image prediction respectively in the first time domain order and the second time domain order based on the at least two reconstructed reference images to determine the first virtual reference image and the second virtual reference image includes: Determine candidate image blocks at the same position in the at least two reconstructed reference images; Traverse each candidate image block in each of the at least two reconstructed reference images in the first temporal order to perform virtual reference image prediction, and determine the first virtual reference image; Traverse each candidate image block in each of the at least two reconstructed reference images in the second temporal order to perform virtual reference image prediction, and determine the second virtual reference image.

12. The method according to claim 11, wherein, The method further includes: Determine each image block in the current image as the candidate image block.

13. The method according to claim 11, wherein, The method further includes: Take some image blocks in the current image as the candidate image blocks.

14. The method according to claim 13, wherein, The method further includes: In the current image, determine at least one image block at a preset interval as the candidate image block.

15. The method according to any one of claims 11 - 14, wherein, The determining the first distortion cost between the first virtual reference image and the current image, and the second distortion cost between the second virtual reference image and the current image includes: Determine the first distortion cost according to the distortion cost between the candidate image block in the first virtual reference image and the image block at the same position in the current image; Determine the second distortion cost according to the distortion cost between the candidate image block in the second virtual reference image and the image block at the same position in the current image.

16. The method according to any one of claims 11-14, wherein, The determining the first distortion cost between the first virtual reference image and the current image, and the second distortion cost between the second virtual reference image and the current image includes: Perform inter-frame prediction coding based on the first virtual reference image to determine first candidate coding information, and determine the first distortion cost by determining the coding cost of the first candidate coding information; Perform inter-frame prediction coding based on the second virtual reference image to determine second candidate coding information, and determine the second distortion cost by determining the coding cost of the second candidate coding information.

17. The method according to claim 10, wherein, The method further includes: In the case where the preset sequence-level syntax identification information is the second value, perform virtual reference image prediction on the at least two reconstructed reference images input in the first temporal order through a prediction network, and determine the target virtual reference image; the second value indicates that the virtual reference image prediction in the second temporal order is not enabled.

18. The method according to claim 10, wherein The method further includes: Encode the preset sequence-level syntax identification information, and write the obtained encoded bits into the bitstream.

19. The method according to claim 18, wherein The determining the image-level syntax identification information according to the temporal order corresponding to the target virtual reference image includes: In the case where the temporal order corresponding to the target virtual reference image is the first temporal order, set the value of the image-level syntax identification information to the third value; In the case where the temporal order corresponding to the target virtual reference image is the second temporal order, set the value of the image-level syntax identification information to the fourth value.

20. The method according to claim 19, wherein The method further includes: Encode the preset image-level syntax identification information, and write the obtained encoded bits into the bitstream.

21. The method according to claim 11, wherein, The traversing each candidate image block in each of the at least two reconstructed reference images in the second temporal order to perform virtual reference image prediction includes: Through the interpolation algorithm module in the prediction network, for the at least two reconstructed reference images input in the second time domain order perform optical flow information prediction on the candidate image blocks, and determine at least two second optical flow feature information; Through the convolution module in the prediction network, extract the feature information of the candidate image blocks in the at least two reconstructed reference images respectively, and determine at least two initial feature information; Through the feature processing module in the prediction network, perform at least one downsampling on the at least two second optical flow feature information and the at least two initial feature information, and extract feature information from the at least two second optical flow feature information and the at least two initial feature information for each downsampling, and determine second intermediate feature information of at least one scale; Through the feature aggregation prediction module in the prediction network, aggregate the second intermediate feature information of at least one scale and perform prediction to determine the second virtual reference image.

22. The method according to any one of claims 8, 9, 11 - 14, 17 - 21, wherein The method further includes: Insert the target virtual reference image into the forward prediction reference image list and the backward prediction reference image list corresponding to the current image, and perform inter-frame prediction on the current image according to the forward prediction reference image list and the backward prediction reference image list to determine the prediction image corresponding to the current image.

23. A decoder, comprising: A parsing part configured to parse the bitstream to determine the image-level syntax identification information corresponding to the current image; A first determination part configured to determine at least two reconstructed reference images corresponding to the current image; A first prediction part configured to determine virtual reference image prediction for the at least two reconstructed reference images in the first time domain order or the second time domain order according to the image-level syntax identification information, and determine the target virtual reference image; The first time domain order represents the original time domain order of the at least two reconstructed reference images; the second time domain order is opposite to the first time domain order; The target virtual reference image is used for inter-frame prediction of the current image.

24. An encoder, comprising: A second determination part configured to determine at least two reconstructed reference images corresponding to the current image; A second prediction part configured to perform virtual reference image prediction on the at least two reconstructed reference images respectively in the first time domain order and the second time domain order to determine a first virtual reference image and a second virtual reference image; The first time domain order represents the original time domain order of the at least two reconstructed reference images; the second time domain order is opposite to the first time domain order; A third determination part configured to determine the target virtual reference image from the first virtual reference image and the second virtual reference image according to the distortion cost, and determine the image-level syntax identification information according to the time domain order corresponding to the target virtual reference image; the target virtual reference image is used for inter-frame prediction of the current image.

25. A decoder, the decoder includes a first memory and a first processor; wherein, The first memory is used to store a computer program that can run on the first processor; The first processor is configured to execute the method according to any one of claims 1 to 7 when running the computer program.

26. An encoder, the encoder comprising a second memory and a second processor; wherein, The second memory is configured to store a computer program that can run on the second processor; The second processor is configured to execute the method according to any one of claims 8 to 22 when running the computer program.

27. A bitstream, which is generated by performing bit encoding on information to be encoded; wherein, The information to be encoded at least includes at least one of the following: Picture-level syntax identification information; Wherein, the picture-level syntax identification information is used to indicate virtual reference picture prediction of at least two reconstructed reference pictures corresponding to the current picture in a first temporal order or a second temporal order to determine a target virtual reference picture.

28. A storage medium, wherein, The storage medium stores a computer program, and when the computer program is executed, it implements the method according to any one of claims 1 to 7, or implements the method according to any one of claims 8 to 22.

Citation Information

Patent Citations

  • Methods and equipment for encoding and decoding images

    CN103297778A

  • Hierarchical structure for neural network based tools in video coding

    CN115486065A

  • Apparatus for removal of fine dust

    KR102187786B1

  • Methods and systems of generating virtual reference picture for video processing

    WO2022104609A1