Inter-frame prediction method, encoding method, decoding method, and their devices

By obtaining and inserting the synthetic frames in the reference frame list, using their correlation with the current encoded frames, designing various application strategies, solving the problem of inflexible synthetic frame applications in the prior art and improving video encoding efficiency.

CN115460412BActive Publication Date: 2025-08-01ZHEJIANG DAHUA TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210952703.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-09
Publication Date
2025-08-01
Estimated Expiration
2042-08-09

AI Technical Summary

Technical Problem

In the existing inter-frame prediction technology based on neural networks, synthetic frames are only used simply to replace reference frames, lacking flexible and effective application strategies, resulting in inefficient encoding efficiency.

Method used

By obtaining the forward and backward reference frame lists of the current coded frame, using the reference frame synthesis network model to generate a composite frame, and insert it into the reference frame list to be updated, various application strategies are designed to improve coding efficiency based on the correlation between the composite frame and the current coded frame.

Benefits of technology

The video encoding efficiency is improved, and through the high correlation between the synthetic frame and the current encoding frame, an application strategy based on synthetic frames is designed to enhance the encoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115460412B_ABST
    Figure CN115460412B_ABST
Patent Text Reader

Abstract

The present application discloses an inter-frame prediction method, an encoding method, a decoding method, and a device thereof. The method includes: obtaining a forward reference frame list and a backward reference frame list of a current encoded frame; obtaining a synthesized frame of the current encoded frame based on an original forward reference frame in the forward reference frame list and an original backward reference frame in the backward reference frame list; inserting the synthesized frame into a reference frame list to be updated, where the reference frame list to be updated includes the forward reference frame list and / or the backward reference frame list; and performing prediction on the current encoded frame based on the reference frame list to be updated into which the synthesized frame is inserted. In the present application, the synthesized frame is inserted into the reference frame list to be updated for predicting the current encoded frame, fully utilizing the high correlation between the synthesized frame and the current encoded frame, thereby being able to improve the encoding efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of inter-frame prediction technology, and particularly relates to an inter-frame prediction method, an encoding method, a decoding method, and their devices. Background Art

[0002] With the development of technology, the amount of video image data is gradually increasing. When video is transmitted wirelessly, it is usually necessary to compress video pixel data, etc. The compressed data is called a video bitstream. The video bitstream is transmitted to the user terminal through a wired or wireless network and then decoded for viewing. The entire video encoding process includes processes such as block partitioning, prediction, transformation, quantization, and encoding.

[0003] In the existing neural network-based inter-frame prediction technology, the forward reference frame and the backward reference frame that are closest to the current encoded frame are used as network inputs, and then a synthetic frame is obtained using a 3D-Unet neural network. And the synthetic frame is used to replace the reference frames in the forward reference frame list and the backward reference frame list. In the prior art, the synthetic frame is only simply used to replace the existing frames in the reference list, lacking a flexible and effective synthetic frame application strategy, which is not conducive to improving the encoding efficiency. Summary of the Invention

[0004] This application proposes an inter-frame prediction method, an encoding method, a decoding method, and their devices to solve the above problems and improve the encoding efficiency of video.

[0005] To solve the above technical problems, a technical solution adopted by this application is: to provide an inter-frame prediction method. The method includes:

[0006] Obtain the forward reference frame list and the backward reference frame list of the current encoded frame; based on the original forward reference frames in the forward reference frame list and the original backward reference frames in the backward reference frame list, obtain the synthetic frame of the current encoded frame; insert the synthetic frame into the reference frame list to be updated, and the reference frame list to be updated includes the forward reference frame list and / or the backward reference frame list; perform prediction on the current encoded frame based on the reference frame list to be updated with the synthetic frame inserted.

[0007] Among them, inserting the synthetic frame into the reference frame list to be updated includes: inserting the synthetic frame at the end of the reference frame list to be updated.

[0008] Among them, inserting the synthetic frame into the reference frame list to be updated includes:

[0009] Obtain the first subjective and objective quality parameters of the current coded frame and the synthesized frame, obtain the second subjective and objective quality parameters of the current coded frame and the forward reference frame, and obtain the third subjective and objective quality parameters of the current coded frame and the backward reference frame; sort the first subjective and objective quality parameters and the second subjective and objective quality parameters according to their magnitude relationship to obtain a first sequence; sort the first subjective and objective quality parameters and the third subjective and objective quality parameters according to their magnitude relationship to obtain a second sequence; insert the synthesized frame into the forward reference frame list based on the position of the first subjective and objective quality parameters in the first sequence; and / or insert the synthesized frame into the backward reference frame list based on the position of the first subjective and objective quality parameters in the second sequence.

[0010] Among them, obtaining the synthesized frame of the current coded frame based on the original forward reference frame in the forward reference frame list and the original backward reference frame in the backward reference frame list includes:

[0011] Adjust the multiple component sizes of the original forward reference frame and the original backward reference frame to the same size threshold; obtain the synthesized frame of the current coded frame by using the size-adjusted original forward reference frame and the size-adjusted original backward reference frame; restore the component size of the synthesized frame to the component size corresponding to the current coded frame.

[0012] Among them, obtaining the synthesized frame of the current coded frame based on the original forward reference frame in the forward reference frame list and the original backward reference frame in the backward reference frame list includes:

[0013] Obtain the first information of the original forward reference frame in the forward reference frame list and the second information of the original backward reference frame in the backward reference frame list; obtain the information of the synthesized frame of the current coded frame based on the first information and the second information; among them, the first information includes the image information of the original forward reference frame, and the second information includes the image information of the original backward reference frame; or, the first information includes the image information of the original forward reference frame and the edge information of the original forward reference frame, and the second information includes the image information of the original backward reference frame and the edge information of the original backward reference frame.

[0014] Among them, the inter-frame prediction method further includes:

[0015] Obtain the first quantity of the original forward reference frame in the forward reference frame list and the second quantity of the backward reference frame in the backward reference frame list; in response to the first quantity being greater than the second quantity, obtain the original forward reference frame located at the head of the forward reference frame list and insert it at the end of the backward reference frame list; in response to the first quantity being less than the second quantity, obtain the backward reference frame located at the head of the backward reference frame list and insert it at the end of the forward reference frame list.

[0016] Among them, the reference frame list to be updated is the forward reference frame list and the backward reference frame list, and predicting the current coded frame based on the reference frame list to be updated with the synthesized frame inserted includes:

[0017] Predict the current encoded frame based on the synthesized frames and at least part of the original forward reference frames in the forward reference frame list, and the synthesized frames and at least part of the original backward reference frames in the backward reference frame list.

[0018] Wherein, the reference frame list to be updated is the forward reference frame list or the backward reference frame list. Predicting the current encoded frame based on the reference frame list to be updated with the inserted synthesized frame includes:

[0019] Predict the current encoded frame based on the synthesized frames and at least part of the original forward reference frames in the forward reference frame list with the inserted synthesized frame, and at least part of the original backward reference frames in the backward reference frame list without the inserted synthesized frame; or, predict the current encoded frame based on the synthesized frames and at least part of the original backward reference frames in the backward reference frame list with the inserted synthesized frame, and at least part of the original forward reference frames in the forward reference frame list without the inserted synthesized frame.

[0020] Wherein, predicting the current encoded frame based on the reference frame list to be updated with the inserted synthesized frame includes:

[0021] Obtain the synthesized frame in the reference frame list to be updated; set the motion vector of the synthesized frame to a zero vector to directly use the synthesized frame as the predicted frame of the current encoded frame.

[0022] To solve the above technical problems, a technical solution adopted by this application is: provide an encoding method, including obtaining the prediction information of the current encoded frame in any of the above items, and encoding the prediction information into the bitstream of the current encoded frame.

[0023] To solve the above technical problems, a technical solution adopted by this application is: provide a decoding method, including obtaining the prediction information of the current encoded frame in any of the above items and the bitstream of the current encoded frame, and decoding the bitstream.

[0024] To solve the above technical problems, a technical solution adopted by this application is: provide an encoder, which includes a processor and a memory; a computer program is stored in the memory, and the processor is used to execute the computer program to implement the above encoding method.

[0025] [[ID=)25]]To solve the above technical problems, a technical solution adopted by this application is: provide a decoder, which includes a processor and a memory; a computer program is stored in the memory, and the processor is used to execute the computer program to implement the above decoding method.

[0026] To solve the above technical problems, a technical solution adopted by this application is: provide an electronic device, which includes a processor and a memory connected to the processor. Program data is stored in the memory, and the processor executes the program data stored in the memory to execute and implement the inter-frame prediction method in any of the above items.

[0027] To solve the above technical problems, another technical solution adopted in this application is: to provide a computer-readable storage medium, which stores program instructions internally, and the program instructions are executed to implement the inter-frame prediction method of any one of the above.

[0028] The beneficial effect of this application is: different from the prior art, this application obtains the original forward reference frame in the forward reference frame list of the current coded frame and the original backward reference frame in the backward reference frame list, uses the original forward reference frame and the original backward reference frame to obtain the synthesized frame of the current coded frame, and inserts the synthesized frame into the reference frame list to be updated. The reference frame list to be updated includes the forward reference frame list and / or the backward reference frame list; and predicts the current coded frame based on the reference frame list to be updated with the synthesized frame inserted; the synthesized frame is obtained by using the original forward reference frames in multiple forward reference frame lists and the original backward reference frames in multiple backward reference frame lists. Therefore, the synthesized frame is relatively similar to the current coded frame, so the synthesized frame has a strong correlation with the current coded frame. When predicting the current coded frame, the synthesized frame has a high probability of being referenced. This application utilizes the correlation between the synthesized frame and the current coded frame, and can design various application strategies based on the synthesized frame, thereby improving the coding efficiency. Description of the Drawings

[0029] Figure 1 It is a schematic diagram of the reference relationship between frames of a group of pictures with a size of 32 frames;

[0030] Figure 2 It is a schematic flowchart of the first embodiment of the inter-frame prediction method of this application;

[0031] Figure 3 It is a schematic structural diagram of the first embodiment of the reference frame synthesis network model of this application;

[0032] Figure 4 It is a schematic structural diagram of the second embodiment of the reference frame synthesis network model of this application;

[0033] Figure 5 It is a schematic structural diagram of an embodiment of the residual neural network of this application;

[0034] Figure 6 It is a schematic structural diagram of an embodiment of the 3D residual neural network of this application;

[0035] Figure 7 It is a schematic diagram of indirect fusion of an embodiment of the fusion module of this application;

[0036] Figure 8 It is a schematic diagram of direct fusion of an embodiment of the fusion module of this application;

[0037] Figure 9It is a schematic structural diagram of an embodiment of the modulation module of the present application;

[0038] Figure 10 It is a schematic structural diagram of the first embodiment of the synthetic frame insertion of the present application;

[0039] Figure 11 It is a schematic structural diagram of the second embodiment of the synthetic frame insertion of the present application;

[0040] Figure 12 is Figure 2 A specific process schematic diagram of step S103 in;

[0041] Figure 13 It is a schematic structural diagram of the third embodiment of the synthetic frame insertion of the present application;

[0042] Figure 14 is the present application Figure 2 A process schematic diagram of the first embodiment of step S102 in;

[0043] Figure 15 It is a reference schematic diagram of the first embodiment of obtaining a synthetic frame by using the original forward reference frame and the original backward reference frame of the present application;

[0044] Figure 16 It is a schematic diagram of an embodiment of adjusting the reference frame component size of the present application;

[0045] Figure 17 It is a schematic diagram of an embodiment of adjusting the synthetic frame component size of the present application;

[0046] Figure 18 is the present application Figure 2 A process schematic diagram of the second embodiment of step S102 in;

[0047] Figure 19 It is a reference schematic diagram of the second embodiment of obtaining a synthetic frame by using the original forward reference frame and the original backward reference frame of the present application;

[0048] Figure 20 It is a matrix schematic diagram of an embodiment of the quantization parameter of the present application;

[0049] Figure 21 It is a reference schematic diagram of the third embodiment of obtaining a synthetic frame by using the original forward reference frame and the original backward reference frame of the present application;

[0050] Figure 22 is the present application Figure 2 A specific process schematic diagram of step S104 in;

[0051] Figure 23 It is a process schematic diagram of the second embodiment of the inter-frame prediction method of the present application;

[0052] Figure 24It is a schematic structural diagram of an encoder according to an embodiment of the present application;

[0053] Figure 25 It is a schematic structural diagram of a decoder according to an embodiment of the present application;

[0054] Figure 26 It is a schematic structural diagram of an electronic device according to an embodiment of the present application;

[0055] Figure 27 It is a schematic structural diagram of a computer-readable storage medium according to an embodiment of the present application. Detailed implementation manners

[0056] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0057] During the video wireless transmission process, it is usually necessary to compress the video pixel data. The compressed data is called a video bitstream. The video bitstream is transmitted to the user side through a wired or wireless network and then decoded for viewing. The entire video encoding process includes processes such as block partitioning, prediction, transformation, quantization, and encoding.

[0058] In video encoding, the most commonly used color encoding methods include YUV, RGB, etc. The color encoding method adopted in the present application is YUV. Y represents the luminance, that is, the grayscale value of the image; U and V (i.e., Cb and Cr) represent the chrominance, and their function is to describe the color and saturation of the image. Each Y luminance block corresponds to a Cb and a Cr chrominance block, and each chrominance block only corresponds to one luminance block. Taking the 4:2:0 sampling format as an example, a block of N*M corresponds to a luminance block of size N*M, and the sizes of the corresponding two chrominance blocks are both (N / 2)*(M / 2), and the chrominance block is 1 / 4 the size of the luminance block. For the 4:4:4 sampling format, the luminance block and the chrominance block have the same size.

[0059] Block partitioning means that during video encoding, the input is individual image frames. However, when encoding a frame of image, it is necessary to divide a frame into several maximum coding units, and then perform recursive coding unit partitioning of different sizes on each coding unit. Video encoding is carried out in units of coding units.

[0060] In video encoding, a video sequence is composed of multiple groups of pictures. A group of pictures can include three types of frames: I frames, P frames, and B frames, which are encoded using intra-frame prediction, intra / uni-directional inter-frame prediction, and intra / bi-directional inter-frame prediction methods respectively.

[0061] Intra prediction utilizes the correlation between pixels to remove spatial redundancy and temporal redundancy. Generally speaking, the luminance and chrominance signal values of adjacent pixel points are relatively close and have strong correlation. If the luminance and chrominance information is directly represented by the number of samples, there is a lot of spatial redundancy in the data. If the redundant data is removed first and then encoded, the average number of bits representing each pixel point will decrease, that is, data compression is performed by reducing spatial redundancy. Intra prediction usually includes processes such as prediction block partitioning, obtaining reference pixels, prediction mode selection, and prediction value filtering.

[0062] Inter prediction obtains the motion information of each block of the current image in the reference image by using the encoded image as the reference image of the current image, and is usually represented by a motion vector and a reference frame index. Generally speaking, the luminance and chrominance signal values of pixel points in adjacent frames in time are relatively close and have strong correlation. Inter prediction searches for the most similar matching block to the current block in the reference frame through methods such as motion search, and records the motion information between the current block and the matching block, such as the motion vector (Motion Vector, MV) and the reference frame index. The motion information is encoded and transmitted to the decoding end. At the decoding end, as long as the decoder parses the MV of the current block through the corresponding syntax elements, it can find the matching block of the current block. Among them, the MV includes two directions: horizontal and vertical. And the pixel values of the matching block are copied to the current block, which is the inter prediction value of the current block. There are various inter prediction techniques in inter prediction, such as the skip mode and the merge mode. In the merge mode, the encoder and decoder construct the MV list according to the same rules, and the list includes multiple MV candidates. Therefore, only the best MV index and the corresponding prediction residual need to be transmitted in the bitstream, rather than directly transmitting the MV. The skip mode is a special case of the merge mode, and only the best MV index is transmitted in the bitstream, without the need for a prediction residual.

[0063] In video coding, there is a hierarchical B-frame structure, which gives the reference, encoding, and display relationships between each frame image during the inter prediction process. Please refer to Figure 1 , Figure 1 is a schematic diagram of the reference relationship between each frame of a group of pictures with a size of 32 frames. Among them, POC is the image display order, and TID is the temporal layer number. During the inter prediction process based on this structure, the images in the higher temporal layer need to refer to the images in the lower temporal layer, so the encoding order of the images is different from the display order. The numbers in the colored boxes in the figure are the encoding order of the images.

[0064] The meaning of a B-frame is a bidirectional prediction frame. Therefore, a B-frame includes a forward reference frame list, a backward reference frame list, and the image display order of the reference frames in the list. In the prior art, the effective lengths of the forward reference list and the backward reference list are usually set to 2. Therefore, there are two reference frames in both the forward and backward directions of each B-frame.

[0065] In the existing inter-frame prediction technology, the forward and backward reference frames closest to the current coded frame are used as the network input, then a synthesized frame is obtained by using a neural network, and finally the quality is enhanced through a filtering neural network and the synthesized frame is output. The synthesized frame of the existing technology is used to replace the reference frames in the forward frame reference list and the backward reference frame list. Therefore, in the existing technology, the synthesized frame is only simply used to replace the reference frames in the forward reference frame list and the backward reference frame list, and the correlation between the synthesized frame and the current coded frame is not utilized, resulting in low coding efficiency.

[0066] To solve the above problems, this application first proposes an inter-frame prediction method. Please refer to Figure 2 , Figure 2 which is a schematic flowchart of the first embodiment of the inter-frame prediction method of this application. As Figure 2 shown, the method specifically includes steps S101 to S104:

[0067] Step S101: Obtain the forward reference frame list and the backward reference frame list of the current coded frame.

[0068] When an electronic device encodes a video sequence, it first obtains the forward reference frame list and the backward reference frame list of the current coded frame. Among them, in order to make more full use of the image information of the encoded frames, the effective lengths of the forward reference frame list and the backward reference frame list of the current coded frame can be set to N, where N is greater than or equal to 1. In this embodiment, the effective length N of the forward reference frame list and the backward reference frame list is not limited to 2.

[0069] Step S102: Obtain the synthesized frame of the current coded frame based on the original forward reference frame in the forward reference frame list and the original backward reference frame in the backward reference frame list.

[0070] After the electronic device obtains the forward reference frame list and the backward reference frame list of the current coded frame, it obtains the original forward reference frame in the forward reference frame list and the original backward reference frame in the backward reference frame list. If the number of the original forward reference frame and the original backward reference frame is different, the reverse reference frame is used as a supplement.

[0071] The obtained original forward reference frame and original backward reference frame are used as inputs and input into the reference frame synthesis network model, and the synthesized frame of the current coded frame is obtained by outputting through the reference frame synthesis network model. Among them, the number of the input original forward reference frame and original backward reference frame is set to N0, where N0 is greater than or equal to 1. In order to make more full use of the image information of the encoded frames, the input of the reference frame synthesis network model can use multiple original forward reference frames and multiple original backward reference frames.

[0072] The input and output information of the reference frame synthesis network model can only include main information, the input main information is the first image information of the original forward reference frame and the second image information of the original backward reference frame in the backward reference frame list, and the output main information is the image information of the synthesized frame.

[0073] The input of the reference frame synthesis network model may only include input main information, and the output includes output main information and output side information. The output side information is a motion mask image of the synthesized frame.

[0074] In other implementations, the input of the reference frame synthesis network model may also include input main information and input side information, where the input side information includes but is not limited to frame type, quantization parameter, time domain distance, and reference direction.

[0075] Step S103: inserting the synthesized frame into a reference frame list to be updated, where the reference frame list to be updated includes a forward reference frame list and / or a backward reference frame list.

[0076] After acquiring the synthesized frame based on the reference frame synthesis network model, the electronic device may insert the synthesized frame into the list of reference frames to be updated.

[0077] If, when performing prediction, the forward prediction needs to refer to the synthesized frame, the reference frame list to be updated can be the forward reference frame list, and the synthesized frame is inserted into the forward reference frame list; if, when performing prediction, the backward prediction needs to refer to the synthesized frame, the reference frame list to be updated can be the backward reference frame list, and the synthesized frame is inserted into the backward reference frame list; if, when performing prediction, both the forward prediction and the backward prediction need to refer to the synthesized frame, the reference frame list to be updated can be the forward reference frame list and the backward reference frame list, and the synthesized frame is inserted into both the forward reference frame list and the backward reference frame list.

[0078] When inserting the synthesized frame into the forward reference frame list and / or the backward reference frame list, the insertion position of the synthesized frame can be determined based on indicators, including but not limited to the sum of absolute error, mean square error, peak signal-to-noise ratio, structural similarity, and the like.

[0079] Step S104: predicting the current coding frame based on the to-be-updated reference frame list inserted into the synthesized frame.

[0080] When forward prediction requires reference to a synthesized frame, when the synthesized frame is inserted into the forward reference frame list, inter-frame prediction is performed on the current coded frame based on the forward reference frame list and the backward reference frame list inserted into the synthesized frame.

[0081] When backward prediction requires reference to a synthesized frame, inserting the synthesized frame into the backward reference frame list is to perform inter-frame prediction on the current coded frame based on the forward reference frame list and the backward reference frame list inserted into the synthesized frame.

[0082] When both forward prediction and backward prediction need to refer to the synthesized frame, when inserting the synthesized frame into the forward reference frame list and the backward reference frame list, the current encoded frame is inter-frame predicted based on the forward reference frame list of the inserted synthesized frame and the backward reference frame list of the inserted synthesized frame.

[0083] Different from the prior art, in this application, the original forward reference frames in the forward reference frame list of the current encoded frame and the original backward reference frames in the backward reference frame list are obtained, the synthesized frame of the current encoded frame is obtained by using the original forward reference frames and the original backward reference frames, and the synthesized frame is inserted into the reference frame list to be updated. The reference frame list to be updated includes the forward reference frame list and / or the backward reference frame list; and the current encoded frame is predicted based on the reference frame list to be updated into which the synthesized frame is inserted; the synthesized frame is obtained by using the original forward reference frames in multiple forward reference frame lists and the original backward reference frames in multiple backward reference frame lists. Therefore, the synthesized frame is relatively similar to the current encoded frame. Therefore, the synthesized frame has a strong correlation with the current encoded frame. When predicting the current encoded frame, the synthesized frame has a high probability of being referenced. This application utilizes the correlation between the synthesized frame and the current encoded frame, and various application strategies based on the synthesized frame can be designed, thereby improving the encoding efficiency.

[0084] In this embodiment, in order to improve the quality of the synthesized frame, a new design is made for the reference frame synthesis network model in this embodiment.

[0085] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of the first embodiment of the reference frame synthesis network model of this application. The reference frame synthesis network model includes five parts: input, input quality enhancement, reference frame synthesis network, output quality enhancement, and output. In other embodiments, the parts of input quality increment and output quality enhancement can be omitted.

[0086] The input information includes input main information and input side information. The input main information is the first image information of the original forward reference frame and the second image information of the original backward reference frame. The input side information includes, but is not limited to, frame type, quantization parameter, temporal distance, and reference direction. The input side information is used to assist the designed reference frame synthesis network to adapt to the reference frame inputs from different temporal layers. Additionally, if the form and size of the input side information are inconsistent with those of the input main information, the input side information needs to be adjusted to be consistent with the input main information. The adjustment includes, but is not limited to, symbol encoding (symbol conversion), padding, interpolation, etc. Here, symbol encoding refers to the process of converting source symbols into digital symbols, and the set of digital symbols used for encoding is not uniquely specified as long as it can correspond to the source symbols one by one. For padding and interpolation, since the required size of the network input is inconsistent with the size of the input information, the corresponding input information needs to be padded or interpolated to the input size required by the network. When the input information is a single symbol, the padding method is adopted; when the input information is multiple symbols, the interpolation method is adopted.

[0087] In this embodiment, the designed reference frame synthesis neural network model can utilize more of the original forward reference frames in the forward reference frame list and the original backward reference frames in the backward reference frame list. And the designed reference frame synthesis neural network model in this embodiment can input side information such as the temporal layer, quantization parameter, and reference direction of the reference frame, enhancing the adaptability of the reference frame synthesis neural network model to different temporal layers, different quantization parameters, and reference directions.

[0088] Input quality enhancement, reference frame synthesis network, and output quality enhancement are the main bodies of the reference frame synthesis network model. Input quality enhancement is used to improve the quality of the first image information of the original forward reference frame and the second image information of the original backward reference frame to better obtain the synthesized frame. The reference synthesis network is used to fuse multiple original forward reference frames and multiple original backward reference frames; output quality enhancement is used to improve the quality of the synthesized frame.

[0089] The reference frame synthesis neural network model designed in this embodiment enhances the quality of the input forward reference frame, backward reference frame, and synthesized frame in two stages respectively, and can more effectively improve the quality of the synthesized frame.

[0090] Please refer to Figure 4 , Figure 4 which is the structural schematic diagram of the second embodiment of the reference frame synthesis network model of this application.

[0091] As Figure 4 shown, the input quality enhancement includes N2 neural networks, where N2 is the total number of input reference frames. The N2 neural networks independently enhance the quality of the N2 inputs. The neural networks used for input quality enhancement include, but are not limited to, residual neural networks, and each neural network can adopt the same network or different networks.

[0092] As Figure 4 shown, the reference synthesis network includes 1 modulation module, S fusion modules, and 1 aggregation module. Among them, the fusion module includes two parts of input. The first part of the input comes from all the outputs of the previous stage, and the second part of the output comes from the modulation module. The fusion module fuses the input of the previous stage in different ways, including but not limited to direct fusion and indirect fusion; the modulation module is used to guide the fusion process of the fusion module, and its input includes but not limited to time-domain distance, reference direction, and quantization parameters, and its network structure is but not limited to a fully connected neural network; the aggregation module further fuses the output of the fusion module, including but not limited to a residual neural network.

[0093] As Figure 4 shown, the output quality enhancement also includes a neural network for improving the quality of the synthesized frame. The neural network used for output quality enhancement includes but not limited to a residual neural network, and the neural network for output quality enhancement can use the same neural network as the input quality enhancement or a different neural network.

[0094] For input quality enhancement and output quality enhancement, a residual neural network is used in this embodiment. Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of an embodiment of the residual neural network of this application. In this embodiment, Figure 5 the left side is the residual neural network, Figure 5 and the right side is the residual block in the residual neural network. The number R of residual blocks in the residual neural network is set to 8. In other embodiments, the number of its residual blocks can be set based on actual needs and will not be elaborated here.

[0095] For obtaining the synthesized frame, two methods of indirect fusion and direct fusion can be adopted in this embodiment. In this embodiment, the fusion module uses a 3D residual neural network.

[0096] For the 3D residual neural network, its structure is the same as that of an ordinary convolutional neural network. Please refer to Figure 6 , Figure ..... which is a schematic structural diagram of an embodiment of the 3D residual neural network of this application. As ​ shown, only 3D convolution is used. The difference between 3D convolution and ordinary convolution is that the 3D convolution layer can move in the channel, while the ordinary convolution layer cannot move in the channel. Therefore, for multi-frame images with time-domain correlation, the movement in the channel means that the generated feature map has time-domain correlation.

[0097] If the fusion module adopts the indirect fusion method, first stack the feature maps of the first image information of the original forward reference frame and the second image information of the original backward reference frame on the channels of the 3D residual neural network, then modulate on the feature channels, and finally perform fusion through the residual neural network.

[0098] Please refer to ​ , ​ which is a schematic diagram of the fusion of an embodiment of the indirect fusion of the fusion module of this application. The dimension of the input feature is W*H*C1,..., W*H*C N2 , C i is the number of feature channels of the i-th input, i is the input code, and i = 1, 2,..., N2. N2 is the total number of input reference frames. After stacking, the feature dimension is W*H*(C1+...+C N2 ). In this embodiment, let C1 = C2 =... = C N2 = N3. Therefore, W*H*(C1+...+C N2 ) = W*H*(N2*N3). Here, the modulation input is that the modulation module assigns a multiplication coefficient to each feature channel.

[0099] If the fusion module adopts the direct fusion method, first weight and combine the feature maps of each input into a feature map according to the channels of the 3D residual neural network, and then perform fusion through the 3D residual neural network. Here, the modulation input is used for weighting in the combination process. Please refer to ​ , ​ which is a schematic diagram of the fusion of an embodiment of the direct fusion of the fusion module of this application. First, the j-th feature maps of each input form the j-th group. C i-j represents the j-th feature map of the i-th input. Then, the modulation input is divided into N3 groups, and the modulation input j represents the weighting coefficients of the feature maps in the j-th group.

[0100] In this embodiment, the modulation module can be implemented by a 3-layer fully connected neural network. Please refer to ​ , ​ which is a schematic diagram of the structure of an embodiment of the modulation module of this application. The input of the modulation module consists of quantization parameters, time domain distance, and reference direction, forming a vector of size N4*1. In this embodiment, N4 = 3, and the value of N4 is related to the number of inputs of this module; in this embodiment, N5 = 64, N6 = 256, and in other embodiments, the values of N5 and N6 can be set arbitrarily; N7 and N8 = N2*N3, and the values of N7 and N8 are equal to the total number of input reference frames and the number of feature channels of the residual neural network. The output S of the adjustment module is used as the modulation input in the fusion module.

[0101] The aggregation module of the reference frame synthesis network first stacks the input feature maps in the channel dimension and then fuses them through a residual neural network.

[0102] Different from the prior art, the designed reference frame synthesis network model in this embodiment does not need to limit the applied temporal layer because the designed reference frame synthesis network model has wide adaptability.

[0103] Optionally, this embodiment can implement step S103 through the following method, which specifically includes the following steps:

[0104] Insert the synthesized frame at the end of the list of reference frames to be updated.

[0105] When the list to be updated is the forward reference frame list, insert the synthesized frame at the end of the forward reference frame list; when the list to be updated is the backward reference frame list, insert the synthesized frame at the end of the backward reference frame list; when the list to be updated is both the forward reference frame list and the backward reference frame list, insert the synthesized frame at the end of both the forward reference frame list and the backward reference frame list.

[0106] In the forward reference frame list, the forward reference frames are sorted according to the distance in the image display order from the current coded frame. The position of the forward coded frame closest to the current coded frame is the head of the forward reference frame list, and the position of the forward coded frame farthest from the current coded frame is the end of the forward reference frame list. The backward reference frame list is the same as the forward reference frame list, and both are sorted based on the distance in the image display order from the current coded frame.

[0107] Please refer to ​ , ​ which is a schematic structural diagram of the first embodiment of inserting the synthesized frame in this application. As ​ shown, in the hierarchical B-frame structure of this embodiment, the numbers of the forward reference frame list and the backward reference frame list are N11 and N12 respectively, and the forward reference index and the backward reference index are idx0 and idx1 respectively, where idx0 ∈ [1, N11] and idx1 ∈ [1, N12].

[0108] During the encoding and decoding process, if the synthesized frame is used in the forward reference frame list and / or the backward reference frame list, the electronic device inserts the synthesized frame at the end of the forward reference frame list and / or the backward reference frame list and adds a new reference index.

[0109] During the encoding or decoding process, the synthesized frame uses a switch syntax to determine whether to insert the synthesized frame into the forward reference frame list and / or the backward reference frame list; and uses an index syntax to identify the position where the synthesized frame is inserted into the forward reference frame list and / or the backward reference frame list.

[0110] Take ​ as an example. In the figure, the forward reference frame list is L1 to L N9 , and the backward reference frame list is R1 to R N10 . L1 is the head of the forward reference frame list, and L N9 is the end of the forward reference frame list; R1 is the head of the backward reference frame list, and R N9 is the end of the backward reference frame list. If the forward reference frame list uses a composite frame, the composite frame is inserted into the L N11+1 th position in the forward reference frame list, and its reference index is N11 + 1. If the backward reference frame list uses a composite frame, the composite frame is inserted into the L N12+1 th position in the backward reference frame list, and its reference index is N12 + 1. If both the forward reference frame list and the backward reference frame list use composite frames, the composite frames are simultaneously inserted into the L N11+1 th position in the forward reference frame list and the L N12+1 th position in the backward reference frame list, and the reference indices N11 + 1 and N12 + 1 are increased.

[0111] When the forward reference frame list of the current encoded frame needs to refer to a composite frame, the reference index syn_idx0 of the composite frame is encoded and transmitted in the bitstream of the current encoded frame. The default value of syn_idx0 is set to 0. When syn_idx0 is 0, it means that the forward reference frame list does not use a composite frame; when syn_idx0 is other values, it means the position where the composite frame is added in the forward reference frame list. The default value of the reference index syn_idx1 of the backward reference frame list is set to 0. When syn_idx1 is 0, it means that the backward reference frame list does not use a composite frame; when syn_idx1 is other values, it means the position where the composite frame is added in the backward reference frame list.

[0112] Please refer to ​ , ​ which is a schematic structural diagram of the second embodiment of the composite frame insertion of the present application. As ​ shown, in the hierarchical B-frame structure of this embodiment, the numbers of the forward reference frame list and the backward reference frame list are N11 and N12 respectively, and the forward reference index and the backward reference index are idx0 and idx1 respectively, where idx0 ∈ [1, N11] and idx1 ∈ [1, N12].

[0113] During the encoding and decoding process, if the forward reference frame list and / or the backward reference frame list use composite frames, the electronic device inserts the composite frame at the end of the forward reference frame list and / or the end of the backward reference frame list, but does not generate additional reference indices, but shares the reference indices N11 and N12 with the last reference frame in the forward reference frame list and / or the backward reference frame list respectively. In the figure, Syn is the composite frame.

[0114] During the encoding and decoding process, when the prediction of an encoding unit needs to refer to a synthesized frame, the synthesized frame flag bit syn_flag is encoded and transmitted in the bitstream of the encoding block, and the default value of syn_flag is set to 0. When syn_flag is 0, it means that the synthesized frame is not used in the current block prediction process; when syn_flag is 1, it means that the synthesized frame is only inserted into the forward reference frame list for the current block; when syn_flag is 2, it means that the synthesized frame is only inserted into the backward reference frame list for the current block; when syn_flag is 3, it means that the synthesized frame is inserted into both the forward reference frame list and the backward reference frame list for the current block.

[0115] Optionally, please refer to ​ , ​ is ​ a specific process schematic diagram of step S103 in ​ This embodiment can implement step S103 through the method shown in

[0116] Step S201: Obtain the first subjective and objective quality parameters of the current encoded frame and the synthesized frame, obtain the second subjective and objective quality parameters of the current encoded frame and the forward reference frame, and obtain the third subjective and objective quality parameters of the current encoded frame and the backward reference frame.

[0117] During the encoding and decoding process, obtain the first subjective and objective quality parameters of the current encoded frame and the synthesized frame, obtain the second subjective and objective quality parameters of the current encoded frame and the forward reference frame, and obtain the third subjective and objective quality parameters of the current encoded frame and the backward reference frame. Among them, the first subjective and objective quality parameters, the second subjective and objective quality parameters, and the third subjective and objective quality parameters include but are not limited to indicators such as sum of absolute errors, mean square error, peak signal-to-noise ratio, and structural similarity.

[0118] Step S202: Sort the first subjective and objective quality parameters and the second subjective and objective quality parameters according to their magnitude relationship to obtain a first sequence.

[0119] The electronic device sorts the first subjective and objective quality parameters and the second subjective and objective quality parameters according to their magnitude relationship to obtain a first sequence.

[0120] Step S203: Sort the first subjective and objective quality parameters and the third subjective and objective quality parameters according to their magnitude relationship to obtain a second sequence.

[0121] The electronic device sorts the first subjective and objective quality parameters and the third subjective and objective quality parameters according to their magnitude relationship to obtain a second sequence.

[0122] Step S204: Insert the synthesized frame into the forward reference frame list based on the position of the first subjective and objective quality parameter in the first sequence; and / or insert the synthesized frame into the backward reference frame list based on the position of the first subjective and objective quality parameter in the second sequence.

[0123] During the encoding and decoding process, if a synthesized frame needs to be inserted into the forward reference frame list, insert the synthesized frame into the forward reference frame list based on the position of the first subjective and objective quality parameter in the first sequence, and move the forward reference frames after the synthesized frame backward by one position.

[0124] If a synthesized frame needs to be inserted into the backward reference frame list, insert the synthesized frame into the backward reference frame list based on the position of the first subjective and objective quality parameter in the second sequence, and move the backward reference frames after the synthesized frame backward by one position.

[0125] If both the forward reference frame list and the backward reference frame list need to insert the synthesized frame, insert the synthesized frame into the forward reference frame list and the backward reference frame list respectively based on the position in the first sequence and the position in the second sequence.

[0126] Please refer to ​ , ​ which is a schematic structural diagram of the third embodiment of the synthesized frame insertion of the present application. As ​ shown, in this embodiment, first calculate the peak signal-to-noise ratio PSNR between the current encoded frame and the synthesized frame Syn syn , then calculate the PSNR between the current frame and the forward reference frames one by one in ascending order idx0 , and compare the PSNR idx0 with the PSNR syn . If the PSNR syn is higher than the PSNR idx0 , then record the current idx0 as idx0_syn, and finally insert the synthesized frame into the forward reference frame table at the position of idx0_syn, and the forward reference frames at idx0_syn and after move backward by one position. The process of inserting the synthesized frame into the backward reference frame list is the same as that of the forward reference frame list. As ​ shown, L idx0_syn and R idx0_syn are the inserted synthesized frames, and the synthesized frame reference indexes in the forward reference frame list and the backward reference frame list are idx0_syn and idx1_syn respectively, and the subsequent frames move backward by one position, N11' = N11 + 1, N12' = N12 + 1.

[0127] Encode and transmit the reference indices idx0_syn and idx1_syn in the bitstream of the current coded frame. idx0_syn and idx1_syn respectively indicate that a synthesized frame is inserted at the idx0_syn and idx1_syn positions in the forward reference frame list and the backward reference frame list, and the reference frames at idx0_syn, idx1_syn and after in the forward reference frame list and the backward reference frame list are moved backward by one position.

[0128] Optionally, the method for obtaining the synthesized frame is as ​ shown ​ This is ​ a schematic flowchart of the first embodiment of step S102 in this application. This embodiment can implement step S102 through the method as ​ shown. The specific implementation steps include steps S301 to S303:

[0129] Step S301: Adjust the multiple component sizes of the original forward reference frame and the original backward reference frame to the same size threshold.

[0130] After the electronic device obtains the original forward reference frame and the original backward reference frame, if it wants to use them as inputs and input them into the reference synthesis frame network model, if the multiple component sizes of the original forward reference frame and the original backward reference frame are different, then the multiple component sizes of the original forward reference frame and the original backward reference frame need to be adjusted to the same size threshold. If the sizes of the components of the original forward reference frame and the original backward reference frame are the same, no adjustment is required.

[0131] Step S302: Obtain the synthesized frame of the current coded frame by using the original forward reference frame with adjusted size and the original backward reference frame with adjusted size.

[0132] The electronic device inputs the original forward reference frame with adjusted size and the original backward reference frame with adjusted size into the reference synthesis frame network model to obtain the synthesized frame of the current coded frame.

[0133] Step S303: Restore the component size of the synthesized frame to the component size corresponding to the current coded frame.

[0134] After the electronic device obtains the synthesized frame, it is necessary to restore the component size of the synthesized frame to the component size corresponding to the current coded frame.

[0135] Please refer to ​ , ​ This is a reference diagram of the first embodiment of this application for obtaining the synthesized frame by using the original forward reference frame and the original backward reference frame. As ​ shown, the network input is the original forward reference frames of the first N0 (N0 >= 1) in the forward reference frame list and the original backward reference frames of the first N0 (N0 >= 1) in the backward reference frame list, and the output of the reference synthesis frame network model is the synthesized frame. Ln Indicates the n-th reference frame in the forward reference frame list, R n Indicates the n-th reference frame in the backward reference frame list, where n ∈ [1, N0].

[0136] For L n or R n , the number of components of the image is N1 (N1 >= 1). If there are inconsistencies in the sizes between the components, the sizes between the components need to be adjusted to the same size threshold.

[0137] In this embodiment, N1 = 3. Please refer to ​ , ​ is a schematic diagram of an embodiment of adjusting the component sizes of the reference frame in this application. As ​ shown are the input images of the forward reference frame and the original backward reference frame in the YUV420 format, and the chrominance components U and V need to be adjusted to the same size as the Y component through interpolation.

[0138] For the synthesized frame, when outputting the synthesized frame, it needs to be adjusted to the corresponding component size of the current coded frame through interpolation. Please refer to ​ , ​ is a schematic diagram of an embodiment of adjusting the component sizes of the synthesized frame in this application. As ​ shown, when outputting, the components of the synthesized frame need to be adjusted to the YUV420 format through interpolation.

[0139] Optionally, ​ is a schematic flowchart of the second embodiment of step S102 in this application. This embodiment can implement step S102 through the method shown in ​ , and the specific implementation steps include steps S401 to S402: ​ Step S401: Obtain the first information of the original forward reference frame in the forward reference frame list and the second information of the original backward reference frame in the backward reference frame list.

[0140] Step S401: Obtain the first information of the original forward reference frame in the forward reference frame list and the second information of the original backward reference frame in the backward reference frame list.

[0141] When the electronic device obtains the synthesized frame by using the original forward reference frame in the forward reference frame list and the original backward reference frame in the backward reference frame list, it can obtain the first information of the original forward reference frame in the forward reference frame list and the second information of the original backward reference frame in the backward reference frame list.

[0142] Among them, the first information includes the image information of the original forward reference frame, and the second information includes the image information of the original backward reference frame; alternatively, the first information includes the image information of the original forward reference frame and the edge information of the original forward reference frame, and the second information includes the image information of the original backward reference frame and the edge information of the original backward reference frame.

[0143] Among them, the side information of the original forward reference frame can be the frame type, quantization parameter, temporal distance, and reference direction of the original forward reference frame; the side information of the original backward reference frame is the frame type, quantization parameter, temporal distance, and reference direction of the original backward reference frame.

[0144] Step S402: Obtain the information of the synthesized frame of the current encoded frame based on the first information and the second information.

[0145] Inputting the first information and the second information into the reference frame synthesis network model can obtain the information of the synthesized frame of the current encoded frame. The information of the synthesized frame of the current encoded frame obtained by using the reference frame synthesis network model in the foregoing can be the image information of the synthesized frame, or the image information of the synthesized frame and the output side information of the synthesized frame.

[0146] In an application scenario, please refer to ​ , ​ is a schematic reference diagram of the second embodiment of the synthesized frame obtained by the present application using the original forward reference frame and the original backward reference frame. As ​ shown, the input of the reference frame synthesis network model is the image information of the first N0 (N0>=1) in the forward reference frame list and the image information of the first N0 (N0>=1) in the backward reference frame list, and the sizes of the respective components of the reference images of each reference frame are the same, and the number of components of each frame of image is N1 (N1>=1). The network output is the image information of the synthesized frame and the output side information. L n represents the nth reference frame in the forward reference frame list, and R n represents the nth reference frame in the backward reference frame list, where n ∈ [1, N0].

[0147] In this embodiment, the output side information of the synthesized frame is a motion mask image, which is used to characterize the motion condition of each pixel point in the synthesized image. Its size is the same as that of the synthesized frame, and each pixel point thereof indicates whether the corresponding pixel point in the synthesized frame is in motion or at rest. A pixel value of 1 indicates that the pixel point is in motion, and a pixel value of 0 indicates that the pixel point is at rest.

[0148] When the electronic device obtains the synthesized frame by using the original forward reference frame in the forward reference frame list and the original backward reference frame in the backward reference frame list, it can obtain the image information and the side information of the original forward reference frame in the forward reference frame list and the image information and the side information of the original backward reference frame in the backward reference frame list.

[0149] After the electronic device obtains the side information of the original forward reference frame, it is necessary to fill the side information of the original forward reference frame into a first matrix having the same size as the image information of the original forward reference frame; similarly, after obtaining the side information of the original backward reference frame, the side information of the original backward reference frame also needs to be filled into a second matrix having the same size as the image information of the original backward reference frame.

[0150] Taking the quantization parameter as an example, please refer to ​ , ​ which is a matrix schematic diagram of an embodiment of the quantization parameter of the present application. As ​ shown, if the quantization parameter is equal to 32, the quantization parameter needs to be filled into a matrix of W*H, where W and H are the width and height of the image information of the original forward reference frame or the image information of the original backward reference frame, respectively.

[0151] The electronic device inputs the image information of the original forward reference frame, the first matrix, the image information of the original backward reference frame, and the second matrix into the reference frame synthesis network model, and can obtain the image information of the synthesized frame of the current coded frame; or based on the image information of the original forward reference frame, the first matrix, the image information of the original backward reference frame, and the second matrix, the image information of the synthesized frame of the current coded frame and the output edge information of the synthesized frame can be obtained by using the reference frame synthesis network model in the foregoing text.

[0152] In an application scenario, please refer to ​ , ​ which is a reference schematic diagram of the third embodiment of obtaining a synthesized frame by using the original forward reference frame and the original backward reference frame of the present application. As ​ shown, the input of the reference frame synthesis network model is the image information and the edge information of the original forward reference frame of the first N0 (N0>=1) of the forward reference frame list and the image information and the edge information of the original backward reference frame of the first N0 (N0>=1) in the backward reference frame list. The input edge information includes frame type, quantization parameter, temporal distance, and reference direction, and the network output is the image information of the synthesized frame. The input and output of the network are shown in the following figure, L' n represents the image information and the edge information of the original forward reference frame of the nth reference frame of the forward reference frame list, and R' n represents the image information and the edge information of the original backward reference frame of the nth reference frame in the backward reference frame list, where n∈[1, N0].

[0153] For each L' n or R' n , it includes M+N1 (M>=1, N1>=1) frame images, M is the number of input edge information, and N1 is the number of image information components. Among them, m∈[1,M+N1], {I1,…,I N1} is the image information, and {I N1+1 ,…,I N1+M} is the adjusted input edge information. In this embodiment, N1 = 1, M = 4, where I2 is the frame type, I3 is the quantization parameter, I4 is the temporal distance, and I5 is the reference direction.

[0154] For quantization parameters, when the quantization parameters are used as inputs, they need to be filled into a matrix with the same size as the image information. The quantization parameters need to be filled into a matrix of W*H, where W and H are the width and height of the image information of the original forward reference frame or the original backward reference frame respectively.

[0155] For the temporal distance, it is defined as the difference in the image display order between the original forward reference frame or the original backward reference frame and the current encoded frame. Similar to the quantization parameters, it needs to be filled into a matrix of W*H.

[0156] For frame types, first, the source symbols need to be encoded into digital symbols and then filled into a matrix of W*H. In this embodiment, the frame types include the source symbol set {I-frame, B-frame, P-frame}, and the digital symbol set can be set as {-1, 0, 1}, that is, the I-frame is encoded as the digital -1, the B-frame is encoded as the digital 0, and the P-frame is encoded as the digital 1. In other embodiments, other numbers can also be used, which are not limited herein.

[0157] For the reference direction, similarly, the source symbols need to be encoded into digital symbols first and then filled into a matrix of W*H. In this embodiment, the reference direction includes the source symbol set {forward, backward}, and the digital symbol set is set as {-1, 1}, the original forward reference frame is encoded as the digital -1, and the original backward reference frame is encoded as the digital 1. In other embodiments, other numbers can also be used, which are not limited herein.

[0158] Optionally, when the reference frame list to be updated is the forward reference frame list and the backward reference frame list, this embodiment can implement step S104 through the following method, which specifically includes the following steps:

[0159] Predict the current encoded frame based on the synthesized frames in the forward reference frame list and at least part of the original forward reference frames, and the synthesized frames in the backward reference frame list and at least part of the original backward reference frames.

[0160] When the reference frame list to be updated is the forward reference frame list and the backward reference frame list, after the electronic device inserts the synthesized frame into both the forward reference frame list and the backward reference frame list, when performing inter-frame prediction on the current encoded frame using the reference frame list to be updated, it can predict the current encoded frame based on the synthesized frames in the forward reference frame list and at least part of the original forward reference frames, and the synthesized frames in the backward reference frame list and at least part of the original backward reference frames.

[0161] In other embodiments, it can also predict the current encoded frame based on the synthesized frames in the forward reference frame list and all the original forward reference frames, and the synthesized frames in the backward reference frame list and all the original backward reference frames.

[0162] Optionally, when the reference frame list to be updated is a forward reference frame list or a backward reference frame list, this embodiment can implement step S104 through the following method, which specifically includes the following steps:

[0163] Predict the current encoded frame based on the synthesized frame and at least part of the original forward reference frames in the forward reference frame list where the synthesized frame is inserted, and at least part of the original backward reference frames in the backward reference frame list where the synthesized frame is not inserted;

[0164] Alternatively, predict the current encoded frame based on the synthesized frame and at least part of the original backward reference frames in the backward reference frame list where the synthesized frame is inserted, and at least part of the original forward reference frames in the forward reference frame list where the synthesized frame is not inserted.

[0165] When the reference frame list to be updated is a forward reference frame list, after the electronic device inserts the synthesized frame into the forward reference frame list, when performing inter-frame prediction on the current encoded frame using the reference frame list to be updated, it can predict the current encoded frame based on the synthesized frame and at least part of the original forward reference frames in the forward reference frame list where the synthesized frame is inserted, and at least part of the original backward reference frames in the backward reference frame list where the synthesized frame is not inserted.

[0166] When the reference frame list to be updated is a backward reference frame list, after the electronic device inserts the synthesized frame into the backward reference frame list, when performing inter-frame prediction on the current encoded frame using the reference frame list to be updated, it can predict the current encoded frame based on the synthesized frame and at least part of the original backward reference frames in the backward reference frame list where the synthesized frame is inserted, and at least part of the original forward reference frames in the forward reference frame list where the synthesized frame is not inserted.

[0167] Optionally, this embodiment can also expand the inter-frame mode through the ​ method shown in the figure. Please refer to ​ , ​ which is a specific flowchart of step S104 in this application. This embodiment can implement step S104 through the ​ method shown in the figure. The specific implementation steps include steps S501 to S502: ​ which is a specific flowchart of step S104 in this application. This embodiment can implement step S104 through the method shown in the figure. The specific implementation steps include steps S501 to S502:

[0168] Step S501: Obtain the synthesized frame in the reference frame list to be updated.

[0169] After the electronic device inserts the synthesized frame into the reference frame list to be updated, where the reference frame list to be updated includes a forward reference frame list and / or a backward reference frame list. The electronic device can directly obtain the synthesized frame of the current encoded frame based on the updated reference frame list to be updated.

[0170] The synthesized frame is obtained by using the original forward reference frames in multiple forward reference frame lists and the original backward reference frames in multiple backward reference frame lists through a reference frame synthesis network model. Therefore, the synthesized frame has a very strong correlation with the current encoded frame and can be regarded as the direct prediction value of the current encoded frame. Thus, the synthesized frame has a high probability of being referenced. Based on the synthesized frame, traditional inter-frame modes can be extended to obtain skip modes based on the synthesized frame, merge modes based on the synthesized frame, etc.

[0171] Step S502: Set the motion vector of the synthesized frame to a zero vector to directly use the synthesized frame as the prediction frame of the current encoded frame.

[0172] When performing the merge mode based on the synthesized frame, similar to the traditional merge mode, it is necessary to construct the MV list according to the same rules at the encoding and decoding ends. Considering that the synthesized frame is very similar to the current frame in content. In this mode, the process of the MV list is as follows. First, fill the zero vector in the first position of the MV list, rather than filling the zero vector at the end of the MV list as in the traditional mode. Then, sort the other candidates in the MV list in ascending order according to the distance. The distance is calculated as the Euclidean distance of the horizontal and vertical MVs.

[0173] When encoding, encode and transmit the merge_syn flag, and merge_syn is default set to 0. merge_syn = 0 indicates that this mode is not enabled; when merge_syn = 1, it indicates that this mode is enabled. Set the motion vector of the synthesized frame to a zero vector to directly use the synthesized frame as the prediction frame of the current encoded frame, and thus adjust the construction of the MV list at the decoding end. In addition, similar to the traditional merge, the MV index and the corresponding residual also need to be transmitted in this mode.

[0174] When performing the skip mode based on the synthesized frame, similar to the merge mode based on the synthesized frame, only the prediction residual does not need to be transmitted in the bitstream. This mode constructs the MV list according to the same rules at the encoding and decoding ends. First, fill the zero vector in the first position of the MV list, and then sort the other candidates in the MV list in ascending order according to the Euclidean distance of the motion vectors.

[0175] When encoding, encode and transmit the skip_syn flag, and skip_syn is default set to 0. When skip_syn = 0, it indicates that this mode is not enabled; when skip_syn = 1, it indicates that this mode is enabled, and thus adjust the construction of the MV list at the decoding end. In addition, similar to the traditional skip, only the MV index needs to be transmitted in this mode.

[0176] In other embodiments, a fast inter-frame mode based on a synthesized frame can also be performed. By referring to the motion mask image output by the reference frame synthesis network model, it can be quickly determined whether the current coding block is a motion block or a stationary block. First, for the current coding block, calculate the probability that it is a motion block. The calculation method is the proportion of motion pixels in the current block, and the expression is as follows:

[0177]

[0178] where W and H are the width and height of the current block respectively, and mask(i,j) is the coordinate value of the motion mask image output by the reference frame synthesis network model.

[0179] If the probability p is greater than the threshold then it is considered a motion block, otherwise it is a stationary block. In this embodiment

[0180] For a stationary block, only a zero vector is used for prediction, and it is also divided into two cases: whether the prediction residual needs to be encoded or not; for a motion block, prediction is performed according to the normal inter-frame mode.

[0181] When encoding, the fast_syn flag is encoded and transmitted, and fast_syn is default set to 0. When fast_syn = 0, it means the current coding block is a motion block, and prediction is performed according to the normal inter-frame mode; when fast_syn = 1, it means the current block is a stationary block, and inter-frame prediction is directly performed according to the zero vector, and the prediction residual needs to be transmitted; when fast_syn = 2, it means the current block is a stationary block, and inter-frame prediction is directly performed according to the zero vector, and no other information needs to be transmitted.

[0182] Optionally, please refer to ​ , ​ which is a schematic flowchart of the second embodiment of the inter-frame prediction method of the present application. When obtaining the original forward reference frame in the forward reference frame list of the current coding frame and the original backward reference frame in the backward reference frame list, as ​ shown, the inter-frame prediction method further includes steps S601 to S603:

[0183] Step S601: Obtain the first quantity of the original forward reference frames in the forward reference frame list and the second quantity of the backward reference frames in the backward reference frame list.

[0184] When the electronic device obtains the original forward reference frame in the forward reference frame list and the original backward reference frame in the backward reference frame list, it first needs to obtain the first quantity of the original forward reference frames in the forward reference frame list and the second quantity of the backward reference frames in the backward reference frame list.

[0185] Step S602: In response to the first quantity being greater than the second quantity, obtain the original forward reference frame located at the head of the forward reference frame list and insert it at the end of the backward reference frame list.

[0186] The electronic device, in response to the first quantity being greater than the second quantity, obtains the original forward reference frame located at the head of the forward reference frame list and inserts it at the end of the backward reference frame list.

[0187] Please refer to ​ , when obtaining the forward reference frame and the backward reference frame with the image display order equal to 5, the first quantity of the forward reference frame is greater than the second quantity. The forward reference frames are the frames with POC = 4, POC = 3, POC = 2, POC = 1, and POC = 0, and the backward reference frames are the frames with POC = 6, POC = 8, POC = 16, and POC = 32 respectively. At this time, using the forward reference frame as a supplement, obtain the forward reference frame located at the head of the forward reference frame list and insert it at the end of the backward reference frame list. At this time, the backward reference frames with the image display order equal to 5 are the frames with POC = 6, POC = 8, POC = 16, POC = 32, and POC = 4.

[0188] Step S603: In response to the first quantity being less than the second quantity, obtain the backward reference frame located at the head of the backward reference frame list and insert it at the end of the forward reference frame list.

[0189] Please refer to ​ , when obtaining the forward reference frame and the backward reference frame with the image display order equal to 1, the first quantity of the forward reference frame is less than the second quantity. The forward reference frame is the frame with POC = 0, and the backward reference frames are the frames with POC = 2, POC = 4, POC = 8, POC = 16, and POC = 32 respectively. At this time, using the backward reference frame as a supplement, obtain the backward reference frame located at the head of the backward reference frame list and insert it at the end of the forward reference frame list. At this time, the forward reference frames with the image display order equal to 1 are the frames with POC = 2, POC = 4, POC = 8, POC = 16, and POC = 0.

[0190] Different from the prior art, in this application, the original forward reference frames in the forward reference frame list of the current coded frame and the original backward reference frames in the backward reference frame list are obtained, and the synthesized frame of the current coded frame is obtained by using the original forward reference frames and the original backward reference frames. The synthesized frame is inserted into the reference frame list to be updated, and the reference frame list to be updated includes the forward reference frame list and / or the backward reference frame list; and the current coded frame is predicted based on the reference frame list to be updated into which the synthesized frame is inserted. The synthesized frame is obtained by using the original forward reference frames in multiple forward reference frame lists and the original backward reference frames in multiple backward reference frame lists. Therefore, the synthesized frame is relatively similar to the current coded frame, and thus the synthesized frame has a strong correlation with the current coded frame. When predicting the current coded frame, the synthesized frame has a high probability of being referenced. This application utilizes the correlation between the synthesized frame and the current coded frame, and various application strategies based on the synthesized frame can be designed, thereby improving the coding efficiency.

[0191] Furthermore, the reference frame synthesis network model designed in this application utilizes more forward reference frames and backward reference frames, and the designed reference frame synthesis network model combines side information such as the time domain layer, quantization parameter, and reference direction of the forward reference frame and the backward reference frame, enhancing the adaptability to different time domain layers, different quantization parameters, and reference directions. Additionally, a designed modulation module further fuses the side information into the reference frame synthesis network model; the reference frame synthesis network model designed in this application has wide adaptability, so there is no need to limit the applied time domain layer; the reference frame synthesis network model designed in this application enhances the quality of the forward reference frame, the backward reference frame, and the synthesized frame in two stages respectively, and can more effectively improve the quality of the synthesized frame.

[0192] This application also makes an independent design for the application of the synthesized frame. By making full use of the high correlation between the synthesized frame and the current coded frame, a skip mode based on the synthesized frame, a merge mode based on the synthesized frame, and a fast inter-frame mode based on the synthesized frame are designed, which can effectively improve the coding efficiency; in addition, by outputting motion estimation information through the network to assist the application of the synthesized frame, the coding efficiency can be further improved.

[0193] Optionally, this application further proposes a coding method, which obtains the prediction information of any one of the above-mentioned current coded frames and encodes the prediction information into the bitstream of the current coded frame. Among them, encoding is to convert the prediction information of the current coded frame into numbers that can be understood by a computer. There are various encoding methods such as arithmetic coding and variable-length coding, which are not limited here.

[0194] Optionally, this application further proposes a decoding method, which includes obtaining the prediction information of any one of the above-mentioned current coded frames and the bitstream of the current coded frame, and decoding the bitstream.

[0195] To implement the encoding method of the above embodiments, the present application proposes an encoder. For details, please refer to ​ , ​ which is a schematic structural diagram of an embodiment of the encoder provided by the present application.

[0196] The encoder 300 includes a memory 31 and a processor 32. Among them, the memory 31 is coupled to the processor 32.

[0197] The memory 31 is used to store computer programs, and the processor 32 is used to execute the computer programs to implement the image encoding method of the above embodiments.

[0198] In this embodiment, the processor 32 can also be referred to as a CPU (Central Processing Unit). The processor 32 may be an integrated circuit chip with signal processing capabilities. The processor 32 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor 32 can also be any conventional processor, etc.

[0199] To implement the above decoding method, the present application proposes a decoder. For details, please refer to ​ , ​ which is a schematic structural diagram of an embodiment of the decoder provided by the present application.

[0200] The decoder 400 includes a memory 41 and a processor 42. Among them, the memory 41 is coupled to the processor 42.

[0201] The memory 41 is used to store computer programs, and the processor 42 is used to execute the computer programs to implement the image decoding method of the above embodiments.

[0202] In this embodiment, the processor 42 can also be referred to as a CPU (Central Processing Unit). The processor 42 may be an integrated circuit chip with signal processing capabilities. The processor 42 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor 42 can also be any conventional processor, etc.

[0203] Optionally, the present application further proposes an electronic device. Please refer to ​ , ​FIG. 0 is a schematic structural diagram of an embodiment of the electronic device of the present application. The electronic device 100 includes a processor 101 and a memory 102 connected to the processor 101.

[0204] The processor 101 may also be referred to as a CPU (Central Processing Unit). The processor 101 may be an integrated circuit chip with signal processing capabilities. The processor 101 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0205] The memory 102 is used to store program data required for the operation of the processor 101.

[0206] The processor 101 is also used to execute the program data stored in the memory 102 to implement the above-mentioned inter-frame prediction method.

[0207] Optionally, the present application further provides a computer-readable storage medium. Please refer to ​ , ​ FIG. 16 is a schematic structural diagram of an embodiment of the computer-readable storage medium of the present application.

[0208] The computer-readable storage medium 200 of the embodiment of the present application stores program instructions 210 therein, and the program instructions 210 are executed to implement the above-mentioned inter-frame prediction method.

[0209] Among them, the program instructions 210 may form a program file and be stored in the above storage medium in the form of a software product, so that an electronic device (which may be a personal computer, a server, or a network device, etc.) or a processor can execute all or part of the steps of the methods of the various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc, or a terminal device such as a computer, a server, a mobile phone, or a tablet.

[0210] The computer-readable storage medium 200 of this embodiment may be, but is not limited to, a USB flash drive, an SD card, a PD optical drive, a mobile hard disk, a large-capacity floppy drive, a flash memory, a multimedia memory card, a server, etc.

[0211] In one embodiment, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions that are stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to cause the electronic device to perform the steps in the foregoing method embodiments.

[0212] In addition, if the above functions are implemented in the form of software functions and sold or used as an independent product, they can be stored in a storage medium readable by a mobile terminal. That is, the present application also provides a storage device storing program data, and the program data can be executed to implement the method of the foregoing embodiments. The storage device can be a USB flash drive, an optical disc, a server, etc. That is to say, the present application can be embodied in the form of a software product, which includes several instructions for causing an intelligent terminal to perform all or part of the steps of the methods described in the various embodiments.

[0213] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0214] Any process or method description shown in a flowchart or described in other ways herein can be understood as representing a mechanism, segment, or part of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed. This should be understood by those skilled in the technical field to which the embodiments of the present application belong.

[0215] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (which can be a personal computer, server, network device, or other system that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then storing it in a computer memory.

[0216] The above are only embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. An inter-frame prediction method, characterized in that Including: Obtain the forward reference frame list and the backward reference frame list of the current coded frame; Based on the original forward reference frames in the forward reference frame list and the original backward reference frames in the backward reference frame list, obtain the synthesized frame of the current coded frame; Insert the synthesized frame into the reference frame list to be updated, where the reference frame list to be updated includes the forward reference frame list and / or the backward reference frame list; Predict the current coded frame based on the reference frame list to be updated into which the synthesized frame is inserted; Wherein, the reference frame list to be updated is the forward reference frame list or the backward reference frame list, and predicting the current coded frame based on the reference frame list to be updated into which the synthesized frame is inserted includes: Predict the current coded frame based on the synthesized frame and at least some of the original forward reference frames in the forward reference frame list into which the synthesized frame is inserted, and at least some of the original backward reference frames in the backward reference frame list into which the synthesized frame is not inserted; Or, predict the current coded frame based on the synthesized frame and at least some of the original backward reference frames in the backward reference frame list into which the synthesized frame is inserted, and at least some of the original forward reference frames in the forward reference frame list into which the synthesized frame is not inserted.

2. The inter-frame prediction method according to claim 1, wherein The inserting the synthesized frame into the reference frame list to be updated includes: Insert the synthesized frame at the end of the reference frame list to be updated.

3. The inter-frame prediction method according to claim 1, wherein The inserting the synthesized frame into the reference frame list to be updated includes: Obtain the first subjective and objective quality parameter between the current coded frame and the synthesized frame, obtain the second subjective and objective quality parameter between the current coded frame and the forward reference frame, and obtain the third subjective and objective quality parameter between the current coded frame and the backward reference frame; Sort the first subjective and objective quality parameter and the second subjective and objective quality parameter according to their magnitude relationship to obtain a first sequence; Sort the first subjective and objective quality parameter and the third subjective and objective quality parameter according to their magnitude relationship to obtain a second sequence; Insert the synthesized frame into the forward reference frame list based on the position of the first subjective and objective quality parameter in the first sequence; and / or insert the synthesized frame into the backward reference frame list based on the position of the first subjective and objective quality parameter in the second sequence.

4. The inter-frame prediction method according to claim 1, wherein The obtaining the synthesized frame of the current coded frame based on the original forward reference frames in the forward reference frame list and the original backward reference frames in the backward reference frame list includes: Adjust the multiple component sizes of the original forward reference frames and the original backward reference frames to the same size threshold; Use the size-adjusted original forward reference frames and the size-adjusted original backward reference frames to obtain the synthesized frame of the current coded frame; Restore the component size of the synthesized frame to the component size corresponding to the current coded frame.

5. The inter-frame prediction method according to claim 1, wherein The obtaining the synthesized frame of the current coded frame based on the original forward reference frames in the forward reference frame list and the original backward reference frames in the backward reference frame list includes: Obtain the first information of the original forward reference frame in the forward reference frame list and the second information of the original backward reference frame in the backward reference frame list; Obtain the information of the synthesized frame of the current coding frame based on the first information and the second information; Wherein, the first information includes the image information of the original forward reference frame, and the second information includes the image information of the original backward reference frame; or, the first information includes the image information of the original forward reference frame and the side information of the original forward reference frame, and the second information includes the image information of the original backward reference frame and the side information of the original backward reference frame.

6. The inter-frame prediction method according to claim 1, wherein Further includes: Obtain the first quantity of the original forward reference frame in the forward reference frame list and the second quantity of the backward reference frame in the backward reference frame list; In response to the first quantity being greater than the second quantity, obtain the original forward reference frame located at the head of the forward reference frame list and insert it at the end of the backward reference frame list; In response to the first quantity being less than the second quantity, obtain the backward reference frame located at the head of the backward reference frame list and insert it at the end of the forward reference frame list.

7. The inter-frame prediction method according to claim 1, wherein The reference frame list to be updated is the forward reference frame list and the backward reference frame list, and the predicting the current coding frame based on the reference frame list to be updated inserted with the synthesized frame includes: Predict the current coding frame based on the synthesized frame and at least part of the original forward reference frames in the forward reference frame list and the synthesized frame and at least part of the original backward reference frames in the backward reference frame list.

8. The inter-frame prediction method according to claim 1, wherein The predicting the current coding frame based on the reference frame list to be updated inserted with the synthesized frame includes: Obtain the synthesized frame in the reference frame list to be updated; Set the motion vector of the synthesized frame to a zero vector to directly use the synthesized frame as the prediction frame of the current coding frame.

9. A coding method, characterized in that, Includes obtaining the prediction information of the current coding frame according to any one of claims 1-8 and encoding the prediction information into the bitstream of the current coding frame.

10. A decoding method, characterized in that, Obtain the prediction information of the current coding frame according to any one of claims 1-8 and the bitstream of the current coding frame according to claim 9, and decode the bitstream.

11. An encoder, characterized in that, The encoder includes a processor and a memory; a computer program is stored in the memory, and the processor is configured to execute the computer program to implement the encoding method as described in claim 9.

12. A decoder, characterized in that, The decoder includes a processor and a memory; a computer program is stored in the memory, and the processor is configured to execute the computer program to implement the decoding method as described in claim 10.

13. An electronic device, characterized in that, The electronic device includes a processor and a memory connected to the processor, wherein a program data is stored in the memory, and the processor executes the program data stored in the memory to execute the inter-frame prediction method according to any one of claims 1-8.

14. A computer-readable storage medium, characterized in that, It internally stores program instructions, and the program instructions are executed by the processor to implement the inter-frame prediction method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Multi-viewpoint video coding method based on time-domain-enhanced viewpoint synthesis prediction

    CN102413332A