Reference image list construction method, video encoding and decoding method, device and system

CN120130071APending Publication Date: 2025-06-10GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380076709.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-01-09
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing video coding and decoding standards fail to fully utilize temporal and spatial domain correlations when compressing multi-view light field videos, resulting in low compression efficiency.

Method used

By constructing a reference image list for multi-view light field video, including temporal reference images and spatial domain reference images, inter-frame prediction is performed using adjacent co-located views in the temporal domain and adjacent views at the same time, and the reference image list is optimized based on spatiotemporal correlation. .

Benefits of technology

The compression efficiency of multi-view light field video is improved and the coding performance is significantly improved, especially in scenarios with strong temporal correlation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120130071A_ABST
    Figure CN120130071A_ABST
Patent Text Reader

Abstract

The invention discloses a reference image list construction method of a multi-view light field video image, a video encoding and decoding method, a video encoding and decoding device and a video encoding and decoding system. When it is determined that a current image has at least one time domain reference image and at least one spatial domain reference image, based on the at least one time domain reference image and the at least one spatial domain reference image of the current image, the reference image list of the multi-view light field video image is constructed; and constructing a reference image list for performing inter-frame prediction on the current image. The time domain reference image is a co-located view in a multi-view array adjacent to the current image in the time domain, and the space domain reference image is an adjacent view in the current multi-view array of the current image.
Need to check novelty before this filing date? Find Prior Art

Description

Reference image list construction method, video encoding and decoding method, device and system

[0001] Technology Neighborhood

[0002] The embodiments of the present disclosure relate to, but are not limited to, video technology, and more specifically, to a method for constructing a reference image list for multi-view light field video images, and a video encoding and decoding method, device, and system for multi-view light field video. Background Art

[0003] Digital video compression technology primarily compresses large amounts of digital video data for easier transmission and storage. Currently popular video codec standards, such as H.266 / Versatile Video Coding (VVC), employ a block-based hybrid coding framework. Each video frame is divided into square largest coding units (LCUs) of equal size (e.g., 128x128, 64x64, etc.). Each LCU can be divided into rectangular coding units (CUs) based on a specific rule. Coding units may also be divided into prediction units (PUs) and transform units (TUs). The hybrid coding framework includes modules such as prediction, transform, quantization, entropy coding, and in-loop filtering. The prediction module includes intra-frame prediction and inter-frame prediction, which are used to reduce or remove inherent redundancy in the video. Inter-frame prediction includes motion estimation and motion compensation. Because adjacent pixels within a video frame are strongly correlated, intra-frame prediction is used in video codecs to eliminate spatial redundancy between adjacent pixels. Because adjacent frames within a video are highly similar, inter-frame prediction is used to eliminate temporal redundancy between them, thereby improving coding efficiency. In contrast to the prediction signal, the residual information is transformed, quantized, and entropy-encoded on a block-by-block basis into a bitstream.

[0004] With the surge in Internet videos and people's increasing demand for video clarity, although existing digital video compression standards can save a lot of video data, there is still a need to pursue better digital video compression technology to reduce the bandwidth and traffic pressure of digital video transmission.

[0005] SUMMARY OF THE INVENTION

[0006] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.

[0007] An embodiment of the present disclosure provides a method for constructing a reference image list for a multi-view light field video image, comprising:

[0008] Determining that a current image has at least one temporal reference image and at least one spatial reference image;

[0009] Constructing a reference image list for inter-frame prediction of the current image based on at least one temporal reference image and at least one spatial reference image of the current image;

[0010] The temporal reference image is a co-located view in a multi-view array adjacent to the current image in the temporal domain, the spatial reference image is an adjacent view of the current image in the current multi-view array, and both the temporal reference image and the spatial reference image are encoded images or decoded images.

[0011] An embodiment of the present disclosure further provides a video decoding method, including:

[0012] Constructing a reference image list for the current image according to the method for constructing a reference image list for a multi-view light field video image according to any embodiment of the present disclosure;

[0013] Decoding obtains motion information of a current block in a current image, and performing inter-frame prediction on the current block according to the motion information and the reference image list.

[0014] An embodiment of the present disclosure further provides a video encoding method, including:

[0015] Constructing a reference image list for the current image according to the method for constructing a reference image list for a multi-view light field video image according to any embodiment of the present disclosure;

[0016] Motion information of a current image in a current image is determined according to the reference image list, inter-frame prediction is performed on the current frame according to the motion information, and the motion information is encoded.

[0017] An embodiment of the present disclosure further provides a code stream, wherein the code stream is generated by the video encoding method as described in any embodiment of the present disclosure.

[0018] An embodiment of the present disclosure further provides a device for constructing a reference image list for multi-view light field video images, comprising a processor and a memory storing a computer program, wherein the processor, when executing the computer program, can implement the method for constructing a reference image list for multi-view light field video images as described in any embodiment of the present disclosure.

[0019] An embodiment of the present disclosure further provides a video decoding device, comprising a processor and a memory storing a computer program, wherein the processor can implement the video decoding method as described in any embodiment of the present disclosure when executing the computer program.

[0020] An embodiment of the present disclosure further provides a video encoding device, including a processor and a memory storing a computer program, wherein the processor can implement the video encoding method as described in any embodiment of the present disclosure when executing the computer program.

[0021] An embodiment of the present disclosure further provides a video encoding and decoding system, which includes the video encoding device described in any embodiment of the present disclosure and the video decoding device described in any embodiment of the present disclosure.

[0022] An embodiment of the present disclosure further provides a non-transitory computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program implements the method described in any embodiment of the present disclosure when executed by a processor.

[0023] An embodiment of the present disclosure further provides a computer program product, including a computer program, wherein when the computer program is executed by a processor, it can implement the method described in any embodiment of the present disclosure.

[0024] Still other aspects will become apparent upon reading and understanding the accompanying drawings and detailed description.

[0025] Summary of the Figures

[0026] The accompanying drawings are used to provide an understanding of the embodiments of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solutions of the present disclosure and do not constitute a limitation to the technical solutions of the present disclosure.

[0027] FIG1A is a schematic diagram of a coding and decoding system according to an embodiment of the present disclosure;

[0028] FIG1B is a framework diagram of an encoding end according to an embodiment of the present disclosure;

[0029] FIG1C is a framework diagram of a decoding end according to an embodiment of the present disclosure;

[0030] FIG2 is a schematic diagram of a group of pictures (GOP) structure of RA;

[0031] FIG3 is a schematic diagram of positions detected when constructing a motion information candidate list of a current block;

[0032] FIG4 is a schematic diagram of derivation of temporal motion information;

[0033] FIG5 is a schematic diagram of three dimensions of a multi-view light field video;

[0034] FIG6 is a schematic diagram of a light field image of an exemplary multi-view array;

[0035] FIG7 is a schematic diagram of collocated views between different multi-view arrays in a multi-view light field video;

[0036] FIG8 is an end-to-end system framework diagram of a dense light field;

[0037] FIG9 is a schematic diagram of two technical routes for dense light field compression;

[0038] FIG10 is an exemplary light field image;

[0039] FIG11 is a flowchart of a method for constructing a reference image list for a multi-view light field video image according to an embodiment of the present disclosure;

[0040] FIG12A is a schematic diagram of progressive scanning;

[0041] 12B and 12C are schematic diagrams of two zigzag scans;

[0042] FIG13 is a schematic diagram of performing a zigzag scan on a multi-view array;

[0043] FIG14 is a schematic diagram of a 5x5 multi-view array;

[0044] FIG15 is a schematic diagram of a temporal reference image and a spatial reference image of an image in a multi-view array at time 2 according to an embodiment of the present disclosure;

[0045] FIG16 is a flow chart of a video decoding method according to an embodiment of the present disclosure;

[0046] FIG17 is a video encoding method and flow chart according to an embodiment of the present disclosure;

[0047] FIG18 and FIG19 are RD curve diagrams of the test results of two sequences according to an embodiment of the present disclosure;

[0048] FIG20 is a schematic diagram of a device for constructing a reference image list for multi-view light field video images according to an embodiment of the present disclosure.

[0049] Details

[0050] The present disclosure describes multiple embodiments, but the description is exemplary rather than restrictive, and it is obvious to those skilled in the art that there may be more embodiments and implementations within the scope of the embodiments described in the present disclosure.

[0051] In the description of the present disclosure, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment described as "exemplary" or "for example" in the present disclosure should not be interpreted as being more preferred or advantageous than other embodiments. "And / or" in this article is a description of the association relationship of associated objects, indicating that there may be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. "Multiple" refers to two or more than two. In addition, in order to facilitate the clear description of the technical solutions of the embodiments of the present disclosure, words such as "first" and "second" are used to distinguish between identical or similar items with basically the same functions and effects. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit them to be different.

[0052] In this document, "including any one or more of the following: Option 1, Option 2,..." or "including any one or more of Option 1, Option 2,..." means including any one of the listed options, or any combination of multiple options. For example, "including any one or more of the following: A, B" or "including any one or more of A and B" means including only A, only B, or both A and B. Another example is: "including any one or more of the following: A, B, C" or "including any one or more of A, B, and C" means including only A, only B, only C, both A and B, both A and C, both B and C, or both A, B, and C. The same applies to options with more options.

[0053] When describing representative exemplary embodiments, the specification may have presented the method and / or process as a specific sequence of steps. However, to the extent that the method or process does not rely on the specific order of the steps described herein, the method or process should not be limited to the steps in the specific order described. As will be understood by those skilled in the art, other sequences of steps are also possible. Therefore, the specific sequence of the steps set forth in the specification should not be interpreted as a limitation to the claims. In addition, the claims for the method and / or process should not be limited to the steps performed in the order written, and those skilled in the art can readily understand that these sequences can vary and still remain within the spirit and scope of the disclosed embodiments.

[0054] The intra-frame prediction method and video coding and decoding method of the embodiments of the present disclosure can be applied to various video coding and decoding standards, such as: H.264 / Advanced Video Coding (AVC), H.265 / High Efficiency Video Coding (HEVC), H.266 / Versatile Video Coding (VVC), AVS (Audio Video Coding Standard), and other standards developed by MPEG (Moving Picture Experts Group), AOM (Alliance for Open Media), JVET (Joint Video Experts Team) and extensions of these standards, or any other customized standards.

[0055] Figure 1A is a block diagram of a video encoding and decoding system that can be used in embodiments of the present disclosure. As shown in the figure, the system is divided into an encoding end 1 and a decoding end 2. The encoding end 1 generates a bitstream. The decoding end 2 can decode the bitstream. The decoding end 2 can receive the bitstream from the encoding end 1 via a link 3. The link 3 includes one or more media or devices capable of moving the bitstream from the encoding end 1 to the decoding end 2. In one example, the link 3 includes one or more communication media that enable the encoding end 1 to send the bitstream directly to the decoding end 2. The encoding end 1 modulates the bitstream according to a communication standard and sends the modulated bitstream to the decoding end 2. The one or more communication media may include wireless and / or wired communication media and may form part of a packet network. In another example, the bitstream can also be output from the output interface 15 to a storage device, and the decoding end 2 can read the stored data from the storage device via streaming or downloading.

[0056] As shown in the figure, encoding end 1 includes a data source 11, a video encoding device 13, and an output interface 15. Data source 11 may include a video capture device (e.g., a camera), an archive containing previously captured data, a feed interface for receiving data from a content provider, a computer graphics system for generating data, or a combination of these sources. Video encoding device 13, also referred to as a video encoding end, encodes the data from data source 11 and outputs it to output interface 15. Output interface 15 may include at least one of a modulator, a modem, and a transmitter. Decoding end 2 includes an input interface 21, a video decoding device 23, and a display device 25. Input interface 21 includes at least one of a receiver and a modem. Input interface 21 may receive a bitstream via link 3 or from a storage device. Video decoding device 23, also referred to as a video decoding end, is used to decode the received bitstream. Display device 25 is used to display the decoded data. Display device 25 may be integrated with other devices in decoding end 2 or provided separately; display device 25 is optional for the decoding end. In other examples, the decoding end may include other devices or equipment for applying the decoded data.

[0057] FIG1B is a block diagram of an exemplary video encoding device that can be used in an embodiment of the present disclosure. As shown in the figure, the video encoding device 10 includes:

[0058] The division unit 101 is configured to cooperate with the prediction unit 100 to divide the received video data into slices, coding tree units (CTUs) or other larger units. The received video data may be a video sequence including video frames such as I frames, P frames or B frames.

[0059] The prediction unit 100 is configured to divide a CTU into coding units (CUs) and perform intra-frame prediction coding or inter-frame prediction coding on the CU. When performing intra-frame prediction and inter-frame prediction on the CU, the CU can be divided into one or more prediction units (PUs).

[0060] The prediction unit 100 includes an inter-frame prediction unit 121 and an intra-frame prediction unit 126 .

[0061] The inter-frame prediction unit 121 is configured to perform inter-frame prediction on the PU and generate prediction data for the PU, wherein the prediction data includes the prediction block of the PU, the motion information of the PU, and various syntax elements. The inter-frame prediction unit 121 may include a motion estimation (ME) unit and a motion compensation (MC) unit. The motion estimation unit can be used to perform motion estimation to generate a motion vector, and the motion compensation unit can be used to obtain or generate a prediction block based on the motion vector.

[0062] The intra prediction unit 126 is configured to perform intra prediction on a PU and generate prediction data for the PU. The prediction data for the PU may include a prediction block and various syntax elements of the PU.

[0063] The residual generating unit 102 (indicated by the circle with a plus sign after the dividing unit 101 in the figure) is configured to subtract the prediction block of the PU into which the CU is divided from the original block of the CU to generate a residual block of the CU.

[0064] The transform processing unit 104 is configured to partition a CU into one or more transform units (TUs). The partitioning of prediction units and transform units may be different. A TU-associated residual block is a sub-block obtained by partitioning the residual block of the CU. A TU-associated coefficient block is generated by applying one or more transforms to the TU-associated residual block.

[0065] The quantization unit 106 is configured to quantize the coefficients in the coefficient block based on a quantization parameter. The quantization degree of the coefficient block can be changed by adjusting the quantization parameter (QP: Quantizer Parameter).

[0066] The inverse quantization unit 108 and the inverse transform unit 110 are configured to apply inverse quantization and inverse transform to the coefficient block, respectively, to obtain a reconstructed residual block associated with the TU.

[0067] The reconstruction unit 112 (represented by the circle with a plus sign after the inverse transform processing unit 110 in the figure) is configured to add the reconstructed residual block and the prediction block generated by the prediction unit 100 to generate a reconstructed image.

[0068] The filter unit 113 is configured to perform loop filtering on the reconstructed image.

[0069] The decoded image buffer 114 is configured to store the reconstructed image after loop filtering. The intra prediction unit 126 can extract reference images of blocks adjacent to the current block from the decoded image buffer 114 to perform intra prediction. The inter prediction unit 121 can use the reference image of the previous frame cached in the decoded image buffer 114 to perform inter prediction on the PU of the current frame image.

[0070] The entropy coding unit 115 is configured to perform entropy coding operations on received data (such as syntax elements, quantized coefficient blocks, motion information, etc.) to generate a video bitstream.

[0071] In other examples, the video encoding apparatus 10 may include more, fewer, or different functional components than those in this example, for example, the transform processing unit 104 and the inverse transform processing unit 110 may be eliminated.

[0072] FIG1C is a block diagram of an exemplary video decoding device that can be used in an embodiment of the present disclosure. As shown in the figure, the video decoding device 15 includes:

[0073] Entropy decoding unit 150 is configured to perform entropy decoding on the received encoded video stream, extracting syntax elements, quantized coefficient blocks, and motion information for PUs. Prediction unit 152, inverse quantization unit 154, inverse transform processing unit 156, reconstruction unit 158, and filter unit 159 may each perform corresponding operations based on the syntax elements extracted from the stream.

[0074] The inverse quantization unit 154 is configured to perform inverse quantization on the coefficient block associated with the quantized TU.

[0075] The inverse transform processing unit 156 is configured to apply one or more inverse transforms to the inverse quantized coefficient block to generate a reconstructed residual block for the TU.

[0076] Prediction unit 152 includes an inter-prediction unit 162 and an intra-prediction unit 164. If the current block is encoded using intra prediction, intra prediction unit 164 determines an intra prediction mode for the PU based on syntax elements decoded from the codestream, and performs intra prediction in conjunction with reconstructed reference information adjacent to the current block obtained from decoded image buffer 160. If the current block is encoded using inter prediction, inter prediction unit 162 determines a reference block for the current block based on motion information of the current block and corresponding syntax elements, and performs inter prediction on the reference block obtained from decoded image buffer 160.

[0077] The reconstruction unit 158 ​​(represented by a circle with a plus sign after the inverse transform processing unit 155 in the figure) is set to perform intra-frame prediction or inter-frame prediction on the current block based on the reconstructed residual block associated with the TU and the prediction unit 152 to obtain a reconstructed image.

[0078] The filter unit 159 is configured to perform loop filtering on the reconstructed image.

[0079] The decoded image buffer 160 is configured to store the reconstructed image after loop filtering as a reference image for subsequent motion compensation, intra-frame prediction, inter-frame prediction, etc. The filtered reconstructed image can also be output as decoded video data for presentation on a display device.

[0080] In other embodiments, the video decoding device 15 may include more, fewer, or different functional components. For example, the inverse transform processing unit 155 may be eliminated in some cases.

[0081] Based on the above-described video encoding and decoding devices, the following basic encoding and decoding process can be performed. On the encoding side, a frame of an image is divided into blocks, or first into multiple slices and then into blocks. Slices within the same image can be processed in parallel. Intra-frame prediction, inter-frame prediction, or other algorithms are performed on the current block to generate a prediction block for the current block. The prediction block is subtracted from the original block of the current block to obtain a residual block. The residual block is transformed and quantized to obtain a quantization coefficient matrix. The quantization coefficient matrix is ​​then entropy encoded to generate a bitstream. On the decoding side, intra-frame prediction or inter-frame prediction is performed on the current block to generate a prediction block for the current block. The quantization coefficient matrix obtained from the decoded bitstream is then inversely quantized and inversely transformed to obtain a residual block. The prediction block and residual block are added together to obtain a reconstructed block. The reconstructed block forms a reconstructed image. The reconstructed image is then subjected to image-based or block-based loop filtering to obtain a decoded image. The encoding side also performs similar operations as the decoding side to obtain a decoded image, also known as a loop-filtered reconstructed image. The loop-filtered reconstructed image can be used as a reference frame for inter-frame prediction of subsequent frames. The block division information, prediction, transform, quantization, entropy coding, loop filtering, and other mode and parameter information determined by the encoder can be written into the bitstream. The decoder determines the block division information, prediction, transform, quantization, entropy coding, loop filtering, and other mode and parameter information used by the encoder by decoding the bitstream or analyzing existing information, thereby ensuring that the decoded image obtained by the encoder and the decoder is the same.

[0082] Although the above example is given of a block-based hybrid coding framework, the embodiments of the present disclosure are not limited thereto. With the development of technology, one or more modules in the framework and one or more steps in the process may be replaced or optimized.

[0083] The embodiments of the present disclosure relate to, but are not limited to, the inter-frame prediction units in the above-mentioned encoding end and decoding end and the corresponding inter-frame prediction methods.

[0084] Videos are composed of images. To ensure a smooth video appearance, each second of video contains dozens or even hundreds of frames. Examples include 24, 30, 50, 60, and 120 frames per second. This results in significant temporal redundancy in the video. In other words, there is a significant amount of temporal correlation. Therefore, inter-frame prediction often uses "motion" to exploit temporal correlation. A very simple "motion" model is that an object is at a certain position in the image corresponding to a certain moment. After a certain amount of time, it moves to another position in the image corresponding to that moment. This is the most basic and commonly used translational motion in video codecs.

[0085] Inter-frame prediction uses motion information to represent "motion". Basic motion information includes information about a reference picture and a motion vector (MV). The reference picture can also be called a reference frame. The codec determines the reference picture based on the information about the reference picture, and determines the coordinates of the reference block based on the information about the motion vector and the coordinates of the current block. The coordinates of the reference block are used to determine the reference block in the reference picture. Motion in a video is not always so simple. Even motion that can be considered as translation can have subtle changes over time, including subtle deformations, changes in brightness, changes in noise, etc.

[0086] More than one reference block can be used to predict the current block to achieve better prediction results. For example, in the commonly used bidirectional prediction, two reference blocks are used to predict the current block. The two reference blocks can use a forward reference block and a backward reference block. Later, both are forward or both are backward. The so-called forward means that the time corresponding to the reference image is before the current frame, and the backward means that the time corresponding to the reference image is after the current frame. In other words, forward means that the position of the reference image in the video is before the current frame, and backward means that the position of the reference image in the video is after the current frame. In other words, forward means that the picture order count (POC) of the reference image is less than the POC of the current frame, and backward means that the POC of the reference image is greater than the POC of the current frame.

[0087] Video coding standards may support the prediction of more reference blocks. A simple way to generate a prediction block using two reference blocks is to average the pixel values ​​at corresponding positions of the two reference blocks to obtain the prediction block. In order to obtain better prediction results, weighted averaging can also be used, such as the BCW (Bi-prediction with CU-level weight) currently used by VVC. The geometric partitioning mode (GPM) in VVC can also be understood as a special bidirectional prediction. In order to use bidirectional prediction, it is naturally necessary to find two reference blocks, so two sets of reference image information and motion vector information are required. Each of them can be understood as a unidirectional motion information.

[0088] Motion in video isn't limited to simple translation; it also includes scaling, rotation, distortion, and various other types of information. Combining these two sets of information creates bidirectional motion information. In practice, unidirectional and bidirectional motion information can use the same data structure, but both reference image information and motion vector information are valid for bidirectional motion information, while one reference image and motion vector information for unidirectional motion information are invalid.

[0089] Reference image (reference frame)

[0090] Without considering parallel processing, video is processed one image at a time. Already encoded and decoded images can be stored in a buffer and used as reference images for subsequent encoded and decoded images. Current codec standards all have a reference image management method to manage reference images. This method manages which frames can be used as reference images for the current frame, the indexes of these reference images, which images need to be cached, and which images can no longer be used as reference images and can be removed from the cache.

[0091] According to the different image encoding and decoding orders, the commonly used scenarios can be divided into two categories: random access (RA) and low delay (LD). The image display order of low delay is the same as the encoding and decoding order, while the image display order of random access can be different from the encoding and decoding order. In layman's terms, low latency means encoding and decoding frame by frame in the order of the video itself, while random access can disrupt the order of the video itself. Some frames are not encoded and decoded first, and some subsequent frames are encoded and decoded first, and then the skipped frames are encoded and decoded again. One advantage of random access is that some frames can refer to both the reference image before it and the reference image after it, so that "motion" can be better utilized to achieve better compression efficiency.

[0092] Figure 2 shows a classic RA group of pictures (GOP) structure. P frames (Predictive Frames) in the figure are frames that can only be predicted in a single forward direction. B frames (Bi-predictive Frames) are frames that can be predicted in a bidirectional direction. This reference relationship restriction can also be applied to sub-image levels, such as by dividing P slices and B slices at the slice level.

[0093] The arrows in Figure 2 indicate reference relationships. I-frames do not require a reference image. After encoding and decoding an I-frame with a POC of 0, encoding and decoding a P-frame with a POC of 4 can refer to the I-frame with a POC of 0. Then, encoding and decoding a B-frame with a POC of 2 can refer to the I-frame with a POC of 0 and the P-frame with a POC of 4. ...

[0094] The codec uses reference picture lists to manage reference images. VVC supports two reference picture lists, denoted as RPL0 and RPL1, where RPL is the abbreviation of Reference Picture List. P slice in VVC can only use RPL0, and B slice can use RPL0 and RPL1. For a slice, there are several reference images in each reference picture list, and the codec finds a reference image through the reference picture index. VVC uses reference picture index and motion vector to represent motion information. For example, for the above-mentioned bidirectional motion information, VVC uses the reference picture index refIdxL0 corresponding to reference picture list 0, and the motion vector mvL0 corresponding to reference picture list 0, the reference picture index refIdxL1 corresponding to reference picture list 1, and the motion vector mvL0 corresponding to reference picture list 1. The reference picture index corresponding to reference picture list 0 and the reference picture index corresponding to reference picture list 1 here can be understood as the information of the above-mentioned reference image.

[0095] Motion information prediction (motion vector prediction or motion information prediction)

[0096] The current block may use the motion information to find a reference block from the reference image and determine a prediction block for the current block based on the reference block.

[0097] The motion information used by the current block can usually be predicted using some relevant information, which can be called motion information prediction or motion vector prediction. For example, the motion information used by the adjacent already-encoded blocks around the current block in the current image (such as a frame or slice) can be used, because there is a strong correlation between adjacent blocks. The motion information used by the already-encoded blocks around the current block in the current image that are not adjacent can also be used, because even if they are not adjacent, there is a certain correlation between the closer areas. This method of using the motion information used by the already-encoded blocks around the current block in the current image to predict motion information is generally called spatial motion information prediction. In addition to the already-encoded blocks around the current block, the motion information of the blocks related to the position of the current block in the already-encoded image can also be used to predict the motion information of the current block, which is generally called temporal motion information prediction.

[0098] The motion information of the encoded and decoded blocks can be maintained in a list in the order of encoding and decoding. The list generally retains the motion information of several different recently encoded and decoded blocks. The motion information in this list is used to predict the motion information of the current block. This is generally called history-based motion information prediction (HBP). It can be simply understood as using the motion information of the same image as the current block to derive spatial motion information, and using the motion information of different images to derive temporal motion information.

[0099] To use spatial motion information prediction, the motion information of the encoded and decoded blocks in the current image (or slice) must be stored. A minimum storage unit is typically set, such as a 4×4 unit, but it can also be 8×8 or other sizes. Each time a block is encoded and decoded, the codec stores motion information for all the corresponding minimum storage units. To find the motion information of a neighboring block, the current block can locate a minimum storage unit based on its coordinates and obtain its motion information. Similarly, to use temporal motion information prediction, the motion information of the encoded and decoded image (or slice) must be stored. A minimum storage unit is typically also set, which can be the same size as or different from the size of the spatial motion information storage unit, depending on the standard. To find the motion information of a block in the image (or slice), the current block can also locate a minimum storage unit based on its coordinates and obtain its motion information. Note that due to storage space limitations or implementation complexity, the current block may only be able to obtain temporal or spatial motion information within a certain coordinate range.

[0100] Motion information prediction can yield one or more motion information. If multiple motion information is obtained, it is usually necessary to specify which one or more motion information to select, or to select one or more motion information according to certain established rules. For example, GPM in VVC requires the use of two motion information sets. For example, sub-block based TMVP (SbTMVP) requires one motion information set for each sub-block, where TMVP refers to temporal motion vector predictor.

[0101] In order to use the predicted motion information, a certain piece of predicted motion information can be directly used as the motion information of the current block. An example is the merge in HEVC. Or, based on a certain piece of predicted motion information, a motion vector difference (MVD) is added to obtain new motion information. From a design point of view, the closer the predicted motion information is to the actual motion information, the better. However, there is no guarantee that the motion information prediction is always accurate. Therefore, in order to obtain more accurate motion information, a motion vector deviation can be used. A new method of representing motion vector deviation has been added to VVC. In merge mode, the method of using motion vector prediction plus this new motion vector deviation can be used. This method is called merge mode with motion vector deviation (MMVD for short). In short, motion information prediction can be used directly or in conjunction with other methods.

[0102] The following is an example of the method for constructing the merge motion information candidate list in VVC.

[0103] The merge motion information candidate list is denoted as mergeCandList. When constructing mergeCandList, first check the prediction of the spatial motion information of positions 1 to 5 in Figure 3, and then check the temporal motion information prediction. When checking the temporal motion information prediction, if position 6 is available, the temporal motion information prediction is derived (driven) based on position 6. This derivation can also be called derivation, that is, the temporal motion information is calculated step by step based on position 6. The availability of position 6 means that position 6 does not exceed the boundary of the image or sub-image and it is in the same CTU row as the current block. For details, please refer to the standard text. If position 6 is not available, the temporal motion information prediction is derived based on position 7. Position 6 can be represented by coordinates (xColBr, yColBr) and position 7 can be represented by coordinates (xColCtr, yColCtr)

[0104] A specific method for determining spatial motion information is as follows. Taking the spatial motion information for position 2 in Figure 3 as an example, position 2 is denoted as B1, or alternatively, the block where position 2 is located is denoted as B1. The following information is derived: the available flag availableFlagB1, the reference image index refIdxLXB1, the reference image list usage flag predFlagLXB1, and the motion vector mvLXB1. The availableFlagB1 can be used to indicate whether the spatial motion information for B1 is available. The reference image index refIdxLXB1, the reference image list usage flag predFlagLXB1, and the motion vector mvLXB1 together constitute the motion information. Where X = 0..1.

[0105] Note that (xCb, yCb) is the coordinate of the upper left corner of the current block relative to the upper left corner of the current image, cbWidth is the width of the current block, and cbHeight is the height of the current block.

[0106] The coordinates (xNbB1, yNbB1) within the adjacent block B1 are set to (xCb + cbWidth - 1, yCb - 1). Determine whether the block containing (xNbB1, yNbB1) is usable. One method of determining this is that if the block containing (xNbB1, yNbB1) has been coded and decoded, and is inter-coded, the block is usable; otherwise, it is unusable. Additional criteria can also be used.

[0107] One understanding is to use the motion information at (xNbB1, yNbB1) as the spatial motion information prediction of position B1, and the spatial motion information prediction of other positions can also be derived in a similar way.

[0108] Temporal motion information derivation

[0109] As shown in Figure 4, currPic is the current image, currCb is the current block, currRef is the reference image for the current block's temporal motion information, colCb is the co-located block, colPic is the image where the co-located block is located, and colRef is the reference image for the motion information used by the co-located block. When deriving temporal motion information, colCb is first found from colPic based on position information, and the motion information of colCb is found. The motion information of colCb shown in the figure is the motion vector from colPic to colRef, as shown by the solid line. Temporal motion information scales the motion vector shown in the solid line to the motion vector from currPic to currRef, as shown by the dotted line. tb is a variable determined based on the difference in the POC from currPic to currRef, and td is a variable determined based on the difference in the POC from colPic to colRef. The motion vector is scaled based on tb and td. The illustration shows unidirectional motion information, or the scaling of a motion vector. It is understood that temporal motion information can use bidirectional motion information.

[0110] A specific method for deriving temporal motion information is as follows. Taking the derivation of temporal motion information of position 6 in Figure 3 as an example, position 6 is recorded as Col, or the block where position 6 is located in the reference image is recorded as Col. Col is the abbreviation of collocated. The block is called a co-located block, and the reference image is called a co-located image. The following information is derived: availableFlagCol, reference image index refIdxLXCol, reference image list usage flag predFlagLXCol, and motion vector mvLXCol. The above availableFlagCol can be used to indicate whether the temporal motion information of Col is available. The reference image index refIdxLXCol, reference image list usage flag predFlagLXCol, and motion vector mvLXCol together constitute the motion information. Where X = 0..1.

[0111] Let the coordinates (xColBr, yColBr) of position 6 be (xCb + cbWidth, yCb + cbHeight). If the coordinates (xColBr, yColBr) meet the requirements, such as not exceeding the range of the current image or sub-image, and not exceeding the range of the CTU row where the current block is located, let the coordinates of the co-located block (xColCb, yColCb) be ((xColBr>>3)<<3, (yColBr>>3)<<3). Let the current block be currCb, and the co-located block on the co-located image ColPic be colCb, where colCb is the block covering (xColCb, yColCb). currPic is the current image. The reason for first shifting right (>>) by 3 bits and then left (<<) by 3 bits when calculating the coordinates is because the motion information in the co-located image in this example is stored in a minimum storage unit of 8×8 (to save cache space, the granularity of the cached reference image motion information can be coarser).

[0112] There are currently many approaches to light field video compression. One method for compressing light field video sequences is to represent light field images as sub-aperture image arrays, referred to as multi-views (typically 5×5, 7×7, 9×9, or larger). The multi-view arrays at each moment are reordered according to certain rules to form a pseudo-video sequence. These pseudo-video sequences at each moment are then concatenated into a complete long video sequence, which can then be encoded and decoded using video encoding tools (such as AVC, HEVC, or VVC).

[0113] Using a multi-view approach to represent the light field, each moment in the light field video can be represented as a multi-view array. Multiple moments in time represent multiple multi-view arrays, as shown in Figure 5. The figure depicts three dimensions: the temporal dimension t and the spatial dimensions xy. Figure 5 shows three multi-view arrays at three moments in time, each consisting of 5x5=25 views. The spatiotemporal correlations between views can be mined from the temporal dimension t and the spatial dimension xy.

[0114] As shown in Figure 6, a multi-view array of a light field image differs significantly from a multi-view array captured using a multi-camera system: the baseline distance between adjacent views in a light field multi-view array is very narrow, and the parallax is very small, making it difficult to detect noticeable changes in parallax with the naked eye. This is one of the characteristics of a light field multi-view array: a narrow baseline. This characteristic can be summarized as follows: at a given moment in a light field video, the spatial correlation between multiple views is generally high, especially between spatially adjacent views.

[0115] Because the compressed video is a light field, in addition to the spatial correlation at a single moment, there is also temporal correlation at different moments. There is a strong temporal correlation between the co-located views in the multi-view array at two moments. The so-called co-located views refer to the views having the same position in the multi-view array. Figure 7 shows some co-located views, and the same type of wireframes are co-located views. For example, in the multi-view array at multiple moments in Figure 7, the views in the i-th row and j-th column are co-located views, i = 1, 2, ..., 5, j = 1, 2, ..., 5.

[0116] The Dense Light Field (DLF) working group within MPEG-I focuses on light field video. Figure 8 shows the end-to-end system architecture of DLF. The overall DLF processing flow includes light field video capture, data conversion, compression, rendering, and final display. Light Field Video Coding (LVC), an AHG working group within WG04, focuses on coding techniques for compressing content captured by light field cameras or multi-camera arrays. LVC corresponds to the encoder and decoder sections in Figure 8.

[0117] There are two technical approaches to DLF compression, as shown in Figure 9. The four icons at the top of the figure capture capture, data format, data format, and display, representing the data representation methods (light field video or multi-view video) during the acquisition phase, pre-decoding phase, post-decoding phase, and final display phase, respectively. The first technical approach is to directly compress light field video (i.e., Lenslet video), as shown in the top half of the figure. Figure 10 shows an example light field image. It can be seen that it is very different from a typical image, with many high-frequency signals, making direct compression difficult.

[0118] The second technical approach is to convert the light field video into a multi-view video (also referred to herein as a multi-view light field video), and then compress the multi-view light field video, as shown in the lower half of the figure. The disclosed embodiments are directed to this technical approach. The encoder and decoder (Encoder & Decoder) can use video codecs such as AVC, HEVC, or VVC.

[0119] However, these video codecs fail to consider the characteristics of light field multi-view arrays during video compression. The reference image list for inter-frame prediction often consists of short-range reference frames. In this case, the video encoder and decoder only look for correlations between different views at a given moment, compressing more of the spatial redundancy between different views at the same moment. As a result, in some cases, when performing inter-frame prediction, the encoder fails to select a more correlated reference image based on the temporal correlation between multiple views. As a result, the selected reference image does not conform to the spatial distribution pattern of the multi-view array, resulting in low compression efficiency.

[0120] The disclosed embodiment performs spatiotemporal predictive coding on the motion features of a multi-view light field video (such as a subaperture array light field video). Before encoding, the multi-view array at each moment of the multi-view light field video is converted into a pseudo video sequence, which is then connected into a complete long video sequence. The long video sequence is encoded using a codec tool. During the inter-frame prediction process, the temporal reference image and the spatial reference image of the current image are added to the reference image list for inter-frame prediction. The temporal reference image selects adjacent co-located frames in the temporal domain, and the spatial reference frame selects adjacent frames on the multi-view array at the same moment. This can take into account both temporal and spatial correlations, thereby improving the compression effect of the multi-view light field video.

[0121] In an embodiment of the present disclosure, the temporal reference image of the current image is a co-located view in a multi-view array adjacent to the current image in the temporal domain, and the spatial reference image of the current image is an adjacent view of the current image in the current multi-view array, and both the temporal reference image and the spatial reference image are encoded images or decoded images.

[0122] In this paper, a view in a multi-view array is referred to as a picture during encoding and decoding, which may also be referred to as a frame, and a reference picture may also be referred to as a reference frame.

[0123] By scanning multiple views in a multi-view array simultaneously, a pseudo-sequence can be generated. There are many ways to scan multiple views in a multi-view array, such as row-by-row scanning (see FIG12A ), zigzag scanning (see FIG12A and FIG12B ), or zigzag scanning ( FIG13 ). The zigzag scanning can start in any direction; FIG13 shows the left direction, but the present disclosure is not limited thereto.

[0124] Under different scanning modes, the spatial reference frames that can be obtained by the current image are different. Referring to Figure 15, when the zigzag scanning mode is adopted, it is assumed that the encoding and decoding (encoding or decoding) of the multi-view array at time 1 has been completed, and the image in the multi-view array at time 2 is being encoded and decoded. According to the order of the zigzag scanning, the first encoded and decoded image in the multi-view array at time 2 is the view marked as 25, and the view marked as n in the figure is hereinafter referred to as view Pn. Because P25 is the first image in the current multi-view array, P25 has no spatial reference image, but P25 has a co-located view, namely P0, in the multi-view array at time 1. When bidirectional prediction is allowed, P25 also allows the co-located view P50 in the multi-view array at the next moment, namely time 3, to be used as another temporal reference image. In other embodiments, the temporal reference image of the current image may be more than 2.

[0125] Assuming that this embodiment only allows unidirectional prediction, as shown in FIG15 , the temporal reference image of P26 is P1, the temporal reference image of P27 is P2, the temporal reference image of P28 is P3, the temporal reference image of P29 is P4, and so on.

[0126] After encoding and decoding P25, when encoding and decoding P26, P26 has a spatial reference image P25 and a temporal reference image P1. After encoding and decoding P26, when encoding P27, the adjacent views of P27 can be different depending on different definitions. Neighborhoods are divided into three categories: 4-neighborhood, diagonal neighborhood, and 8-neighborhood. For a nine-square grid centered on the current image, the four views covered by a "plus sign" are views in the 4-neighborhood of the current image; the views in the four corners of the nine-square grid are views in the diagonal neighborhood of the current image; and all eight views surrounding the nine-square grid are views in the 8-neighborhood of the current image.

[0127] When the views in the 4-neighborhood of the current image are used as the neighboring views of the current image, for P27, the neighboring views are P26, P36, P38, and P28. Of these, only P26 has been coded and decoded. In this case, P27 has only one spatial reference image, P26. In another embodiment, the views in the 8-neighborhood of a view are used as the neighboring views of that view. For P27, the neighboring views are P26, P35, P36, P37, P38, P39, P28, and P25. Of these, only P25 and P26 have been coded and decoded. In this case, the spatial reference images for P27 are P25 and P26. The spatial reference images for subsequent images can be determined in a similar manner and will not be further described.

[0128] As described above, in the example shown in Figure 15, the center view of the multi-view array is the first view to be coded at the current moment. There is no spatial reference image, and only the temporally co-located view serves as the temporal reference image. While only one spatial reference image is allowed, the other views all have one temporal reference image and one spatial reference image, indicated by solid and dashed lines, respectively.

[0129] In one embodiment, as shown in FIG15 , when views in an 8-neighborhood of the current image are used as neighboring views of the current image, the spatial reference images for P32 are P26, P25, P30, and P31. In this case, views P25 and P31 in the 4-neighborhood of P32 are preferentially added to the reference image list. After P25 and P31 have been added to the reference image list, if spatial reference images are still needed, spatial reference images P26 and P30 in the diagonal neighborhood of P32 are added.

[0130] The embodiments of the present disclosure do not limit the tools used to encode and decode long video sequences, and the tools may be any one of VVC, HEVC, or AVS, or other video encoding and decoding tools.

[0131] An embodiment of the present disclosure provides a method for constructing a reference image list for a multi-view light field video image, as shown in FIG11 , including:

[0132] Step S110, determining that the current image has at least one temporal reference image and at least one spatial reference image;

[0133] Step S120, constructing a reference image list for inter-frame prediction of the current image based on at least one temporal reference image and at least one spatial reference image of the current image;

[0134] The temporal reference image is a co-located view in a multi-view array adjacent to the current image in the temporal domain, the spatial reference image is an adjacent view of the current image in the current multi-view array, and both the temporal reference image and the spatial reference image are encoded images or decoded images.

[0135] In this document, the current multi-view array is the multi-view array where the current image is located. The multi-view array adjacent to the current image in the time domain refers to the multi-view array adjacent to the current multi-view array at the corresponding moment. For example, when the current image is the image in the multi-view array corresponding to time 2 in Figure 15, the multi-view array adjacent to the current image in the time domain includes the multi-view array corresponding to time 1 (i.e., the multi-view array at the previous moment) and the multi-view array corresponding to time 3 (i.e., the multi-view array at the next moment). In other embodiments, the multi-view array corresponding to time 4 may also be included. This depends on whether the hardware platform supports it. If it does not support it, even the co-located view in the coded multi-view array cannot be used as the time domain reference view of the current view. For example, in some cases, the image in the multi-view array corresponding to time 3 cannot use the co-located view in the multi-view array corresponding to time 1 as the time domain reference image.

[0136] The disclosed embodiment adds short-range spatial reference images (with strong spatial correlation) and long-range reference images (temporal reference frames) to the reference image list, providing reference images in both time and space dimensions, which can improve the compression effect of multi-view light field video.

[0137] In an exemplary embodiment of the present disclosure, the method further includes:

[0138] When it is determined that the current image has only a spatial reference image, the reference image list is constructed based on the determined at least one spatial reference image. For example, for an image in the first multi-view array of a multi-view light field video, when only unidirectional prediction is allowed and there is no temporal reference image, one or more spatial reference images can be added to the reference image list.

[0139] When it is determined that the current image has only a temporal reference image, the reference image list is constructed based on the determined at least one temporal reference image. For example, for the first image scanned in the multi-view array, there is only a temporal reference image and no spatial reference image. In this case, one or more temporal reference images may be added to the reference image list of the current image.

[0140] For the first image that is encoded and decoded first in the first multi-view array in a multi-view light field video, its reference image list can be empty.

[0141] In an exemplary embodiment of the present disclosure, the adjacent views are views in a 4-neighborhood of the current image in the current multi-view array, or the adjacent views are views in an 8-neighborhood of the current image in the current multi-view array. This has been explained above and will not be repeated here.

[0142] In an exemplary embodiment of the present disclosure, the co-located view includes one or more of the following: the co-located view of the current image in the multi-view array at the previous moment, and the co-located view of the current image in the multi-view array at the next moment. Here, the "multi-view array at the previous moment" refers to the multi-view array at the previous moment corresponding to the corresponding moment of the current multi-view array, and the "multi-view array at the next moment" refers to the multi-view array at the next moment corresponding to the corresponding moment of the current multi-view array. Please refer to the above description. In one example, during uni-directional prediction, the co-located view of the current image in the multi-view array at the previous moment can be used as the temporal reference image of the current image, and during bi-directional prediction, the co-located views of the current image in the multi-view arrays at the previous and next moments can be used as the temporal reference images of the current image.

[0143] In an exemplary embodiment of the present disclosure, constructing a reference image list for inter-frame prediction of the current image based on at least one temporal reference image and at least one spatial reference image of the current image includes: when the current image has M temporal reference images and N spatial reference images, and M≥m, N≥n, adding m temporal reference images and n spatial reference images of the current image to the reference image list; where m and n are set values, and M, N, m, and n are all greater than or equal to 1.

[0144] In this embodiment, m is the maximum number of temporal reference images that can be added to the reference image list, and n is the maximum number of spatial reference images that can be added to the reference image list. When both the temporal reference images and spatial reference images of the current image exceed their respective maximum numbers, only m temporal reference images and n spatial reference images of the current image are allowed to be added to the reference image list. For example, when m = 1, n = 2, and the current image has 2 temporal reference images and 3 spatial reference images, only 1 temporal reference image and 2 spatial reference images of the current image are allowed to be added to the reference image list.

[0145] In one example of this embodiment, constructing a reference image list for inter-frame prediction of the current image based on at least one temporal reference image and at least one spatial reference image of the current image further includes:

[0146] When the current image has M temporal reference images and N spatial reference images, and M<m, N≥n, adding M temporal reference images and n spatial reference images of the current image to the reference image list;

[0147] When the current image has M temporal reference images and N spatial reference images, and M≥m, N<n, adding m temporal reference images and N spatial reference images of the current image to the reference image list;

[0148] When there are M temporal reference images and N spatial reference images in the current image, and M < m, N < n, add the M temporal reference images and N spatial reference images of the current image to the reference image list.

[0149] This example illustrates the processing when one or both of the temporal reference image and the spatial reference image of the current image do not exceed the set maximum number. When the numbers of the temporal reference image and the spatial reference image of the current image do not exceed their respective maximum numbers, add both the temporal reference image and the spatial reference image of the current image to the reference image list. If the number of the temporal reference image or the spatial reference image of the current image exceeds the corresponding maximum number, only the set maximum number of temporal reference images or spatial reference images can be added to the reference image list.

[0150] In this embodiment, m = n = 1; or, m = 1, n = 2; or, m = 1, n = 3; or, m = 1, n = 4; or, m = 1, n = 5; or, m = n = 2; or, m = 2, n = 3; or, m = 2, n = 4; or, m = 3, n = 3; where m + n = x, x ≥ 2, and x is the maximum number of reference frames that can be added to the reference image list, which can be a set value. In one example, x is 2, 3, 4, 5, or 6. This embodiment gives some typical set values of m and n, but the embodiments of the present disclosure are not limited to the numbers listed here.

[0151] In one example of this embodiment, the encoder selects one temporal reference image and one spatial reference image (i.e., the case of m = n = 1) to construct the reference image list. During the predictive coding process, the encoder will adaptively select a frame with higher correlation as the reference image for the current block.

[0152] In an exemplary embodiment of the present disclosure, constructing a reference image list for inter-frame prediction of a current image based on at least one temporal reference image and at least one spatial reference image of the current image includes:

[0153] When there are M temporal reference images and N spatial reference images in the current image, and M + N ≤ x, add the M temporal reference images and N spatial reference images of the current image to the reference image list, where M, N ≥ 1, x ≥ 2, and x is the maximum number of reference frames that can be added to the reference image list.

[0154] The disclosed embodiment does not limit the number of temporal reference images and spatial reference images respectively, but only limits the number of temporal reference images and spatial reference images added to the list according to the length x of the reference image list (that is, the maximum number of reference images that can be added). When the total number of temporal reference images and spatial reference images of the current image does not exceed the length of the reference image list, all the temporal reference images and spatial reference images of the current image are added to the reference image list.

[0155] In an exemplary embodiment of the present disclosure, constructing a reference image list for inter-frame prediction of the current image based on at least one temporal reference image and at least one spatial reference image of the current image includes:

[0156] When the current image has M temporal reference images and N spatial reference images, and M+N>x, if M=1, add 1 temporal reference image and x-1 spatial reference images to the reference image list; if N=1, add x-1 temporal reference images and 1 spatial reference image to the reference image list; wherein M, N≥1, x≥2, and x is the maximum number of reference frames that can be added to the reference image list.

[0157] In this embodiment, if the total number of temporal reference images and spatial reference images of the current image exceeds the length of the reference image list, and the number of either the temporal reference image or the spatial reference image is 1, the temporal reference image or the spatial reference image with the number 1 is added to the reference image list. Then, multiple spatial reference images or temporal reference images with the number not being 1 are added to fill the list.

[0158] In an exemplary embodiment of the present disclosure, constructing a reference image list for inter-frame prediction of the current image based on at least one temporal reference image and at least one spatial reference image of the current image includes:

[0159] When the current image has M temporal reference images and N spatial reference images, M+N>x and both M and N are greater than 1, determining a temporal correlation between the temporal reference images and the current image, as well as a spatial correlation between the spatial reference images and the current image, and determining the number of temporal reference images and the number of spatial reference images to be added to the reference image list based on the temporal correlation and the spatial correlation;

[0160] Wherein, x≥2, and x is the maximum number of reference frames that can be added to the reference image list.

[0161] For a video, the degree of change in adjacent images due to motion in the video content can be large or small, which is referred to in this article as the video's motion signature. When the motion amplitude is large, the correlation between adjacent images in the temporal domain decreases, while the spatial correlation is higher. When the motion amplitude is small, or even when the scene is nearly static, the correlation between adjacent images in the temporal domain is even higher. This embodiment determines the video's motion signature by calculating the spatiotemporal correlation between images, and then determines the reference image allocation strategy. This ensures that reference images that are more relevant to the current image are more likely to be added to the reference image list.

[0162] Specifically, in this embodiment, when the total number of temporal reference images and spatial reference images of the current image exceeds the length of the reference image list, and the number of temporal reference images and spatial reference images is not 1, a corresponding number of temporal reference frames and spatial reference frames are adaptively allocated according to the current motion characteristics of the video, so that the reference images added to the reference image list are images with a high similarity to the current image, and the current block in the current image can have a greater probability of finding a reference block with a higher matching degree, which can improve the coding efficiency.

[0163] In an example of this embodiment, the temporal correlation between the temporal reference image and the current image is determined based on a first similarity between at least one temporal reference image of the current image and the current image, wherein a higher the first similarity, a greater the temporal correlation;

[0164] The spatial correlation between the spatial reference image and the current image is determined based on a second similarity between at least one spatial reference image of the current image and the current image. The higher the second similarity, the greater the spatial correlation.

[0165] For example, for P32 in Figure 15, the temporal correlation can be calculated based on the similarity between P7 and P32, and the spatial correlation can be calculated based on the similarity between P32 and P31. Alternatively, the spatial correlation can be obtained based on the average of the similarities between P32 and P31 and P25.

[0166] In an example of this embodiment, the similarity between the temporal reference image and the current image, and the similarity between the spatial reference image and the current image are represented by the following parameters:

[0167] Mean-square error (MSE)

[0168] Peak Signal to Noise Ratio (PSNR)

[0169] Cosine similarity;

[0170] Hash similarity;

[0171] Mutual information, which is calculated through image entropy.

[0172] In addition to the above parameters, parameters such as sum of absolute difference (SAD) and sum of absolute transformed difference (SATD) after Hadamard transform can also be used for characterization. The smaller the SAD and SATD are, the greater the similarity.

[0173] In an example of this embodiment, determining the number of temporal reference images and the number of spatial reference images added to the reference image list according to the temporal correlation and the spatial correlation includes:

[0174] In the case of p < q,

[0175] If a·p > q holds, determine m = 2, n = x - m;

[0176] If b·p > q holds, determine m = 1, n = x - m;

[0177] Where p is the value of the temporal correlation, q is the value of the spatial correlation, a is the first coefficient, b is the second coefficient, 1 < a < b, m is the number of temporal reference images added to the reference image list, n is the number of spatial reference images added to the reference image list, and x ≥ 3.

[0178] For example, assume p < q, and the reference frame list has a total of 4 frames (x = 4). Gradually judge the following conditions. When 1.5p > q holds: Determine that the number of temporal reference images and the number of spatial reference images added to the reference image list are both 2, that is, m = n = 2; and when 2p > q holds: Determine that the number of temporal reference images added to the reference image list is 1, that is, m = 1, and the number of spatial reference images added to the reference image list is 3, that is, n = 3 frames.

[0179] In an example of this embodiment, determining the number of temporal reference images and the number of spatial reference images added to the reference image list according to the temporal correlation and the spatial correlation includes:

[0180] In the case of q < p,

[0181] If a·q > p holds, determine n = 2, m = x - n;

[0182] If b·q > p holds, determine n = 1, m = x - n;

[0183] Where p is the value of temporal correlation, q is the value of spatial correlation, a is the first coefficient, b is the second coefficient, 1 < a < b, m is the number of temporal reference images added to the reference image list, n is the number of spatial reference images added to the reference image list, and x ≥ 3.

[0184] For example, assume q < p and the reference frame list has 4 frames (x = 4). Gradually judge the following conditions: when 1.5q > p holds, it is determined that the number of temporal reference images and spatial reference images added to the reference image list are both 2, i.e., m = n = 2; when 2p > q, it is determined that the number of spatial reference images added to the reference image list is 1, i.e., n = 1, and the number of temporal reference images added to the reference image list is 3, i.e., m = 3 frames.

[0185] In the above example, regardless of the relationship between p and q, at least one spatial reference image and one temporal reference image need to be added to the reference image list.

[0186] In an exemplary embodiment of the present disclosure, the adjacent views refer to the views in the 8-neighborhood of the current image in the current multi-view array;

[0187] When constructing a reference image list for inter-frame prediction of the current image based on the at least one temporal reference image and at least one spatial reference image, the views in the 4-neighborhood of the current image are preferentially added to the reference image list. When all the views in the available 4-neighborhood of the current image have been added to the reference image list and there is still a need to add spatial reference images of the current image, the views in the diagonal neighborhood of the current image are then added to the reference image list.

[0188] In this embodiment, the views in the 4-neighborhood are preferentially added to the reference image list. As already exemplified above, it will not be elaborated here.

[0189] In an exemplary embodiment of the present disclosure, there is one reference image list; or, there are two reference image lists, which are respectively used to add temporal reference images and spatial reference images;

[0190] The reference images in the reference image list belong to one GOP, or the reference images in the reference image list belong to multiple GOPs.

[0191] The embodiment of the present disclosure divides reference images into two categories: time domain reference images and spatial domain reference images. The time domain reference images are derived from the co-located views of the adjacent multi-view matrix in the time domain, and the spatial domain reference images are derived from the adjacent views in the multi-view matrix at the same moment. The total number of reference images is fixed, some of which are time domain reference images and some of which are spatial domain reference images. In one embodiment of the present disclosure, more reference images can be allocated to the party with stronger correlation based on the current spatiotemporal correlation. Suppose the total number of reference images is x, the number of time domain reference images is m, and the number of time domain reference images is n, then x=m+n. The time domain correlation is p, and the spatial domain correlation is q. According to the relationship between p and q, the sizes of m and n are allocated.

[0192] An embodiment of the present disclosure further proposes a video decoding method for a multi-view light field video, as shown in FIG16 , comprising:

[0193] Step S210: constructing a reference image list for the current image according to the method for constructing a reference image list for a multi-view light field video image according to any embodiment of the present disclosure;

[0194] Step S220 , decoding to obtain motion information of a current block in the current image, and performing inter-frame prediction on the current block according to the motion information and the reference image list.

[0195] The decoding method of the disclosed embodiment can make greater use of the spatiotemporal correlation of multi-view light field videos while taking into account both temporal and spatial correlations.

[0196] In this article, the current block can be the current coding unit (CU) or the current prediction unit (PU), or other coding units. The current image is the image where the current block is located.

[0197] In an exemplary embodiment of the present disclosure, the method further includes: setting a GOP size (GOPsize) based on the number of views in the multi-view array, wherein the larger the number of views in a multi-view array, the larger the GOP size. For example, when the size of the multi-view array is 5x5, the GOPsize (the number of images in a GOP) is smaller than the GOPsize when the size of the multi-view array is 7x7. This embodiment sets the GOP size based on the number of views in the multi-view array, such as making the GOPsize proportional to the multi-view scale.

[0198] By considering the scale of multiple views when designing the reference relationship, for example, each frame must reference 25 frames forward and 1 frame forward. In this case, GOPSize can be set to any value because the reference method is the same for each frame, and the period can be any value. This can also achieve good results.

[0199] An embodiment of the present disclosure further proposes a video encoding method for a multi-view light field video, as shown in FIG17 , including:

[0200] Step S310: constructing a reference image list for the current image according to the method for constructing a reference image list for a multi-view light field video image according to any embodiment of the present disclosure;

[0201] Step S320 , determining motion information of a current image in a current image according to the reference image list, performing inter-frame prediction on the current frame according to the motion information, and encoding the motion information.

[0202] The encoding method of the disclosed embodiment can utilize the spatiotemporal correlation of multi-view light field videos for encoding to a greater extent, thereby improving compression efficiency.

[0203] When the encoder performs inter-frame encoding on the current block, it can search for reference images from the reference image list according to the motion search method supported by the encoder, find the reference block with the greatest similarity to the current block from one of the reference images, and determine the motion information of the current block, which includes the reference image index and motion vector information. The reference image index is used to indicate the position of the selected reference image in the reference image list, and the motion vector information is used to indicate the position offset of the reference block relative to the current block. Because the embodiment of the present disclosure adds spatial reference images and temporal reference images to the reference image list, it fully utilizes the spatial correlation between adjacent views in the same view array in the multi-view light field video, as well as the temporal correlation between co-located views at similar times, which helps to find a reference block with a high degree of match with the current block, thereby improving compression efficiency.

[0204] The solution of the embodiment of the present disclosure was tested on two sequences, where Sequence 1 has strong spatial correlation and Sequence 2 has strong temporal correlation. The experimental results show that the solution can effectively improve compression performance.

[0205] The test platform is the official H.266 / VVC (Versatile Video Coding) test model VTM-17.0. Sequence 1 is from the LVC test sequence Nagoya train 2, the color multi-view version Tunnel_Train_2c_MV, from moments 0-3. Each moment contains 25 views, totaling 100 frames, with a single-view resolution of 920×880. Sequence 2 is from moments 25-28 of the LVC test sequence Nagoya Origami, from moments 25-28. Each moment contains 25 views, totaling 100 frames, with a single-view resolution of 912×880.

[0206] Anchor uses the official LD_P (lowdelay_P) profile of VTM-17.0, and adjusts the following parameters: Profile is set to main_10, Level is set to 4.1, and encoding bit depth is set to 8.

[0207] In the LD configuration, IntraPeriod is -1, meaning only the first frame in the sequence is an I-frame, and the rest are B or P-frames. In the LD_P configuration, the rest are all P-frames.

[0208] The embodiment of the present disclosure mainly makes modifications to the reference image list. The reference image list of each image contains two images, namely the first forward frame (spatial reference frame) and the 25th forward frame (temporal reference frame).

[0209] FIG18 (taking Sequence 1 as the test sequence) and FIG19 (taking Sequence 2 as the test sequence) are RD curves of the above two configurations under the two sequences.

[0210] The RD curve shows bitrate on the horizontal axis (lower values ​​are better), and PSNR on the vertical axis (higher values ​​are better). It represents the PSNR of the algorithm's compression results at a specific bitrate, with the curve moving toward the upper left indicating better performance. We can see that on Sequence 1, which has strong spatial correlation, our ts method still achieves a certain degree of performance improvement, even though temporal correlation is weaker. The RD curve is slightly above the anchor. On Sequence 2, which has stronger temporal correlation, the ts method shows significant performance improvement, with the RD curve clearly above the anchor.

[0211] An embodiment of the present disclosure further provides a device for constructing a reference image list for a multi-view light field video image, as shown in FIG20 . The device includes a processor 71 and a memory 73 storing a computer program. When the processor 71 executes the computer program, the method for constructing a reference image list for a multi-view light field video image as described in any embodiment of the present disclosure can be implemented.

[0212] An embodiment of the present disclosure further provides a video decoding device, comprising a processor and a memory storing a computer program, wherein the processor can implement the video decoding method as described in any embodiment of the present disclosure when executing the computer program.

[0213] An embodiment of the present disclosure further provides a video encoding device, including a processor and a memory storing a computer program, wherein the processor can implement the video encoding method as described in any embodiment of the present disclosure when executing the computer program.

[0214] The processor of the above-mentioned embodiment of the present disclosure may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a microprocessor, etc., or other conventional processors; the processor may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), discrete logic or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or other equivalent integrated or discrete logic circuits, or a combination of the above devices. That is, the processor of the above-mentioned embodiment may be any processing device or device combination that implements the various methods, steps and logic block diagrams disclosed in the embodiments of the present disclosure. If the embodiments of the present disclosure are partially implemented in software, the instructions for the software may be stored in a suitable non-volatile computer-readable storage medium, and one or more processors may be used to execute the instructions in hardware to implement the methods of the embodiments of the present disclosure. The term "processor" used herein may refer to the above-mentioned structure or any other structure suitable for implementing the technology described herein.

[0215] An embodiment of the present disclosure further provides a video encoding and decoding system, which includes the video encoding device described in any embodiment of the present disclosure and the video decoding device described in any embodiment of the present disclosure.

[0216] An embodiment of the present disclosure further provides a non-transitory computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, can implement the method described in any embodiment of the present disclosure.

[0217] An embodiment of the present disclosure further provides a computer program product, including a computer program, wherein when the computer program is executed by a processor, the method described in any embodiment of the present disclosure can be implemented.

[0218] In one or more exemplary embodiments above, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored as one or more instructions or codes on a computer-readable medium or transmitted via a computer-readable medium and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium that facilitates the transfer of a computer program from one place to another, such as according to a communication protocol. In this way, a computer-readable medium may generally correspond to a non-transitory tangible computer-readable storage medium or a communication medium such as a signal or carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures for implementing the technology described in this disclosure. A computer program product may include a computer-readable medium.

[0219] By way of example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Furthermore, any connection may also be referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwaves, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwaves are included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient (transient) media, but rather refer to non-transient tangible storage media. As used herein, disk and optical disk include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, or Blu-ray disc, among others, where disks typically reproduce data magnetically, while optical discs use lasers to reproduce data optically. Combinations of the above should also be included within the scope of computer-readable media.

[0220] In some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. Furthermore, the techniques may be fully implemented in one or more circuits or logic elements.

[0221] The technical solutions of the embodiments of the present disclosure can be implemented in a wide variety of devices or equipment, including wireless mobile phones, integrated circuits (ICs), or a group of ICs (e.g., a chipset). Various components, modules, or units are described in the embodiments of the present disclosure to emphasize the functional aspects of the devices configured to perform the described techniques, but they do not necessarily need to be implemented by different hardware units. Instead, as described above, the various units can be combined in the codec-side hardware unit or provided by a collection of interoperable hardware units (including one or more processors as described above) in combination with appropriate software and / or firmware.

Claims

1. A method for constructing a reference image list of multi-view light field video images, comprising: Determine that there is at least one temporal reference image and at least one spatial reference image for the current image; Based on at least one temporal reference image and at least one spatial reference image of the current image, construct a reference image list for the current image for inter-frame prediction; Wherein, the temporal reference image is the co-located view in the multi-view array adjacent to the current image in the time domain, the spatial reference image is the adjacent view of the current image in the current multi-view array, and both the temporal reference image and the spatial reference image are encoded images or decoded images.

2. The method according to claim 1, wherein: The method further comprises: When it is determined that the current image has only spatial reference images, construct the reference image list based on the determined at least one spatial reference image; When it is determined that the current image has only temporal reference images, construct the reference image list based on the determined at least one temporal reference image.

3. The method according to claim 1, wherein: The adjacent view is the view in the 4-neighborhood of the current image in the current multi-view array, or the adjacent view is the view in the 8-neighborhood of the current image in the current multi-view array.

4. The method according to claim 1, wherein: The co-located view includes one or more of the following: the co-located view of the current image in the multi-view array at the previous moment, the co-located view of the current image in the multi-view array at the next moment.

5. The method according to claim 1, wherein: The constructing a reference image list for the current image for inter-frame prediction based on at least one temporal reference image and at least one spatial reference image of the current image includes: When the current image has M temporal reference images and N spatial reference images, and M≥m, N≥n, add m temporal reference images and n spatial reference images of the current image to the reference image list; wherein, m, n are set values, and M, N, m, n are all greater than or equal to 1.

6. The method according to claim 5, wherein: m = n = 1; Or, m = 1, n = 2; Or, m = 1, n = 3; Or, m = 1, n = 4; Or, m = 1, n = 5; Or, m = n = 2; Or, m = 2, n = 3; Or, m = 2, n = 4; Or, m = 3, n = 3; Wherein, m + n = x, x≥2, and x is the maximum number of reference frames that can be added to the reference image list.

7. The method according to claim 5 or 6, wherein: The constructing a reference image list for the current image for inter-frame prediction based on at least one temporal reference image and at least one spatial reference image of the current image further includes: When the current image has M temporal reference images and N spatial reference images, and M < m, N≥n, add M temporal reference images and n spatial reference images of the current image to the reference image list; When there are M temporal reference images and N spatial reference images in the current image, and M≥m, N<n, add m temporal reference images and N spatial reference images of the current image to the reference image list; When there are M temporal reference images and N spatial reference images in the current image, and M<m, N<n, add M temporal reference images and N spatial reference images of the current image to the reference image list.

8. The method according to claim 1, wherein: Constructing a reference image list for inter-frame prediction of the current image based on at least one temporal reference image and at least one spatial reference image of the current image includes: When there are M temporal reference images and N spatial reference images in the current image, and M+N≤x, add M temporal reference images and N spatial reference images of the current image to the reference image list, where M, N≥1, x≥2, and x is the maximum number of reference frames that can be added to the reference image list.

9. The method according to claim 1, wherein: Constructing a reference image list for inter-frame prediction of the current image based on at least one temporal reference image and at least one spatial reference image of the current image includes: When there are M temporal reference images and N spatial reference images in the current image, and M+N>x, if M = 1, add 1 temporal reference image and x-1 spatial reference images to the reference image list; if N = 1, add x-1 temporal reference images and 1 spatial reference image to the reference image list; where M, N≥l, x≥2, and x is the maximum number of reference frames that can be added to the reference image list.

10. The method according to claim 1, wherein: Constructing a reference image list for inter-frame prediction of the current image based on at least one temporal reference image and at least one spatial reference image of the current image includes: When there are M temporal reference images and N spatial reference images in the current image, and M+N>x and both M and N are greater than 1, determine the temporal correlation between the temporal reference image and the current image, and the spatial correlation between the spatial reference image and the current image, and determine the number of temporal reference images and the number of spatial reference images added to the reference image list according to the temporal correlation and the spatial correlation; where x≥2, and x is the maximum number of reference frames that can be added to the reference image list.

11. The method according to claim 10, wherein: The temporal correlation between the temporal reference image and the current image is determined according to the first similarity between at least one temporal reference image of the current image and the current image, and the higher the first similarity, the greater the temporal correlation; The spatial correlation between the spatial reference image and the current image is determined according to the second similarity between at least one spatial reference image of the current image and the current image, and the higher the second similarity, the greater the spatial correlation.

12. The method according to claim 11, wherein: The similarities between the time-domain reference image and the current image, and between the spatial-domain reference image and the current image are represented by the following parameters: Mean squared error; Peak signal-to-noise ratio; Cosine similarity; Hash similarity; Mutual information, which is calculated through image entropy.

13. The method according to claim 10, wherein: Determining the number of time-domain reference images and the number of spatial-domain reference images to be added to the reference image list according to the time-domain correlation and the spatial-domain correlation includes: When p < q, If a·p > q holds, determine m = 2, n = x - m; If b·p > q holds, determine m = 1, n = x - m; Where, p is the value of the time-domain correlation, q is the value of the spatial-domain correlation, a is the first coefficient, b is the second coefficient, 1 < a < b, m is the number of time-domain reference images to be added to the reference image list, n is the number of spatial-domain reference images to be added to the reference image list, and x ≥ 3.

14. The method according to claim 10, wherein: Determining the number of time-domain reference images and the number of spatial-domain reference images to be added to the reference image list according to the time-domain correlation and the spatial-domain correlation includes: When q < p, If a·q > p holds, determine n = 2, m = x - n; If b·q > p holds, determine n = 1, m = x - n; Where, p is the value of the time-domain correlation, q is the value of the spatial-domain correlation, a is the first coefficient, b is the second coefficient, 1 < a < b, m is the number of time-domain reference images to be added to the reference image list, n is the number of spatial-domain reference images to be added to the reference image list, and x ≥ 3.

15. The method according to claim 1, wherein: The adjacent views refer to the views in the 8-neighborhood of the current image in the current multi-view array; When constructing the reference image list for inter-frame prediction of the current image based on the at least one time-domain reference image and the at least one spatial-domain reference image, the views in the 4-neighborhood of the current image are preferentially added to the reference image list. When all the views in the available 4-neighborhood of the current image have been added to the reference image list and there is still a need to add spatial-domain reference images for the current image, the views in the diagonal neighborhood of the current image are then added to the reference image list.

16. The method according to claim 1, wherein: There is one reference image list; or, there are two reference image lists, which are respectively used to add time-domain reference images and spatial-domain reference images; The reference images in the reference image list belong to one group of pictures (GOP), or the reference images in the reference image list belong to multiple GOPs.

17. A video decoding method for multi-view light field video, comprising: Constructing the reference image list of the current image according to the method described in any one of claims 1 to 16; Decoding to obtain the motion information of the current block in the current image, and performing inter-frame prediction on the current block according to the motion information and the reference image list.

18. The method according to claim 1, wherein: The method further includes setting a size of the GOP according to the number of views in the multi-view array, wherein the larger the number of views in the multi-view array, the larger the size of the GOP.

19. A video encoding method for a multi-view light field video, comprising: Constructing a reference picture list for the current picture according to the method of any one of claims 1 to 16; Motion information of a current image in a current image is determined according to the reference image list, inter-frame prediction is performed on the current frame according to the motion information, and the motion information is encoded.

20. A code stream, characterized in that The code stream is generated based on the multi-view light field video coding method according to claim 19.

21. A device for constructing a reference image list for a multi-view light field video image, comprising a processor and a memory storing a computer program, wherein: When the processor executes the computer program, it is capable of implementing the method for constructing a reference image list for multi-view light field video images according to any one of claims 1 to 16.

22. A video decoding device comprising a processor and a memory storing a computer program, wherein: When the processor executes the computer program, the video decoding method for the multi-view light field video according to claim 17 or 18 can be implemented.

23. A video encoding device comprising a processor and a memory storing a computer program, wherein: When the processor executes the computer program, it is capable of implementing the video encoding method for multi-view light field video according to claim 19.

24. A video encoding and decoding system, wherein: The method comprises the video encoding device according to claim 23 and the video decoding device according to claim 22.

25. A non-transitory computer-readable storage medium storing a computer program, wherein: When the computer program is executed by a processor, it can implement the method according to any one of claims 1 to 19.

26. A computer program product comprising a computer program, wherein When the computer program is executed by a processor, it can implement the method according to any one of claims 1 to 19.