Video coding method, video decoding method, computer-readable medium, and electronic device
Patent Information
- Application Number
- PCT/CN2026/085913
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-28
- Filing Date
- 2026-03-25
- Publication Date
- 2026-10-01
Smart Images

Figure CN2026085913_01102026_PF_FP_ABST
Abstract
Description
Video encoding and decoding methods, computer-readable media and electronic devices
[0001] This application claims priority to Chinese Patent Application No. 202510392104.X, filed on March 28, 2025, entitled "Video Decoding Method, Computer-Readable Medium and Electronic Device". Technical Field
[0002] This application relates to the field of video encoding and decoding technology, and more specifically, to a video encoding and decoding method, apparatus, computer-readable medium, and electronic device. Background Technology
[0003] In video encoding and decoding technology, a virtual reference frame (VRFrame) refers to an additional reference frame that is generated during the video encoding process using a specified algorithm or technique and does not directly exist in the original video sequence. Although the concept of VRFrame has been proposed in related technologies, current video coding standards lack support for VRFrame. Summary of the Invention
[0004] The embodiments of this application provide a video encoding and decoding method, apparatus, computer-readable medium, and electronic device, which can flexibly adjust the decoding strategy according to the actual encoding situation to achieve support for virtual reference frames, and can enhance the accuracy of inter-frame prediction through virtual reference frames, thereby improving the overall encoding and decoding performance.
[0005] This application provides a video decoding method, including: decoding a video stream to determine whether a current slice uses a virtual reference frame; if it is determined that the current slice uses a virtual reference frame, constructing a reference image list corresponding to the current slice based on the generated virtual reference frame; and decoding the current slice based on the constructed reference image list.
[0006] This application provides a video encoding method, including: determining whether a current stripe uses a virtual reference frame; if it is determined that the current stripe uses a virtual reference frame, constructing a reference image list corresponding to the current stripe based on the generated virtual reference frame; and encoding the current stripe based on the constructed reference image list.
[0007] This application provides a video encoding method, including: generating a video stream and storing the video stream, wherein generating the video stream includes: determining whether the current slice uses a virtual reference frame; if the current slice uses a virtual reference frame, constructing a reference image list for the current slice based on the generated virtual reference frame; and encoding the current slice based on the constructed reference image list.
[0008] This application provides a video decoding apparatus, including: a decoding unit configured to decode a video stream to determine whether a current slice uses a virtual reference frame; a generation unit configured to, if it is determined that the current slice uses a virtual reference frame, construct a reference image list corresponding to the current slice based on the generated virtual reference frame; and a processing unit configured to decode the current slice based on the constructed reference image list.
[0009] This application provides a video encoding apparatus, including: a determining unit configured to determine whether a current stripe uses a virtual reference frame; a generating unit configured to, if it is determined that the current stripe uses a virtual reference frame, construct a reference image list corresponding to the current stripe based on the generated virtual reference frame; and a processing unit configured to perform encoding processing on the current stripe based on the constructed reference image list.
[0010] This application provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the video decoding method or video encoding method as described in the above embodiments.
[0011] This application provides an electronic device, including: one or more processors; and a storage device for storing one or more computer programs, which, when executed by the one or more processors, cause the electronic device to implement the video decoding method or video encoding method as described in the above embodiments.
[0012] This application provides a computer program product comprising a computer program stored in a computer-readable storage medium. An electronic device's processor reads and executes the computer program from the computer-readable storage medium, causing the electronic device to perform the video decoding or video encoding methods provided in the various embodiments described above.
[0013] This application provides a method for storing video streams, wherein the video streams are decoded according to the video decoding method described in the above embodiments, or the video streams are generated according to the video encoding method described in the above embodiments.
[0014] In some embodiments of this application, the video stream can be decoded to determine whether a virtual reference frame is used in the current slice. If a virtual reference frame is used, a reference image list corresponding to the current slice is constructed based on the generated virtual reference frame, and then the current slice is decoded based on the constructed reference image list. Therefore, the technical solution of this application can dynamically detect whether a virtual reference frame is used for encoding in each slice during the decoding process of the video stream. This allows the decoder to flexibly adjust the decoding strategy according to the actual encoding situation, adapting to the application requirements of different encoding standards and technologies. Furthermore, when a virtual reference frame is used in the current slice, constructing a corresponding reference image list based on the generated virtual reference frame not only enriches the available reference frames but also improves the quality of the reference frames. This reduces prediction errors, enhances the accuracy of inter-frame prediction, and improves compression efficiency, overall encoding / decoding performance, and video quality.
[0015] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0016] Figure 1 illustrates a schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of this application can be applied;
[0017] Figure 2 shows a schematic diagram of the placement of the video encoding device and the video decoding device in the streaming system;
[0018] Figure 3 shows a basic flowchart of a video encoder;
[0019] Figure 4 shows a schematic diagram of a block partitioning structure in the HEVC standard;
[0020] Figure 5 shows a schematic diagram of a block partitioning structure in the AVS3 standard;
[0021] Figure 6 shows a schematic diagram of a block partitioning structure in the VVC standard;
[0022] Figure 7 shows a schematic diagram of the partitioning effect using multiple nested partitioning structures;
[0023] Figure 8 shows a schematic diagram of inter-frame prediction;
[0024] Figure 9 shows a flowchart of a video decoding method according to some embodiments of this application;
[0025] Figure 10 shows a flowchart of a video encoding method according to some embodiments of this application;
[0026] Figure 11 shows a block diagram of a video decoding apparatus according to some embodiments of this application;
[0027] Figure 12 shows a block diagram of a video encoding apparatus according to some embodiments of the present application;
[0028] Figure 13 shows a schematic diagram of the structure of a computer system suitable for implementing the computer device of the present application. Detailed Implementation
[0029] Exemplary embodiments will now be described in a more comprehensive manner with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to these examples; rather, these embodiments are provided so that this application will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art.
[0030] Furthermore, the features, structures, or characteristics described in this application can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to provide a full understanding of the embodiments of this application. However, those skilled in the art will recognize that when implementing the technical solutions of this application, not all the detailed features in the embodiments may be used, one or more specific details may be omitted, or other methods, elements, devices, steps, etc., may be employed.
[0031] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0032] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. For example, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0033] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0034] In this article, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0035] Figure 1 shows a schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of this application can be applied.
[0036] As shown in Figure 1, system architecture 100 includes multiple terminal devices that can communicate with each other via, for example, a network 150. For instance, system architecture 100 may include a first terminal device 110 and a second terminal device 120 interconnected via network 150. In the embodiment of Figure 1, the first terminal device 110 and the second terminal device 120 perform unidirectional data transmission.
[0037] For example, the first terminal device 110 may encode video data (e.g., a video image stream captured by the first terminal device 110) for transmission to the second terminal device 120 via network 150. The encoded video data is transmitted in the form of one or more encoded video streams. The second terminal device 120 may receive the encoded video data from network 150, decode the encoded video data to recover the video data, and display video images based on the recovered video data.
[0038] System architecture 100 may include a third terminal device 130 and a fourth terminal device 140 that perform bidirectional transmission of encoded video data, such as during a video conference. For bidirectional data transmission, each of the third terminal device 130 and the fourth terminal device 140 may encode video data (e.g., a video image stream captured by the terminal device) for transmission over network 150 to the other terminal device. Each of the third terminal device 130 and the fourth terminal device 140 may also receive encoded video data transmitted by the other terminal device, decode the encoded video data to recover the video data, and display the video images on an accessible display device based on the recovered video data.
[0039] In the embodiment shown in FIG1, the first terminal device 110, the second terminal device 120, the third terminal device 130 and the fourth terminal device 140 may be servers or terminals, but the principles disclosed in this application are not limited to these.
[0040] Servers can be standalone physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminals can be smartphones, tablets, laptops, desktop computers, smart speakers, smart voice interaction devices, smartwatches, smart home appliances, in-vehicle terminals, aircraft, etc., but are not limited to these.
[0041] The network 150 shown in Figure 1 represents any number of networks, including, for example, wired and / or wireless communication networks, that transmit encoded video data between the first terminal device 110, the second terminal device 120, the third terminal device 130, and the fourth terminal device 140. The communication network 150 may exchange data in circuit-switched and / or packet-switched channels. This network may include telecommunications networks, local area networks (LANs), wide area networks (WANs), and / or the Internet. For the purposes of this application, unless explained below, the architecture and topology of network 150 may be irrelevant to the operation of the disclosure herein.
[0042] Figure 2 illustrates the placement of the video encoding and decoding devices in a streaming environment. The subject matter disclosed in this application is equally applicable to other video-enabled applications, including, for example, video conferencing, digital television (TV), and storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0043] The streaming system may include an acquisition subsystem 213, which may include a video source 201 such as a digital camera, which creates an uncompressed video image stream 202. In an embodiment, the video image stream 202 includes samples captured by a digital camera. The video image stream 202 is depicted as a thick line to emphasize the high data volume of the video image stream compared to encoded video data 204 (or encoded video bitstream 204). The video image stream 202 may be processed by an electronic device 220, which includes a video encoding device 203 coupled to the video source 201. The video encoding device 203 may include hardware, software, or a combination of hardware and software to implement or enforce aspects of the disclosed subject matter as described in more detail below. Compared to the video image stream 202, the encoded video data 204 (or encoded video bitstream 204) is depicted as a thin line to emphasize the lower data volume of the encoded video data 204 (or encoded video bitstream 204), which may be stored on a streaming server 205 for future use. One or more streaming client subsystems, such as client subsystems 206 and 208 in FIG. 2, can access streaming server 205 to retrieve copies 207 and 209 of encoded video data 204. Client subsystem 206 may include, for example, a video decoding device 210 in electronic device 230. Video decoding device 210 decodes the incoming copy 207 of the encoded video data and produces an output video picture stream 211 that can be displayed on display 212 (e.g., a screen) or another presentation device. In some streaming systems, the encoded video data 204, video data 207, and video data 209 (e.g., video stream) may be encoded according to certain video encoding / compression standards.
[0044] It should be noted that electronic devices 220 and 230 may include other components not shown in the figures. For example, electronic device 220 may include a video decoding device, and electronic device 230 may also include a video encoding device.
[0045] In some embodiments of this application, taking High Efficiency Video Coding (HEVC) and Versatile Video Coding (VVC) from international video coding standards, as well as the Chinese national video coding standard AVS, as examples, after an input video frame image, the video frame image is divided into several non-overlapping processing units according to a block size. Each processing unit performs a similar compression operation. This processing unit is called a Coding Tree Unit (CTU) or Largest Coding Unit (LCU). The CTU can be further subdivided into more refined units to obtain one or more basic Coding Units (CUs). The CU is the most basic element in a coding process.
[0046] In other embodiments, this processing unit may also be called a tile, which is a rectangular area of a multimedia data frame that can be independently decoded and encoded. In the Alliance for Open Media Video 1 (AV1) standard, the tile can be further subdivided into one or more superblocks (SBs). The SB is the starting point for block partitioning and can be further divided into multiple subblocks. The superblocks are then further subdivided into one or more blocks. Each block is the most basic element in a coding process. For example, an SB can contain several blocks.
[0047] The above method of dividing video frame images can be called a block partition structure. The following introduces some concepts in the encoding process:
[0048] Predictive coding includes intra-frame prediction and inter-frame prediction. The original video signal is predicted from a selected reconstructed video signal to obtain a residual video signal. The encoder needs to decide which predictive coding mode to choose for the current coding unit (or coding block) and inform the decoder. Intra-frame prediction refers to the predicted signal coming from a region within the same image that has already been encoded and reconstructed; inter-frame prediction refers to the predicted signal coming from another encoded image (called a reference image) that is different from the current image.
[0049] Transform and Quantization: After the residual video signal undergoes transformation operations such as Discrete Fourier Transform (DFT) and Discrete Cosine Transform (DCT), the signal is transformed into the transform domain, and these are called transform coefficients. The transform coefficients are then subjected to lossy quantization, losing some information to make the quantized signal more suitable for compression. In some video coding standards, there may be more than one transform method to choose from. Therefore, the encoder needs to select one of the transform methods for the current coding unit (or coding block) and inform the decoder. The fineness of quantization is usually determined by the quantization parameter (QP). A larger QP value means that coefficients with a wider range of values will be quantized into the same output, which usually leads to greater distortion and a lower bit rate. Conversely, a smaller QP value means that coefficients with a smaller range of values will be quantized into the same output, which usually leads to less distortion and a higher bit rate.
[0050] Entropy coding, or statistical coding, involves statistically compressing the quantized transform-domain signal based on the frequency of each value, ultimately outputting a binary (0 or 1) compressed bitstream. Simultaneously, other information generated during encoding, such as the selected coding mode and motion vector data, also requires entropy coding to reduce the bit rate. Statistical coding is a lossless coding method that effectively reduces the bit rate required to represent the same signal. Common statistical coding methods include Variable Length Coding (VLC) and Context Adaptive Binary Arithmetic Coding (CABAC).
[0051] Context-Based Binary Arithmetic Coding (CABAC) primarily involves three steps: binarization, context modeling, and binary arithmetic coding. After binarizing the input syntax elements, the binary data can be encoded using either a regular coding mode or a bypass coding mode. The bypass coding mode eliminates the need to assign a specific probability model to each binary bit; the input binary bit bin value is directly encoded using a simple bypass encoder, thus accelerating the overall encoding and decoding speed. Generally, different syntax elements are not completely independent, and even identical syntax elements possess a certain degree of memory. Therefore, according to conditional entropy theory, using other encoded syntax elements for conditional coding can further improve coding performance compared to independent coding or memoryless coding. This encoded symbol information used as conditions is called the context. In the regular coding mode, the binary bits of the syntax elements sequentially enter the context modeler. The encoder assigns an appropriate probability model to each input binary bit based on the values of previously encoded syntax elements or binary bits; this process is called context modeling. The context model corresponding to a grammatical element can be located using the context index increment (ctxIdxInc) and the context index start (ctxIdxStart). After the bin value and the assigned probability model are fed into the binary arithmetic encoder for encoding, the context model needs to be updated based on the bin value, which is the adaptive process in encoding.
[0052] Loop Filtering: The transformed and quantized signal undergoes inverse quantization, inverse transform, and prediction compensation to obtain a reconstructed image. Due to the effects of quantization, the reconstructed image differs from the original image in some aspects, resulting in distortion. Therefore, filtering operations can be performed on the reconstructed image. Filters such as deblocking filters (DB), Sample Adaptive Offset (SAO), and Adaptive Loop Filters (ALF) can effectively reduce the distortion caused by quantization. Since these filtered reconstructed images will be used as a reference for subsequent coded images to predict future image signals, the above filtering operations are also called loop filtering, i.e., filtering operations within the coding loop.
[0053] Figure 3 shows a basic flowchart of a video encoder, illustrating the process using intra-frame prediction as an example. The original image signal s... k [x,y] and the predicted image signal Perform the difference operation to obtain the residual signal u. k [x,y]. Residual signal u k After transformation and quantization, [x,y] is transformed to obtain quantization coefficients. These coefficients are then used to obtain the encoded bitstream through entropy coding, and to obtain the reconstructed residual signal u′ through inverse quantization and inverse transform. k [x,y]. Predict the image signal. With the reconstructed residual signal u′ k [x,y] superimposed to generate image signals Image signal On one hand, the signal is input to the intra-frame mode decision module and the intra-frame prediction module for intra-frame prediction processing; on the other hand, the reconstructed image signal s′ is output through loop filtering. k [x,y], reconstruct the image signal s′ k [x,y] can be used as a reference image for the next frame for motion estimation and motion compensation prediction. Then, based on the motion compensation prediction result s′ r [x+m x ,y+m y ] and intra-frame prediction results Obtain the predicted image signal for the next frame. And continue repeating the above process until the coding is complete.
[0054] Based on the above encoding process, at the decoding end, for each coding unit (or coding block), after acquiring the compressed bitstream (i.e., bitstream), entropy decoding is performed to obtain various mode information and quantization coefficients. Then, the quantization coefficients undergo inverse quantization and inverse transform processing to obtain the residual signal. On the other hand, based on the known coding mode information, the prediction signal corresponding to that coding unit (or coding block) can be obtained. Then, the residual signal and the prediction signal are added together to obtain the reconstructed signal. The reconstructed signal then undergoes loop filtering and other operations to produce the final output signal.
[0055] Currently, mainstream video coding standards, such as the international video coding standards High Efficiency Video Coding (HEVC) and Versatile Video Coding (VVC), as well as the Chinese national video coding standard AVS3, all adopt a block-based hybrid coding framework. Specifically, the original video data is divided into a series of coding blocks, and video coding methods such as prediction, transform, and entropy coding are combined to achieve video data compression.
[0056] In a block-based hybrid coding framework, video images are divided into several non-overlapping processing units for video compression. This processing unit is called a Coding Tree Unit (CTU). A CTU can be further subdivided into one or more basic coding units, called Coding Units (CUs). Each CU is the most basic element in a coding process, and each CU can choose a different coding mode. Figure 4 shows a block partitioning structure diagram in the HEVC standard; a CTU can be partitioned downwards using a 4-tree approach.
[0057] The AVS3 standard adopts a basic block partitioning structure of Quad-Tree (QT) + Binary-Tree (BT) + Extended Quad-Tree (EQT). For example, the representation of the QT+BT+EQT basic block partitioning structure in the bitstream in AVS3 is shown in Figure 5. For a CU, the first step is to determine whether to use QT for partitioning. If QT is used, QT partitioning is performed directly; if not, the next step is to determine whether not to partition. If partitioning is required, it is necessary to determine whether to use EQT or BT. Furthermore, regardless of whether EQT or BT is used, it is necessary to determine whether to partition horizontally or vertically. Block partitioning is a recursive decision-making process starting from the LCU and proceeding from top to bottom. During the recursive process, the optimal partitioning method and encoding mode are determined through optimization at the encoding end.
[0058] The VVC standard employs multiple block partitioning structures, including QT, BT, and Triple Tree (TT), as shown in Figure 6. Figure 7 illustrates the division of the CTU into multiple CUs using QT and various nested partitioning structures, where bold block edges represent QT partitions, and the remaining edges represent partitions of other tree types. These multiple block partitioning structures form a content adaptive coding tree structure composed of CUs.
[0059] In the field of coding technology, inter-frame prediction is a commonly used predictive coding technique. As shown in Figure 8, inter-frame prediction utilizes the correlation in the temporal domain of video, using pixels from neighboring encoded images to predict pixels in the current image, thereby effectively removing temporal redundancy and saving bits of coding residual data. Here, P represents the current frame, Pr represents the reference frame, B represents the current coding block, and Br represents the reference block of B. The coordinates of B' in the reference frame are the same as the coordinates of B in the current frame, and the coordinates of Br are (x... r ,y rThe coordinates of B' are (x, y). The displacement between the current coded block and its reference block is called the motion vector (MV), where MV = (x, y). r -x,y r -y). In other words, inter-frame prediction refers to the process of searching for a reference block in neighboring encoded images (i.e., reference frames) based on the current block to be encoded in the current frame, with the aim of removing temporal redundancy in the video signal.
[0060] In the VVC standard, for a unidirectional predictive slice (P slice), a reference block can be determined from one reference frame, allowing the inter-frame prediction value of the current block to be derived. For a bidirectional predictive slice (B slice), inter-frame prediction values can be derived from two reference frames. For each coding unit of inter-frame prediction, the motion parameters of the inter-frame prediction include motion vectors, reference picture indices, and reference picture list indices. Motion parameters can be indicated explicitly or implicitly, including using Skip mode, Merge mode, or Advanced Motion Vector Prediction (AMVP) mode. When encoding a CU using Skip mode, the CU does not need to transmit residual coefficients, motion vector difference (MVD), or reference picture indices. When using Merge mode, one approach is to obtain motion parameters from neighboring CUs (including spatial and temporal adjacency); another approach is to explicitly transmit motion parameters, which requires transmitting the motion information needed for each CU, including motion vectors, the corresponding reference image index for each reference image list, the reference image list usage flag, and other necessary information.
[0061] In the VVC standard, each slice requires a corresponding Reference Picture List (RPL). Each slice includes two reference lists, RPL 0 and RPL 1. A P slice uses only RPL 0, a B slice can use both RPL 0 and RPL 1, and an intra slice (I slice) does not require an RPL (e.g., the RPL list is empty). In addition, the VVC standard includes a bitstream conformance check on the RPLs.
[0062] The following describes the syntax elements related to RPL construction included in the VVC standard in the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), and Slice Header.
[0063] Table 1 shows the syntax elements related to RPL construction in SPS:
[0064] Table 1
[0065] In Table 1, sps_long_term_ref_pics_flag is a 1-bit unsigned integer used to indicate whether the current sequence can use a Long-Term Reference Picture (LTRP), where a value of 0 indicates that it is not available and a value of 1 indicates that it is allowed.
[0066] sps_inter_layer_prediction_enabled_flag is a 1-bit unsigned integer used to indicate whether inter-layer prediction can be used for the current sequence. A value of 0 indicates that it is not available, and a value of 1 indicates that it is allowed.
[0067] `sps_idr_rpl_present_flag` is a 1-bit unsigned integer used to indicate whether a slice of Network Abstraction Layer (NAL) cell equal to IDR_N_LP or IDR_W_RADL contains RPL syntax elements. A value of 1 indicates that it contains RPL, and a value of 0 indicates that it does not.
[0068] The value of sps_rpl1_same_as_rpl0_flag indicates whether RPL 1 is the same as RPL 0. If it is equal to 1, it means that the configuration of RPL 1 is exactly the same as that of RPL 0, and no separate syntax elements for RPL 1 need to be defined. If it is equal to 0, it means that RPL 1 is not the same as RPL 0.
[0069] sps_num_ref_pic_lists[i] is an unsigned exponential Golomb code that indicates the number of ref_pic_list_struct(listIdx,rplsIdx) in SPS where listIdx is i (i is 0 or 1), that is, the number of reference images in RPL 0 or RPL1. The value of sps_num_ref_pic_lists[i] is in the range [0, 64].
[0070] Table 2 shows the syntax elements related to RPL construction in PPS:
[0071] Table 2
[0072] In Table 2, the value of pps_num_ref_idx_default_active_minus1[i] plus 1 is used to determine the default number of active reference frames in each reference image list, where i is 0 or 1. When i equals 0, the default value of NumRefIdxActive[0] for P slice or B slice with sh_num_ref_idx_active_override_flag equal to 0 is specified; when i equals 1, the default value of NumRefIdxActive[1] for B slice with sh_num_ref_idx_active_override_flag equal to 0 is specified. The value of pps_num_ref_idx_default_active_minus1[i] should be in the range [0, 14].
[0073] The value of pps_rpl1_idx_present_flag indicates whether rpl_sps_flag[1] and rpl_idx[1] exist in the image header syntax structure or the slice header referencing the PPS image. A value of 0 indicates that these syntaxes are not allowed, and a value of 1 indicates that they are allowed.
[0074] A value of 1 for pps_rpl_info_in_ph_flag indicates that the RPL information exists in the Picture Header (PH) syntax structure, but not in the slice header of a PPS that does not contain a PH syntax structure. A value of 0 for pps_rpl_info_in_ph_flag indicates that the RPL information does not exist in the PH syntax structure, but may exist in the slice header of a PPS that references it.
[0075] Table 3 shows the syntax elements related to RPL construction in the Slice header:
[0076] Table 3
[0077] In Table 3, if pps_rpl_info_in_ph_flag is 0 (indicating that the RPL information is not in the image header but may be in the slice header), or if the current NAL unit type is not IDR_W_RADL or IDR_N_LP (e.g., a normal P frame or B frame), or if sps_idr_rpl_present_flag is 1 (even if the current frame is an IDR frame, the sequence parameter set indicates that these frames should contain RPL information), then the condition is met, and the ref_pic_lists() function is called to construct or update the reference image list.
[0078] Next, if the current slice is not an I slice (e.g., sh_slice_type != I) and there is more than one reference frame in RPL 0 (e.g., num_ref_entries[0][RplsIdx[0]]>1), or the current slice is a B slice (e.g., sh_slice_type == B) and there is more than one reference frame in RPL 1 (e.g., num_ref_entries[1][RplsIdx[1]]>1), then proceed to the next step of processing.
[0079] Parse sh_num_ref_idx_active_override_flag. If this flag is set to 1, it means that the default number of active reference images will be overridden, and further parsing of sh_num_ref_idx_active_minus1[i] is required.
[0080] Based on the state of sh_num_ref_idx_active_override_flag, iterate through all relevant reference image lists: for non-ISlice, if it is a B Slice, consider both RPL0 and RPL1 (e.g., the loop variable i ranges from 0 to 1); if it is a PS Slice, consider only RPL0 (e.g., the loop variable i ranges only 0). For each reference list i, if it contains more than 1 reference image (i.e., the syntax element num_ref_entries[i][RplsIdx[i]]>1), then sh_num_ref_idx_active_minus1[i] is parsed.
[0081] sh_num_ref_idx_active_minus1[i] is used to deduce the number of active reference images NumRefIdxActive[i] actually used in the reference image list i. The value of sh_num_ref_idx_active_minus1[i] should be in the range [0, 14].
[0082] If i equals 0 or 1, the current slice is B, sh_num_ref_idx_active_override_flag equals 1, and sh_num_ref_idx_active_minus1[i] does not exist, then it is inferred that sh_num_ref_idx_active_minus1[i] equals 0.
[0083] If the current slice is P, sh_num_ref_idx_active_override_flag is equal to 1, and sh_num_ref_idx_active_minus1[0] does not exist, then it can be inferred that sh_num_ref_idx_active_minus1[0] is equal to 0.
[0084] The value of the NumRefIdxActive[i] variable is obtained according to the method shown in Table 4:
[0085] Table 4
[0086] In Table 4, if the current slice is a B slice, then two reference image lists (RPL 0 and RPL 1) need to be processed, for example, i=0 and i=1. If the current slice is a P slice, then only RPL 0 (i=0) needs to be processed, while the NumRefIdxActive[1] of RPL 1 (i=1) will be set to 0. For I slices (intra-coded), the NumRefIdxActive[i] of both reference image lists will be set to 0, indicating that the reference image lists are not used to decode the slice.
[0087] The `sh_num_ref_idx_active_override_flag` flag indicates whether the default number of active reference frames needs to be overridden. If the value is 1, it means that the number of active reference frames specific to the current slice needs to be transmitted; if the value is 0, the default value is used.
[0088] When sh_num_ref_idx_active_override_flag is 1, the value of sh_num_ref_idx_active_minus1[i] plus 1 is used directly as NumRefIdxActive[i].
[0089] When sh_num_ref_idx_active_override_flag is 0, check the relationship between the number of available reference images in the current reference image list i and the default maximum number of active reference images. If the number of available reference images in the current reference image list is greater than or equal to the default maximum number of active reference images, then use the default maximum value. Otherwise, use the number of available reference images in the current reference image list.
[0090] ref_pic_lists can exist in the PH syntax structure or the slice header. Their main function is to build and configure the list of reference images for the current slice or image, such as RPL0 and RPL1.
[0091] Furthermore, the construction of the reference image list must satisfy the bitstream consistency check. For example, for RPL I, num_ref_entries[i][RplsIdx[i]] must be greater than or equal to NumRefIdxActive[i]. The reference image must exist in the Decoded Picture Buffer (DPB), and its Temporal ID (TID) must be less than or equal to the TID of the current image. The reference image cannot be the current image, and ph_non_ref_pic_flag (used to indicate whether it is a non-reference image; if set to 1, it indicates a non-reference image) must be 0. Short-Term Reference Picture (STRP) and Long-Term Reference Picture (LTRP) of the same image cannot reference the same image. LTRP entries are not allowed, and the difference between the Picture Order Count (POC) value of the current image and the POC value of the image referenced by the entry must be greater than or equal to 224. The number of reference images in the same layer as the current image must be less than or equal to the maximum DPB value - 1, and the number of reference images in the same layer as the current image for all slices must be consistent.
[0092] After decoding the Slice Header of each image and completing the RPL construction, a reference image marking process is required. This process dynamically updates the marking status of reference images in the Decoded Image Buffer (DPB) based on the current image's RPL 0 and RPL 1, including: Unused for Reference, STRP, and LTRP. For Clean Layered Video Stream Segments (CLVSS) images, such as Instantaneous Decoder Refresh (IDR) frames and Clean Random Access (CRA) frames, all reference images in the DPB that are at the same layer as the current image (with the same nuh_layer_id) are marked as "Unused for Reference". For non-CLVSS images, the LTRP entries in the reference list are first traversed. If a reference image is currently marked as STRP and is at the same layer as the current image, it is remarked as LTRP. Then, a cleanup operation for unreferenced reference images is performed. All reference images in the DPB at the same layer as the current image are marked as "Not Used for Reference" if they are not referenced by any entry in the reference list (RPL 0 or RPL 1). Images referenced by all inter-layer reference image entries are marked as LTRP. It is important to note that the same reference image can only be marked as STRP, LTRP, or Not Used for Reference at any given time.
[0093] In the exploration of next-generation video coding standards, Neural Network Based Video Coding (NNVC) has been extensively studied. One proposed technique involves inputting adjacent reference frames into a neural network to derive high-quality virtual reference frames, which are then inserted into a list of reference images. This improves inter-frame prediction accuracy and compression performance. However, current video coding standards lack support for virtual reference frames, leading to encoding and decoding errors in these inserted frames.
[0094] To address the aforementioned technical issues, this application provides a novel video encoding / decoding technology that dynamically detects whether a virtual reference frame is used for encoding in the current slice as decoding progresses. This allows the decoder to flexibly adjust its decoding strategy based on the actual encoding situation, adapting to the application requirements of different encoding standards and technologies. Furthermore, when a virtual reference frame is used in the current slice, constructing a corresponding list of reference images based on the generated virtual reference frames not only enriches the available reference frames but also improves their quality. This reduces prediction errors, enhances the accuracy of inter-frame prediction, and improves compression efficiency, overall encoding / decoding performance, and video quality.
[0095] The implementation details of the technical solutions in the embodiments of this application are described in detail below:
[0096] Figure 9 shows a flowchart of a video decoding method according to some embodiments of this application. This video decoding method can be executed by a device with computing processing capabilities, such as a terminal device or a server. Referring to Figure 9, the video decoding method includes at least steps S910 to S930, which are described in detail below:
[0097] In S910, the video stream is decoded to determine whether the current stripe uses a virtual reference frame.
[0098] In some embodiments, the video contains a sequence of video image frames. The video image frame sequence includes a series of images, each of which can be further divided into slices. The current slice refers to the slice that needs to be decoded at the current time.
[0099] In some embodiments, whether the current slice uses a virtual reference frame can be determined based on one or more of the following flag bits: a first flag bit in the sequence header, a second flag bit in the image header, and a third flag bit in the slice header. The first flag bit indicates whether the current sequence containing the current slice uses a virtual reference frame. If the first flag bit is 1, it indicates that the current sequence uses a virtual reference frame; if the first flag bit is 0, it indicates that the current sequence does not use a virtual reference frame. The second flag bit indicates whether the current image containing the current slice uses a virtual reference frame. If the second flag bit is 1, it indicates that the current image uses a virtual reference frame; if the second flag bit is 0, it indicates that the current image does not use a virtual reference frame. The third flag bit indicates whether the current slice uses a virtual reference frame. If the third flag bit is 1, it indicates that the current slice uses a virtual reference frame; if the third flag bit is 0, it indicates that the current slice does not use a virtual reference frame.
[0100] For example, whether the current slice uses a virtual reference frame can be determined by combining the first flag in the sequence header, the second flag in the image header, and the third flag in the slice header. If the first flag in the sequence header indicates that the current sequence uses a virtual reference frame, the second flag in the image header indicates that the current image uses a virtual reference frame, and the third flag in the slice header indicates that the current slice uses a virtual reference frame, then it can be determined that the current slice uses a virtual reference frame.
[0101] For example, whether the current slice uses a virtual reference frame can be determined based on the first flag bit in the sequence header and the second flag bit in the image header. If the first flag bit in the sequence header indicates that the current sequence uses a virtual reference frame, and the second flag bit in the image header indicates that the current image uses a virtual reference frame, then it can be determined that the current slice uses a virtual reference frame.
[0102] For example, the first flag in the sequence header can be used to determine whether the current stripe uses a virtual reference frame. If the first flag in the sequence header indicates that the current sequence uses a virtual reference frame, then it can be determined that the current stripe uses a virtual reference frame.
[0103] In some embodiments, if it is determined at least based on a first flag bit whether the current stripe uses a virtual reference frame, then the first flag bit is decoded again if at least one of the following conditions is met:
[0104] 1.1 The first parameter in the sequence parameter set indicates that the current sequence does not use long-term reference frames (e.g., the value of sps_long_term_ref_pics_flag in SPS is 0).
[0105] 1.2 The second parameter in the sequence parameter set indicates that inter-layer prediction is not used for the current sequence (e.g., the value of sps_inter_layer_prediction_enabled_flag in SPS is 0).
[0106] 1.3 The third parameter in the sequence parameter set indicates that the two lists of reference images used (e.g., RPL 0 and RPL 1) are not the same (e.g., the value of sps_rpl1_same_as_rpl0_flag in SPS is 0).
[0107] 1.4 The fourth parameter in the sequence parameter set indicates whether the number of reference images contained in the two reference image lists used (such as RPL 0 and RPL 1) is greater than or equal to a quantity threshold, or whether the number of reference images contained in one reference image list used (such as RPL 0 or RPL 1) is greater than or equal to a quantity threshold. The quantity threshold can be a positive integer such as 1 or 2. This quantity threshold can be a first quantity threshold. The fourth parameter is, for example, sps_num_ref_pic_lists[i].
[0108] 1.5. A reference image with a target image sequence count (POC) exists in the reference image list corresponding to the current sequence, and the target POC is in the set POC list. For example, a reference image with a target POC exists in RPL 0 corresponding to the current sequence, and the target POC is in the set POC list; or a reference image with a target POC exists in RPL 1 corresponding to the current sequence, and the target POC is in the set POC list; or both RPL 0 and RPL 1 corresponding to the current sequence contain reference images with target POCs, and the target POC is in the set POC list. This target POC can be the first target POC.
[0109] 1.6 The time level identifier (TID) of the reference image list corresponding to the current sequence is greater than or equal to a set threshold. For example, the TID of RPL 0 corresponding to the current sequence is greater than or equal to the set threshold; or the TID of RPL 1 corresponding to the current sequence is greater than or equal to the set threshold; or the TIDs of both RPL 0 and RPL 1 corresponding to the current sequence are greater than or equal to the set threshold. For example, the set threshold can be 2, 3, 4, etc. The set threshold can be a first set threshold.
[0110] In some embodiments, if it is determined at least based on the second flag bit whether the current stripe uses a virtual reference frame, then the second flag bit is decoded again if at least one of the following conditions is met:
[0111] 2.1 The fifth parameter in the sequence parameter set indicates that the current sequence uses a virtual reference frame (as the value of the first flag bit mentioned above indicates that the current sequence uses a virtual reference frame).
[0112] 2.2 The frame type of the current image is the specified type (e.g., the specified type is B frame).
[0113] 2.3 The sixth parameter in the image parameter set indicates whether the number of reference images contained in the two reference image lists used (such as RPL 0 and RPL 1) is greater than or equal to a quantity threshold, or whether the number of reference images contained in one reference image list used (such as RPL 0 or RPL 1) is greater than or equal to a quantity threshold. The quantity threshold can be a positive integer such as 1 or 2. This quantity threshold can be a second quantity threshold. The second quantity threshold can be the same as or different from the first quantity threshold.
[0114] 2.4. A reference image with a target POC exists in the reference image list corresponding to the current image, and the target POC is in the set POC list. For example, a reference image with a target POC exists in RPL 0 corresponding to the current image, and the target POC is in the set POC list; or a reference image with a target POC exists in RPL 1 corresponding to the current image, and the target POC is in the set POC list; or both RPL 0 and RPL 1 corresponding to the current image contain reference images with target POCs, and the target POC is in the set POC list. This target POC can be a second target POC. The second target POC can be the same as or different from the first target POC.
[0115] 2.5 The TID of the reference image list corresponding to the current image is greater than or equal to a set threshold. For example, the TID of RPL 0 corresponding to the current image is greater than or equal to the set threshold; or the TID of RPL 1 corresponding to the current image is greater than or equal to the set threshold; or the TIDs of both RPL 0 and RPL 1 corresponding to the current image are greater than or equal to the set threshold. The set threshold can be 2, 3, 4, etc. The set threshold can be a second set threshold, which may be the same as or different from the first set threshold.
[0116] In some embodiments, when one or more of the conditions in 2.1 to 2.5 above are met, it can be implicitly determined directly that the current image needs to use a virtual reference frame without decoding the second flag bit (in which case the bitstream may not contain the second flag bit).
[0117] In some embodiments, if it is determined at least based on a third flag bit whether the current stripe uses a virtual reference frame, then the third flag bit is decoded again if at least one of the following conditions is met:
[0118] 3.1 The fifth parameter in the sequence parameter set indicates that the current sequence uses a virtual reference frame (as the value of the first flag bit above indicates that the current sequence uses a virtual reference frame).
[0119] 3.2 The seventh parameter in the image parameter set indicates that the current image uses a virtual reference frame (as the value of the second flag bit mentioned above indicates that the current sequence uses a virtual reference frame).
[0120] 3.3 The current image frame type is the specified type (e.g., the specified type is B frame).
[0121] 3.4 The eighth parameter in the strip parameter set indicates whether the number of reference images contained in the two reference image lists used (such as RPL 0 and RPL 1) is greater than or equal to a quantity threshold, or whether the number of reference images contained in one reference image list used (such as RPL 0 or RPL 1) is greater than or equal to a quantity threshold. The quantity threshold can be a positive integer such as 1 or 2. This quantity threshold can be a third quantity threshold. The third quantity threshold can be the same as or different from the first and second quantity thresholds.
[0122] 3.5. A reference image with a target POC exists in the reference image list corresponding to the current strip, and the target POC is in the set POC list. For example, a reference image with a target POC exists in RPL 0 corresponding to the current strip, and the target POC is in the set POC list; or a reference image with a target POC exists in RPL 1 corresponding to the current strip, and the target POC is in the set POC list; or both RPL 0 and RPL 1 corresponding to the current strip contain reference images with target POCs, and the target POC is in the set POC list. This target POC can be a third target POC. The third target POC can be the same as or different from the first and second target POCs.
[0123] 3.6 The TID of the reference image list corresponding to the current strip is greater than or equal to a set threshold. For example, the TID of RPL 0 corresponding to the current strip is greater than or equal to the set threshold; or the TID of RPL 1 corresponding to the current strip is greater than or equal to the set threshold; or the TIDs of both RPL 0 and RPL 1 corresponding to the current strip are greater than or equal to the set threshold. The set threshold can be 2, 3, 4, etc. The set threshold can be a third set threshold, which may be the same as or different from the first and second set thresholds.
[0124] In some embodiments, when one or more of the conditions in 3.1 to 3.6 above are met, it can be implicitly determined directly that the current stripe needs to use a virtual reference frame without decoding the third flag bit (in which case the bitstream may not contain the third flag bit).
[0125] Referring to Figure 9, in S920, if it is determined that the current strip uses a virtual reference frame, a list of reference images corresponding to the current strip is constructed based on the generated virtual reference frame.
[0126] In some embodiments, constructing a reference image list corresponding to the current strip based on the generated virtual reference frame can be achieved by inserting the generated virtual reference frame into a predetermined position in the reference image list corresponding to the current strip.
[0127] For example, the generated first virtual reference frame can be inserted into a first predetermined position in the first reference image list (e.g., RPL 0) corresponding to the current slice; or the generated second virtual reference frame can be inserted into a second predetermined position in the second reference image list (e.g., RPL 1) corresponding to the current slice; or the generated first virtual reference frame can be inserted into the first predetermined position in the first reference image list corresponding to the current slice, and the generated second virtual reference frame can be inserted into the second predetermined position in the second reference image list corresponding to the current slice. The first virtual reference frame and the second virtual reference frame may be the same or different, and the first predetermined position and the second predetermined position may be the same or different. In other words, the generated virtual reference frame can be inserted into both reference image lists corresponding to the current slice, or it can be inserted into only one of the reference image lists.
[0128] In some embodiments, if a virtual reference frame is inserted into the reference image list corresponding to the current slice, the length of the reference image list needs to be adjusted. For example, the length of the reference image list corresponding to the current slice can be directly increased by a first preset value, which is the number of virtual reference frames to be inserted. For instance, the length of a preset reference image list corresponding to the current slice can be increased by the first preset value. For example, if one virtual reference frame needs to be inserted, the length of the reference image list corresponding to the current slice can be increased by 1. When directly increasing the length of the reference image list corresponding to the current slice by the first preset value, the lengths of both reference image lists (such as RPL 0 and RPL 1) corresponding to the current slice can be increased by the first preset value, or only the length of one of the reference image lists (such as RPL 0 or RPL 1) can be increased by the first preset value. In these embodiments, the insertion of virtual reference frames is supported by adjusting the length of the reference image list. In other embodiments of this application, the insertion of virtual reference frames can also be supported by adjusting the decoding process of the bitstream, as detailed below.
[0129] In some embodiments, if the current slice is not an intra-frame slice (e.g., it could be a B slice or a P slice) and the number of available reference images in the first reference image list (e.g., RPL 0) corresponding to the current slice is greater than 1, or if the current slice is a bidirectional prediction slice (e.g., a B slice) and the sum of the number of available reference images in the second reference image list (e.g., RPL 1) corresponding to the current slice and a first set value is greater than 1, then the fourth flag bit (e.g., the fourth flag bit could be sh_num_ref_idx_active_override_flag) is decoded. The value of the fourth flag bit is used to indicate whether the default number of active reference images should be overridden; for example, a value of 1 indicates that the default number of active reference images needs to be overridden; a value of 0 indicates that the default number of active reference images does not need to be overridden.
[0130] If the value of the fourth flag indicates the number of active reference images that override the default, then the number of active reference images actually used in the first reference image list is updated if the current slice is a unidirectional prediction slice (e.g., P slice) and the sum of the number of reference images available in the first reference image list (e.g., RPL 0) and the first set value is greater than 1. For example, the number of active reference images actually used in the first reference image list NumRefIdxActive[0] is updated by decoding sh_num_ref_idx_active_minus1[0].
[0131] If the value of the fourth flag indicates the number of active reference images that overrides the default, then when the current slice is a bidirectional prediction slice (e.g., B slice), the following process is performed: if the sum of the number of available reference images in the first reference image list (e.g., RPL 0) and the first set value is greater than 1, then the number of active reference images actually used in the first reference image list is updated, such as by decoding sh_num_ref_idx_active_minus1[0] to update the number of active reference images actually used in the first reference image list NumRefIdxActive[0]; if the sum of the number of available reference images in the second reference image list (e.g., RPL 1) and the first set value is greater than 1, then the number of active reference images actually used in the second reference image list is updated, such as by decoding sh_num_ref_idx_active_minus1[1] to update the number of active reference images actually used in the second reference image list NumRefIdxActive[1].
[0132] If the value of the fourth flag indicates the number of active reference frames to cover by the default, then for the P slice, RPL 0 needs to be processed; for the B slice, both RPL 0 and RPL 1 need to be processed. If the current slice does not use virtual reference frames, the above first setting value is 0; if the current slice uses virtual reference frames, the above first setting value is the number of virtual reference frames used. For example, if one virtual reference frame is used, then the first setting value can be 1.
[0133] In some embodiments, the actual number of active reference images used in the reference image list can be updated in the following manner: if the number of available reference images in reference image list i (reference image list i represents at least one of the first reference image list RPL 0 and the second reference image list RPL 1) is greater than or equal to the sum of the default length of reference image list i and 1 and a first set value, then the actual number of active reference images used in reference image list i is set to the sum of the default length of reference image list i and 1 and the first set value; if the number of available reference images in reference image list i is less than the sum of the default length of reference image list i and 1 and the first set value, then the actual number of active reference images used in reference image list i is set to the sum of the number of available reference images in reference image list i and the first set value.
[0134] In some embodiments, the actual number of active reference images used in the reference image list can also be updated as follows: if the number of available reference images in reference image list i (reference image list i represents at least one of the first reference image list RPL 0 and the second reference image list RPL 1) is greater than or equal to the sum of the default length of reference image list i and 1, then the actual number of active reference images used in reference image list i is set to the sum of the default length of reference image list i and 1; if the number of available reference images in reference image list i is less than the sum of the default length of reference image list i and 1, then the actual number of active reference images used in reference image list i is set to the sum of the number of available reference images in reference image list i and a first preset value. In these embodiments, the default length of reference image list i is increased by a first preset value based on the original length. Wherein, if the current strip does not use virtual reference frames, the above-mentioned first preset value is 0; if the current strip uses virtual reference frames, the above-mentioned first preset value is the number of virtual reference frames used. For example, if one virtual reference frame is used, then the first preset value can be 1.
[0135] In some embodiments, besides adjusting the length of the reference image list, the length of the reference image list can also be maintained by deleting reference images from the list. For example, a reference image at a set position in the reference image list corresponding to the current strip can be deleted, and then a generated virtual reference frame can be inserted into the set position in the reference image list corresponding to the current strip. In other words, the virtual reference frame can replace the reference image at the set position in the reference image list. This ensures that the length of the reference image list remains unchanged, thus eliminating the need to adjust the length of the reference image list. The reference image list can be one or both of the two reference image lists corresponding to the current strip.
[0136] In some embodiments, after inserting the generated virtual reference frame into a predetermined position in the reference image list corresponding to the current slice, a specified reference image can be removed from the reference image list corresponding to the current slice to maintain the same length. This ensures that the length of the reference image list remains unchanged, eliminating the need to adjust its length. The reference image list can be one or both of the two reference image lists corresponding to the current slice.
[0137] In some embodiments, after inserting the generated virtual reference frame into a predetermined position in the reference image list corresponding to the current strip, if the number of reference images contained in the updated reference image list exceeds a reference image number threshold (exemplarily, the reference image number threshold may be less than or equal to the length of the reference image list), then a specified reference image can be removed from the updated reference image list. For example, the reference image list may be one or both of the two reference image lists corresponding to the current strip.
[0138] For example, the specified reference image to be removed may include one or more of the following reference images: reference images of a specified frame type (such as I-frames) in the reference image list corresponding to the current slice, reference images in the reference image list corresponding to the current slice whose time level identifier is less than or equal to a first threshold, and reference images listed later in the reference image list corresponding to the current slice. For example, when removing reference images listed later in the reference image list corresponding to the current slice, the removal can be performed in reverse order from back to front, depending on the number of reference images to be removed. For instance, if only one reference image needs to be removed, the last reference image in the reference image list can be removed.
[0139] In some embodiments, after constructing a list of reference images corresponding to the current strip based on the generated virtual reference frame, the virtual reference frame can be set as a short-term reference frame.
[0140] In S930, the current strip is decoded based on the constructed list of reference images.
[0141] The process of decoding the current strip based on the constructed reference image list can be to select a reference image from the reference image list and then perform decoding through inter-frame prediction. For details, please refer to the relevant content in the foregoing embodiments.
[0142] In some embodiments, the reference image list corresponding to the current strip may not be constructed based on the generated virtual reference frame, i.e., the scheme of inserting virtual reference frames may not be adopted, when one or more of the following conditions are met: the number of reference images in the reference image list corresponding to the current strip (such as one or more of RPL 0 and RPL1) is less than a second threshold, and the number of reference images in the reference image list corresponding to the current strip (such as one or more of RPL 0 and RPL1) is greater than a third threshold.
[0143] In the embodiments of this application, virtual reference frames are additional reference frames generated during the video encoding process using specific algorithms or techniques that do not directly exist in the original video sequence. These frames can be used to improve the accuracy of inter-frame prediction, thereby improving compression efficiency and video quality. Virtual reference frames can be generated through one or more of the following methods:
[0144] (1) Generation based on motion compensation interpolation. For example, by analyzing the motion vectors between two reference frames, motion estimation and compensation techniques can be used to interpolate in time to generate an intermediate frame, which can then be used as a virtual reference frame.
[0145] (2) By estimating the motion trajectory of pixels in the image sequence, a virtual reference frame is generated based on this.
[0146] (3) Generate new reference frames by weighted averaging of multiple reference frames in the time dimension or by other forms of filtering operations, so as to serve as virtual reference frames.
[0147] (4) Using the content of multiple reference frames, a new virtual reference frame is generated through a specific synthesis algorithm (such as content copying, patching, color adjustment, etc.).
[0148] (5) Generate virtual reference frames using neural networks or deep learning. For example, input the reference frames adjacent to the current frame into the neural network to derive high-quality virtual reference frames; or first use traditional motion estimation methods to generate preliminary virtual reference frames, and then refine and optimize them using neural networks.
[0149] (6) Use a neural network-based image and video compression method to generate a virtual reference frame. For example, input the current image to be encoded, or the current image to be encoded and its reference image into a neural network for compression, and export the decoded and reconstructed image as the virtual reference frame of the current image.
[0150] Figure 9 illustrates the technical solution of the embodiment of this application from the perspective of video decoding. The technical solution of the embodiment of this application will be described again below from the perspective of video encoding with reference to Figure 10.
[0151] Figure 10 shows a flowchart of a video encoding method according to some embodiments of this application. This video encoding method can be executed by a device with computing processing capabilities, such as a terminal device or a server. Referring to Figure 7, the video encoding method includes at least steps S1010 to S1030, which are described in detail below:
[0152] In S1010, it is determined whether the current stripe uses a virtual reference frame.
[0153] In S1020, if the current strip uses a virtual reference frame, a list of reference images corresponding to the current strip is constructed based on the generated virtual reference frame.
[0154] In S1030, the current strip is encoded based on the constructed list of reference images.
[0155] This application also provides a video encoding method, including:
[0156] Generate a video stream and store the video stream, wherein generating the video stream includes:
[0157] Determine whether the current strip uses a virtual reference frame.
[0158] If the current strip uses a virtual reference frame, then a list of reference images for the current strip is constructed based on the generated virtual reference frame.
[0159] The current strip is encoded based on a constructed list of reference images.
[0160] The processing at the video encoding end is similar to that at the video decoding end. For details, please refer to the aforementioned processing procedures at the decoding end, which will not be repeated here.
[0161] The technical solution of this application can support the insertion of virtual reference frames into a reference image list and perform corresponding validity verification, which helps to improve video encoding efficiency and video quality. The following uses a reference frame generated by a neural network as an example to illustrate the technical solution of this application:
[0162] In some embodiments, it can be first determined whether to use a virtual reference frame generated based on a neural network, based on the bitstream.
[0163] For example, a flag bit (first flag bit) (such as sps_nn_inter_flag) can be added to the sequence header to indicate whether a neural network-based virtual reference frame generation technique is used. The specific semantic structure is shown in Table 5.
[0164] Table 5
[0165] In some embodiments, sps_nn_inter_flag may be decoded only if one or more of the following conditions are met:
[0166] (1) When long-term reference frames are not used (e.g., the value of sps_long_term_ref_pics_flag is 0);
[0167] (2) When inter-layer prediction is not used (e.g., the value of sps_inter_layer_prediction_enabled_flag is 0);
[0168] (3) RPL 0 and RPL 1 are not the same (for example, the value of sps_rpl1_same_as_rpl0_flag is 0);
[0169] (4) The number of reference images contained in one or more of RPL 0 and RPL 1 is greater than or equal to a set threshold TH_NUM_RPL, for example, TH_NUM_RPL is a positive integer greater than or equal to 1;
[0170] (5) The ref_pic_list_struct(i,j) obtained by decoding in the sequence header satisfies at least one of the following conditions:
[0171] The POC of the reference image contained in RPL 0 or RPL 1 meets the requirements;
[0172] The TID of the reference frame contained in RPL 0 or RPL 1 meets the requirement that it is greater than or equal to the threshold TH_TID.
[0173] For example, a target POC list can be set up, which contains one or more target POCs. If a reference frame with the target POC exists in RPL 0 or RPL 1, then the POC of the reference frame contained in RPL 0 or RPL 1 meets the requirements.
[0174] For example, TH_TID is a positive integer greater than 1, such as TH_TID = 2, 3 or 4.
[0175] In this embodiment, `ref_pic_list_struct(i,j)` is used to define the contents of the reference image list, where `i` indicates the reference image list to be constructed. For example, `i = 0` indicates constructing reference image list 0 (RPL 0); `i = 1` indicates constructing reference image list 2 (RPL 2). `j` represents the index of the predefined reference image list configuration in the Sequence Parameter Set (SPS), meaning that if the encoder chooses to use a predefined configuration in the SPS, it will select the specific configuration based on this index.
[0176] For example, a flag bit (second flag bit) (such as ph_nn_inter_flag) can be added to the image header (PH) to indicate whether a neural network-based virtual reference frame generation technique is used. The specific semantic structure is shown in Table 6.
[0177] Table 6
[0178] In some embodiments, ph_nn_inter_flag may be decoded only when one or more of the following conditions are met; alternatively, the use of nn_inter may be implicitly determined directly when one or more of the following conditions are met, without decoding the syntax element ph_nn_inter_flag:
[0179] (1) If a high-level syntax element such as sps_nn_inter_flag exists, then the value of sps_nn_inter_flag is 1, that is, the high-level syntax element indicates that the nn_inter mode is used;
[0180] (2) The frame type of the current image is the specified type, such as B frame;
[0181] (3) The reference image list for the current image satisfies one or more of the following conditions:
[0182] The number of reference images contained in RPL 0 or RPL 1 meets the requirements, such as being greater than or equal to a set threshold TH_NUM_RPL, for example, TH_NUM_RPL is a positive integer greater than or equal to 1;
[0183] The TID of the reference image contained in RPL 0 or RPL 1 meets the requirement that it is greater than or equal to the threshold TH_TID. For example, TH_TID is a positive integer greater than 1, such as TH_TID = 2, 3 or 4, etc.
[0184] The POC of the reference image contained in RPL 0 or RPL 1 meets the requirements.
[0185] For example, a list of target POCs can be set up, which contains one or more target POCs. If a reference frame with the target POC exists in RPL 0 or RPL 1, then the POC of the reference image contained in RPL 0 or RPL 1 meets the requirements.
[0186] For example, the target POC list can be {poc_cur-1, poc_cur-2, ..., poc_cur-K}. Here, poc_cur represents the POC value of the current image, and K is a positive integer greater than or equal to 1. For instance, if RPL 0 or RPL 1 contains a reference frame with poc_cur-1, or a reference frame with both poc_cur-1 and poc_cur-2, then the POC of the reference images contained in RPL 0 or RPL 1 meets the requirements.
[0187] For example, the target POC list can also be derived based on the TID of the current slice. For instance, the target POC list could be {poc_cur – pow(2, G – tid_cur), poc_cur + pow(2, G – tid_cur)}, where pow() represents a power function, tid_cur is the TID of the current slice, and G is a value associated with the current Group of Pictures (GOP). For example, G = log(2, gop_size). For instance, when gop_size is 32, the value of G is equal to 5.
[0188] For example, a flag (a third flag) (such as slice_nn_inter_flag) can be added to the slice header to indicate whether a neural network-based virtual reference frame generation technique is used. The specific semantic structure is shown in Table 7.
[0189] Table 7
[0190] In some embodiments, slice_nn_inter_flag may be decoded only when one or more of the following conditions are met; alternatively, the use of nn_inter may be implicitly determined directly when one or more of the following conditions are met, without decoding the syntax element slice_nn_inter_flag:
[0191] (1) If there are high-level syntax elements such as sps_nn_inter_flag, ph_nn_inter_flag, etc., and the high-level syntax elements indicate the use of nn_inter mode;
[0192] (2) The frame type of the current image is the specified type, such as the current image being a B frame;
[0193] (3) The reference image list of the current slice satisfies one or more of the following conditions:
[0194] The number of reference images contained in RPL 0 or RPL 1 meets the requirements, such as being greater than or equal to a set threshold TH_NUM_RPL, for example, TH_NUM_RPL is a positive integer greater than or equal to 1;
[0195] The TID of the reference image contained in RPL 0 or RPL 1 meets the requirement that it is greater than or equal to the threshold TH_TID. For example, TH_TID is a positive integer greater than 1, such as TH_TID = 2, 3 or 4, etc.
[0196] The POC of the reference image contained in RPL 0 or RPL 1 meets the requirements.
[0197] For example, a target POC list can be set up, which contains one or more target POCs. If a reference frame with the target POC exists in RPL 0 or RPL 1, then the POC of the reference image contained in RPL 0 or RPL 1 meets the requirements.
[0198] For example, the target POC list can be {poc_cur-1, poc_cur-2, ..., poc_cur-K}. Here, poc_cur represents the POC value of the current image, and K is a positive integer greater than or equal to 1. For instance, if RPL 0 or RPL 1 contains a reference frame with poc_cur-1, or a reference frame with both poc_cur-1 and poc_cur-2, then the POC of the reference images contained in RPL 0 or RPL 1 meets the requirements.
[0199] For example, the target POC list can also be derived based on the TID of the current slice. For instance, the target POC list could be {poc_cur – pow(2, G – tid_cur), poc_cur + pow(2, G – tid_cur)}. Here, pow() represents the power function; tid_cur is the TID of the current slice; and G is a value associated with the current Group of Pictures (GOP). For example, G = log(2, gop_size). For instance, when gop_size is 32, the value of G is equal to 5.
[0200] In some embodiments, when it is determined that a virtual reference frame generated based on a neural network will be used, the list of reference images can be adjusted accordingly to support the virtual reference frame. For example, the length of the list of reference images can be adjusted, or the length of the list of reference images can be kept constant by adjusting the reference images in the list of reference images, which will be explained below.
[0201] In some embodiments, when it is necessary to adjust the length of the reference image list, the correct reference image list length can be decoded by modifying the decoding conditions of the flag bits (e.g., sh_num_ref_idx_active_override_flag) in the bitstream that indicate whether the default number of active reference images is overridden. For example, the modifications to the decoding conditions of sh_num_ref_idx_active_override_flag and related flag bits can be shown in Table 8:
[0202] Table 8
[0203] Referring to Table 8, num_ref_nn_inter is determined based on the decoding information in the bitstream. For example, if nn_inter is used, then num_ref_nn_inter is the number of inserted virtual reference frames; if nn_inter is not used, then num_ref_nn_inter is 0.
[0204] If the current slice is not an I slice (e.g., sh_slice_type != I) and the number of reference frames in RPL 0 plus num_ref_nn_inter is greater than 1 (e.g., num_ref_entries[0][RplsIdx[0]]+num_ref_nn_inter>1); or the current slice is a B slice (e.g., sh_slice_type == B) and the number of reference frames in RPL 1 plus num_ref_nn_inter is greater than 1 (e.g., num_ref_entries[1][RplsIdx[1]]+num_ref_nn_inter>1), then proceed to the next step of processing.
[0205] Parse sh_num_ref_idx_active_override_flag. If this flag is set to 1, it means that the default number of active reference images will be overridden, and further parsing of sh_num_ref_idx_active_minus1[i] is required.
[0206] Based on the state of sh_num_ref_idx_active_override_flag, iterate through all relevant reference image lists and perform the following operations. For non-I Slices, if it is a B Slice, both RPL 0 and RPL 1 are considered (e.g., the loop variable i ranges from 0 to 1); if it is a P Slice, only RPL 0 is considered (i.e., the loop variable i ranges only 0). For each reference list i, if the number of reference images it contains + num_ref_nn_inter is greater than 1 (e.g., num_ref_entries[i][RplsIdx[i]] + num_ref_nn_inter>1), then sh_num_ref_idx_active_minus1[i]. sh_num_ref_idx_active_minus1[i] is used to deduce the actual number of active reference images NumRefIdxActive[i] used in reference image list i.
[0207] In some embodiments, based on the modified decoding conditions described above, the default length of the reference image list (i.e., pps_num_ref_idx_default_active_minus1[i]) can be increased by num_ref_nn_inter, and correspondingly, the maximum value of the reference image list can also be increased by num_ref_nn_inter.
[0208] In some embodiments, based on the modified decoding conditions described above, the method for deriving NumRefIdxActive[i] (which indicates the number of active reference images actually used in reference image list i) can be modified, as shown in Table 9 below:
[0209] Table 9
[0210] Referring to Table 9, num_ref_nn_inter is determined based on the decoding information in the bitstream. For example, if nn_inter is used, then num_ref_nn_inter is the number of inserted virtual reference frames; if nn_inter is not used, then num_ref_nn_inter is 0.
[0211] Where the number of reference images available in RPL i (i = 0 or 1) (e.g., num_ref_entries[i][RplsIdx[i]]) is greater than or equal to the default length of RPL i (e.g., pps_num_ref_idx_default_active_minus1[i]) plus 1 and num_ref_nn_inter, then the number of active reference images actually used by RPL i (e.g., NumRefIdxActive[i]) is set to the default length of RPL i (i.e., pps_num_ref_idx_default_active_minus1[i]) plus 1 and num_ref_nn_inter. Otherwise, if the number of reference images available in RPL i is less than the default length of reference image list i plus 1 and num_ref_nn_inter, then the number of active reference images actually used by RPL i (e.g., NumRefIdxActive[i]) is set to the sum of the number of reference images available in RPL i (e.g., num_ref_entries[i][RplsIdx[i]]) and num_ref_nn_inter.
[0212] In some embodiments, based on the modified decoding conditions described above, and the processing of increasing num_ref_nn_inter by the default length of the reference image list (e.g., pps_num_ref_idx_default_active_minus1[i]), the method for deriving NumRefIdxActive[i] (which indicates the number of active reference images actually used in reference image list i) can also be modified as follows:
[0213] As shown in Table 10:
[0214] Table 10
[0215] Referring to Table 10, num_ref_nn_inter is determined based on the decoding information in the bitstream. For example, if nn_inter is used, then num_ref_nn_inter is the number of inserted virtual reference frames; if nn_inter is not used, then num_ref_nn_inter is 0.
[0216] If the number of available reference images in RPL i (i = 0 or 1) (e.g., num_ref_entries[i][RplsIdx[i]]) is greater than or equal to the sum of the default length of RPL i (e.g., pps_num_ref_idx_default_active_minus1[i], the default length having been increased by num_ref_nn_inter) and 1, then the number of active reference images actually used by RPL i (e.g., NumRefIdxActive[i]) is set to the sum of the default length of RPL i (e.g., pps_num_ref_idx_default_active_minus1[i]) and 1. Otherwise, if the number of available reference images in RPL i is less than the sum of the default length of reference image list i and 1, then the number of active reference images actually used by RPL i (e.g., NumRefIdxActive[i]) is set to the sum of the number of available reference images in RPL i (e.g., num_ref_entries[i][RplsIdx[i]]) and num_ref_nn_inter.
[0217] In some embodiments, when it is necessary to adjust the length of the reference image list, the decoding conditions of the relevant syntax elements in the bitstream can be changed without modifying them; instead, the length of the reference image list can be directly increased by num_ref_nn_inter. Here, if nn_inter is used, num_ref_nn_inter represents the number of inserted virtual reference frames; if nn_inter is not used, num_ref_nn_inter is 0.
[0218] For example, when increasing the length of the reference image list, the length of RPL 0 can be increased, the length of RPL 1 can be increased, or the lengths of both RPL 0 and RPL 1 can be increased simultaneously.
[0219] In some embodiments, when inserting a virtual reference frame into the reference image list, the generated virtual reference frame can be inserted at the Y-th position in the reference image list. Exemplarily, a virtual reference frame can be inserted into only RPL 0 or RPL 1, or it can be inserted into both RPL 0 and RPL 1 simultaneously. Exemplarily, the virtual reference frames inserted into RPL 0 and RPL 1 can be the same or different, and the insertion positions can be the same or different.
[0220] For example, a virtual reference frame can be inserted into the position at index 1 in RPL 0 and RPL 1 (index values start from 0).
[0221] For example, different positions can be inserted based on whether the current image has a backward reference frame. For instance, if the current image does not have a backward reference frame, the position with index 1 is inserted; if the current image has a backward reference frame, the position with index 2 is inserted.
[0222] In some embodiments, if the length of the reference image list can be maintained by adjusting the reference images in the reference image list, then a virtual reference frame can be used to directly replace the reference images in the reference image list. That is, the reference image at the position where the virtual reference frame is to be inserted can be deleted, thus ensuring that the length of the reference image list remains unchanged.
[0223] For example, a virtual reference frame can be used to replace the reference image in RPL 0 or RPL 1, or the reference images in both RPL 0 and RPL 1 can be replaced simultaneously. For example, the virtual reference frames used to replace the reference images in RPL 0 and RPL 1 can be the same or different, and the positions of the replaced reference images can be the same or different. For example, certain types of reference images may not be replaced, such as I-frames, or reference images with a TID less than or equal to a specific threshold (e.g., less than or equal to 2).
[0224] In some embodiments, if the length of the reference image list is maintained by adjusting the reference images in the reference image list, then after inserting virtual reference frames into the reference image list, some reference images in the reference image list can be deleted to maintain the length of the reference image list. Exemplarily, this method can insert virtual reference frames only into RPL 0 or RPL 1, or it can insert virtual reference frames into both RPL 0 and RPL 1 simultaneously. Exemplarily, the virtual reference frames inserted into RPL 0 and RPL 1 can be the same or different, and the insertion positions can be the same or different.
[0225] In some embodiments, after inserting a virtual reference frame into the reference image list, if the number of reference images in the resulting reference image list exceeds a set threshold (exemplarily, this threshold may be less than or equal to the length of the reference image list), then a specified reference image can be deleted from the reference image list. Exemplarily, this method can be applied only to RPL 0 or RPL 1, or it can be applied to both RPL 0 and RPL 1 simultaneously. The technical solution of this embodiment allows for flexible control over the number of reference images contained in the reference image list after inserting the virtual reference frame.
[0226] For example, in the above embodiments, when deleting a specified reference image, reference images that are ranked late in the reference image list can be deleted. For instance, if a virtual reference frame is inserted, the last reference image in the reference image list can be deleted. If two virtual reference frames are inserted, the last two reference images in the reference image list can be deleted. For example, specific types of reference images can also be deleted, such as I-frames, or reference images with a TID less than or equal to a specific threshold (e.g., less than or equal to 2).
[0227] In some embodiments, virtual reference frames may not be added to the reference image list when certain conditions are met. In this case, an inter-frame prediction scheme that does not include virtual reference frames can be used. The certain conditions may include one or more of the following:
[0228] The number of reference images in the reference image list is less than the second threshold;
[0229] The number of reference images in the reference image list is greater than the third threshold.
[0230] If the number of reference images in the reference image list is less than the second threshold, for example, if there is only one reference image in RPL 0, the virtual reference frame will not be added to the reference image list.
[0231] If the number of reference images in the reference image list is greater than the third threshold, for example, if RPL 0 already has 5 reference images, the virtual reference frame will not be added to the reference image list.
[0232] In some embodiments, after a virtual reference frame is added to a list of reference images, it can be set as a short-term reference frame.
[0233] In the above embodiments, the list of reference images to which the virtual reference frame is added in any way can be referred to as a predetermined list of reference images. After adding the virtual reference frame to the predetermined list of reference images, the final constructed list of reference images is obtained.
[0234] The above embodiments illustrate the use of a virtual reference frame generated via a neural network. In other embodiments of this application, the virtual reference frame can also be generated in other ways, such as through motion-compensated interpolation, or by estimating the motion trajectory of pixels in an image sequence and generating the virtual reference frame accordingly.
[0235] The technical solution of this application improves the way the reference image list is exported, and can support the insertion of virtual reference frames into the reference image list. By inserting high-quality virtual reference frames, the video encoding performance can be effectively improved, and the robustness of the codec can also be improved.
[0236] The technical solutions of the various embodiments shown above can be used individually or in combination. Furthermore, the technical solutions of the embodiments of this application can be applied to related products such as video codecs or video compression.
[0237] The following describes an apparatus embodiment of this application, which can be used to perform the methods described in the above embodiments of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the method embodiments described above.
[0238] Figure 11 shows a block diagram of a video decoding apparatus according to some embodiments of the present application. The video decoding apparatus can be installed in a device with computing processing capabilities, such as a terminal device or a server.
[0239] Referring to FIG11, a video decoding apparatus 1100 according to some embodiments of the present application includes: a decoding unit 1102, a generating unit 1104, and a processing unit 1106.
[0240] The decoding unit 1102 is configured to decode the video stream to determine whether the current stripe uses a virtual reference frame; the generation unit 1104 is configured to construct a reference image list corresponding to the current stripe based on the generated virtual reference frame if it is determined that the current stripe uses a virtual reference frame; and the processing unit 1106 is configured to decode the current stripe based on the constructed reference image list.
[0241] In some embodiments of this application, based on the foregoing scheme, it is determined whether the current stripe uses a virtual reference frame according to one or more of the following flag bits: a first flag bit in the sequence header, a second flag bit in the image header, and a third flag bit in the stripe header;
[0242] Wherein, the first flag bit is used to indicate whether the current sequence uses a virtual reference frame, the second flag bit is used to indicate whether the current image where the current strip is located uses a virtual reference frame, and the third flag bit is used to indicate whether the current strip where the current strip is located uses a virtual reference frame.
[0243] In some embodiments of this application, based on the foregoing scheme, if it is determined at least according to the first flag bit whether the current stripe uses a virtual reference frame, then the first flag bit is decoded when at least one of the following conditions is met:
[0244] The first parameter in the sequence parameter set indicates that the current sequence does not use a long-term reference frame;
[0245] The second parameter in the sequence parameter set indicates that inter-layer prediction is not used for the current sequence;
[0246] The third parameter in the sequence parameter set indicates that the two lists of reference images used are different;
[0247] The fourth parameter in the sequence parameter set indicates that the number of reference images contained in the two reference image lists used is greater than or equal to the first quantity threshold, or indicates that the number of reference images contained in one reference image list used is greater than or equal to the first quantity threshold.
[0248] The current sequence has a reference image with a first target image sequence count in the reference image list, and the first target image sequence count is in the set image sequence count list;
[0249] The time level identifier of the reference image list corresponding to the current sequence is greater than or equal to the first set threshold.
[0250] In some embodiments of this application, based on the foregoing scheme, if it is determined at least according to the second flag bit whether the current strip uses a virtual reference frame, then the second flag bit is decoded when at least one of the following conditions is met; or the decoding process of the second flag bit is skipped when at least one of the following conditions is met, and it is determined that the current image needs to use a virtual reference frame:
[0251] The fifth parameter in the sequence parameter set indicates that the current sequence uses a virtual reference frame;
[0252] The frame type of the current image is the specified type;
[0253] The sixth parameter in the image parameter set indicates that the number of reference images contained in the two reference image lists used is greater than or equal to the second quantity threshold, or indicates that the number of reference images contained in one reference image list used is greater than or equal to the second quantity threshold.
[0254] The current image has a reference image with a second target image sequence count in the reference image list, and the second target image sequence count is in the set image sequence count list;
[0255] The time level identifier of the reference image list corresponding to the current image is greater than or equal to the second set threshold.
[0256] In some embodiments of this application, based on the foregoing scheme, if it is determined at least according to the third flag bit whether the current stripe uses a virtual reference frame, then the third flag bit is decoded when at least one of the following conditions is met; or the decoding process of the third flag bit is skipped when at least one of the following conditions is met, and it is determined that the current stripe needs to use a virtual reference frame:
[0257] The fifth parameter in the sequence parameter set indicates that the current sequence uses a virtual reference frame;
[0258] The seventh parameter in the image parameter set indicates that the current image uses a virtual reference frame;
[0259] The frame type of the current image is the specified type;
[0260] The eighth parameter in the strip parameter set indicates that the number of reference images contained in the two reference image lists used is greater than or equal to the third quantity threshold, or indicates that the number of reference images contained in one reference image list used is greater than or equal to the third quantity threshold.
[0261] The current stripe has a reference image with a third target image sequence count in the reference image list, and the third target image sequence count is in the set image sequence count list;
[0262] The time level identifier of the reference image list corresponding to the current strip is greater than or equal to the third set threshold.
[0263] In some embodiments of this application, based on the foregoing scheme, the set image sequence counting list is determined according to the time level identifier of the current stripe.
[0264] In some embodiments of this application, based on the foregoing scheme, the generation unit 1104 is configured to insert the generated virtual reference frame into a set position in the reference image list of the current strip.
[0265] In some embodiments of this application, based on the foregoing scheme, the generated virtual reference frame is inserted into a set position in the reference image list of the current stripe, including at least one of the following methods:
[0266] The generated first virtual reference frame is inserted into the first set position in the first reference image list of the current strip;
[0267] The generated second virtual reference frame is inserted into the second predetermined position in the second reference image list of the current strip;
[0268] The first virtual reference frame may be the same as or different from the second virtual reference frame, and the first set position may be the same as or different from the second set position.
[0269] In some embodiments of this application, based on the foregoing scheme, the processing unit 1106 is further configured to: if the current stripe is not an intra-frame stripe and the number of available reference images in the first reference image list of the current stripe is greater than 1, or if the current stripe is a bidirectional prediction stripe and the sum of the number of available reference images in the second reference image list of the current stripe and a first set value is greater than 1, then decode a fourth flag bit, the value of which is used to indicate whether the default number of active reference images is overridden; if the value of the fourth flag bit indicates that the default number of active reference images is overridden, then when the current stripe is a unidirectional prediction stripe and the sum of the number of available reference images in the first reference image list and the first set value is greater than 1, update the number of actually used active reference images in the first reference image list;
[0270] Wherein, if the current strip does not use virtual reference frames, the first setting value is 0; if the current strip uses virtual reference frames, the first setting value is the number of virtual reference frames used.
[0271] In some embodiments of this application, based on the foregoing scheme, the processing unit 1106 is further configured to: if the value of the fourth flag indicates the number of active reference images covering the default, then when the current strip is a bidirectional prediction strip, perform the following process: if the sum of the number of available reference images in the first reference image list and the first set value is greater than 1, then update the number of active reference images actually used in the first reference image list; if the sum of the number of available reference images in the second reference image list and the first set value is greater than 1, then update the number of active reference images actually used in the second reference image list.
[0272] In some embodiments of this application, based on the foregoing scheme, updating the actual number of active reference images used includes: if the number of available reference images in reference image list i is greater than or equal to the sum of the default length of reference image list i, 1, and the first set value, then the actual number of active reference images used in reference image list i is set to the sum of the default length of reference image list i, 1, and the first set value.
[0273] If the number of available reference images in reference image list i is less than the sum of the default length of reference image list i, 1, and the first setting value, then the actual number of active reference images used by reference image list i is set to the sum of the number of available reference images in reference image list i and the first setting value.
[0274] Wherein, reference image list i represents at least one of the first reference image list and the second reference image list.
[0275] In some embodiments of this application, based on the foregoing scheme, updating the actual number of active reference images used includes: if the number of available reference images in reference image list i is greater than or equal to the sum of the default length of reference image list i and 1, then the actual number of active reference images used in reference image list i is set to the sum of the default length of reference image list i and 1; if the number of available reference images in reference image list i is less than the sum of the default length of reference image list i and 1, then the actual number of active reference images used in reference image list i is set to the sum of the number of available reference images in reference image list i and the first set value.
[0276] Wherein, reference image list i represents at least one of the first reference image list and the second reference image list; and the default length of reference image list i is the original length plus the first set value.
[0277] In some embodiments of this application, based on the aforementioned scheme, the length of the reference image list corresponding to the current strip (preset reference image list length) is increased by a first preset value, where the first preset value is the number of virtual reference frames to be inserted.
[0278] In some embodiments of this application, based on the aforementioned scheme, inserting the generated virtual reference frame into a set position in the reference image list corresponding to the current stripe includes: deleting a predetermined reference image at the set position and inserting the generated virtual reference frame into the set position.
[0279] In some embodiments of this application, based on the foregoing scheme, inserting the generated virtual reference frame into a set position in the reference image list corresponding to the current strip includes: inserting the generated virtual reference frame into the set position and removing a specified reference image from the reference image list corresponding to the current strip, so that the length of the reference image list corresponding to the current strip remains unchanged.
[0280] In some embodiments of this application, based on the foregoing scheme, inserting the generated virtual reference frame into a set position in the reference image list corresponding to the current stripe includes: inserting the generated virtual reference frame into the set position to obtain an updated reference image list; if the number of reference images contained in the updated reference image list exceeds a reference image number threshold, then removing a specified reference image from the updated reference image list.
[0281] In some embodiments of this application, based on the foregoing scheme, the designated reference image includes one or more of the following reference images: a reference image of a specified frame type in the reference image list corresponding to the current strip, a reference image in the reference image list corresponding to the current strip whose time level identifier is less than or equal to a first threshold, and a reference image arranged later in the reference image list corresponding to the current strip.
[0282] In some embodiments of this application, based on the foregoing scheme, the construction of the reference image list for the current stripe is rejected when one or more of the following conditions are met: the number of reference images in the reference image list for the current stripe is less than a second threshold, or the number of reference images in the reference image list for the current stripe is greater than a third threshold.
[0283] In some embodiments of this application, based on the foregoing scheme, the processing unit 1106 is further configured to: set the virtual reference frame as a short-term reference frame.
[0284] Figure 12 shows a block diagram of a video encoding apparatus according to some embodiments of the present application. The video encoding apparatus can be installed in a device with computing processing capabilities, such as a terminal device or a server.
[0285] Referring to FIG12, a video encoding apparatus 1200 according to some embodiments of the present application includes: a determining unit 1202, a generating unit 1204, and a processing unit 1206.
[0286] The determining unit 1202 is configured to determine whether the current strip uses a virtual reference frame; the generating unit 1204 is configured to construct a reference image list corresponding to the current strip based on the generated virtual reference frame if the current strip uses a virtual reference frame; and the processing unit 1206 is configured to perform encoding processing on the current strip based on the constructed reference image list.
[0287] Figure 13 shows a schematic diagram of the structure of a computer system suitable for implementing the computer device of the present application. The computer device may be a video encoding device or a video decoding device in the foregoing embodiments.
[0288] It should be noted that the computer system 1300 of the computer device shown in Figure 13 is only an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0289] As shown in Figure 13, the computer system 1300 may include a Central Processing Unit (CPU) 1301, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 1302 or programs loaded from storage portion 1308 into Random Access Memory (RAM) 1303, such as performing the methods described in the above embodiments. The RAM 1303 also stores various programs and data required for system operation. The CPU 1301, ROM 1302, and RAM 1303 are interconnected via a bus 1304. An Input / Output (I / O) interface 1305 is also connected to the bus 1304.
[0290] The following components can be connected to I / O interface 1305: an input section 1306 including a keyboard, mouse, etc.; an output section 1307 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1308 including a hard disk, etc.; and a communication section 1309 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 1309 performs communication processing via a network such as the Internet. A drive 1310 is also connected to I / O interface 1305 as needed. Removable media 1311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 1310 as needed so that computer programs read from them can be installed into storage section 1308 as needed.
[0291] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1309, and / or installed from removable medium 1311. When the computer program is executed by central processing unit (CPU) 1301, it performs various functions defined in the system of this application.
[0292] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a computer program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.
[0293] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and a computer program.
[0294] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.
[0295] In another aspect, this application also provides a computer-readable medium, which may be included in the computer device described in the above embodiments; or it may exist independently and not assembled into the computer device. The computer-readable medium carries one or more computer programs, which, when executed by the computer device, cause the computer device to perform the methods described in the above embodiments.
[0296] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0297] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, and includes several instructions to cause a computer device to execute the method according to the embodiments of this application.
[0298] For example, a computer device can be a video decoding device, in which case the video decoding device can perform the video decoding method shown in Figure 9; or, for instance, a computer device can be a video encoding device, in which case the video encoding device can perform the video encoding method shown in Figure 10.
[0299] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A video decoding method, characterized in that, include: The video stream is decoded to determine whether the current strip uses a virtual reference frame; If it is determined that the current strip uses a virtual reference frame, then a list of reference images for the current strip is constructed based on the generated virtual reference frame; The current strip is decoded based on the constructed list of reference images.
2. The video decoding method according to claim 1, characterized in that, Whether the current stripe uses a virtual reference frame is determined based on one or more of the following flags: a first flag in the sequence header, a second flag in the image header, and a third flag in the stripe header; Wherein, the first flag bit is used to indicate whether the current sequence in which the current strip is located uses a virtual reference frame, the second flag bit is used to indicate whether the current image in which the current strip is located uses a virtual reference frame, and the third flag bit is used to indicate whether the current strip uses a virtual reference frame.
3. The video decoding method according to claim 2, characterized in that, The first flag bit is decoded when at least one of the following conditions is met: The first parameter in the sequence parameter set indicates that the current sequence does not use a long-term reference frame; The second parameter in the sequence parameter set indicates that the current sequence does not use inter-layer prediction; The third parameter in the sequence parameter set indicates that the two lists of reference images used are different; The fourth parameter in the sequence parameter set indicates that the number of reference images contained in the two reference image lists used is greater than or equal to the first number threshold, or indicates that the number of reference images contained in one reference image list used is greater than or equal to the first number threshold. The current sequence has a reference image in the reference image list that has a first target image sequence count, and the first target image sequence count is in the set image sequence count list; The time-level identifier of the reference image list in the current sequence is greater than or equal to a first set threshold.
4. The video decoding method according to claim 2 or 3, characterized in that, The second flag bit is decoded when at least one of the following conditions is met; or the decoding process for the second flag bit is skipped when at least one of the following conditions is met, and it is determined that the current image needs to use a virtual reference frame: The fifth parameter in the sequence parameter set indicates that the current sequence uses a virtual reference frame; The frame type of the current image is a specified type; The sixth parameter in the image parameter set indicates that the number of reference images contained in the two reference image lists used is greater than or equal to the second quantity threshold, or indicates that the number of reference images contained in one reference image list used is greater than or equal to the second quantity threshold. The current image's reference image list contains a reference image with a second target image sequence count, and the second target image sequence count is located in the set image sequence count list; The time-level identifier of the reference image list of the current image is greater than or equal to a second set threshold.
5. The video decoding method according to any one of claims 2 to 4, characterized in that, The third flag is decoded when at least one of the following conditions is met; or the decoding process for the third flag is skipped and it is determined that the current stripe requires the use of a virtual reference frame when at least one of the following conditions is met: The fifth parameter in the sequence parameter set indicates that the current sequence uses a virtual reference frame; The seventh parameter in the image parameter set indicates that the current image uses a virtual reference frame; The frame type of the current image is a specified type; The eighth parameter in the strip parameter set indicates that the number of reference images contained in the two reference image lists used is greater than or equal to the third quantity threshold, or indicates that the number of reference images contained in one reference image list used is greater than or equal to the third quantity threshold. The current strip's reference image list contains a reference image with a third target image sequence count, and the third target image sequence count is located in a set image sequence count list; The time-level identifier of the reference image list of the current strip is greater than or equal to a third preset threshold.
6. The video decoding method according to claim 4 or 5, characterized in that, The set image sequence count list is determined based on the time level identifier of the current stripe.
7. The video decoding method according to any one of claims 1 to 6, characterized in that, The reference image list for the current stripe is constructed based on the generated virtual reference frame, including: The generated virtual reference frame is inserted into a set position in the reference image list of the current strip.
8. The video decoding method according to claim 7, characterized in that, Inserting the generated virtual reference frame into the designated position in the reference image list of the current strip includes at least one of the following methods: The generated first virtual reference frame is inserted into the first set position in the first reference image list of the current strip; The generated second virtual reference frame is inserted into the second predetermined position in the second reference image list of the current strip; The first virtual reference frame may be the same as or different from the second virtual reference frame, and the first set position may be the same as or different from the second set position.
9. The video decoding method according to claim 7 or 8, characterized in that, The video decoding method further includes: If the current slice is not an intra-frame slice and the number of available reference images in the first reference image list of the current slice is greater than 1, or if the current slice is a bidirectional prediction slice and the sum of the number of available reference images in the second reference image list of the current slice and the first set value is greater than 1, then the fourth flag bit is decoded. The value of the fourth flag bit is used to indicate whether the default number of active reference images is overridden. If the value of the fourth flag indicates the number of active reference images that cover the default, then when the current strip is a unidirectional prediction strip and the sum of the number of reference images available in the first reference image list and the first set value is greater than 1, the number of active reference images actually used in the first reference image list is updated. Wherein, if the current strip does not use virtual reference frames, the first setting value is 0; if the current strip uses virtual reference frames, the first setting value is the number of virtual reference frames used.
10. The video decoding method according to claim 9, characterized in that, The video decoding method further includes: If the value of the fourth flag indicates the number of active reference images covered by the default, then when the current strip is a bidirectional prediction strip, the following procedure is performed: If the sum of the number of available reference images in the first reference image list and the first set value is greater than 1, then update the number of active reference images actually used in the first reference image list. If the sum of the number of available reference images in the second reference image list and the first set value is greater than 1, then the number of active reference images actually used in the second reference image list is updated.
11. The video decoding method according to claim 9 or 10, characterized in that, Update the actual number of active reference images used, including: If the number of available reference images in reference image list i is greater than or equal to the sum of the default length of reference image list i, 1, and the first set value, then the actual number of active reference images used in reference image list i is set to the sum of the default length of reference image list i, 1, and the first set value. If the number of available reference images in reference image list i is less than the sum of the default length of reference image list i, 1, and the first setting value, then the actual number of active reference images used by reference image list i is set to the sum of the number of available reference images in reference image list i and the first setting value. Wherein, reference image list i represents at least one of the first reference image list and the second reference image list.
12. The video decoding method according to claim 9 or 10, characterized in that, Update the actual number of active reference images used, including: If the number of available reference images in reference image list i is greater than or equal to the sum of the default length of reference image list i and 1, then the actual number of active reference images used in reference image list i is set to the sum of the default length of reference image list i and 1. If the number of available reference images in reference image list i is less than the sum of the default length of reference image list i and 1, then the actual number of active reference images used by reference image list i is set to the sum of the number of available reference images in reference image list i and the first set value. Wherein, reference image list i represents at least one of the first reference image list and the second reference image list; and the default length of reference image list i is the original length plus the first set value.
13. The video decoding method according to any one of claims 7 to 12, characterized in that, Increase the length of the preset reference image list corresponding to the current strip by a first set value, where the first set value is the number of virtual reference frames to be inserted.
14. The video decoding method according to any one of claims 7 to 12, characterized in that, Inserting the generated virtual reference frame into the designated position includes: Delete the predetermined reference image at the set position, and insert the generated virtual reference frame into the set position; or Insert the generated virtual reference frame into the set position, and remove the specified reference image from the reference image list corresponding to the current strip; or The generated virtual reference frame is inserted into the set position to obtain an updated list of reference images; If the number of reference images in the updated reference image list exceeds the reference image number threshold, then the specified reference image in the updated reference image list is removed.
15. The video decoding method according to claim 14, characterized in that, The designated reference image includes one or more of the following reference images: The reference image of the specified frame type, the reference image whose time level identifier is less than or equal to the first threshold, and the reference image that is listed later in the reference image list corresponding to the current strip.
16. The video decoding method according to any one of claims 1 to 15, characterized in that, The video decoding method further includes: The construction of the reference image list for the current stripe based on the generated virtual reference frame is rejected if one or more of the following conditions are met: the number of reference images in the reference image list for the current stripe is less than a second threshold, or the number of reference images in the reference image list for the current stripe is greater than a third threshold.
17. A video encoding method, characterized in that, include: Determine whether the current stripe uses a virtual reference frame; If it is determined that the current strip uses a virtual reference frame, then a list of reference images for the current strip is constructed based on the generated virtual reference frame; The current strip is encoded based on a constructed list of reference images.
18. A video encoding method, characterized in that, include: Generate a video stream and store the video stream, wherein generating the video stream includes: Determine whether the current stripe uses a virtual reference frame; If it is determined that the current strip uses a virtual reference frame, then a list of reference images for the current strip is constructed based on the generated virtual reference frame; The current strip is encoded based on a constructed list of reference images.
19. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 18.
20. A computer device, characterized in that, include: One or more processors; A memory for storing one or more computer programs that, when executed by one or more processors, cause the computer device to perform the method of any one of claims 1 to 18.