Encoder, decoder, and corresponding methods
By using flags to reorder leading and non-leading pictures in VVC systems, interlaced video coding is efficiently integrated, addressing resource inefficiencies and maintaining VVC constraints, thereby enhancing coding efficiency and reducing resource usage.
Patent Information
- Application Number
- JP2025199151
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-06-21
- Filing Date
- 2025-11-19
- Publication Date
- 2026-01-29
- Estimated Expiration
- 2040-04-02
AI Technical Summary
Existing video coding systems, such as VVC, do not efficiently support interlaced video coding, violating constraints by requiring leading pictures to follow Intra Random Access Point (IRAP) pictures in decoding order, leading to inefficiencies in processor, memory, and network resource usage.
Implementing flags in the bitstream to allow for flexible ordering of leading and non-leading pictures, enabling interlaced video coding within VVC systems by positioning non-leading pictures before or after leading pictures, as indicated by the flag values, thus adhering to VVC constraints while supporting interlaced video.
Enhances coding efficiency, reduces resource usage, and allows for seamless integration of interlaced video coding with VVC systems without increasing streaming bandwidth.
Smart Images

Figure 2026015540000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This patent application is incorporated herein by reference in its entirety. ss Point And Leading Pictures In Video Coding” by FNU Hendry et al., 201 No. 62 / 828,875, filed April 3, 1999, and U.S. Provisional Patent Application ... "Access Point And Leading Pictures In Video Coding" by FNU Hendry et al. This application claims the benefit of U.S. Provisional Patent Application No. 62 / 864,958, filed June 21, 2019.
[0002] This disclosure relates generally to video coding, and more particularly to interlaced video coding. This document relates to coding leading pictures in the context of video coding. [Background technology]
[0003] The amount of video data required to render even a relatively short video is quite large. This is because data is streamed over a communications network with limited bandwidth capacity. This can cause difficulties when the data is to be sent or otherwise communicated. Video data is typically compressed before being transmitted over modern telecommunications networks. Memory resources may be limited, so if the video is stored on a storage device, The size of the video can also be an issue when it is being compressed. Video compression devices often using software and / or hardware in the necessary to code the video data and thereby represent a digital video image The compressed data is then passed to a video decoder that decodes the video data. The data is received at the destination by a network resource decompression device. Limited video resources and ever-increasing demand for higher video quality This allows for improved compression ratios with little or no sacrifice in image quality. A compression and decompression technique that utilizes the Summary of the Invention [Means for solving the problem]
[0004] In an embodiment, the present disclosure includes a method implemented in a decoder, the method comprising: The decoder receiver will then decode the flag and the Intra Random Access Point (IRAP) picture. A plurality of codecs including one or more non-leading pictures associated with an IRAP picture and an IRAP picture. receiving a bitstream comprising a coded picture; When the flag is set to a first value, the processor associates the IRAP picture with Any leading picture that is used must be in decoding order with all the pictures associated with the IRAP picture. determining that the picture is before a non-leading picture, and when the flag is set to a second value, The processor associates the non-leading pictures with the IRAP picture in decoding order. determining that the flag is set to a first value and that the flag is set to a first value; The processor determines whether the IRAP picture is an IRAP picture or an IRAP picture based on whether the IRAP picture is set to a first value or a second value. any leading pictures associated with the IRAP picture, and one or more leading pictures associated with the IRAP picture. decoding the non-leading pictures in decoding order; Translate one or more decoded pictures for display as part of a decoded video sequence. and a step of transmitting the
[0005] Versatile Video Coding (VVC) video systems use IRAP pictures, leading pictures, and In some examples, a bitstream containing leading pictures may be utilized. A leading picture may also be called a trailing picture. An IRAP picture is a picture that is A leading picture is an intra-predictive coded picture that serves as the starting point of a sequence. The picture precedes the IRAP picture in presentation order but is located after the IRAP picture in coding order. A non-leading / trailing picture is a picture coded after a presentation A picture that comes after the IRAP picture in both sequence and coding order. The video coding system may be configured such that the leading picture immediately follows the IRAP picture in decoding order. requires that there be a leading picture and that all non-leading pictures are after the leading picture. Interlace video coding provides a perceived high quality image without increasing streaming bandwidth. Interlaced video coding is a mechanism for increasing the frame rate of a video. A frame is divided into two fields. The horizontal lines for the first field of the frame are , is captured at a first time and coded in the first picture. The horizontal lines for the second field are captured at a second time and are placed immediately after the first picture. In this way, the resulting frame is , a slice from a first picture at a first time and a slice from a second picture at a second time. VVC systems use interlaced video, which enhances the sense of motion. For example, interlaced frames may not be supported by the device. It uses IRAP pictures and adjacent intra-predictive coded pictures to perform Intra-predictive coded pictures are non-leading / trailing pictures. Furthermore, when leading pictures are used, they are considered to be adjacent to It is positioned after the adjacent intra-predictive coded picture. The picture immediately follows the IRAP picture in decoding order, and all non-leading pictures This violates the VVC constraint that the leading picture must be after the leading picture. In the VVC system Contains flags that can be used to implement stray video coding. When set to the first value, such as , the leading picture, if any, is However, there may be a single non-preceding picture between the IRAP picture and any preceding picture. To indicate to the decoder that a row picture is being positioned, the encoder sets the flag to 1, etc. In one example, the non-leading pictures are positioned between the leading pictures. The flag may be included in the Sequence Parameter Set (SPS). This example may be applied to the entire sequence of pictures, not just the preceding picture. It is expected that multi-channel and interlaced video may be implemented together in the same bitstream. Contains flags that enhance the functionality of the encoder and / or decoder by allowing Additionally, this example demonstrates that leading pictures and interlaced video can be implemented together. By allowing for this, the coding efficiency of the resulting bitstream is increased. Thus, this example illustrates the processor resources in the encoder and / or decoder, The use of memory and / or network resources may be reduced.
[0006] Optionally, in any of the preceding aspects, another implementation of the aspect is further characterized in that the flag is a second When set to a value of , the first leading picture and the last leading picture in decoding order are determining that the leading picture is not positioned between the leading picture and the subsequent leading picture; This stipulates that:
[0007] Optionally, in any of the preceding aspects, another implementation of the aspect comprises: Specifies that the group contains an SPS and that the flags are obtained from the SPS.
[0008] Optionally, in any of the preceding aspects, another implementation of the aspect is It specifies that it is a sequential field flag (field_seq_flag).
[0009] Optionally, in any of the preceding aspects, another implementation of the aspect comprises coding field_seq_ to indicate that the video sequence contains pictures representing fields. flag is set to 1 and the coded video sequence contains a picture representing a frame. When indicating that it is included, it is specified that field_seq_flag is set to 0.
[0010] Optionally, in any of the preceding aspects, another implementation of the aspect is contains the first field of the frame, and the non-leading pictures before the first leading picture are Specifies that the second field of the frame is included.
[0011] Optionally, in any of the preceding aspects, another implementation of the aspect is and decoding one or more non-leading pictures from the first field from the IRAP picture. and the second field from the non-leading picture before the first leading picture. This includes lacing together to create a single frame.
[0012] In an embodiment, the present disclosure includes a method implemented in an encoder, the method comprising: The encoder processor generates an IRAP picture and one associated with the IRAP picture. Coding for a video sequence comprising a plurality of pictures including at least one non-leading picture determining a coding order; and marking the flags into the bitstream by a processor. a step of encoding, wherein any preceding picture associated with the IRAP picture is encoded; It must precede all non-leading pictures associated with the IRAP picture in the streaming order. When the flag is set to a first value, the non-leading picture is an IRAP picture in coding order. The flag is set to the second value when the frame is before the first leading picture associated with the frame. and a processor for generating an IRAP picture and any associated IRAP picture. Any leading picture and one or more non-leading pictures associated with the IRAP picture. encoding the data into a bitstream in a loading order; and storing the bitstream for communication to a decoder by a memory provided with the bitstream; It is equipped with:
[0013] The VVC video system handles video including IRAP pictures, leading pictures, and non-leading pictures. In some examples, the non-leading pictures are also called trailing pictures. An IRAP picture is an intermediate picture that serves as the start of a coded video sequence. The leading picture is a predictively coded picture. The leading picture is an IRAP picture in presentation order. A picture that comes before an IRAP picture but is coded after the IRAP picture in coding order. Non-leading / trailing pictures are pictures in both presentation order and coding order. Some video coding systems use the IRAP picture as the first picture. The row picture must immediately follow the IRAP picture in decoding order and all non-leading pictures must be Interlaced video coding requires that the next picture follows the previous picture. It is a mechanism to increase the perceived frame rate without increasing the streaming bandwidth. In interlaced video coding, a video frame is divided into two fields. The horizontal lines for the first field of the frame are captured at the first time and The horizontal lines for the second field of the frame are coded in the first picture. , captured at a second time and coded in a second picture immediately adjacent to the first picture. In this way, the resulting frame is the first picture at the first time. and a slice from the second picture in the second, which gives a sense of motion. VVC systems are not designed to support interlaced video. For example, an interlaced frame must be connected to an IRAP picture and its neighbors in order to function. Intra-predictive coded pictures may be used to predict the quality of the image. The loaded picture is considered to be a non-leading / trailing picture. When a picture is used, the preceding picture is used in conjunction with its adjacent intra-predictive coded picture. This means that the preceding picture is positioned after the IRAP picture in decoding order. immediately after the leading picture, and all non-leading pictures are after the leading picture. This example violates the constraints of C. Perform source video coding When the flag is set to a first value, such as 0, If there is a picture, it precedes all of the non-leading pictures. The default is to position a single non-leading picture between the picture and any leading picture. To indicate this to the coder, the encoder can set the flag to a second value, such as 1. In some examples, non-leading pictures may not be positioned between leading pictures. S and may be applied to the entire sequence of pictures. An example is when leading pictures and interlaced video are mixed together in the same bitstream. By allowing the implementation to Additionally, this example shows how leading pictures and interlaced video are handled together. The bitstream coding efficiency gained by allowing Therefore, this example is a Reduce processor, memory, and / or network resource usage. obtain.
[0014] Optionally, in any of the preceding aspects, another implementation of the aspect is further characterized in that the flag is a second When set to a value of , the first and last leading pictures in coding order are This specifies that no leading pictures are positioned between the current and previous frames.
[0015] Optionally, in any of the preceding aspects, another implementation of the aspect comprises: This specifies that the frame contains an SPS and that flags are encoded into the SPS.
[0016] Optionally, in any of the preceding aspects, another implementation of the aspect is It is specified that it is d_seq_flag.
[0017] Optionally, in any of the preceding aspects, another implementation of the aspect comprises coding field_seq_ to indicate that the video sequence contains pictures representing fields. flag is set to 1 and the coded video sequence contains a picture representing a frame. When indicating that it is included, it is specified that field_seq_flag is set to 0.
[0018] Optionally, in any of the preceding aspects, another implementation of the aspect is contains the first field of the frame, and the non-leading pictures before the first leading picture are Specifies that the second field of the frame is included.
[0019] Optionally, in any of the preceding aspects, another implementation of the aspect is the first field from the first leading picture and the second field from the non-leading picture before the first leading picture A field contains alternating lines of video data that represent a single interlaced video frame. It stipulates that
[0020] In an embodiment, the present disclosure provides a method for transmitting a signal to a processor, a receiver coupled to the processor, and a method for transmitting a signal to a processor. a memory coupled to the processor; and a transmitter coupled to the processor. a receiving device, the processor, the receiver, the memory, and the transmitter being as described in the preceding aspect. It is configured to perform either method.
[0021] In some embodiments, the present disclosure provides a computer program for use by a video coding device. a non-transitory computer-readable medium having a computer program product thereon, The program product, when executed by a processor, precedes a video coding device. A computer program stored on a non-transitory computer-readable medium that performs any of the methods of the present invention. It comprises computer-executable instructions.
[0022] In an embodiment, the present disclosure provides a method for identifying a flag and an IRAP picture and an associated IRAP picture. a plurality of coded pictures including one or more non-leading pictures, receiving means for receiving the bitstream; and when the flag is set to a first value, an IRA Any leading picture associated with a P picture must be an IRAP picture in decoding order. The flag is set to a second value. When a non-leading picture is associated with an IRAP picture, the non-leading picture is the first picture in decoding order to be associated with the IRAP picture. determining whether the flag is set to a first value; The second value is set to the IRAP picture, and any associated IRAP picture. and one or more non-leading pictures associated with the IRAP picture. decoding means for decoding in order and displaying as part of the decoded video sequence; and a forwarding means for forwarding the one or more decoded pictures to Includes Da.
[0023] Optionally, in any of the preceding aspects, another implementation of the aspect is It is provided that the method is further configured to perform any of the methods of the aspects described above.
[0024] In an embodiment, the present disclosure provides an IRAP picture and an IRAP picture associated with the IRAP picture. Coding for a video sequence comprising a plurality of pictures including at least one non-leading picture A decision means for determining the order of IRAP packets and a flag encoding method for encoding the flags into the bitstream and Any preceding picture associated with the picture must be an IRAP picture in coding order. The flag is set to the first value when the frame is before all non-leading pictures associated with the frame. A non-leading picture is associated with the IRAP picture in coding order. When it is before the first leading picture, the flag is set to the second value, and it is an IRAP picture, IRAP Any preceding pictures associated with the picture, and the IRAP picture associated with the In order to encode one or more non-leading pictures into a bitstream in coding order, encoding means for encoding the bitstream; and storage means for storing the bitstream for communication to a decoder. The encoder includes a stage.
[0025] Optionally, in any of the preceding aspects, another implementation of the aspect is further characterized in that the encoder It is provided that the device is further configured to perform the method of any of the preceding aspects.
[0026] For clarity, any one of the above embodiments may be considered a new implementation within the scope of this disclosure. The present invention may be combined with any one or more of the other aforementioned embodiments to create an embodiment.
[0027] These and other features will become more apparent from the following detailed description taken in conjunction with the accompanying drawings and claims. This will be more clearly understood.
[0028] For a more complete understanding of the present disclosure, reference is now made to the accompanying drawings and detailed description, in which: Reference is made to the following brief description, in which like reference numerals represent like parts. [Brief explanation of the drawings]
[0029] [Figure 1] 1 is a flowchart of an exemplary method for coding a video signal. [Figure 2] 1 is a schematic diagram of an example encoding and decoding (codec) system for video coding. [Figure 3] FIG. 1 is a schematic diagram illustrating an exemplary video encoder. [Figure 4] FIG. 1 is a schematic diagram illustrating an exemplary video decoder. [Figure 5] FIG. 1 is a schematic diagram illustrating an example coded video sequence with leading pictures. [Figure 6A] 1A-1C are schematic diagrams collectively illustrating examples of interlaced video coding. [Figure 6B] 1A-1C are schematic diagrams collectively illustrating examples of interlaced video coding. [Figure 6C]1A-1C are schematic diagrams collectively illustrating examples of interlaced video coding. [Figure 7] 1 is a schematic diagram illustrating an example coded video sequence utilizing both interlaced video coding and leading pictures. [Figure 8] FIG. 1 is a schematic diagram illustrating an example bitstream configured to include both interlaced video coding and leading pictures. [Figure 9] 1 is a schematic diagram of an exemplary video coding device. [Figure 10] 1 is a flowchart of an example method for encoding a video sequence and leading pictures into a bitstream involving interlaced video coding. [Figure 11] 1 is a flowchart of an example method for decoding a video sequence and leading pictures from a bitstream involving interlaced video coding. [Figure 12] 1 is a schematic diagram of an example system for coding a video sequence and leading pictures into a bitstream involving interlaced video coding. DETAILED DESCRIPTION OF THE INVENTION
[0030] Illustrative implementations of one or more embodiments are provided below, but the disclosed system Any system and / or method, whether currently known or existing, It should be understood at the outset that this disclosure may be implemented using any number of techniques. Even in this case, the following, including the exemplary designs and implementations illustrated and described herein, should not be limited to the illustrative implementations, diagrams, and techniques illustrated therein, This invention may be modified within the scope of the appended claims, along with their full range of equivalents.
[0031] The following terms are defined as follows, unless used herein to the contrary: Specifically, the following definitions are intended to facilitate a more clear understanding of the present disclosure: However, the terms may be explained differently in different contexts. Therefore, the following definitions The definitions given for such terms in this specification should be considered supplementary. These definitions should not be construed as limiting other definitions of the descriptions provided.
[0032] A bitstream is a video data stream that is compressed for transmission between an encoder and a decoder. The encoder converts the video data into a bitstream. and a device configured to utilize an encoding process to compress the performs the decoding process to reconstruct the video data from the bitstream for display. The flag is set by the encoder during encoding. The bits or bits coded into the bitstream signal the mechanism used. Since is a group of bits, the mechanism to be utilized by the decoder during decoding is Intra prediction refers to accurately reconstructing video data from a picture stream. A picture can be reconstructed without reference to other pictures by referencing itself. Inter prediction is a mechanism for coding a picture by referencing one or more other pictures. Intra random access points are used to code pictures. Intraprediction (IRAP) pictures are coded according to intra prediction, and the coded video A leading picture is a picture that serves as the starting point for a sequence. It is coded after the associated IRAP picture in order, but not in output order. A non-leading picture, also called a trailing picture, is a picture that precedes the associated IRAP picture. The picture is the picture that comes after the IRAP picture in both coding order and output order. Interlaced video coding is the first time in the first picture. coding a first field of video data and coding the first field of video data at a second time in a second picture; Coding a second field of video data in the image to give the impression of an improved frame rate To give the first field and the second field, we present them in a single frame. A frame is a video coding mechanism that combines A complete image intended for complete or partial display to the user at the corresponding moment. The pictures are interlaced video, where the pictures are fields of a frame. The parameter set is used to determine the time course of a coded video. Contains data such as flags and other parameters for the corresponding section of the sequence. The part of the bitstream that signals the sequential field flags (f field_seq_flag) is used for interlaced video and in coding order Signals when a non-leading picture is positioned between the IRAP picture and a leading picture. This is a flag.
[0033] The following acronyms are used in this document: Coding Tree Block (CTB), Coding Tree Block (CTB), Coding Tree Block (CTB), Coding Tree Block (CTB), Coding Tree Block (CTB ... Coding Tree Unit (CTU), Coding Unit (CU), Coded Video Video Sequence (CVS), Joint Video Experts Team (JVET), Motion Constrained Tileset Maximum Transmission Unit (MCTS), Maximum Transmission Unit (MTU), Network Abstraction Layer (NAL), Picture Order Count (POC), Raw Byte Sequence Payload (RBSP), Sequence Parameter Set (SPS), and and Working Draft (WD).
[0034] Many video formats are used to reduce the size of video files while minimizing data loss. For example, video compression techniques can be used to compress the spatial (e.g., image) performing intrapicture prediction and / or temporal (e.g., interpicture) prediction, This may include reducing or removing data redundancy in a video sequence. For the base video coding, a video slice (e.g., a video picture or A video image (part of a video picture) may be partitioned into video blocks, which are called tree blocks. block, coding tree block (CTB), coding tree unit (CTU), code They may also be called coding units (CUs) and / or coding nodes. Video blocks in a track-coded (I) slice are coded relative to neighboring blocks in the same picture. It is coded using spatial prediction with respect to reference samples within the block. Video in an inter-coded unidirectionally predicted (P) or bidirectionally predicted (B) slice of A block can be spatially predicted with respect to reference samples in neighboring blocks of the same picture, or Coding is achieved by utilizing temporal prediction with respect to reference samples in other reference pictures. A picture may be called a frame and / or an image, and a reference picture may be used. The frames are sometimes called reference frames and / or reference pictures. The prediction results in a predicted block that represents the image block. The residual data is used to represent the original image block. It represents the pixel difference between the inter-coded block and the predicted block. A block is associated with a motion vector pointing to a block of reference samples that form the predicted block. , and residual data indicating the difference between the coded block and the predicted block. Intra-coded blocks are coded according to the intra-coding mode and For further compression, the residual data is coded in the pixel domain. These result in residual transform coefficients that can be quantized. The quantized transform coefficients may first be arranged in a two-dimensional array. The entropy coding is performed by: Such video compression techniques can be applied to achieve further compression. It will be discussed in detail.
[0035] To ensure that the encoded video can be decoded correctly, the corresponding video Video is encoded and decoded according to a video coding standard. The standard is International Telecommunication Union (ITU) Standardization Sector (ITU-T) H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Motion Picture Experts Group (MPEG)-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, ITU-T H.264 or Advanced Video Coding (AVC), also known as ISO / IEC MPEG-4 Part 10, and High Efficiency Video Coding, also known as ITU-T H.265 or MPEG-H Part 2 AVC includes Scalable Video Coding (SVC), Multiview Video Coding (MVC), and Multiview Video Coding (MVC) and Multiview Video Coding plus Depth (MVC+D), as well as three HEVC includes extensions such as 3D AVC (3D-AVC). HEVC also includes Scalable HEVC (SHVC), Multiview It includes extensions such as Multi-Voltage HEVC (MV-HEVC), and 3D HEVC (3D-HEVC). The Joint Video Experts Team (JVET) is a leading provider of Versatile Video Coding (VVC) and The development of a video coding standard called VVC was initiated. VVC is included in the Working Draft (WD). This includes JVET-M1001-v7.
[0036] Video coding systems utilize IRAP and non-IRAP pictures. The IRAP pictures are random access pictures for a video sequence. A picture coded according to inter prediction that serves as an access point In intra prediction, blocks of a picture are predicted by other blocks in the same picture. This is the same as a non-IRAP picture that uses inter prediction. In inter prediction, the blocks in the current picture are predicted by a different A block is coded by reference to other blocks in a different reference picture. Since an IRAP picture is coded without reference to other pictures, the first IRAP picture is coded without reference to other pictures. Therefore, the decoder can decode any IRAP picture without decoding the picture. In contrast, non-IRAP pictures can start decoding the video sequence at Since pictures are coded with reference to other pictures, decoders generally Decoding of a video sequence cannot begin at this point. The data picture buffer (DPB) can be refreshed. This is because the IRAP picture is coded The beginning of a video sequence (CVS) is the beginning of a video sequence, and a picture in the CVS is a picture in the previous CVS. Therefore, IRAP pictures do not refer to inter prediction related codes. It can also stop the occurrence of framing errors, because such errors are propagated through the IRAP picture. However, IRAP pictures are not large in terms of data size. Therefore, video sequences are generally coded To balance coding efficiency and functionality, fewer non-IRAP pictures are used, along with a larger number of non-IRAP pictures. For example, a 60-frame CVS contains one IRAP picture and and 59 non-IRAP pictures. Therefore, IRAP pictures are not included in the bitstream. Furthermore, the presence of IRAP pictures in the bitstream reduces the compression efficiency of the video stream. This penalty to compression efficiency is due to the fact that the Therefore, intra prediction requires many more bits than inter prediction. This is partly caused by the fact that IRAP pictures are used in decoding processes. This can refresh the process and remove the reference picture from the DPB. Reference pictures available for inter prediction when coding subsequent pictures This reduces the number of frames, temporarily reducing the efficiency of the inter prediction process.
[0037] Video coding systems may also utilize leading pictures. It is placed after the IRAP picture in the moving order and before the IRAP picture in the presentation order. A leading picture is a picture that is determined by the corresponding picture being efficiently moved from the IRAP picture. When a picture can be predicted reasonably well, the corresponding picture should be presented before the IRAP picture. Such pictures can be used as references for IRAP pictures for inter prediction. In order to allow it to be used as an IRAP picture, The decoder then uses the presentation order to create different presentation orders. The order of leading pictures and IRAP pictures can be swapped before the Random Access Skip Ahead (RASL) and Random Access Decode Ahead (RADL) pictures An IRAP picture may also depend on a picture before the IRAP picture. This means that IRAP pictures are skipped when they are used as random access points. This means that no other pictures are decoded and therefore decoding starts from the IRAP picture. This is because such other reference pictures are not available when the RADL picture is used. For reference, use the IRAP picture or other pictures between the RADL picture and the IRAP picture. Therefore, RADL pictures depend only on IRAP as a random access point. It is decoded even when used because coding starts at an IRAP picture. Is it guaranteed that every picture that a RADL picture can reference is decoded even when The video coding system uses the direct ordering of the IRAP pictures it references in decoding order. Any relevant trailing edge may then need to be positioned after the leading edge. The picture comes after the preceding picture in decoding order.
[0038] Video coding uses a wide range of mechanisms, for example, interlace coding. Coding involves coding a frame into more than one field and more than one picture. For example, a frame may be divided into an even field and an odd field. The even field of a star-cross frame contains samples from the even-numbered horizontal lines of the frame. The odd field of an interlaced frame contains the odd-numbered horizontal lines of the frame. As a specific example, the odd field contains the samples captured at the first time. The odd field can then be captured and stored in the first picture. The first field can be captured and stored in a second picture. Interlaced coding therefore enhances the sense of motion. It creates the impression of increased frame rate without increasing the sequence bandwidth. Race coding is natively supported by a standardized coding system However, interlaced coding is not always possible. video to indicate that the program is an interlaced coded bitstream. By utilizing syntax elements in the VUI, Such syntax elements are field_seq_flag, general_fr It may include the ame_only_constraint_flag.
[0039] Standardized video coding systems that utilize leading pictures are interlaced. Not configured to support video coding, e.g., VVC and HEVC. Coding order requires that the leading picture, if any, comes after the IRAP picture Then, after the leading pictures, there are non-leading / trailing pictures. Such an order ensures that non-leading pictures are positioned between the IRAP picture and the associated leading picture. However, in the context of interlaced video coding, the IRAP frame The frame is divided between two fields in two pictures. The first field is The first picture is coded as an IRAP picture. The second picture with the second field The picture is coded as a non-leading / trailing picture instead of an IRAP picture, This is because the second picture cannot be used as a random access point. This means that both pictures are needed to start decoding, so the first picture The two pictures that make up an IRAP frame are effectively split into two parts, one for each frame and the other for each picture. They should be positioned next to each other for efficient coding. , a non-leading picture with a second IRAP field next to an IRAP picture with a first IRAP field Positioning the channel violates the coding order of VVC and HEVC. This is because positioning such as places a non-leading picture before every leading picture.
[0040] Disclosed herein is a method for encoding interlaced video using advance picture encoding. This is a mechanism for constructing a video coding system that uses To implement interlaced video coding in a VVC system that uses row pictures, A flag may be used to indicate whether a picture is to be framed between the IRAP picture and any preceding picture. This can be used to signal to the decoder whether there are any non-leading pictures. The coder reads the flag and sets it to support interlaced video coding. You can adjust the order as desired. When the flag is set to the first value, such as 0, , if there is a leading picture, it precedes all of the non-leading pictures. The encoder ensures that a single non-leading picture is located between the IRAP picture and any leading picture. A flag can be set to a second value, such as 1, to indicate to the decoder that the In some instances, the non-leading pictures may not be positioned between leading pictures. For example, the sequential field flag (field_seq_flag) can be used for this purpose. This flag may be included in the Sequence Parameter Set (SPS) and is used to In the context of interlaced video, a frame may be multiple frames. Note that the number of pictures (e.g., two) may be included. Outside the context of racing video, the term frame is used because a frame contains a single picture. The terms frame and picture may be used interchangeably. The following use of the term interlaced coding is not in the context of interlaced coding. should not be considered limiting.
[0041] FIG. 1 is a flow chart of an exemplary operational method 100 of coding a video signal. Specifically, the video signal is encoded in an encoder. The encoding process includes the following steps: Compress the video signal by utilizing various mechanisms to reduce the video file size. The smaller the file size, the more compressed the video file will be sent to the user. The decoder then decodes the compressed data while reducing the associated bandwidth overhead. Decodes the compressed video file and restores the original video signal for display to the end user. The decoding process generally allows the decoder to reconstruct the video signal reliably. The encoding process is mirrored to allow for
[0042] In step 101, a video signal is input to an encoder. For example, may be an uncompressed video file stored in memory. Video files are files captured by a video capture device such as a video camera and used to record video. The video file may be encoded to support live streaming of the video. It may contain both audio and video components. The video components provide a visual experience when viewed in sequence. It contains a series of image frames that give the effect of motion. The frames are divided into luma components (or luma samples) A pixel is represented in terms of light, referred to herein as a pixel, and a chroma component (or color component). In some cases, the image contains pixels that are described in terms of colors, called frame samples. The frame may also include depth values to support three-dimensional viewing.
[0043] In step 103, the video is partitioned into blocks. The partitions are made up of the pictures of each frame. including subdividing cells into square and / or rectangular blocks for compression For example, High Efficiency Video Coding (HEVC) (also known as H.265 and MPEG-H Part 2) In a coding tree, a frame can first be divided into coding tree units (CTUs). The CTU is divided into blocks of a predetermined size (e.g., 64 pixels by 64 pixels). A CTU contains both luma and chroma samples. The coding tree is into blocks and then divide the blocks until a configuration is achieved that supports further encoding. For example, the luma component of a frame can be divided into individual The blocks can be further subdivided until they contain relatively uniform illumination values. The color components can be subdivided until each block contains relatively uniform color values. The segmentation scheme varies depending on the content of the video frame.
[0044] In step 105, the image block partitioned in step 103 is compressed. Various compression mechanisms are used, for example, inter-prediction and / or intra-prediction. Inter prediction can be used to predict when objects in a common scene appear in consecutive frames. It is designed to take advantage of the fact that objects in a frame of reference tend to The blocks depicting the body do not need to be repeated in adjacent frames. In other words, an object such as a table may remain in a constant position over multiple frames. Thus, the table is written once and adjacent frames can refer to the reference frame. A pattern matching mechanism can be used to match objects across multiple frames. Furthermore, a moving object may be captured in multiple frames, e.g., due to the object's motion or the camera's motion. As a specific example, a video may be represented across multiple frames. A motion vector may represent such a movement. The motion vectors can be used to describe the coordinates of the object in the frame. It is a two-dimensional vector that gives the offset to the coordinates of the object in the frame. Therefore, inter prediction uses motion vectors that indicate an offset from the corresponding block in the reference frame. We can encode the image blocks in the current frame as a set of vectors. .
[0045] Intra prediction encodes blocks within a common frame. It takes advantage of the fact that the chromatic and chromatic components tend to be densely packed in a frame. For example, green spots in a piece of wood tend to be positioned next to similar green spots. Intra prediction supports multiple directional prediction modes (e.g., 33 in HEVC), planar modes, and Direct Current (DC) mode. Directional mode uses the current block in the corresponding direction. The planar mode indicates that the sample in the neighboring block is similar / same as the sample in the row / column ( For example, a series of blocks along a plane are interpolated based on neighboring blocks at the ends of the rows. The planar mode can be essentially a relatively constant gradient by varying the value. By utilizing the DC mode, it shows smooth transitions of light / color across rows / columns. All neighbors associated with the angular direction of the directional prediction mode are used for field smoothing. Indicates that the block is similar / same as the mean value associated with the block's samples. Therefore, intra-predicted blocks are coded as various related prediction mode values instead of actual values. Furthermore, inter-predicted blocks can be represented by a In either case, the predicted block can be represented as a motion vector value. The blocks may not exactly represent the image blocks in some cases. are stored in the residual block. To further compress the file, the residual block is may be applicable.
[0046] In step 107, various filtering techniques can be applied. The data is applied according to the in-loop filtering method. Prediction can result in the creation of blocky images at the decoder. The prediction scheme encodes a block and then stores it for later use as a reference block. The in-loop filtering scheme can reconstruct the coded block. filter, deblocking filter, adaptive loop filter, and sample adaptive offset (S AO) filters are applied to the block / frame iteratively. These filters To eliminate such blocking artifacts, the file can be accurately reconstructed. Furthermore, these filters reduce artifacts in the reconstructed reference block. Since the artifacts are reduced, the coding is based on the reconstructed reference block. are less likely to produce additional artifacts in subsequent blocks where .
[0047] Once the video signal has been segmented, compressed and filtered, in step 109: The resulting data is encoded in a bitstream, which is the same as the bitstream discussed above. to support the encoded data and reconstruction of the appropriate video signal at the decoder. For example, such data may be data, prediction data, residual blocks, and coding instructions to the decoder. The bitstream may contain various flags. The bitstream can be broadcast to multiple decoders and / or stored in a It can also be multicast. Creating a bitstream is an iterative process. Therefore, steps 101, 103, 105, 107, and 109 are performed for a number of frames and blocks. The sequence shown in Figure 1 may occur sequentially and / or simultaneously across blocks. is presented for ease of discussion and to illustrate the video coding process. No particular order is intended to be limiting.
[0048] In step 111, the decoder receives the bitstream and begins the decoding process. Specifically, the decoder uses an entropy decoding method to decode the bitstream. In step 111, the decoding is performed by converting the received data into the corresponding syntax and video data. The decoder uses syntax data from the bitstream to determine the division into frames. This division must match the block division result in step 103. The entropy encoding / decoding as utilized in step 111 is not described herein. The encoder selects several possible values based on their spatial location in the input image. Many choices are made during the compression process, such as choosing a block partitioning scheme from a variety of options. Signaling of strict selection may utilize multiple bins. Herein, a bin is defined as: A binary value that is treated as a variable (e.g., a bit value that can change depending on the situation). Political coding eliminates all options that are not clearly feasible for a particular case. Allow the encoder to discard, leaving a set of acceptable choices. Then, for each The length of the codeword is the number of allowable choices. based on (e.g., one bin for two choices, two for three to four choices, etc. The encoder then encodes the codeword for the selected choice. This scheme reduces the size of the codeword, which gives a large probability of finding all possible choices. Rather than uniquely representing a choice from the set, it represents a choice from a small subset of acceptable choices. This is because the codeword is as large as desired to uniquely represent the choice of The set of acceptable choices is then determined in a manner similar to that of the encoder. By determining the set of allowable choices, the decoder decodes the codeword can be read to determine the selection made by the encoder.
[0049] In step 113, the decoder performs block decoding. Specifically, the decoder , and uses the inverse transform to generate a residual block. The decoder then The corresponding predicted block is used to reconstruct the image block according to the partition. The block is an intra-predicted block as generated by the encoder in step 105 and an intra-predicted block. The reconstructed image block may then be processed in step 111. The segment data determined in step (a) is positioned in the frame of the reconstructed video signal. The syntax for step 113 also uses the entropy as discussed above. The information can be signaled in the bitstream via bit coding.
[0050] In step 115, the encoder reconstructs the image in a similar manner to step 107. Filtering is performed on frames of the video signal. For example, noise suppression filters The blocking filter, deblocking filter, adaptive loop filter, and SAO filter This filter can be applied to a frame to remove noise artifacts. Once ringed, the video signal is transmitted in step 117 for viewing by the end user. and output to a display.
[0051] FIG. 2 illustrates an exemplary coding and decoding (codec) system for video coding. 1 is a schematic diagram of a codec system 200. Specifically, the codec system 200 is adapted to perform the method of operation 100. The codec system 200 provides the functionality to support the encoder and decoder. It is generalized to describe components used in both The lock system 200 operates as discussed with respect to steps 101 and 103 in the method of operation 100. The codec receives and segments a video signal, which results in a segmented video signal 201. The network system 200 then performs the steps 105, 107, and 109 of the method 100. When operating as such an encoder, the segmented video signal 201 is converted into coded video. When operating as a decoder, the codec system 200 , bits as discussed with respect to steps 111, 113, 115, and 117 of the method of operation 100. The codec system 200 generates an output video signal from the stream. Component 211, Transform Scaling and Quantization Component 213, Intra Picture The estimation component 215, the intra-picture prediction component 217, the motion compensation component component 219, motion estimation component 221, scaling and inverse transform component 229, The filter control analysis component 227, the in-loop filter component 225, the decoded picture Header buffer component 223, and header formatting and context It includes an adaptive binary arithmetic coding (CABAC) component 231. The components are combined as shown. In Figure 2, the black lines represent the movement of the data to be encoded / decoded. The dashed lines indicate the movement of control data that controls the operation of other components. The components of the clock system 200 may all reside in the encoder. , may include a subset of the components of the codec system 200. For example, The data includes an intra-picture prediction component 217, a motion compensation component 219, a scaling component 220, and a the decoding and inverse transform component 229, the in-loop filter component 225, and the These components may include a signal picture buffer component 223. It will be revealed.
[0052] The segmented video signal 201 is segmented into blocks of pixels by a coding tree. The coding tree is a representation of the various segments of a captured video sequence. Use the divide mode to subdivide blocks of pixels into smaller blocks of pixels. These blocks can then be further subdivided into smaller blocks. A link can be called a node on the coding tree. A larger parent node has a smaller The number of times a node is subdivided depends on the number of nodes in the node / coding tree. In some cases, the divided blocks are contained in a coding unit (CU). For example, a CU may contain a luma block, a red differential chroma (Cr) block, and a blue differential chroma (Cr) block. The lower part of the CTU contains the roma (Cb) block, along with the corresponding syntax instructions for the CU. The split mode can be two, three, or four, which change shape depending on the split mode used. or binary tree (BT), which is used to partition each node into four child nodes. The segmented video signal 201 may include a quadtree (TT), a quadtree (QT), and a quadtree (TT). a general purpose coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, a filter control analysis component 227, and The motion estimation component 221 then forwards the motion estimation signal to the motion estimation component 221.
[0053] The generic coder control component 211 performs the coding of the video sequence according to the application constraints. It is configured to make decisions regarding the coding of images into a bitstream. For example, the general coder control component 211 may be configured to handle bitrate / bitstream size versus reconstruction. Such decisions depend on storage space / bandwidth availability and image quality optimization. The general coder control component 211 also controls the buffer size. To mitigate the problem of underrun and overrun of the data, the To manage these issues, a general purpose coder control component is used. The component 211 manages the segmentation, prediction, and filtering by the other components. For example, the general coder control component 211 dynamically increases the compression complexity to improve resolution. Increase the compression complexity to improve bandwidth utilization, or decrease the resolution and bandwidth utilization. Therefore, the general coder control component 211 Controls other components of the Stem 200 to adjust the quality and bit rate of the video signal reconstruction. The general purpose coder control component 211 creates the control data and It controls the behavior of other components. The control data is used for header formatting and and the CABAC component 231 to generate the parameters for decoding at the decoder. is coded in the bitstream to signal the
[0054] The segmented video signal 201 also includes a motion estimation component 221 and a and motion compensation component 219. Each slice may be divided into multiple video blocks. and motion compensation component 219 for one or more blocks in one or more reference frames. Perform inter-predictive coding of the received video block relative to the time The codec system 200 performs multiple coding passes, for example For example, an appropriate coding mode may be selected for each block of video data.
[0055] The motion estimation component 221 and the motion compensation component 219 may be highly integrated. are shown separately for conceptual purposes. Motion estimation is the process of generating motion vectors that estimate the motion of video blocks. The vector is, for example, the coded object relative to the predicted block. The predicted block should be coded with respect to pixel differences. The predicted block is the block that is found to be a good match for the reference block. Such pixel differences can be calculated using the sum of absolute differences (SAD), sum of squared differences (SSD), or other differential measures. HEVC uses CTU, coding tree block ( It utilizes several coded objects, including CTB, and CU. For example, a CTU can be split into CTBs, and then the CTBs can be split into CBs for inclusion in a CU. A CU may be a prediction unit (PU) containing prediction data and / or a transform for the CU. The motion estimation component can be coded as a transform unit (TU) containing the residual data. By using rate-distortion analysis as part of the rate-distortion optimization process, For example, the motion estimation component 221 , determining multiple reference blocks, multiple motion vectors, etc. for the current block / frame. and selects the reference block, motion vector, etc. that has the best rate-distortion performance. The best rate-distortion performance is achieved by the quality of the video reconstruction (e.g., the data loss due to compression). The objective of this paper is to balance the amount of loss (e.g., the amount of loss) and coding efficiency (e.g., the size of the final encoding).
[0056] In some examples, the codec system 200 includes a decoded picture buffer component 2 23. For example, For example, the video codec system 200 may use quarter pixel positions, eighth pixel positions, or or other fractional pixel positions of the reference picture. Component 221 performs motion search for integer pixel positions and fractional pixel positions, The motion estimation component 221 may output a motion vector with fractional pixel accuracy. Inter-coding is performed by comparing the position of the predicted block in the reference picture with the position of the predicted block in the reference picture. The motion estimation controller calculates motion vectors for the PUs of the video blocks in the selected slice. The component 221 stores the calculated motion vectors in a header as motion data for encoding. Formatting and CABAC component 231, and motion compensation component 232. The output is sent to port 219.
[0057] The motion compensation performed by the motion compensation component 219 is 21 to fetch or generate a prediction block based on the motion vector determined by Again, the motion estimation component 221 and the motion compensation component 219 may involve: In some cases, the functionality may be integrated. Upon receiving the motion vector, the motion compensation component 219 calculates the predicted block to which the motion vector points. The residual video block may then be used to locate the current video being coded. Subtract the pixel values of the predicted block from the pixel values of the block to form pixel difference values Generally, the motion estimation component 221 performs the motion estimation for the luma component. The motion compensation component 219 performs motion estimation for both the chroma and luma components. The motion vector is calculated based on the luma component of the predicted block and the residual. The blocks are forwarded to the transform scaling and quantization component 213 .
[0058] The segmented video signal 201 includes an intra-picture estimation component 215 and an intra-picture estimation component 216. It is also sent to the picture prediction component 217. The intra-picture estimation component 215 and the intra-picture compensation component 219 The picture prediction component 217 may be highly integrated, but is shown separately for conceptual purposes. The intra-picture estimation component 215 and the intra-picture prediction component The motion estimation component 221 and the motion estimation component 222 are used to estimate the motion between frames, as described above. As an alternative to inter-prediction performed by the compensation component 219, The current block is intra-predicted with respect to the blocks in the frame. The picture estimation component 215 determines the image to be used to encode the current block. In some examples, the intra-picture estimation component determines the intra-prediction mode. 215 selects a mode for encoding the current block from a plurality of tested intra-prediction modes. An appropriate intra-prediction mode is selected. The selected intra-prediction mode is then used for encoding. The header is then forwarded to the header formatting and CABAC component 231 for further processing.
[0059] For example, the intra-picture estimation component 215 may use various tested intra-picture predictions. Rate-distortion analysis for the measured modes is used to calculate the rate-distortion values for the modes tested. The intra prediction mode with the best rate-distortion performance is selected from the set of intra prediction modes. In general, the encoded block and the original data that was encoded to produce the encoded block are the amount of distortion (or error) between the coded block and the uncoded block, Determine the bit rate (e.g., number of bits) used to produce the resulting block. The intra picture estimation component 215 determines which intra prediction mode is used for the block. The quantization of the various coded blocks is performed to determine which one gives the best rate-distortion value. In addition, the intra-picture estimation component 21 calculates the ratio from the distortion and rate. 5 uses depth modeling mode (DMM) based on rate-distortion optimization (RDO). The JPEG 2000 may be configured to code the depth block of the frame.
[0060] The intra-picture prediction component 217, when implemented on an encoder, The selected intra-prediction mode is determined by the picture estimation component 215. Generate a residual block from the predicted block based on , the residual block may be read from the bitstream. The residual block is represented as a matrix. The residual block contains the difference in values between the predicted block and the original block, which is then transformed. The image is forwarded to the transform scaling and quantization component 213. The intra-picture prediction component 215 and the intra-picture prediction component 217 are used to predict the luma and chroma components. It can work for both minutes.
[0061] The transform scaling and quantization component 213 further compresses the residual block. The transform scaling and quantization component 213 is configured to A transform such as the discrete cosine transform (DCT), discrete sine transform (DST), or a conceptually similar transform, is added to the residual block. to produce a video block with residual transform coefficient values. A digital transform, a subband transform, or other types of transforms may also be used. It may be transformed from the pixel value domain to a transform domain such as the frequency domain. The scaling component 213 also scales the transformed residual information, for example, based on frequency. Such scaling is configured to scale different frequency information differently. This involves applying a scale factor to the residual information so that it is quantized to a granularity. Transform scaling and quantization controls can affect the final visual quality of the constructed video. Component 213 also includes a quantization function for quantizing the transform coefficients to further reduce the bit rate. The quantization process involves determining the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting the quantization parameter. In some examples, the transform scaling and quantization component 213 then The quantized transform coefficients may be scanned through a matrix containing the quantized transform coefficients. Formatting and CABAC component 231 for bitstream is encoded as follows:
[0062] The scaling and inverse transform component 229 performs the transformation to support motion estimation. Apply the inverse operation of the transform scaling and quantization component 213. and inverse transform component 229 applies inverse scaling, transformation, and / or quantization. as a reference block that can be a predictive block for another current block, for example. Reconstruct the residual block in the pixel domain for later use. The motion compensation component 221 and / or the motion compensation component 219 compensate for the motion of subsequent blocks / frames. By adding the residual block back to the corresponding predicted block for use in the estimation The reference block can be calculated using the A filter is applied to the reconstructed reference block to reduce artifacts Such artifacts would otherwise occur when subsequent blocks are predicted. This can lead to inaccurate predictions (and further artifacts) in
[0063] The filter control analysis component 227 and the in-loop filter component 225 Apply filters to the residual blocks and / or reconstructed image blocks. For example, The transformed residual block from the scaling and inverse transform component 229 is In order to reconstruct the image blocks, an intra-picture prediction component 217 and / or The motion compensation component 219 may combine the corresponding predicted block with the corresponding predicted block. The filter may then be applied to the reconstructed image block. In some examples, the filter Instead, it can be applied to the residual block. Like the other components in Figure 2, the filter The control analysis component 227 and the in-loop filter component 225 are highly integrated. , may be implemented together, but are shown separately for conceptual purposes. The filters applied to the block are applied to a specific spatial region, and how such filters The filter control analysis component contains several parameters to adjust how the filter is applied. Component 227 is restructured to determine where such filters should be applied. The received reference block is analyzed and the corresponding parameters are set. Header formatting and CABAC components are used as filter control data for The in-loop filter component 225 then forwards the filter control data to the in-loop filter component 231. The filters are deblocking filters, noise suppression filters, and so on. Such filters may include a control filter, a SAO filter, and an adaptive loop filter. in the spatial / pixel domain (e.g. on the reconstructed pixel block), depending on the example , or may be applied in the frequency domain.
[0064] When acting as an encoder, the filtered reconstructed image blocks, residual The difference block and / or the predicted block are used later in the motion estimation as discussed above. The decoded picture is stored in the decoded picture buffer component 223 for use by the decoder and When operating as a decoded picture buffer component 223, a portion of the output video signal is The reconstructed and filtered blocks are stored and transferred to the display as The decoded picture buffer component 223 stores the predicted blocks, residual blocks, and / or can be any memory device capable of storing reconstructed image blocks.
[0065] The header formatting and CABAC component 231 is part of the codec system 200. receives data from various components of the The data is encoded into a coded bitstream. The 231 general control data and filter control data It generates various headers to encode control data such as intra-frame data. It includes prediction and motion data, as well as residual data in the form of quantized transform coefficient data. All prediction data is coded in the bitstream. The system performs all the steps desired by the decoder to reconstruct the original segmented video signal 201. Such information includes the intra-prediction mode index table (codeword (also called mapping table), definition of coding context for various blocks , an indication of the most probable intra prediction mode, an indication of partition information, etc. Such data can be encoded using entropy coding. For example, information can be provided by Context-Adaptive Variable Length Coding (CAVLC), CABAC, and Syntax-Based Context-adaptive binary arithmetic coding (SBAC), probability interval partitioned entropy (PIPE) coding Entropy coding is performed using a coding scheme or other entropy coding technique. Following entropy coding, the coded bitstream may be It may be transmitted to another device (e.g., a video decoder) or for later transmission. or may be archived for retrieval.
[0066] 3 is a block diagram illustrating an example video encoder 300. to implement the encoding functionality of the codec system 200 and / or the operating method 10. 0 may be utilized to implement steps 101, 103, 105, 107, and / or 109. The encoder 300 segments the input video signal and generates a video signal substantially similar to the segmented video signal 201. This results in a segmented video signal 301. The segmented video signal 301 is then compressed , are encoded into a bitstream by components of the encoder 300.
[0067] Specifically, the partitioned video signal 301 is subjected to intra-picture prediction for intra prediction. The intra-picture prediction component 317 then forwards the intra-picture prediction data to the intra-picture prediction component 317. The sub-picture estimation component 215 and the intra-picture prediction component 217 The segmented video signal 301 may also be stored in a decoded picture buffer component. The motion compensation component 321 performs inter-prediction based on a reference block in the motion compensation component 323. The motion compensation component 321 is a component of the motion estimation component 221 and the motion compensation component 322. The intra-picture prediction component 3 may be substantially similar to the compensation component 219. The prediction block and the residual block from the motion compensation component 321 are The image is forwarded to the transform and quantization component 313 for transformation and quantization of the image. The transform and quantization component 313 is a transform scaling and quantization component 2 13. The transformed and quantized residual block and the corresponding prediction block The block (together with associated control data) is then extracted for coding into the bitstream. The entropy coding component 331 then forwards the resulting image to the entropy coding component 331. Component 331 is essentially the same as Header Formatting and CABAC Component 231. It can be similar to.
[0068] The transformed and quantized residual block and / or the corresponding prediction block are motion compensated. Transformation and quantization for reconstruction into a reference block used by component 321 The transform component 313 also forwards the image to the inverse transform and quantization component 329. and the quantization component 329 is substantially the same as the scaling and inverse transform component 229. The in-loop filter in the in-loop filter component 325 may be Depending on the example, it may also be applied to the residual block and / or the reconstructed reference block. The in-loop filter component 325 controls the filter control analysis component 227 and the loop The in-loop filter component may be substantially similar to the in-loop filter component 225. The in-loop filter component 325 may include multiple filters such as those discussed with respect to the in-loop filter component 225. The filtered block may then be passed to the motion compensation component 321. The decoded picture buffer component 323 stores the image data for future use as a reference block. The decoded picture buffer component 323 stores the decoded picture buffer The component 223 may be substantially similar to the component 223.
[0069] 4 is a block diagram illustrating an example video decoder 400. To implement the decoding functionality of the codec system 200 and / or the steps of the operating method 100 The decoder 400 may be utilized to implement steps 111, 113, 115, and / or 117. receives the bitstream, for example from the encoder 300, and displays it to the end user. To do this, a reconstructed output video signal is generated based on the bitstream.
[0070] The bitstream is received by the entropy decoding component 433. The Tropey Decoding Component 433 supports CAVLC, CABAC, SBAC, PIPE coding, or other configured to implement an entropy decoding scheme, such as the entropy coding technique of For example, the entropy decoding component 433 may decode the encoded data in the bitstream. To provide a context for interpreting additional data that is coded as words, The decoded information may include general control data, filter control data, Segment information, motion information, prediction data, and quantized transform coefficients from the residual block. The quantized transform coefficients include any desired information for decoding the video signal. The residual blocks are then forwarded to the inverse transform and quantization component 429 for reconstruction into residual blocks. The inverse transform and quantization component 429 is similar to the inverse transform and quantization component 329. It could be.
[0071] The reconstructed residual block and / or predicted block are based on an intra prediction operation. and forwarded to the intra-picture prediction component 417 for reconstruction into image blocks. The intra-picture prediction component 417 is an intra-picture estimation component. The components may be similar to the intra-picture prediction component 215 and the intra-picture prediction component 217. The intrapicture prediction component 417 locates reference blocks within a frame. To do this, we use a prediction mode and apply the residual block to the result to generate an intra predicted image block. The reconstructed intra-predicted image block and / or residual block are The clock and corresponding inter prediction data are fed to the in-loop filter component 425. These are transferred to the decoded picture buffer component 423 via the buffer component 223 and in-loop filter component 225, respectively, The in-loop filter component 425 filters the reconstructed image blocks, The residual block and / or the predicted block are filtered, and such information is decoded. The decoded picture is stored in the picture buffer component 423. The reconstructed image block from the frame 423 is used for the motion compensation component for inter prediction. The motion compensation component 421 is connected to the motion estimation component 221 and / or or may be substantially similar to the motion compensation component 219. Specifically, the motion compensation component The component 421 generates a predicted block using a motion vector from a reference block. , and apply the residual block to the result to reconstruct the image block. The loop also passes the picture data to the decoded picture buffer component via the in-loop filter component 425. The decoded picture buffer component 423 can then forward the additional reconstruction It continues to store the image blocks that have been extracted, which can be reconstructed into frames via the segmentation information. Such frames may also be arranged in sequences. The sequences are then reconstructed. The output video signal is then output to the display.
[0072] FIG. 5 is a schematic diagram illustrating an exemplary CVS 500 with leading pictures. For example, the CVS 500 , an encoder, such as the codec system 200 and / or the encoder 300, according to the method 100. Furthermore, the CVS 500 may be used in conjunction with the codec system 200 and / or the decoder. The CVS 500 can be decoded by a decoder such as decoder 400. The CVS 500 is coded in decoding order 508. The decoding order 508 indicates how pictures are positioned in the bitstream. The pictures in CVS 500 are then output in presentation order 510. Presentation order 510 is , the picture should be displayed by the decoder to properly display the resulting video. For example, pictures in a CVS 500 may generally be positioned in presentation order 510. However, similar pictures may be placed closer together, for example to support inter-prediction. By placing some pictures in different positions, the coding efficiency is improved. Moving such a picture in this way results in a decoding order 508. In the example shown, the pictures are indexed in decoding order 508 from 0 to 4. In presentation order 510, the pictures at index 2 and index 3 are It has been moved before the picture at index 0.
[0073] The CVS 500 includes an IRAP picture 502. The IRAP picture 502 is a random access picture for the CVS 500. A picture coded according to intra prediction that serves as a starting point. Specifically, the blocks of the IRAP picture 502 have references to other blocks of the IRAP picture 502. An IRAP picture 502 is coded without reference to other pictures. Since the IRAP picture 502 is decoded without first decoding any other pictures, Therefore, the decoder starts decoding the CVS 500 at the IRAP picture 502. Furthermore, the IRAP picture 502 allows the DPB to be refreshed. For example, most pictures presented after the IRAP picture 502 are inter-predicted. Therefore, the IRAP picture 502 does not depend on the picture before it (for example, picture index 0). Therefore, the picture buffer is refreshed when the IRAP picture 502 is decoded. This has the effect of stopping any inter-prediction related coding errors. Is it possible that such an error cannot propagate through the IRAP picture 502? The IRAP picture 502 may include various types of pictures. The architecture is implemented as Instantaneous Decoder Refresh (IDR) or Clean Random Access (CRA). The IDR can be coded to start a new CVS500 and refresh the picture buffer. CRA launches new CVS500 Random access points can be accessed without having to refresh the picture buffer or In this way, the CRA The leading picture 504 associated with the CRA may refer to a picture before the IDR. The associated leading picture 504 may not refer to a picture before the IDR.
[0074] CVS 500 also includes various non-IRAP pictures. These are leading pictures 504 and trailing pictures. The leading picture 504 is positioned after the IRAP picture 502 in decoding order 508. 502, but positioned before the IRAP picture 502 in presentation order 510. The trailing picture 506 is positioned after the IRAP picture 502 in both decoding order 508 and presentation order 510. Both the leading picture 504 and the trailing picture 506 are most often interleaved. The last picture 506 is coded according to the prediction. The picture is coded with reference to a picture positioned after picture 502. The edge picture 506 can always be decoded once the IRAP picture 502 is decoded. The row picture 504 is a random access skip ahead (RASL) picture and a random access The IRAP picture 502 may include a Reverse Decodable Leading (RADL) picture. The IRAP picture 502 is coded by reference to the IRAP picture 502, but at a position after the IRAP picture 502. Since an IRAP picture depends on the previous picture, it is loaded into the IRAP picture 502. When the decoder starts decoding in this case, it cannot decode RASL pictures. Therefore, the RASL picture uses the IRAP picture 502 as a random access point. However, if the decoder detects a random access point, it is skipped and not decoded. When using the previous IRAP picture (before index 0 and not shown) as the RADL pictures are decoded and displayed. RADL pictures are IRAP pictures 502 and / or is coded with reference to a picture after IRAP picture 502, but in presentation order The RADL picture is positioned before the RAP picture 502. The RADL picture is the picture before the IRAP picture 502. Therefore, when the IRAP picture 502 is a random access point, the RADL picture The chat can be decrypted and displayed.
[0075] 6A-6C are schematic diagrams collectively illustrating an example of interlaced video coding. Interlaced video coding involves coding a first picture 6 as shown in Figures 6A and 6B. 6C. From the first picture 601 and the second picture 602, an interlaced video frame 604 as shown in FIG. 6C is generated. For example, interlaced video coding produces 0. When encoding a video containing frame 600 as part of method 100, codec system 20 0 and / or encoders such as encoder 300. The clock system 200 and / or a decoder such as the decoder 400 may be configured to process interlaced video frames. In addition, the interlaced video frame 600 may be decoded as follows: It may be encoded in a CVS such as CVS 500, as discussed in more detail below with respect to FIG.
[0076] When performing interlaced video coding, as shown in FIG. 6A, A field 610 is captured at a first time and encoded into a first picture 601 . The first field 610 contains a horizontal line of video data. The horizontal line of video data in the first picture 601 extends from the left boundary of the first picture 601 to the right boundary of the first picture 601. However, the first field 610 omits alternating rows of video data. In one example implementation, the first field 610 is captured by a video capture at a first time. Contains half of the video data captured by the capture device.
[0077] As shown in FIG. 6B, a second field 612 is captured at a second time, For example, the second time is a frame for a video There can be only a value set based on the rate set immediately after the first time. For example, 15 frames. For videos that are set to display at a frame rate called frames per second (FPS), The interval may be 1 / 15th of a second after the first time. As shown, the second field 612 6 includes a horizontal line of video data that complements a horizontal line of the first field 610 of picture 601. Specifically, the horizontal line of video data in the second field 612 is aligned with the left side of the second picture 602. The second field 612 extends from the right boundary of the first picture 602 to the right boundary of the second picture 602. In addition, the second field 612 includes the horizontal lines omitted by the first field 610. The horizontal lines contained in field 610 are omitted.
[0078] A first field 610 of a first picture 601 and a second field 612 of a second picture 602 is represented at the decoder as an interlaced video frame 600 as shown in FIG. 6C. Specifically, interlaced video frame 600 may be synthesized to show a first time 6. The first field 610 of the first picture 601 captured at a second time 6 includes a second field 612 of a second picture 602 captured at the same time. has visual effects that emphasize and / or exaggerate the movement. Once the sequence of interlaced video frames 600 is complete, the additional frames are actually encoded. This creates the impression that the video is encoded at an increased frame rate without the need for In this way, an interlaced video frame 600 is used. Video coding is a method for effectively increasing the video frame rate without increasing the video data size. Interlaced video coding can therefore , may improve the coding efficiency of the encoded video sequence.
[0079] FIG. 7 shows an example of an interlaced video frame 600. 7 is a schematic diagram illustrating an exemplary CVS 700 that utilizes both decoding and leading pictures. The CVS700 is quite similar to the CVS500, but it has some differences, such as the first picture 601 and the second picture 602. Modified to preserve leading pictures while encoding pictures with any field. For example, the CVS 700 may be implemented as a codec system 200 and / or an encoder according to the method 100. The CVS 700 may be encoded by an encoder such as the codec system 300. The encoded data may be decoded by a decoder such as the system 200 and / or the decoder 400.
[0080] CVS 700 has a decoding order 708 and a presentation order 710, which are the same as decoding order 508 and presentation order 510, respectively. The CVS 700 also operates in a manner very similar to the IRAP picture 702, the preceding picture The IRAP picture 502, the leading picture 504, and the trailing picture 706 are included. , and trailing picture 506. The difference is that IRAP picture 702, leading picture 704, and The first and last pictures 706 are all generated in the first field as described with respect to FIGS. 6A-6C. The computer can be implemented by utilizing fields 610 and 612 in a manner very similar to the first field. Each frame therefore contains two pictures. Therefore, CVS700 contains twice as many pictures as CVS500. Since each picture omits half the frame, it contains roughly the same amount of data as CVS500.
[0081] The problem with CVS700 is that the first field of intra-predictive coded data Then, the IRAP picture 702 is coded by including the A second field of predictively coded data is contained in the non-leading picture 703. The non-leading picture 703 is not an IRAP picture 702 because the decoder 3, because it is not possible to start decoding CVS700. This is because half of the frames associated with the VVC-based video stream are omitted. The decoding system may select the preceding picture immediately following the IRAP picture 702 in decoding order 708. This creates a problem as the user may be constrained to position 704.
[0082] This disclosure allows the CVS 700 to be utilized with VVC systems. Specifically, IRAP A single non-leading picture 703 is positioned between the picture 702 and the leading picture 704. A flag can be signaled to indicate when this is allowed. However, the non-leading pictures 703 and / or trailing pictures 706 are positioned between the leading pictures 704. Therefore, this flag can be constrained to prevent the decoding order 708 from being P-picture 702, a single non-leading picture 703, any leading picture 704 (e.g., leading picture (Image 704 is optional and may be omitted in some instances), then one or more trailing edge pictures. This flag may therefore indicate whether CVS 500 should be expected. , or CVS700. , field_seq_flag in the SPS can be used for purposes discussed below.
[0083] FIG. 8 is configured to include both interlaced video coding and leading pictures. 8 is a schematic diagram illustrating an exemplary bitstream 800 that may be used. is for decoding by the codec system 200 and / or decoder 400 according to the method 100. For this purpose, the codec system 200 and / or the encoder 300 may generate the In addition, the bitstream 800 may contain CVS 500 and / or 700. The stream 800 is a first stream that can be combined to create the interlaced video frame 600. 6. The bitstream 800 may include a first picture 601 and a second picture 602. It may include pictures 504 and / or 704.
[0084] The bitstream 800 includes an SPS 810, multiple picture parameter sets (PPSs) 811, and multiple The SPS 810 includes a slice header 815 and image data 820. A sequence common to all pictures in a coded video sequence Such data includes picture size, bit depth, coding tools, It may contain parameters, bit rate limits, etc. PPS 811 is a parameter that applies to the entire picture. Therefore, each picture in a video sequence can refer to a PPS 811. Each picture references a PPS 811, but in some instances a single PPS 811 may refer to multiple pictures. Note that the image may contain data for multiple similar pictures. , may be coded according to similar parameters. In such a case, a single PPS 811 PPS 811 may contain data for such similar pictures. Shows available coding tools, quantization parameters, offsets, etc. for each slice in the image. The slice header 815 contains parameters specific to each slice in the picture. Thus, there is one slice header 815 for each slice in the video sequence. The slice header 815 contains slice type information, a picture order count (POC), a reference Reference picture list, prediction weights, tile entry points, deblocking parameters, etc. The slice header 815 may also be referred to as a tile group header in some contexts. Note that it can be called.
[0085] The image data 820 is a video signal that is coded according to inter-prediction and / or intra-prediction. The video data includes the corresponding transformed and quantized residual data. The sequence includes a number of frames 821. The frames 821 correspond to corresponding frames in the video sequence. It is a complete image intended for complete or partial display to the user at a given moment. A frame 821 may contain one or more pictures 823. In most contexts, a frame 821 is In such a case, the pictures contained in a single access unit (AU) are Picture 823 image / frame 821. However, in the context of interlaced video, The pixel 823 is a pixel of the horizontal line included in an AU, such as the first field 610 or the second field 612. Therefore, frame 821 is a field. When used, a picture 823 can be generated from two pictures 823. A picture 823 can be divided into one or more slices 82 5. A slice 825 is exclusively contained in a single Network Abstraction Layer (NAL) unit. An integer number of complete tiles or an integer number of consecutive complete codings of a picture 823 to be included A tree unit (CTU) can be defined as a row (e.g., within a tile). The channel 725 is further divided into CTUs and / or coding tree blocks (CTBs). The CTU / CTB is further divided into coding blocks based on the coding tree. The coding block can then be encoded / decoded according to a prediction mechanism.
[0086] The bitstream 800 may include a field_seq_flag 827. The field_seq_flag 827 is a CVS500 As shown in FIG. 1, any preceding picture associated with an IRAP picture is coded When it precedes in order all non-leading pictures associated with the IRAP picture, This flag may be set to a first value, indicating that a non-leading picture, as indicated in CVS 700, In coding order, the first leading picture associated with the IRAP picture This means that there are no leading pictures between the first and last leading pictures in decoding order. When not positioned, it can be set to a second value, in which case the IRAP picture is The non-leading picture that contains the first field and precedes the first leading picture is the second leading picture of the frame. In the example shown, field_seq_flag 827 may be included in SPS 810. As a typical example, field_seq_flag 827 specifies a picture 823 representing a field of frame 821. It may be set to 1 to indicate that the loaded video sequence contains represents a coded video sequence with pictures 823 each representing a complete frame 821. It may be set to 0 to indicate that the field_seq_flag8 27 and decoding the IRAP picture and one or more non-leading pictures. The first field from the picture and the second field from the non-leading picture before the first leading picture. This should include when to interlace fields to create a single frame. Therefore, the field_seq_flag 827 determines whether the preceding picture is related to the This allows interlaced video coding to be used. Utilizing d_seq_flag 827 enhances the functionality of the encoder and / or decoder. Furthermore, using field_seq_flag827 is a This allows for an effective increase in frame rate without significantly increasing the amount of data required for By doing so, the coding efficiency of the bitstream 800 can be improved. Utilizing the field_seq_flag 827 allows for the processing of This may reduce processor, memory, and / or network transmission resource usage.
[0087] Now, the above information will be explained in more detail later in this specification. IRAP pictures are Although it provides various useful features, it also creates a penalty in compression efficiency. The presence of this noise can cause a spike in the bit rate. This penalty to compression efficiency can be For example, an IRAP picture is an intra-predicted picture. Therefore, IRAP pictures have more features to represent compared to inter-predicted pictures. Furthermore, the presence of IRAP pictures can destroy temporal prediction. because the decoder can refresh the decoding process when it receives an IRAP picture. , which results in the removal of previous reference pictures in the DPB. This allows for inter prediction. Such a system has access to fewer reference pictures when performing coding. Therefore, the coding of pictures that come after the IRAP picture in decoding order is less efficient. It could be.
[0088] Among the picture types used as IRAP pictures, IDR pictures are the only ones that are not part of other pictures. It may utilize different signaling and derivations compared to the IEEE 802.11a type. Some of the differences are: When signaling and / or deriving the POC value for an IDR picture, the highest The most significant bits (MSB) part may not be derived from the previous key picture. Instead, The MSB may be set equal to 0. Furthermore, the slice header of an IDR picture may contain the reference picture May not contain information to assist the decoder in performing the management. CRA, back end, and other picture types such as Temporal Sub-Layer Access (TSA) reference pictures Information such as the reference picture set (RPS) or reference picture list is included in the slice header. The picture marking process can be used for the picture marking process. as being either used for or not used for reference , is used to determine the status of a reference picture in the DPB. However, the IDR For pictures, such information is not used for reference purposes. The presence of an IDR indicates that the decoding process should simply mark all reference pictures. Therefore, it may not be signaled.
[0089] Additionally, leading pictures may be associated with an IRAP. Leading pictures are, in decoding order, Pictures that come after its associated IRAP picture but before the IRAP picture in output order Depending on the coding structure and picture reference structure, the leading picture may be further divided into two Two types of pictures can be distinguished: the first type of picture known as a RASL picture may not be decoded correctly when the decoding process starts at the associated IRAP picture. This means that the picture that precedes the IRAP picture in decoding order is a leading picture. This is possible because these leading pictures are coded with reference to the RADL picture. The second type of picture, known as a picture, is an IRAP picture, to which the decoding process is related. This is a leading picture that will be correctly decoded even if it starts in the , these leading pictures directly or indirectly precede the IRAP picture in decoding order. This is possible because the image is coded without reference to any other picture. In this video coding system, the RASL picture associated with an IRAP picture is It is constrained to be before the RADL picture associated with the same IRAP picture in the input order. can be.
[0090] IRAP pictures and leading pictures are generated by the system level application. The NAL unit types may be given different NAL unit types so that they can be easily identified by the The video splicer interprets detailed syntax elements in the coded bitstream. It is possible to understand the coded picture type without having to consider: for example, a splice , including determining the RASL picture and the RADL picture from the trailing edge picture, It may be necessary to identify the IRAP picture from the picture and identify the leading picture. A trailing picture is associated with an IRAP picture and is located after the IRAP picture in output order. The current picture is the picture that is being played back if the current picture is an IRAP picture in decoding order. and before any other IRAP picture in decoding order, Therefore, the NAL units corresponding to the IRAP pictures and the preceding pictures are related. Providing types supports the functionality of such applications.
[0091] In some video coding systems, the IRAP picture and the preceding picture NAL unit types may include: Broken Link Access with Leading Picture A BLA picture (BLA_W_LP) can be followed by one or more preceding pictures in decoding order. A BLA with RADL (BLA_W_RADL) is a NAL unit for a single A BLA picture may be followed by one or more RADL pictures but may not be followed by a RASL picture. A BLA without leading pictures (BLA_N_LP) is an NAL unit for This is a NAL unit of a BLA picture that is not followed by a preceding picture. IDR with RADL (IDR_W_RADL) may be followed in decoding order by one or more RADL pictures, but not by any RASL pictures. An IDR without a leading picture (IDR_N_LP) is an NAL unit of an IDR picture. , is an NAL unit of an IDR picture that is not followed by a preceding picture in decoding order. A CRA picture that may be followed by leading pictures, including ASL and / or RADL pictures RADL is the NAL unit of a RADL picture. RASL is the NAL unit of a RASL picture. It is a NAL unit.
[0092] Other video coding systems use the following NAL units for IRAP and leading pictures: IDR_W_RADL is used when one or more RADL pictures follow in decoding order. It is an NAL unit of an IDR picture that may be in the frame but may not be followed by an IDR picture. R_N_LP is a NAL unit of an IDR picture that is not followed by a preceding picture in decoding order. A CRA may be followed by a preceding picture such as a RASL picture and / or a RADL picture. RADL is the NAL unit of a RADL picture. RASL is the NAL unit of a RASL picture. It is a NAL unit of Kucha.
[0093] For bitstream conformance, some constraints are imposed, e.g., on HEVC and / or VVC systems. The constraints may be applied to the leading pictures in the stem. Each picture in the bitstream other than the first picture in the bitstream shall be A picture can be considered to be related to the preceding IRAP picture. When a picture is an IR picture, it shall be an RADL or RASL picture. When it is the last picture of an AP picture, the picture is not a RADL or RASL picture. If a picture is a leading picture of an IRAP picture, the picture is In order, it precedes all trailing pictures associated with the same IRAP picture. RASL pictures shall not be associated with IDR pictures. RADL pictures shall not be associated with IDR pictures. _N_LP shall not be associated with an IDR picture with nal_unit_type equal to IRAP. Random access is achieved by discarding all access units before the access unit. Note that the access can be performed at the position of the IRAP access unit. Such random access occurs at the IRAP picture and all subsequent non-IRAP pictures in decoding order. Each parameter set available can result in the correct decoding of the picture. Assuming that such a parameter set should be activated, bit Such data can be generated either in the stream or by external means such as user input. Furthermore, random access can be performed using the IRAP picture that precedes the IRAP picture in decoding order. Any picture that precedes an IRAP picture in output order and It shall precede any RADL picture associated with the picture. Any RASL picture that is associated with a CRA picture in output order Any RADL picture associated with a CRA picture shall precede any RADL picture. The CRA picture is the first IRAP picture in output order that precedes the CRA picture in decoding order. It shall be later.
[0094] Therefore, the bitstream conformance constraint on leading pictures as described above is , may conflict with interlaced video coding schemes. The conflicts are as follows: When trace coding is used, the two fields of an IRAP picture are both IRAP Instead, only the first field is marked as an IRAP picture. The second field of the picture is marked as the trailing picture. The interlaced last picture containing the field is an interlaced IRAP picture in decoding order. This applies to interlaced IRAP pictures and interlaced This is because the trailing picture in the sequence forms a complete frame. If it is after an IRAP picture, then if the picture is a preceding picture of an IRAP picture , a picture is a sequence of all trailing pictures associated with the same IRAP picture in decoding order. This violates the constraint that states that the IRAP and Whether there are any leading pictures associated with it and whether all leading pictures have been considered This may assist an external entity, such as a video splicer, in efficiently determining whether However, these constraints cannot simply be removed. Such external entities include: Starting from the IRAP picture, the picture immediately following the IRAP picture is the trailing picture. If the IRAP picture is a picture with no leading pictures associated with it, the external entity Therefore, it can be determined that all preceding pictures associated with the IRAP picture To find the structure, an external entity must use the IRAP Without the above constraints, the external encoder can find the first trailing picture after the picture. The entity then performs the following steps to find all preceding pictures associated with the IRAP picture: It may be necessary to seek up to the IRAP picture.
[0095] Generally, this disclosure describes a method for handling leading pictures associated with an IRAP picture. More specifically, the present disclosure provides a method for efficiently extracting leading pictures associated with an IRAP picture. Supports efficient coding of interlaced video content while efficiently locating and identifying This description of the technique is based on the VVET standard of ITU-T and ISO / IEC. C standard, however, the techniques are applicable to other video codec standards as well. It can be used.
[0096] To solve the problems enumerated above, the present disclosure includes the following aspects, which individually: For example, the preceding picture associated with the IRAP picture may be applied to the The pictures may be positioned consecutively in decoding order with no non-leading pictures between them. In addition, the following constraints apply for bitstream adaptation of IRAP pictures and leading pictures: Let picA and picB be the first preceding pictures associated with an IRAP picture, respectively. In such a case, let picA be the last leading picture and picB be the last leading picture in decoding order. Assume that there are no pictures that are after and before picB that are not leading pictures.
[0097] The following constraints may also apply: field_seq_flag is set equal to 0 and the current picture is I If the current picture is a leading picture associated with a RAP picture, the current picture is In this case, it precedes all non-leading pictures associated with the same IRAP picture. If field_seq_flag is set equal to 1, then picA and picB are The first and last leading pictures associated with the IRAP picture are In such a case, there is at most one non-leading picture before picA in decoding order. picA comes after picA in decoding order and comes before picB in decoding order. Assume that there are no non-leading pictures.
[0098] The following constraints may also be applied: If general_frame_only_constraint_flag is equal to 1 and the If this picture is the leading picture associated with the IRAP picture, the current picture is , in decoding order, before all non-leading pictures associated with the same IRAP picture. Otherwise, if general_frame_only_constraint_flag is equal to 0, , picA and picB are the first and second IRAP pictures associated with an IRAP picture in decoding order, respectively. In such a case, there are many pictures before picA in decoding order. Assume that there is only one non-leading picture at a time, and that it is after picA in decoding order. Assume that there are no non-leading pictures prior to picB in the order.
[0099] In one example, the NAL unit type of an IRAP picture is the destination NAL unit type associated with the IRAP picture. It provides enough information to determine whether a row picture is present. The following methods can be used: The NAL unit type CRA_NUT indicates that the leading picture is a CRA picture. Replaced by CRA_W_LP to indicate that it is associated with the leading picture and / or It may be replaced by CRA_N_LP to indicate that the image is not associated with a CRA picture. In the example, the NAL unit types IDR_W_RADL, IDR_N_LP, and CRA_NUT indicate that the leading pictures Replaced by IRAP_W_LP to indicate that it is associated with an IRAP picture, and It may be replaced with IRAP_N_LP to indicate that the pixel is not associated with an IRAP picture.
[0100] In one example, the following applies to CRA_W_LP, CRA_N_LP, IDR_W_RADL, and IDR_N_LP: An IDR picture with NalUnitType equal to IDR_N_LP is included in the bitstream. Not associated with any leading picture present. NalUnitType equal to IDR_W_RADL is not associated with any RASL picture present in the bitstream, It can be associated with a RADL picture in the bitstream. NalUnitTyp equal to CRA_N_LP A CRA picture with e is not associated with any leading pictures present in the bitstream. A CRA picture with a NalUnitType equal to CRA_W_LP is the preceding picture in the bitstream. It can be associated with Kucha.
[0101] In one example, the mapping of the above NAL unit types to Stream Access Point (SAP) types is The mapping is as follows: IDR_N_LP and CRA_N_LP are associated with SAP type 1, and IDR_W_RA DL is associated with SAP type 2 and CRA_W_LP is associated with SAP type 3.
[0102] In one example, the following may apply to IRAP_W_LP and IRAP_N_LP: IRAP pictures with equal NalUnitType are the leading pictures present in the bitstream. An IRAP picture with NalUnitType equal to IRAP_W_LP is not associated with It may be associated with a leading picture in the bit stream.
[0103] In one example, the mapping of the above NAL unit types to SAP types is as follows: IRAP_N_LP is associated with SAP type 1 and IRAP_W_LP is associated with SAP type 3.
[0104] In one example, to determine whether there is a leading picture associated with the IRAP: The device may check the NAL unit type of the IRAP picture. When an IRAP picture can be associated with one or more preceding pictures, the following steps are performed: The device can be used to find all leading pictures associated with the IRAP. The picture immediately following the IRAP picture in decoding order is the non-preceding picture. If it is a non-preceding picture immediately following an IRAP picture, the picture may be ignored. The presence of a picture indicates that the bitstream is an interlaced video coding bitstream. Note that the next picture may be a leading picture. The process checks for the next picture until it encounters the first non-leading picture. can continue.
[0105] 9 is a schematic diagram of an example video coding device 900. The device 900 is suitable for implementing the disclosed examples / embodiments as described herein. The video coding device 900 transmits data upstream over a network. and / or a transmitter and / or a receiver for communicating the , downstream port 920, upstream port 950, and / or transceiver The video coding device 900 also includes a transceiver unit (Tx / Rx) 910. a processor 930 including a logic unit and / or central processing unit (CPU) for and memory 932 for storing the data. Video coding device 900 also includes an electrical Components, Optical-Electrical (OE) Components, Electro-Optical (EO) Components, and / or or telecommunications networks, optical communications networks, or wireless communications networks Upstream port 950 and / or downstream port 960 for communication of data over the network The video codec may include a wireless communication component coupled to a wireless port 920. The operating device 900 may also include input and / or output ports for communicating data to and from a user. The I / O device 960 may include a display for displaying video data. may include output devices such as a display, a speaker for outputting audio data, etc. The I / O device 960 also includes input devices such as a keyboard, a mouse, a trackball, and and / or a corresponding interface for interacting with such output devices. It can be seen.
[0106] The processor 930 is implemented in hardware and software. A processor consists of one or more CPU chips, cores (for example, as a multi-core processor), and field processors. Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), and Digital Signal Processors The processor 930 may be implemented as a digital signal processor (DSP). , Tx / Rx 910, upstream port 950, and memory 932. 930 includes a coding module 914. The coding module 914 includes a CVS 500, an Utilizing Interlaced Video Frame 600, CVS 700, and / or Bitstream 800 The disclosed implementations described herein, such as methods 100, 1000, and 1100, The coding module 914 may implement any other implementation described herein. Furthermore, the coding module 914 may also implement the method / mechanism of the codec system 200, encoder 300, and / or decoder 400. For example, The module 914 determines when there are non-leading pictures between the IRAP picture and the set of leading pictures. A flag can be set in the SPS to indicate whether it is positioned. , the coding module 914 may include additional functions and and / or coding efficiency provided by video coding device 900. Thus, coding module 914 improves the functionality of video coding device 900. It addresses the problems inherent in video coding techniques as well as the advantages of video coding. Transformation module 914 performs transformations of video coding device 900 into different states. Alternatively, the coding module 914 may be stored in a memory 932 and executed by a processor 930. as instructions to be executed (e.g., computer program products stored on non-transitory media) It may be implemented as a product.
[0107] The memory 932 may be a disk, tape drive, solid state drive, read-only memory, or ROM, Random Access Memory (RAM), Flash Memory, Ternary Content Addressable Memory (TCAM), It comprises one or more memory types, such as static random access memory (SRAM). The memory 932 stores programs when they are selected for execution. and for storing instructions and data that are read during program execution. It can be used as an overflow data storage device.
[0108] FIG. 10 illustrates an example of an interlaced video coding system, such as an interlaced video frame 600. A video sequence such as CVS500 and / or 700 with a leading picture is 10 is a flowchart of an exemplary method 1000 for encoding a video signal into a bitstream such as the stream 800. The method 1000 includes a codec system 200, an encoder 300, and a , and / or may be utilized by an encoder such as video coding device 900. .
[0109] The method 1000 includes an encoder receiving a video sequence including a plurality of pictures, For example, the system may decide to encode the video sequence into a bitstream based on user input. In step 1001, the encoder generates a The video sequence is divided into IRAP pictures and IRAP pictures. The video sequence comprises a plurality of pictures, including one or more non-leading pictures associated with the video sequence. The sequence also optionally includes one or more (e.g., a group) of preceding pictures. obtain.
[0110] In step 1003, the encoder encodes the flag into the bitstream. The flag indicates that any preceding pictures associated with the IRAP picture are coded In order, all non-leading pictures associated with the IRAP picture as in CVS500 This may be set to a first value when the video sequence is interleaved. The flag also indicates that non-leading pictures do not contain source video. In CVS700, before the first leading picture associated with the IRAP picture When the flag is set to the second value, the bitstream The frame also has a gap between the first and last leading picture in coding order. This means that the video sequence may be constrained so that leading pictures are not positioned in the In one particular example, an encoder may encode an SPS as a binary The data can be encoded into the bitstream and the flags can be encoded into the SPS. In this example, the flag is field_seq_flag. For example, field_seq_flag specifies whether a field is Set to 1 to indicate that the coded video sequence contains a picture representing Furthermore, the field_seq_flag may be used to indicate that a picture representing a frame is It can be set to 0 to indicate that the interlaced video sequence contains A flag is set to indicate that decoding is used in the bitstream. Therefore, when an IRAP picture contains the first field of a frame, and and the non-leading picture before the first leading picture contains the second field of the frame. For example, the first field and first The second field from the non-leading picture that precedes the leading picture is shown in FIG. 6A-6C. Alternating lines of video data representing a single interlaced video frame, as shown may include:
[0111] In step 1005, the encoder generates an IRAP picture, which is associated with the IRAP picture. Any leading picture and one or more non-leading pictures associated with the IRAP picture , can be encoded into a bitstream in coding order. The method then stores the bitstream for communication to a decoder in step 1007. It is possible.
[0112] FIG. 11 illustrates an example of an interlaced video coding system, such as an interlaced video frame 600. A video sequence such as CVS500 and / or 700 with a leading picture is coded by bit coding. 11 is a flowchart of an exemplary method 1100 for decoding from a bitstream such as stream 800. The method 1100 is a method for implementing the method 100, which is implemented by the codec system 200, the decoder 400, and the like. , and / or may be utilized by a decoder such as video coding device 900.
[0113] The method 1100 may, for example, result in a coding representation of a video sequence as a result of the method 1000. Step 1000 may begin when the decoder begins receiving the bitstream of encoded data. In step 1101, a decoder generates flags, IRAP pictures, and frames associated with the IRAP pictures. and a plurality of coded pictures including one or more non-leading pictures to be coded. The video sequence also receives one or more of the preceding pictures (e.g. For example, a group).
[0114] In step 1103, the decoder determines whether the flag is set to a first value, as indicated in CVS 500. When the IRAP picture is set, any leading pictures associated with the IRAP picture are , it can be determined that the IRAP picture is before all non-leading pictures associated with it. This indicates that the video sequence does not contain interlaced video. In step 1105, the decoder determines whether the flag is set to a second value, as shown in CVS 700. When a non-leading picture is the first leading picture associated with an IRAP picture in decoding order, When the flag is set to the second value, the decoded image can be determined to be in front of the picture. The reader also determines the first and last leading pictures in coding order. It can be decided that no leading pictures are positioned between the video sequences. A specific example would be a bitstream containing interlaced video. The system may include an SPS, and the flag may be obtained from the SPS. field_seq_flag, e.g., a video coded picture representing a field field_seq_flag may be set to 1 to indicate that the sequence contains field_ to indicate that the coded video sequence contains pictures that represent seq_flag can be set to 0. Therefore, if interlaced video coding is enabled, the bit A flag may be set to indicate that it is used in the stream. An IRAP picture contains the first field of a frame and is a non-leading picture that precedes the first leading picture. When a row picture contains the second field of a frame, a flag may be set.
[0115] In step 1107, the decoder determines whether the IRAP picture is an IRAP picture or not based on the flag. any leading pictures associated with the IRAP picture, and one or more leading pictures associated with the IRAP picture Decode non-leading pictures in decoding order. For example, IRAP pictures, leading pictures ( Decoding the first picture (if any) and one or more non-leading pictures is described with respect to FIGS. 6A-6C. The first and second fields from the IRAP picture are used to create a single frame to be and the second field from the non-leading picture before the early leading picture. In step 1109, the decoder may include downloading the decoded video sequence. The one or more decoded pixels resulting from step 1107 are then displayed as part of the The chat can be forwarded.
[0116] FIG. 12 illustrates an example of an interlaced video coding system, such as an interlaced video frame 600. A video sequence such as CVS500 and / or CVS700 with a preceding picture is An exemplary system for coding a The system 1200 is a schematic diagram of the codec system 200, the encoder 300, the decoder 300, and the decoder 300. encoder and decoder, such as decoder 400, and / or video coding device 900. Additionally, system 1200 may be implemented by performing methods 100, 1000, and / or 1100. It can be used when carrying out
[0117] The system 1200 includes a video encoder 1202. The video encoder 1202 encodes IRAP pictures. and one or more non-leading pictures associated with the IRAP picture. a decision module 1201 for determining a coding order for a video sequence comprising: The video encoder 1202 further comprises a flag encoding unit 1206 for encoding the flag into the bitstream. and a coding module 1203 for coding any preceding picture associated with the IRAP picture. , of all non-leading pictures associated with the IRAP picture in coding order When it is before, the flag is set to the first value and the non-leading picture is and before the first leading picture associated with the IRAP picture, the flag is set to The encoding module 1203 further encodes the IRAP picture, the IRAP picture and the associated Any leading picture attached to the IRAP picture, and one or more non-leading pictures associated with the IRAP picture For encoding pictures into a bitstream in coding order. The video encoder 1202 also stores the bitstream for communication to the decoder. The video encoder 1202 further comprises a storage module 1205 for storing the bitstream. a transmission module 1207 for transmitting the video encoding stream to a video decoder 1210. Reader 1202 may be further configured to perform any of the steps of method 1000.
[0118] The system 1200 also includes a video decoder 1210. The video decoder 1210 receives flags and IRAP signals. A set of multiple codes including one or more non-leading pictures associated with an IRAP picture and an IRAP picture. a receiving module for receiving a bitstream, the receiving module comprising: a picture encoded in the bitstream; The video decoder 1210 further includes an IRAP 1211 when the flag is set to a first value. Any leading pictures associated with the picture must be related to the IRAP picture in decoding order. a determining module 1213 for determining that the picture is before all non-leading pictures associated with the picture; The determination module 1213 further determines whether the flag is set to a second value. The picture is located in decoding order before the first leading picture associated with the IRAP picture. The video decoder 1210 further determines the decoding order based on the flag. In the introduction, an IRAP picture, any preceding pictures associated with the IRAP picture, and a decoding module for decoding one or more non-leading pictures associated with an IRAP picture; The video decoder 1210 further includes a video decoder 1215 as part of the decoded video sequence. a transfer module 1217 for transferring one or more decoded pictures for display by The video decoder 1210 is further configured to perform any of the steps of the method 1100. It can be configured.
[0119] Except for the wire, wiring, or other medium between the first and second components When there are no intervening components, the first component directly connects to the second component. A wire, wiring, or other connection is made between the first and second components. When there is an intervening component other than the medium, the first component The term "coupled" and its variations refer to a directly coupled component. The use of the term "about" includes both indirect and indirect binding. Unless otherwise specified, the range means a range including ±10% of the number that follows.
[0120] The steps of the exemplary methods described herein may not necessarily be performed in the order described. It is understood that no particular order is required and that the order of steps in such methods is merely exemplary. It should also be understood that additional steps may be included in such a method. Optionally, some steps may be omitted or completed in a manner consistent with various embodiments of the present disclosure. They may be combined.
[0121] Although several embodiments have been provided in this disclosure, the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of this disclosure. It can be understood that the present examples are illustrative rather than limiting. and the intention is not to be limited to the details given herein. For example, in another system, various elements or components are combined or integrated. or some features may be omitted or not implemented.
[0122] Additionally, various embodiments may be described or illustrated as separate or distinct. The techniques, systems, subsystems, and methods described herein may be modified without departing from the scope of this disclosure. , may be combined or integrated with other systems, components, techniques, or methods. Other examples of variations, substitutions, and alterations are ascertainable by one of ordinary skill in the art and are disclosed herein. may be made without departing from the spirit and scope thereof. [Explanation of symbols]
[0123] 200 Codec System 201 segmented video signal 211 General-Purpose Coder Control Component 213 Transform Scaling and Quantization Components 215 Intra-picture Estimation Component 217 Intra-picture Prediction Component 219 Motion Compensation Component 221 Motion Estimation Component 223 Decoded Picture Buffer Component 225 In-Loop Filter Components 227 Filter Control Analysis Component 229 Scaling and Inverse Transformation Components 231 Header Formatting and CABAC Components 300 Encoder 301 Segmented Video Signal 313 Transform and Quantize Components 317 Intra-picture Prediction Component 321 Motion Compensation Component 323 Decoded Picture Buffer Component 325 In-Loop Filter Components 329 Inverse Transform and Quantization Components 331 Entropy Coding Component 400 decoder 417 Intra-picture Prediction Component 421 Motion Compensation Component 423 Decoded Picture Buffer Component 425 In-Loop Filter Components 429 Inverse Transform and Quantization Components 433 Entropy Decoding Component 500 CVS 502 IRAP Picture 504 Leading Picture 506 Rear Picture 508 Decoding Order 510 Presentation order 600 interlaced video frames 601 First Picture 602 Second Picture 610 First Field 612 Second Field 700 CVS 702 IRAP Picture 703 Non-leading pictures 704 Leading Picture 706 Rear Picture 708 Decoding Order 710 Presentation order 725 slices 810 SPS 811 Picture Parameter Set (PPS) 815 slice header 820 image data 821 frames 823 Pictures 825 slices 900 Video Coding Device 910 Transceiver Unit (Tx / Rx) 914 Coding Module 920 downstream ports 930 processor 932 memory 950 upstream ports 960 I / O devices 1200 System 1201 Decision Module 1202 Video Encoder 1203 Encoding Module 1205 Memory Module 1207 Transmitter 1210 Video Decoder 1211 Receiver, receiving module 1213 Decision Module 1215 Decryption Module 1217 Transfer Module
Claims
1. 1. A method implemented in a decoder, comprising: The decoder receiver receives a flag and an Intra Random Access Point (IRAP) a plurality of pictures including one or more non-leading pictures associated with the IRAP picture; receiving a bitstream comprising the coded pictures of When the flag is set to a first value, the IRA Any leading pictures associated with a P picture are in decoding order adjacent to the IRAP picture. determining that the picture is before all non-leading pictures associated with the picture; When the flag is set to a second value, the processor , which in decoding order precedes the first leading picture associated with said IRAP picture and determining based on whether the flag is set to the first value or the second value, The processor selects the IRAP picture, any preceding picture associated with the IRAP picture, a picture, and the one or more non-leading pictures associated with the IRAP picture; and decoding in code order.
2. When the flag is set to the second value, the processor and a leading picture is positioned between the first leading picture and the last leading picture. The method of claim 1 , further comprising determining that the
3. The bitstream includes a sequence parameter set (SPS), and the flag is The method according to any one of claims 1 to 2, wherein the protein is obtained from PS.
4. Claims 1 to 3, wherein the flag is a sequential field flag (field_seq_flag).
10. The method according to any one of the preceding claims.
5. When the coded video sequence includes pictures representing fields, the fi eld_seq_flag is set to 1 and the coded video sequence represents a frame.
5. Any of claims 1 to 4, wherein the field_seq_flag is set to 0 when the picture contains a 10. The method according to claim 1.
6. The IRAP picture includes the first field of a frame and is located before the first leading picture.
6. The method of claim 1, wherein the non-leading picture in the frame comprises a second field of the frame. The method according to any one of claims 1 to 4.
7. The step of decoding the IRAP picture and the one or more non-leading pictures comprises: the first field from a P picture and the non-leading picture that precedes the first leading picture. and the second field from the picture to create a single frame. The method according to claim 1 , comprising the step of:
8. 1. A method implemented in an encoder, comprising: The encoder processor generates an Intra Random Access Point (IRAP) picture. a plurality of pictures including one or more non-leading pictures associated with said IRAP picture; determining a coding order for a video sequence comprising the structure; encoding, by the processor, the flag into a bitstream; , any preceding picture associated with said IRAP picture is , when it is before all non-leading pictures associated with said IRAP picture, the leading picture is set to a first value, and the non-leading picture is located in coding order after the IRAP picture. the flag is set to a second value when the picture is before the first leading picture associated with the The steps are: The processor selects the IRAP picture, any associated IRAP picture, and the leading picture and the one or more non-leading pictures associated with the IRAP picture. encoding the data into the bitstream in coding order; A memory coupled to the processor may be configured to transmit the bitstream for communication to a decoder. and storing the stream.
9. When the flag is set to the second value, the first 9. The method according to claim 8, wherein no leading picture is positioned between the row picture and the last leading picture. The method described.
10. The bitstream includes a sequence parameter set (SPS), and the flag is 10. The method according to claim 8, wherein the data is encoded into PS.
11. 8 to 1, wherein the flag is a sequential field flag (field_seq_flag).
10. The method of any one of claims 1 to 0.
12. When the coded video sequence includes pictures representing fields, the fi eld_seq_flag is set to 1 and the coded video sequence represents a frame.
12. Any of claims 8 to 11, wherein the field_seq_flag is set to 0 when the picture contains a 10. The method according to claim 1.
13. The IRAP picture includes the first field of a frame and is located before the first leading picture.
13. The method of claim 8, wherein the non-leading picture in the frame comprises a second field of the frame. The method according to any one of claims 1 to 4.
14. the first field from the IRAP picture and the first preceding picture The second field from the non-leading picture is a single interlaced video frame.
14. A method according to any one of claims 8 to 13, wherein the video data includes alternating lines of video data representing a game. 。
15. A processor, a receiver coupled to the processor, and a memory coupled to the processor. a transmitter coupled to the processor, and a transmitter configured to perform the method according to any one of claims 1 to 14. A video coding device that can be used with HDMI.
16. a computer program product for use by a video coding device; a non-transitory computer-readable medium, the computer program product comprising: When executed by the processor, the video coding device is A computer readable medium for executing the method of any one of claims 1 to 5, A non-transitory computer-readable medium comprising computer-executable instructions.
17. Flag and Intra Random Access Point (IRAP) picture and said IRAP picture a plurality of coded pictures including one or more non-leading pictures associated with the coded pictures; receiving means for receiving a bitstream, When the flag is set to a first value, any IRAP picture associated with the IRAP picture The leading pictures are all non-leading pictures associated with the IRAP picture in decoding order. Determined to be in front of the picture, When the flag is set to a second value, the non-leading picture is the preceding picture in decoding order. determine that the IRAP picture is before the first preceding picture associated with the IRAP picture; and a determining means for determining whether based on whether the flag is set to the first value or the second value, an IRAP picture, any preceding pictures associated with said IRAP picture, and said IRAP a decoding order for decoding the one or more non-leading pictures associated with the picture; and a decoding means for Decode one or more decoded pictures for display as part of a decoded video sequence. and a transfer means for transferring the decoder.
18. The decoder is further configured to perform the method according to any one of claims 1 to 7.
18. The decoder of claim 17, wherein
19. Intra Random Access Point (IRAP) pictures and associated for a video sequence comprising a plurality of pictures including one or more non-leading pictures that are determining means for determining the coding order of the encoding a flag into the bitstream associated with the IRAP picture; Any preceding picture that is associated with the IRAP picture in coding order When the flag is set to a first value, the flag is set to a second value when the non-leading picture is in front of all non-leading pictures in the frame. The leading picture is the first picture in coding order associated with the IRAP picture. encoding, wherein the flag is set to a second value when it is before a leading picture; the IRAP picture, any preceding pictures associated with the IRAP picture, and the one or more non-leading pictures associated with the IRAP picture in coding order encoding the bitstream in and encoding means for performing and storage means for storing said bitstream for communication to a decoder. , encoder.
20. The encoder is further configured to perform the method of any one of claims 8 to 14.
20. The encoder of claim 19, configured as follows:
Citation Information
Patent Citations
Method and System for Processing HEVC Coded Video in Broadcast and Streaming Applications
US20160234527A1
Methods and systems of coding a predictive random access picture using a background picture
US20170105004A1
Encoding device, encoding method, transmission device, decoding device, decoding method, and reception device
WO2015025747A1
Image processing device and method
WO2018150934A1