Enhanced signaling of missing or corrupted samples in media files
By refining the damage signaling of NAL units in media files, the problem of unclear distinction between NAL unit damage in existing technologies is solved, achieving more accurate video decoding and rendering.
Patent Information
- Application Number
- CN202380069025.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-05
- Filing Date
- 2023-09-25
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2043-09-25
AI Technical Summary
In existing technologies, the signaling design for lost or corrupted samples in media files is not detailed enough, making it impossible to effectively distinguish different types of NAL unit damage, which affects the accuracy and consistency of video decoding and rendering.
By introducing a new signaling mechanism, it is clearly indicated that the NAL unit, which is a similar parameter set required for decoding the bitstream in the media unit, is corrupted or lost. It distinguishes between SEI messages that affect HRD consistency and unnecessary SEI messages, and clarifies the damage status of the NAL unit header, strip header, and picture header, as well as the damage or loss of the strip reference picture.
It improves the accuracy and consistency of the video decoding process, ensuring that the decoder can correctly handle damaged or missing NAL units, thereby enhancing the reliability and quality of video rendering.
Smart Images

Figure CN119948881B_ABST
Abstract
Description
[0001] Cross-referencing of related patent applications
[0002] This patent application claims the benefit of U.S. Provisional Patent Application No. 63 / 410,509, filed September 27, 2022, and U.S. Provisional Patent Application No. 63 / 413,551, filed October 5, 2022, the teachings and disclosures of which are incorporated herein by reference in their entirety. Technical Field
[0003] This disclosure relates to the generation, storage, and consumption of digital audio and video media information in file formats. Background Technology
[0004] Digital video consumes the largest share of bandwidth on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demands of digital video are likely to continue to grow. Summary of the Invention
[0005] The first aspect relates to a method for processing video data, comprising: determining one or more indications of media units, wherein a first indication indicates that one or more network abstraction layer (NAL) units of a similar set of parameters required for decoding a bitstream in associated data are corrupted; and performing a conversion between visual media data and a visual media data file based on the first indication.
[0006] The second aspect relates to an apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein, when executed by the processor, the instructions cause the processor to perform the method described in any of the preceding aspects.
[0007] The third aspect relates to a non-transitory computer-readable medium including a computer program product for use by a video codec apparatus, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium, such that when the computer-executable instructions are executed by a processor, the video codec apparatus performs the method described in any of the preceding aspects.
[0008] The fourth aspect relates to a non-transitory computer-readable recording medium storing a bitstream of video generated by a video processing apparatus performing a method, wherein the method includes: determining one or more indications of media units, wherein a first indication indicates that one or more network abstraction layer (NAL) units of a similar set of parameters required for decoding the bitstream in associated data are corrupted; and generating the bitstream based on the determination.
[0009] The fifth aspect relates to a method for storing a bitstream of video, comprising: determining one or more indications of media units, wherein a first indication indicates that a network abstraction layer (NAL) unit of one or more similar parameter sets required for decoding the bitstream in associated data is corrupted; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium. Attached Figure Description
[0010] To gain a more complete understanding of this disclosure, reference is now made to the following brief description in conjunction with the accompanying drawings and specific embodiments, wherein like reference numerals denote like parts.
[0011] Figure 1 This is a block diagram illustrating an exemplary video processing system.
[0012] Figure 2 This is a block diagram of an exemplary video processing apparatus.
[0013] Figure 3 This is a flowchart of an exemplary method for video processing.
[0014] Figure 4 This is a block diagram illustrating an exemplary video encoding and decoding system.
[0015] Figure 5 This is a block diagram illustrating an exemplary encoder.
[0016] Figure 6 This is a block diagram illustrating an exemplary decoder.
[0017] Figure 7 This is a schematic diagram of an exemplary encoder.
[0018] Figure 8 This is a flowchart of an exemplary method for video processing. Detailed Implementation
[0019] It should be understood from the outset that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods can be implemented using any number of techniques, whether currently known or not yet developed. This disclosure should not be limited in any way to the illustrative implementations, drawings, and techniques shown below, including the exemplary designs and implementations shown and described herein, but modifications can be made within the full scope of the appended claims and their equivalents.
[0020] Chapter headings are used in this disclosure for ease of understanding and not to limit the applicability of the techniques and embodiments disclosed in each chapter to that chapter. Furthermore, the use of H.266 terminology in some descriptions is merely for ease of understanding and not to limit the scope of the disclosed techniques. Therefore, the techniques described herein are also applicable to other video codec protocols and designs. In this disclosure, edits to text are shown in bold italics to indicate canceled text and bold text to indicate added text, in contrast to the Multi-Functional Video Codec (VVC) specification and / or the International Organization for Standardization (ISO) Basic Media File Format (ISOBMFF) standard.
[0021] 1. Preliminary Discussion
[0022] This disclosure relates to media file formats. Specifically, this disclosure relates to signaling of lost or corrupted samples in media files. For media file formats, such as ISOBMFF or its extensions, such as Network Abstraction Layer (NAL) unit structured video in ISOBMFF, these examples can be applied individually or in various combinations.
[0023] 2. Abbreviation
[0024] Adaptive Color Transformation (ACT), Adaptive Loop Filter (ALF), Adaptive Motion Vector Resolution (AMVR), Adaptive Parameter Set (APS), Access Unit (AU), Access Unit Delimiter (AUD), Advanced Video Coding (AVC) as described in Rec.ITU-TH.264|ISO / IEC14496-10, Bidirectional Prediction (B), Bidirectional Prediction with CU-level Weights (BCW), Bidirectional Optical Flow (BDOF), Block-based Incremental Pulse Code Modulation (BDPCM), Buffer Period (BP), Context-based Adaptive Binary Arithmetic Coding (CABAC), Code Block (CB), Constant Bit Rate (CBR), Cross-Component Adaptive Loop Filter (CCALF), and Image Buffering for Coding / Decoding. Clean Random Access Buffer (CPB), Clean Random Access (CRA), Cyclic Redundancy Check (CRC), Code-Decoder Tree Block (CTB), Code-Decoder Tree Unit (CTU), Code-Decoder Unit (CU), Code-Decoder Video Sequence (CVS), Decoder Picture Buffer (DPB), Decoder Capability Information (DCI), Dependent Random Access Point (DRAP), Decoder Unit (DU), Decoder Unit Information (DUI), Exponential Golomb Code (EG), k-order Exponential Golomb Code (EGk), Bit Stream End (EOS), Padding Data (FD), First-In-First-Out (FIFO), Fixed Length (FL), Green, Blue and Red (GBR), General Constraint Information (GCI), Gradual Decoder Refresh (GDR), Geometric Partitioning Mode (GPM), as shown in Rec.ITU-T H.265|ISO / IEC 23008-2 describes the following: High-Efficiency Video Coding and Decoding (HEVC), Hypothetical Reference Decoder (HRD), Hypothetical Stream Scheduler (HSS), Intra-Frame (I), Intra-Frame Block Copy (IBC), Instantaneous Decode Refresh (IDR), Inter-Layer Reference Picture (ILRP), Intra-Frame Random Access Point (IRAP), Low-Frequency Inseparable Transform (LFNST), Minimum Probability Symbol (LPS), Least Significant Bit (LSB), Long-Term Reference Picture (LTRP), Luminance Map with Chroma Scaling (LMCS), Matrix-Based Intra-Frame Prediction (MIP), Maximum Probability Symbol (MPS), Most Significant Bit (MSB), Multiple Transform Selection (MTS), Motion Vector Prediction (MVP), Network Abstraction Layer (NAL), Output Layer Set (OLS), Operation Point (OP), Operation Point Information (OPI), Prediction (P), Picture Header (PH), and Picture Sequence Counting. (POC), Picture Parameter Set (PPS), Optical Flow Prediction Refinement (PROF), Picture Timing (PT), Picture Unit (PU), Quantization Parameter (QP), Random Access Decodable Preamble (RADL) Picture, Random Access Skip Preamble (RASL) Picture, Raw Byte Sequence Payload (RBSP), Red, Green, and Blue (RGB), Reference Picture List (RPL), Sample Adaptive Offset (SAO), Sample Aspect Ratio (SAR), Supplemental Enhancement Information (SEI), Strip Header (SH), Subpicture Level Information (SLI), Data Bit String (SODB), Sequence Parameter Set (SPS), Short-Term Reference Picture (STRP), Stepped Temporal Sublayer Access (STSA), Truncated Rice Code (TR), Variable Bit Rate (VBR), Video Coding Layer (VCL), Video Parameter Set (VPS), such as Rec.ITU-T Multifunctional Supplemental Enhancement Information (VSEI), Video Availability Information (VUI), as described in H.274|ISO / IEC 23002-7, and Multifunctional Video Coding and Decoding (VVC) as described in Rec.ITU-T H.266|ISO / IEC 23090-3.
[0025] 3. Further discussion
[0026] 3.1 Video codec standards
[0027] Video coding standards have evolved primarily through the development of standards by the International Telecommunication Union (ITU) Telecommunication Standardization Sector (ITU-T) and the International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC). ITU-T developed H.261 and H.263, ISO / IEC developed Moving Picture Experts Group (MPEG)-1 and MPEG-4 video, and the two organizations jointly developed the H.262 / MPEG-2 video and H.264 / MPEG-4 Advanced Video Codec (AVC) and H.265 / High Efficiency Video Codec (HEVC) standards [1]. Since H.262, video coding standards have been based on a hybrid video coding structure, which utilizes time prediction plus transform coding. Recently, the Multi-Functional Video Coding (VVC) standard (ITU-T H.266|ISO / IEC 23090-3) and its associated Multi-Functional Supplemental Enhancement Information (VSEI) standard (ITU-T H.274|ISO / IEC 23002-7) have been designed for a wide range of applications, including traditional uses such as television broadcasting, video conferencing, or playback from storage media, as well as newer and more advanced use cases such as adaptive bitrate streaming, video region extraction, synthesis and merging of content from multiple codec video bitstreams, multi-view video, scalable layered coding and decoding, and viewport-adaptive 360° immersive media. The Basic Video Coding (EVC) standard (ISO / IEC 23094-1) is another video coding standard recently developed by MPEG.
[0028] 3.2 File Format Standards
[0029] Streaming applications are typically based on Internet Protocol (IP), Transmission Control Protocol (TCP), and Hypertext Transfer Protocol (HTTP) delivery methods and often rely on file formats such as ISOBMFF[5]. One such streaming system is HTTP-based Dynamic Adaptive Streaming (DASH)[6]. To use video formats with ISOBMFF and DASH, a file format specification specific to that video format, also known as the Network Abstraction Layer File Format (NALFF)[7], will be required. This includes file format specifications for all NAL unit-based video codecs (such as AVC, HEVC, VVC, and their extensions) for encapsulating video content in ISOBMFF tracks and DASH representations and segments. Important information about the video bitstream (e.g., profiles, hierarchies and levels, and many other details) will need to be exposed as file format level metadata and / or DASH Media Presentation Description (MPD) for content selection purposes, such as for selecting appropriate media segments, for initialization at the start of a streaming session, and for stream adaptation during a streaming session. Similarly, in order to use an image format with ISOBMFF, a file format specification specific to that image format will be required, such as the AVC image file format and the HEVC image file format in [8].
[0030] 3.3 Parameter Set
[0031] AVC, HEVC, and VVC define parameter sets. Parameter set types include SPS, PPS, APS, and VPS. AVC, HEVC, and VVC all support SPS and PPS. HEVC introduces VPS, and both HEVC and VVC include VPS. AVC and HEVC do not include APS, but VVC includes VPS.
[0032] SPS is designed to carry sequence-level header information, while PPS is designed to carry infrequently changing image-level header information. Using SPS and PPS, infrequently changing information does not need to be repeated for each sequence or image, thus avoiding redundant signaling. Furthermore, the use of SPS and PPS enables out-of-band transmission of important header information, thereby not only avoiding the need for redundant transmission but also improving error recovery capabilities.
[0033] The purpose of introducing a VPS is to carry sequence-level header information common to all layers in a multi-layer bitstream.
[0034] The APS is introduced to carry this type of image-level or strip-level information, which requires a considerable number of bits for encoding and decoding. The APS can be shared by multiple images and can have a considerable number of different variables in a sequence.
[0035] 3.4VVC's APS, PH, DCI, and OPI NAL units
[0036] Several additional types of NAL units are introduced into VVC, including APS, PH, DCI, and OPI NAL units.
[0037] 3.4.1 Adaptive Parameter Set (APS)
[0038] The Adaptive Parameter Set (APS) transmits image and / or strip-level information that can be shared across multiple strips of an image and / or strips of different images. However, this information can change frequently across images, and the total number of variables can be high, making it unsuitable for inclusion in the PPS. The APS includes three types of parameters: Adaptive Loop Filter (ALF) parameters, Luminance Map with Chroma Scaling (LMCS) parameters, and scaling list parameters. The APS can be carried in two different NAL unit types, placed as a prefix or suffix before or after the associated strip. The latter can be helpful in ultra-low latency scenarios, for example, allowing the encoder to send the image strips before generating ALF parameters based on the image, which will then be used by subsequent images in the decoding order.
[0039] 3.4.2 Image Header (PH)
[0040] Each PU has a Picture Header (PH) structure. The PH exists in a separate PH NAL unit or is included in a stripe header (SH). If the PU consists of only one stripe, the PH can only be included in the SH. For design simplicity, in the codec layer video sequence (CLVS), the PH can only be entirely in the PH NAL unit or entirely in the SH. When the PH is in the SH, there is no PH NAL unit in the CLVS.
[0041] The PH (Programmable Message) is designed for two purposes. First, it helps reduce the signaling overhead of the SH (Search Engine) for each image containing multiple stripes by carrying all parameters with the same values for all stripes of the image, thus preventing the repetition of identical parameters in each SH. These include IRAP / GDR image indication, inter-frame / intra-frame stripe allow flags, and information related to POC (Programmable Context), RPL (Randomized Frame), deblocking filter, SAO (Single Aspect), ALF (Alternating Field), LMCS (Lapse Scale Controller), scaling list, QP increment, weighted prediction, codec block splitting, virtual boundary, co-occurring image, etc. Second, it helps the decoder identify the first stripe of each codec image containing multiple stripes. Since there is one and only one PH for each PU (Programmable Component), when the decoder receives a PH NAL unit, it easily knows that the next VCLNAL unit is the first stripe of the image.
[0042] 3.4.3 Decoding Capability Information (DCI)
[0043] A DCI NAL unit includes Bitstream Level Profile Level (PTL) information. A DCI NAL unit includes one or more PTL syntax structures that can be used during session negotiation between the sender and receiver of a VVC bitstream. When a DCI NAL unit is present in a VVC bitstream, each Output Layer Set (OLS) in the bitstream's CVS should conform to the PTL information carried in at least one PTL structure within the DCI NAL unit.
[0044] In AVC and HEVC, PTL information for session negotiation is available in the SPS (for HEVC and AVC) and VPS (for HEVC layered extensions). This design of transmitting PTL information for session negotiation in HEVC and AVC has drawbacks because the scope of the SPS and VPS is within the CVS, not the entire bitstream. Therefore, sender-receiver session initiation may be re-initiated during bitstream streaming at each new CVS. DCI solves this problem because it carries bitstream-level information, thus guaranteeing adherence to the indicated decoding capabilities until the end of the bitstream.
[0045] 3.4.4 Operation Point Information (OPI)
[0046] Both HEVC and VVC decoding processes have similar input variables to set the decoding operation point via the decoder's application programming interface (API), such as the target OLS and highest sublayer of the bitstream to be decoded. However, if layers and / or sublayers of the bitstream are removed during transmission or the device does not expose the decoder API to the application, it is possible that the decoder cannot be correctly informed of the operation point for processing a given bitstream. Therefore, the decoder may be unable to infer the properties of images in the bitstream, such as the appropriate buffer allocation for decoding images and whether to output separate images. To address this issue, VVC adds a pattern indicating these two variables in the bitstream through a newly introduced Operation Point Information (OPI) NAL unit. In the initial AU of the bitstream and in each CVS of the AU, the OPI NAL unit informs the decoder about the target OLS and highest sublayer of the bitstream to be decoded.
[0047] When an OPI NAL unit exists and the operation point is also provided to the decoder via decoder API information (e.g., the application may have more updated information about the target OLS and sublayers), the decoder API information takes precedence. In the absence of a decoder API and any OPI NAL units in the bitstream, appropriate fallback options are specified in the VVC to allow correct decoder operation.
[0048] 3.5 VUI and SEI messages
[0049] The VUI is a syntax structure sent as part of the SPS (and possibly in the HEVC VPS). The information carried by the VUI does not affect the standard decoding process, but it may be important for the correct rendering of the encoded video.
[0050] SEI assists in processes related to decoding, display, or other purposes. Like VUI, SEI does not affect standard decoding processes. SEI is carried within SEI messages. Decoder support for SEI messages is optional. However, SEI messages do affect bitstream consistency (e.g., if the syntax of SEI messages in the bitstream does not conform to the specification, the bitstream is not compliant with the specification), and some SEI messages are required in the HRD specification.
[0051] The VUI syntax and most SEI messages used by VVC are not specified in the VVC specification, but rather in the VSEI specification. The SEI messages required for HRD conformance testing are specified in the VVC specification. VVC v1 defines five SEI messages related to HRD conformance testing, and VSEI v1 specifies 20 additional SEI messages. The SEI messages carried in the VSEI specification do not directly affect the behavior of the conformance decoder and have been defined to be used in a codec-agnostic manner, thus allowing VSEI to be used with other video codec standards besides VVC in the future. The VSEI specification does not specifically mention VVC syntax element names, but rather variables, whose values are set in the VVC specification.
[0052] Compared to HEVC, VVC's VUI syntax structure focuses solely on information relevant to the correct rendering of the image and does not include any timing information or bitstream limitation indications. In VVC, the VUI is transmitted via signaling within the SPS, which includes a length field preceding the VUI syntax structure to signal the length of the VUI payload in bytes. This allows the decoder to easily skip information, and more importantly, similar to how SEI message syntax is extended, it allows for convenient future VUI syntax extensions by adding new syntax elements directly to the end of the VUI syntax structure.
[0053] The VUI syntax structure contains the following information:
[0054] ●The content is interwoven or progressive;
[0055] ●Does the content include frame-encapsulated stereoscopic video or projected omnidirectional video?
[0056] ●Aspect ratio of the sample;
[0057] ●Is the content suitable for overscan display?
[0058] ● Color description, including color primary colors, matrix, and transmission characteristics, is especially important for the ability to transmit ultra-high definition (UHD) and high definition (HD) color spaces and high dynamic range (HDR) through signals;
[0059] ● Chromaticity position relative to luminance (signaling is clarified for progressive content compared to HEVC).
[0060] When the SPS does not contain any VUI, the information is considered unspecified, and if the content of the bitstream is intended to be rendered on a display, the information must be transmitted via external means or specified by the application.
[0061] Table 1 lists the SEI messages specified for VVC v1, along with the specification containing their syntax and semantics. Of the 20 SEI messages specified in the VSEI specification, many are inherited from HEVC (e.g., the Padding Payload and two User Data SEI messages). Some SEI messages are essential for the proper processing or rendering of encoded or decoded video content. This is the case, for example, for SEI messages controlling display color capacity, content light level information, or alternative delivery characteristics, which are particularly relevant to HDR content. Other examples include isometric projection, spherical rotation, zone-by-zone encapsulation, or omnidirectional viewport SEI messages, all of which are related to signaling and processing of 360° video content.
[0062] Table 1: SEI Message List in VVC v1
[0063]
[0064]
[0065] The SEI messages used for VVC v1 are specified to include frame field information SEI messages, sample aspect ratio information SEI messages, and sub-picture level information SEI messages.
[0066] The Frame Field Information (SEI) message contains information indicating how associated pictures should be displayed (such as field parity or frame repeat period), the source scan type of the associated picture, and whether the associated picture is a copy of a previous picture. In previous video codec standards, this information was transmitted via signaling in the Picture Timing SEI message along with the timing information of the associated picture. However, it was observed that frame field information and timing information are two different types of information and do not necessarily need to be transmitted via signaling together. A typical example involves transmitting system-level timing information via signaling, but transmitting frame field information via signaling within the bitstream. Therefore, it was decided to remove frame field information from the Picture Timing SEI message and transmit it via signaling within a dedicated SEI message. This change allows the syntax of the frame field information to be modified to send additional and clearer instructions to the display, such as pairing fields together or providing more values for frame repeating.
[0067] The Sample Aspect Ratio (SEI) message can transmit different sample aspect ratios of different images within the same sequence via signaling, while the corresponding information contained in the VUI applies to the entire sequence. This can be relevant when using reference image resampling features with a scaling factor that causes different images in the same sequence to have different sample aspect ratios.
[0068] The Sub-Picture Level Information (SEI) message provides level information for a sequence of sub-pictures.
[0069] 3.6 Signaling for lost or damaged samples in ISOBMFF
[0070] Section 1 of the proposed ISOBMFF technical document [9] includes a design for signaling of lost or damaged samples using the following specifications:
[0071] ■ Category: CorruptedSampleInfoEntry()
[0072] Extend SampleGroupDescriptionEntry('corr')
[0073]
[0074] The value "corrupted" indicates the state of corruption of the associated data. A value of 0 indicates that all data is lost, and the associated data size (sample size or NAL size) should be 0. A value of 1 indicates that the data is corrupted, with no additional information about the corruption. A value of 2 indicates that the data is corrupted, with codec-specific information about the corruption. A value of 3 is reserved.
[0075] The `codec_specific_param` directive indicates codec-specific information about the corrupted codec. The codec format is one of the samples associated with this sample group description.
[0076] Note: The codec_specific_param information depends on the codec format. Each time the codec format is changed between samples, the file writer may need to add different CorruptedSampleInfoEntry() entries and associate them with the samples.
[0077] If the data is not associated with a CorruptedSampleInfoEntry, or if the data is associated with a description_group_index=0 via a sample group with grouping_type'corr', it means the data is not corrupted.
[0078] The handling of samples with corrupted values of 1 or 2 depends on the specific context and implementation.
[0079] For the NAL element (NALUFF) in ISOBMFF, please note:
[0080] For NALU-based codecs, we propose the following semantics:
[0081] For NALU-based video formats, the codec_specific_param field of CorruptedSampleInfoEntry is defined as a bitmask of the following flags, with the most significant bit at the beginning:
[0082] ●ParameterSetCorruptedFlag (value 0x00000001): Indicates that one or more parameter sets (DCI, VPS, SPS, PPS, APS, OPI?) in the associated data are corrupted;
[0083] ●SEICorruptedFlag (value 0x00000002): Indicates that one or more SEI messages in the associated data have been corrupted;
[0084] ●SliceHeaderCorruptedFlag (value 0x00000004): Indicates that one or more slice headers or image headers in the associated data are corrupted;
[0085] ●VCLCorruptedFlag (value 0x00000008): Indicates that the VCL data of one or more stripes in the associated data is corrupted;
[0086] ●OtherNonVCLNALCorruptedFlag (value 0x00000010): Indicates that one or more NAL cells of a different type than those described above are corrupted in the associated data.
[0087] A value of 0 for codec_specific_param indicates that there is no information available to describe the corruption.
[0088] Using grouping_type_parameter 'corr', CorruptedSampleInfoEntry can be used with sample groups of grouping_type 'nalm' and NALUMapEntry. The groupID of the NALUMapEntry mapping entry indicates the index starting from 1 in the sample group description of CorruptedSampleInfoEntry. A groupID of 0 indicates that no entry is associated (the identified data exists and has not been corrupted).
[0089] 4. Technical problems solved through publicly available technical solutions
[0090] In the following text, NAL units with similar terminology to parameter sets are collectively referred to as parameter set NAL units, DCI NAL units, and OPINAL units.
[0091] An exemplary design for corrupted signaling in media data includes five scenarios as part of the CorruptedSampleInfoEntry in NALFF: 1) corruption of NAL cells with similar parameter sets, 2) corruption of SEI messages, 3) corruption of stripe headers or picture headers, 4) corruption of VCL data for one or more stripes, and 5) corruption of other types of data. However, the following issues exist:
[0092] First, among NAL units with similar parameter sets of different types, some affect bitstream decoding, such as SPS and PPS, while others do not, such as DCI. Distinguishing between these NAL units is necessary when indicating damage to NAL units with similar parameter sets via signal transmission.
[0093] Second, SEI messages are contained within SEI NAL units, and each SEI NAL unit can contain one or more SEI messages. In addition to the contained SEI messages, each SEI NAL unit also contains header data. However, indications of corruption of the header data within SEI NAL units are not currently covered.
[0094] Third, some SEI messages, such as buffered SEI messages and image timing SEI messages, affect bitstream consistency by assuming the reference decoder (HRD) operation, while other SEI messages do not affect HRD consistency. It is necessary to distinguish between these SEI messages when indicating corruption of SEI messages via signal transmission.
[0095] Fourth, SEI messages can be divided into essential messages and non-essential messages. Damage to the former may render the video unusable, while damage to the latter can usually be ignored. Therefore, it is beneficial to distinguish between essential and non-essential messages when indicating SEI message corruption via signal transmission.
[0096] Fifth, for each VCL NAL unit, in addition to the stripe header or image header and stripe data, there is also a NAL unit header. However, currently there is no indication of NAL unit header corruption.
[0097] Sixth, the term VCL data is unclear and needs to be clearly specified.
[0098] Seventh, in the exemplary implementation, there is a lack of indication of damage to one or more reference images of the stripes in the sample.
[0099] Eighth, in the exemplary implementation, there is a lack of indication of corruption or loss of NAL units of one or more similar parameter sets required for decoding stripes in the sample.
[0100] Ninth, the exemplary implementation does not yet cover indications of the loss of some stripes in the sample.
[0101] Tenth, in the exemplary implementation, indications of missing NAL units with some similar parameter sets in the sample have not yet been covered.
[0102] 5. List of solutions and implementation examples
[0103] To address the aforementioned issues, a method outlined below is disclosed. The aspects should be considered as examples for explaining general concepts and should not be interpreted in a narrow manner. Furthermore, these examples can be applied individually or in combination in any way.
[0104] In the following text, NAL units with similar terminology to parameter sets are collectively referred to as parameter set NAL units, DCI NAL units, and OPINAL units. In the following text, a media unit may refer to a group of samples, a sample, a subsample, a set of samples, or a set of subsamples belonging to a sample group.
[0105] Example 1
[0106] To solve the first problem, one or more of the following are used by the signal transmission media unit.
[0107] An indication that one or more NAL units, which are required to decode the bitstream in the signal transmission media unit, are corrupted.
[0108] An indication that one or more NAL units with similar parameter sets that are not needed in the decoding bitstream of the signal transmission media unit are corrupted.
[0109] Example 2
[0110] To resolve issues two through four, one or more of the following are used by the signal transmission media unit.
[0111] An indication that one or more SEI NAL units containing SEI messages that affect HRD consistency are corrupted via a signal transmission media unit.
[0112] An indication that one or more SEINAL units, which contain necessary SEI messages that do not affect HRD consistency, are corrupted via a signal transmission media unit.
[0113] An indication that one or more SEI NAL units containing non-essential SEI messages that do not affect HRD consistency are corrupted via a signal transmission media unit.
[0114] An indication that one or more SEI NAL units containing the necessary SEI messages in the signal transmission media unit are corrupted.
[0115] An indication that one or more SEI NAL units containing unnecessary SEI messages are corrupted via a signal transmission media unit.
[0116] Example 3
[0117] To solve the fifth problem, one or more of the following are used by the signal transmission media unit.
[0118] Indication that one or more NAL unit headers, strip headers, or picture headers in a signal transmission media unit are damaged.
[0119] An indication that one or more NAL unit headers in the signal transmission media unit are damaged.
[0120] Indication that one or more strip heads in the signal transmission media unit are damaged.
[0121] Indication that one or more image headers in the media unit are damaged is transmitted via signal.
[0122] Example 4
[0123] To solve the sixth problem, one or more of the following are used by the signal transmission media unit.
[0124] An indication that the VCL data of one or more stripes in the signal transmission media unit is corrupted, wherein the VCL data refers to the data in the VCL NAL unit excluding the NAL unit header, strip header, and picture header (if any).
[0125] Example 5
[0126] To solve problems seven through ten, one or more of the following are used by the signal transmission media unit.
[0127] Indication that one or more reference images of a strip in a signal transmission media unit are damaged.
[0128] An indication that one or more NAL units, which are required to decode a stripe in a signal transmission media unit, are corrupted.
[0129] An indication that the reference data (motion or sample) of one or more reference images used in the strip of the signal transmission media unit has been corrupted.
[0130] Indication that one or more reference images of a strip in a signal transmission media unit are damaged or missing.
[0131] An indication that reference data (motion or sample) of one or more reference images used in the strip of the signal transmission media unit has been damaged or lost.
[0132] Indication of the loss of one or more reference images in the strip of the signal transmission media unit.
[0133] Indication of loss of reference data (motion or sample) for one or more reference images used in the strip of the signal transmission media unit.
[0134] An indication that one or more NAL units with similar parameter sets required by the signal transmission decoding media unit are corrupted or missing.
[0135] Indication of the loss of one or more NAL units with similar parameter sets required by the signal transmission decoding media unit.
[0136] Indication of the loss of one or more strips in the signal transmission media unit.
[0137] Indication of the loss of one or more NAL units with similar parameter sets in the signal transmission media unit.
[0138] Indication that one or more strips in the signal transmission media unit are damaged or lost.
[0139] An indication that one or more NAL units with similar parameter sets in the signal transmission media unit are damaged or lost.
[0140] 6. Example
[0141] The following are some exemplary embodiments of the aspects summarized in Section 5. Most of the relevant parts that have been added or modified are shown in bold, and some deleted parts are shown in italic bold. There may be other editable changes, which are not highlighted.
[0142] 6.1 First Embodiment
[0143] The following is a first embodiment, which is used in Examples 1, 2, 3, 4 and 5 summarized above. The text changes shown are related to the design of using the sample group in [9] to handle lost or damaged samples.
[0144] ■ Category: CorruptedSampleInfoEntry()
[0145] Extend SampleGroupDescriptionEntry('corr')
[0146]
[0147] The value "corrupted" indicates the state of corruption of the associated data. A value of 0 indicates that all data is lost, and the associated data size (sample size or NAL size) should be 0. A value of 1 indicates that the data is corrupted, with no additional information about the corruption. A value of 2 indicates that the data is corrupted, with codec-specific information about the corruption. A value of 3 is reserved.
[0148] The `codec_specific_param` directive indicates codec-specific information about the corrupted codec. The codec format is one of the samples associated with this sample group description.
[0149] Note: The codec_specific_param information depends on the codec format. Each time the codec format is changed between samples, the file writer may need to add different CorruptedSampleInfoEntry() entries and associate them with the samples.
[0150] If the data is not associated with a CorruptedSampleInfoEntry, or if the data is associated with a description_group_index=0 via a sample group with grouping_type'corr', it means the data is not corrupted.
[0151] The handling of samples with corrupted values of 1 or 2 depends on the specific context and implementation.
[0152] Regarding NALUFF:
[0153] For NALU-based codecs, we propose the following semantics:
[0154] In the following text, NAL units with similar terminology to parameter sets are collectively referred to as parameter set NAL units, DCI NAL units, and OPINAL units.
[0155] For NALU-based video formats, the codec_specific_param field of CorruptedSampleInfoEntry is defined as a bitmask of the following flags, with the most significant bit at the beginning:
[0156] -NonDecodingParameterSetCorruptedFlag (value 0x00000001): Indicates that the NAL unit of one or more parameter sets (DCI, VPS, SPS, PPS, APS, OPI?) required for decoding the bitstream in the associated data is corrupted.
[0157] -DecodingParameterSetCorruptedFlag (value 0x00000002): Indicates that one or more NAL units of a similar parameter set that are not needed in the decoded bitstream of the associated data have been corrupted.
[0158] -SEIConfSeiCorruptedFlag (value 0x00000002 0x00000004): Indicates that one or more SEI message NAL units containing SEI messages that affect the consistency of the bitstream's HRD are corrupted in the associated data.
[0159] -EsntSeiCorruptedFlag (value 0x00000008): Indicates that one or more SEI NAL cells in the associated data containing necessary SEI messages that do not affect the consistency of the bitstream in the HRD.
[0160] -NesnSeiCorruptedFlag (value 0x00000010): Indicates that one or more SEI NAL cells in the associated data containing non-essential SEI messages that do not affect the consistency of the bitstream have been corrupted.
[0161] -SliceVclHeaderCorruptedFlag (value 0x000000004 0x00000020): Indicates that one or more NAL cell headers, stripe headers, or picture headers in the associated VCL NAL cell are corrupted.
[0162] -VCLVclDataCorruptedFlag (value 0x00000008 0x00000040): Indicates that the VCL data of one or more stripes in the associated data is corrupted, where VCL data refers to the data in the VCL NAL unit that does not include the NAL unit header, strip header, and picture header (if any).
[0163] -OtherNonVCLNALVclNalCorruptedFlag (value 0x00000010 0x00000080): Indicates that one or more non-VCL NAL cells in the associated data of a type different from the above types have been corrupted. These non-VCL NAL cells are not NAL cells of a similar parameter set, nor are they SEI NAL cells.
[0164] -RefPicCorruptedFlag (value 0x00000100): Indicates that one or more reference images of the stripe in the associated data are corrupted.
[0165] -RefDecParamSetCorruptedFlag (value 0x00000200): Indicates that one or more NAL cells of a similar parameter set required for decoding the associated data stripe are corrupted.
[0166] A value of 0 for codec_specific_param indicates that there is no information available to describe the corruption.
[0167] Using grouping_type_parameter 'corr', CorruptedSampleInfoEntry can be used with sample groups of grouping_type 'nalm' and NALUMapEntry. The groupID of the NALUMapEntry mapping entry indicates the index starting from 1 in the sample group description of CorruptedSampleInfoEntry. A groupID of 0 indicates that no entry is associated (the identified data exists and has not been corrupted).
[0168] 6. References
[0169] [1]ITU-T and ISO / IEC, “High efficiency video coding”, Rec.ITU-T H.265|ISO / IEC 23008-2 (effective version).
[0170] [2] J.Chen, E.Alshina, G.J.Sullivan, J.-R.Ohm, J.Boyce, “Algorithm description of Joint Exploration Test Model 7 (JEM7),” JVET-G1001, August 2017.
[0171] [3] Rec.ITU-T H.266|ISO / IEC 23090-3, “Versatile Video Coding”.
[0172] [4] Rec.ITU-T Rec.H.274|ISO / IEC 23002-7, “Versatile Supplemental Enhancement Information Messages for Coded Video Bitstreams”.
[0173] [5] ISO / IEC 14496-12: “Information technology - Coding of audio-visual objects - Part 12: ISO base media file format”.
[0174] [6] ISO / IEC 23009-1: “Information technology - Dynamic adaptive streaming over HTTP (DASH) - Part 1: Media presentation description and segment formats”.
[0175] [7] ISO / IEC 14496-15: “Information technology - Coding of audio-visual objects - Part 15: Carriage of network abstraction layer (NAL) unit structured video in the ISO base media file format”.
[0176] [8]ISO / IEC 23008-12: "Information technology-High efficiency coding and media delivery in heterogeneous environments-Part 12:Image File Format".
[0177] [9] MPEG WG 03 Output Document N0598, “Technologies under Consideration for ISO / IEC 14496-12”, July 2022.
[0178] Figure 1 This is a block diagram illustrating an exemplary video processing system 4000, in which various techniques disclosed herein can be implemented. Various implementations may include some or all of the components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8-bit or 10-bit multi-component pixel values, or it may be received in a compressed or encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.
[0179] System 4000 may include an encoding / decoding component 4004, which can implement the various encoding / decoding or coding methods described in this disclosure. Encoding / decoding component 4004 can reduce the average bit rate of the video from input 4002 to the output of encoding / decoding component 4004 to generate an encoded / decoded representation of the video. Therefore, encoding / decoding techniques are sometimes referred to as video compression or video codec techniques. The output of encoding / decoding component 4004 can be stored or transmitted via connected communication (as represented by component 4006). Component 4008 can use the stored or communicated bitstream (or encoded / decoded) representation of the video received at input 4002 to generate pixel values or displayable video to be sent to display interface 4010. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “encoding / decoding” operations or tools, it should be understood that encoding / decoding tools or operations are used at the encoder, and corresponding decoding tools or operations, the opposite of the encoding / decoding result, will be performed by the decoder.
[0180] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), Peripheral Component Interconnect (PCI), Integrated Drive Electronic Devices (IDE), etc. The technologies described in this disclosure can be embodied in a variety of electronic devices, such as mobile phones, laptop computers, smartphones, or other devices capable of performing digital data processing and / or video display.
[0181] Figure 2 This is a block diagram of an exemplary video processing apparatus 4100. Apparatus 4100 can be used to implement one or more of the methods described herein. Apparatus 4100 can be embodied in a smartphone, tablet computer, computer, Internet of Things (IoT) receiver, etc. Apparatus 4100 may include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. The processors(multiple) processors 4102 may be configured to implement one or more methods described in this disclosure. The memories(multiple) memories 4104 may be used to store data and code for implementing the methods and techniques described herein. The video processing circuitry 4106 may be used to implement some of the techniques described in this disclosure in hardware circuitry. In some embodiments, the video processing circuitry 4106 may be at least partially included in the processor 4102, such as a graphics coprocessor.
[0182] Figure 3 This is a flowchart of an exemplary method 4200 for video processing. In step 4202, method 4200 determines an indication of a media unit, wherein the indication indicates lost or corrupted data associated with the media unit. In step 4204, a conversion between visual media data and a visual media data file is performed based on the indication.
[0183] It should be noted that method 4200 can be implemented in an apparatus for processing video data, including a processor and a non-transitory memory having instructions thereon, such as a video encoder 4400, a video decoder 4500, and / or an encoder 4600. In such cases, the instructions, upon execution by the processor, cause the processor to perform method 4200. Furthermore, method 4200 can be executed by a non-transitory computer-readable medium comprising a computer program product for use by a video codec device. The computer program product includes computer-executable instructions stored on the non-transitory computer-readable medium, causing the video codec device to perform method 4200 when executed by a processor.
[0184] Figure 4This is a block diagram illustrating an exemplary video encoding / decoding system 4300 that can utilize the techniques disclosed herein. The video encoding / decoding system 4300 may include a source device 4310 and a target device 4320. The source device 4310 generates encoded video data and may be referred to as a video encoding device. The target device 4320 can decode the encoded video data generated by the source device 4310 and may be referred to as a video decoding device.
[0185] Source device 4310 may include video source 4312, video encoder 4314, and input / output (I / O) interface 4316. Video source 4312 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems for generating video data, or combinations of such sources. Video data may include one or more pictures. Video encoder 4314 encodes the video data from video source 4312 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded representation of the video data. The bitstream may include encoded / decoded pictures and associated data. The encoded / decoded pictures are encoded / decoded representations of the pictures. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 4316 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to target device 4320 via network 4330 through I / O interface 4316. Encoded video data may also be stored on storage medium / server 4340 for access by target device 4320.
[0186] Target device 4320 may include I / O interface 4326, video decoder 4324, and display device 4322. I / O interface 4326 may include a receiver and / or a modem. I / O interface 4326 may acquire encoded video data from source device 4310 or storage medium / server 4340. Video decoder 4324 may decode the encoded video data. Display device 4322 may display the decoded video data to a user. Display device 4322 may be integrated with target device 4320 or may be external to target device 4320, which may be configured to interface with an external display device.
[0187] The video encoder 4314 and the video decoder 4324 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other current and / or further standards.
[0188] Figure 5 This is a block diagram illustrating an example of a video encoder 4400, which can be... Figure 4The system 4300 shown includes a video encoder 4314. The video encoder 4400 can be configured to perform any or all of the techniques disclosed herein. The video encoder 4400 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 4400. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0189] The functional components of the video encoder 4400 may include a segmentation unit 4401; a prediction unit 4402, which may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405, and an intra-frame prediction unit 4406; a residual generation unit 4407; a transform processing unit 4408; a quantization unit 4409; an inverse quantization unit 4410; an inverse transform unit 4411; a reconstruction unit 4412; a buffer 4413; and an entropy coding unit 4414.
[0190] In other examples, the video encoder 4400 may include more, fewer, or different functional components. In one example, the prediction unit 4402 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture containing the current video block.
[0191] Furthermore, some components such as the motion estimation unit 4404 and the motion compensation unit 4405 can be highly integrated, but for illustrative purposes, they are shown separately in the example of the video encoder 4400.
[0192] The segmentation unit 4401 can segment an image into one or more video blocks. The video encoder 4400 and the video decoder 4500 can support various video block sizes.
[0193] The mode selection unit 4403 can, for example, select one of the encoding / decoding modes (intra-frame or inter-frame) based on the error result, and provide the resulting intra-frame or inter-frame encoded / decoded block to the residual generation unit 4407 to generate residual block data, and provide it to the reconstruction unit 4412 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 4403 can select an intra-frame and inter-frame prediction combined (CIIP) mode, where prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 4403 can also select the resolution of the motion vector for the block (e.g., sub-pixel or integer pixel precision).
[0194] To perform inter-frame prediction on the current video block, motion estimation unit 4404 can generate motion information for the current video block by comparing one or more reference frames from buffer 4413 with the current video block. Motion compensation unit 4405 can determine the predicted video block for the current video block based on motion information from images other than those associated with the current video block and decoded samples from buffer 4413.
[0195] The motion estimation unit 4404 and the motion compensation unit 4405 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-band, P-band, or B-band.
[0196] In some instances, motion estimation unit 4404 can perform unidirectional prediction on the current video block, and can search for a reference video block for the current video block in the reference images of list 0 or list 1. Motion estimation unit 4404 can then generate a reference index and a motion vector, whereby the reference index indicates a reference image in list 0 or list 1 that includes the reference video block, and the motion vector indicates the spatial displacement between the current video block and the reference video block. Motion estimation unit 4404 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 4405 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.
[0197] In other examples, motion estimation unit 4404 can perform bidirectional prediction on the current video block. Motion estimation unit 4404 can search for a reference video block for the current video block in the reference images in list 0, and can also search for another reference video block for the current video block in the reference images in list 1. Motion estimation unit 4404 can then generate a reference index and a motion vector. The reference index indicates the reference images in lists 0 and 1 that include the reference video blocks, and the motion vector indicates the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 4404 can output the reference index and motion vector of the current video block as the motion information of the current video block. Motion compensation unit 4405 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.
[0198] In some examples, the motion estimation unit 4404 may output the complete set of motion information for decoding processing by the decoder. In some examples, the motion estimation unit 4404 may not output the complete set of motion information for the current video. Instead, the motion estimation unit 4404 may reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 4404 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0199] In one example, the motion estimation unit 4404 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 4500 that the current video block has the same motion information as another video block.
[0200] In another example, motion estimation unit 4404 may identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 4500 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0201] As described above, the video encoder 4400 can predictively transmit motion vectors via signals. Two examples of predictive signaling techniques that can be implemented by the video encoder 4400 include Advanced Motion Vector Prediction (AMVP) and merge mode signaling.
[0202] The intra-prediction unit 4406 can perform intra-prediction on the current video block. When the intra-prediction unit 4406 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0203] The residual generation unit 4407 can generate residual data for the current video block by subtracting the predicted video block from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.
[0204] In other examples, the current video block may not have residual data for the current video block, such as in skip mode, and the residual generation unit 4407 may not perform the subtraction operation.
[0205] The transform processing unit 4408 can generate one or more transform coefficient video blocks of the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0206] After the transform processing unit 4408 generates a transform coefficient video block associated with the current video block, the quantization unit 4409 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0207] The inverse quantization unit 4410 and the inverse transform unit 4411 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 4412 can add the reconstructed residual video block to the corresponding sample from one or more predicted video blocks generated by the prediction unit 4402 to generate a reconstructed video block associated with the current block, which is then stored in the buffer 4413.
[0208] After the video block is reconstructed by the reconstruction unit 4412, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0209] The entropy encoding unit 4414 can receive data from other functional components of the video encoder 4400. When the entropy encoding unit 4414 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0210] Figure 6 This is a block diagram illustrating an example of a video decoder 4500, which can be... Figure 4 The system 4300 shown includes a video decoder 4324. The video decoder 4500 can be configured to perform any or all of the techniques disclosed herein. In the illustrated example, the video decoder 4500 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 4500. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0211] In the example shown, the video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra-frame prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, the video decoder 4500 can perform a decoding process that is generally the inverse of the encoding process described in the reference video encoder 4400.
[0212] The entropy decoding unit 4501 can retrieve the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded video data blocks). The entropy decoding unit 4501 can decode the entropy-coded video data, and based on the entropy-coded video data, the motion compensation unit 4502 can determine motion information including motion vectors, motion vector precision, reference image list indexes, and other motion information. For example, the motion compensation unit 4502 can determine such information by executing AMVP and merge modes.
[0213] The motion compensation unit 4502 can generate motion-compensated blocks, possibly by performing interpolation based on an interpolation filter. The syntax elements can include identifiers of the interpolation filters to be used at sub-pixel precision.
[0214] The motion compensation unit 4502 can use interpolation filters, such as those used by the video encoder 4400 during the encoding of a video block, to calculate the sub-integer pixel interpolation of the reference block. The motion compensation unit 4502 can determine the interpolation filter used by the video encoder 4400 based on the received syntax information, and use the interpolation filter to generate the prediction block.
[0215] The motion compensation unit 4502 may use some syntax information to determine the size of the blocks used to encode the encoded video sequence (multiple) frames and / or (multiple) stripes, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame codec block, and other information used to decode the encoded video sequence.
[0216] Intra-prediction unit 4503 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 4504 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 4501. Inverse transform unit 4505 applies the inverse transform.
[0217] The reconstruction unit 4506 can add the residual block to the corresponding prediction block generated by the motion compensation unit 4502 or the intra-frame prediction unit 4503 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 4507, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates the decoded video for presentation on the display device.
[0218] Figure 7This is a schematic diagram of an exemplary encoder 4600. Encoder 4600 is suitable for implementing VVC techniques. Encoder 4600 includes three loop filters: a deblocking filter (DF) 4602, a sample adaptive offset (SAO) 4604, and an adaptive loop filter (ALF) 4606. Unlike DF 4602, which uses predefined filters, SAO 4604 and ALF 4606 utilize the original samples of the current image and reduce the mean square error between the original and reconstructed samples by adding an offset and applying a finite impulse response (FIR) filter, respectively. Information from the encoder / decoder side is transmitted via the offset and filter coefficients. ALF 4606 is located in the last processing stage of each image and can be considered as a tool attempting to capture and repair artifacts generated in previous stages.
[0219] The encoder 4600 also includes an intra-frame prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610 configured to receive input video. The intra-frame prediction component 4608 is configured to perform intra-frame prediction, while the ME / MC component 4610 is configured to perform inter-frame prediction using a reference picture obtained from a reference picture buffer 4612. Residual blocks from inter-frame or intra-frame prediction are fed to a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, which are then fed to an entropy encoder / decoder component 4618. The entropy encoder / decoder component 4618 entropy-encodes and decodes the prediction results and quantized transform coefficients and sends them to a video decoder (not shown). The quantization component output from the quantization component 4616 can be fed to an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. REC component 4624 can output images to DF 4602, SAO 4604 and ALF 4606 for filtering before these images are stored in reference image buffer 4612.
[0220] Figure 8 This is a flowchart of an exemplary method 4700 for video processing. Method 4700 includes determining one or more indications for media units in step 4702, wherein a first indication indicates that one or more Network Abstraction Layer (NAL) units with a similar set of parameters required for decoding the bitstream in the associated data are corrupted. In step 4704, based on the first indication, a conversion between visual media data and the bitstream is performed.
[0221] It should be noted that method 4700 can be implemented in an apparatus for processing video data, including a processor and a non-transitory memory having instructions thereon, such as a video encoder 4400, a video decoder 4500, and / or an encoder 4600. In such cases, the instructions, upon execution by the processor, cause the processor to perform method 4200. Furthermore, method 4700 can be executed by a non-transitory computer-readable medium comprising a computer program product for use by a video encoding / decoding device. The computer program product includes computer-executable instructions stored on the non-transitory computer-readable medium, causing the video encoding / decoding device to perform method 4700 when executed by a processor.
[0222] The following is a list of some preferred solutions.
[0223] The following solutions illustrate examples of the techniques discussed in this disclosure.
[0224] 1. A method for processing media data, comprising: determining an indication of a media unit, wherein the indication indicates lost or corrupted data associated with the media unit; and performing a conversion between visual media data and a visual media data file based on the indication.
[0225] 2. The method of Solution 1, wherein the instruction indicates that one or more of the Network Abstraction Layer (NAL) units in the media unit, which are similar to the set of parameters required to decode the bitstream, are corrupted.
[0226] 3. The method of any one of solutions 1-2, wherein the instruction indicates that one or more of the NAL units of a similar set of parameters not required for decoding the bitstream in the media unit are corrupted.
[0227] 4. The method of any of solutions 1-3, wherein the instruction indicates that one or more SEI NAL units containing supplemental enhancement information (SEI) messages in the media unit are corrupted, and the SEI messages affect the consistency of the hypothetical reference decoder (HRD).
[0228] 5. The method of any of solutions 1-4, wherein the instruction indicates that one or more SEI NAL units containing a necessary SEI message in the media unit are corrupted, and the necessary SEI message does not affect HRD consistency.
[0229] 6. The method of any of solutions 1-5, wherein the instruction indicates that one or more SEI NAL units containing a non-essential SEI message in the media unit are corrupted, and the non-essential SEI message does not affect HRD consistency.
[0230] 7. The method of any of solutions 1-6, wherein the instruction indicates that one or more SEI NAL units containing the necessary SEI message in the media unit are corrupted.
[0231] 8. The method of any of solutions 1-7, wherein the instruction indicates that one or more SEI NAL units containing unnecessary SEI messages in the media unit are corrupted.
[0232] 9. The method of any of solutions 1-8, wherein the instruction indicates that one or more NAL unit headers, strip headers or picture headers in a media unit are corrupted.
[0233] 10. The method of any of solutions 1-9, wherein the instruction indicates that one or more NAL unit headers in the media unit are damaged.
[0234] 11. The method of any of solutions 1-10, wherein the instruction indicates that one or more strip heads in the media unit are damaged.
[0235] 12. The method of any of solutions 1-11, wherein the instruction indicates that one or more picture headers in the media unit are corrupted.
[0236] 13. The method of any one of solutions 1-12, wherein the instruction indicates that the video codec layer (VCL) data of one or more stripes in a media unit is corrupted, wherein the VCL data includes data in the VCL NAL unit other than the NAL unit header, strip header and picture header.
[0237] 14. The method of any of solutions 1-13, wherein the instruction indicates that one or more reference pictures of a strip in a media unit are damaged.
[0238] 15. The method of any of solutions 1-14, wherein the instruction indicates that one or more NAL units of a similar set of parameters required for decoding a stripe in a media unit are corrupted.
[0239] 16. The method of any of solutions 1-15, wherein the instruction indicates that reference data of one or more reference pictures used by the strip in the media unit is corrupted.
[0240] 17. The method of any of solutions 1-16, wherein the instruction indicates that one or more reference pictures of a strip in a media unit are damaged or missing.
[0241] 18. The method of any of solutions 1-17, wherein the instruction indicates that reference data of one or more reference pictures used by the strip in the media unit is corrupted or lost.
[0242] 19. The method of any of solutions 1-18, wherein the instruction indicates that one or more reference pictures of a strip in a media unit are missing.
[0243] 20. The method of any of solutions 1-19, wherein the instruction indicates that reference data of one or more reference pictures used by the strip in the media unit is lost.
[0244] 21. The method of any of solutions 1-20, wherein the instruction indicates that one or more NAL units of a similar set of parameters required to decode the media unit are corrupted or lost.
[0245] 22. The method of any of solutions 1-21, wherein the instruction indicates that one or more NAL units of a similar set of parameters required for decoding the media unit are lost.
[0246] 23. The method of any of solutions 1-22, wherein the instruction indicates that one or more stripes in the media unit are lost.
[0247] 24. The method of any of solutions 1-23, wherein the instruction indicates that one or more NAL units of a similar parameter set in the media unit are missing.
[0248] 25. The method of any of solutions 1-24, wherein the instruction indicates that one or more strips in the media unit are damaged or missing.
[0249] 26. The method of any of solutions 1-25, wherein the instruction indicates that one or more NAL units of a similar parameter set in the media unit are corrupted or lost.
[0250] 27. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any of solutions 1-26.
[0251] 28. A non-transitory computer-readable medium comprising a computer program product for use by a video codec apparatus, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, such that, when executed by a processor, the video codec apparatus performs the method of any of solutions 1-26.
[0252] 29. A non-transitory computer-readable recording medium storing a bitstream of video generated by a video processing apparatus performing a method, wherein the method includes: determining an indication of a media unit, wherein the indication indicates lost or corrupted data associated with the media unit; and generating a bitstream based on the determination.
[0253] 30. A method for storing a bitstream of video, comprising: determining an indication of a media unit, wherein the indication indicates lost or corrupted data associated with the media unit; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0254] 31. A method, apparatus or system described in this disclosure.
[0255] The following solutions demonstrate other examples of the techniques discussed in this article.
[0256] 1. A method for processing media data, comprising: determining one or more indications of a media unit, wherein a first indication indicates that a network abstraction layer (NAL) unit of one or more similar parameter sets required for decoding a bitstream in associated data is corrupted; and performing a conversion between visual media data and a visual media data file based on the first indication.
[0257] 2. Solution 1 method, wherein the first indicator is the DecodingParameterSetCorruptedFlag with a value of 0x00000001.
[0258] 3. The method of any of solutions 1-2, wherein the second indication indicates that one or more NAL units of a similar set of parameters not required for decoding the bitstream in the associated data are corrupted.
[0259] 4. Solution 3, wherein the second indicator is the NonDecodingParameterSetCorruptedFlag with a value of 0x00000002.
[0260] 5. The method of any of solutions 1-4, wherein the third indication indicates that one or more supplementary enhancement information (SEI) NAL units containing supplementary enhancement information (SEI) messages in the associated data are corrupted, and the SEI messages affect the hypothetical reference decoder (HRD) consistency of the bitstream.
[0261] 6. Solution 5, wherein the third indicator is a ConformanceSeiCorruptedFlag with a value of 0x00000004.
[0262] 7. The method of any of solutions 1-6, wherein the fourth indication indicates that one or more SEI NAL cells containing a necessary SEI message in the associated data are corrupted, and the necessary SEI message does not affect the HRD consistency of the bitstream.
[0263] 8. The method of Solution 7, wherein the fourth indicator is the Essential SEI Corrupted Flag with a value of 0x00000008.
[0264] 9. The method of any of solutions 1-8, wherein the fifth instruction indicates that one or more SEI NAL cells containing a non-essential SEI message in the associated data are corrupted, and the non-essential SEI message does not affect the HRD consistency of the bitstream.
[0265] 10. Solution 9, wherein the fifth indicator is the Nonessential SeiCorrupted Flag with a value of 0x00000010.
[0266] 11. The method of any of solutions 1-10, wherein the sixth instruction indicates that one or more NAL unit headers, strip headers or picture headers of the video codec layer (VCL) NAL units in the associated data are corrupted.
[0267] 12. Solution 11, wherein the sixth indicator is the VCL Header Corrupted Flag with a value of 0x00000020.
[0268] 13. The method of any of solutions 1-12, wherein the seventh instruction indicates that the VCL data of one or more stripes in the associated data is corrupted, and wherein the VCL data refers to the data in the VCL NAL unit other than the NAL unit header, strip header and picture header.
[0269] 14. The method of Solution 13, wherein the seventh indicator is the VCL data corruption flag (VclDataCorruptedFlag) with a value of 0x00000040.
[0270] 15. The method of any of solutions 1-14, wherein the eighth instruction indicates that one or more reference images of the stripe in the associated data are corrupted.
[0271] 16. The method of Solution 15, wherein the eighth indicator is the Reference Picture Corrupted Flag with a value of 0x00000100.
[0272] 17. The method of any of solutions 1-16, wherein the ninth instruction indicates that one or more NAL units of a similar set of parameters required for decoding the associated data strip are corrupted.
[0273] 18. The method of Solution 17, wherein the ninth indicator is the Reference Image Decoding Parameter Set Corrupted Flag (RefDecParamSetCorruptedFlag) with a value of 0x00000200.
[0274] 19. The method of any of solutions 1-18, wherein at least one instruction indicates that one or more SEI NAL units containing necessary SEI messages in the media unit are corrupted.
[0275] 20. The method of any of solutions 1-19, wherein at least one instruction indicates that one or more SEI NAL units containing unnecessary SEI messages in the media unit are corrupted.
[0276] 21. The method of any one of solutions 1-20, wherein at least one indicator indicates that one or more NAL unit headers in the media unit are damaged.
[0277] 22. The method of any of solutions 1-21, wherein at least one indicator indicates that one or more strip heads in the media unit are damaged.
[0278] 23. The method of any of solutions 1-22, wherein at least one indicator indicates that one or more picture headers in a media unit are corrupted.
[0279] 24. The method of any of solutions 1-23, wherein at least one indicator indicates that reference data of one or more reference pictures used by the strip in the media unit is corrupted.
[0280] 25. The method of any of solutions 1-24, wherein at least one indicator indicates that one or more reference pictures of a strip in a media unit are damaged or lost.
[0281] 26. The method of any of solutions 1-25, wherein at least one indicator indicates that reference data of one or more reference pictures used by the strip in the media unit is damaged or lost.
[0282] 27. The method of any of solutions 1-26, wherein at least one indicator indicates that one or more reference pictures of a strip in a media unit are missing.
[0283] 28. The method of any of solutions 1-27, wherein at least one indicator indicates that reference data of one or more reference pictures used by the strip in the media unit is lost.
[0284] 29. The method of any of solutions 1-28, wherein at least one NAL unit indicating that one or more similar parameter sets required to decode the media unit is corrupted or lost.
[0285] 30. The method according to any one of claims 1-29, wherein at least one indicator indicates the loss of one or more NAL units of a similar set of parameters required for decoding the media unit.
[0286] 31. The method of any of solutions 1-30, wherein at least one indicator indicates that one or more stripes in a media unit are lost.
[0287] 32. The method of any of solutions 1-31, wherein at least one NAL unit of a similar parameter set in a media unit is lost.
[0288] 33. The method of any of solutions 1-32, wherein at least one indicator indicates that one or more strips in the media unit are damaged or lost.
[0289] 34. The method of any of solutions 1-33, wherein at least one NAL unit indicating one or more similar parameter sets in the media unit is corrupted or lost.
[0290] 35. The method of any of solutions 1-34, wherein the conversion includes encoding visual media data into a visual media data file.
[0291] 36. The method of any of solutions 1-34, wherein the conversion includes decoding visual media data from a visual media data file.
[0292] 37. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any of solutions 1-36.
[0293] 38. A non-transitory computer-readable medium comprising: a computer program product for use by a video codec device, wherein the computer program product includes computer-executable instructions stored on the non-transitory computer-readable medium, such that, when executed by a processor, the video codec device performs the method of any of solutions 1-36.
[0294] 39. A non-transitory computer-readable recording medium storing a bitstream of video generated by a video processing apparatus performing a method, wherein the method includes: determining one or more indications of media units, wherein a first indication indicates that one or more network abstraction layer (NAL) units of a similar set of parameters required for decoding the bitstream in associated data are corrupted; and generating the bitstream based on the determination.
[0295] 40. A method for storing a bitstream of video, comprising: determining one or more indications of media units, wherein a first indication indicates that a network abstraction layer (NAL) unit of one or more similar parameter sets required for decoding the bitstream in associated data is corrupted; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0296] In the solution described in this disclosure, the encoder can conform to the format rules by generating a codec representation based on those rules. In the solution described herein, the decoder can use the format rules to parse the syntax elements in the codec representation, determining the presence or absence of these elements to generate the decoded video.
[0297] In this disclosure, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. As defined by the syntax, the bitstream representation of the current video block can, for example, correspond to bits that are co-located or scattered at different locations within the bitstream. For example, a macroblock can be encoded based on the error residual values after transformation and encoding / decoding, and also using bits from the header and other fields in the bitstream. Furthermore, during the conversion, the decoder can, based on this determination, parse the bitstream knowing that some fields may or may not be present, as described in the solutions above. Similarly, the encoder can determine whether certain syntax fields are included or excluded, and generate the encoding / decoding representation accordingly by including or excluding syntax fields in the encoding / decoding representation.
[0298] The disclosed and other solutions, examples, embodiments, modules, and functional operations described herein can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures and their structural equivalents in this disclosure, or in combinations thereof. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or control of the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of substances that implement machine-readable propagating signals, or a combination thereof. The term "data processing apparatus" includes all means, devices, and machines for processing data, such as a programmable processor, a computer, or a plurality of processors or computers. In addition to hardware, the apparatus may include code that creates an execution environment for the computer program in discussion, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination thereof. Propagating signals are artificially generated signals, such as machine-generated electrical, optical, or electromagnetic signals, generated to encode information for transmission to a suitable receiver device.
[0299] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), as a single file dedicated to the program in question, or as multiple coordinating files (e.g., files storing portions of one or more modules, subroutines, or code). Computer programs can be deployed to execute on one or more computers located at a single site or distributed across multiple sites and interconnected via a communications network.
[0300] The processes and logic flows described in this disclosure can be executed by one or more programmable processors that execute one or more computer programs, thereby performing functions by manipulating input data and generating outputs. The processes and logic flows can also be executed by special-purpose logic circuitry, and the apparatus can also be implemented as special-purpose logic circuitry, such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs).
[0301] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or operatively coupled to receive data from or transfer data to one or more mass storage devices, or both. However, a computer does not necessarily need to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as: semiconductor memory devices, such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and compact optical disc read-only memory (CD ROM) and digital versatile optical disc read-only memory (DVD-ROM). Processors and memory may be complemented or incorporated into dedicated logic circuitry.
[0302] While this disclosure contains numerous details, these should not be construed as limiting any subject matter or potentially claimed scope, but rather as descriptions of features characteristic of particular embodiments of a particular art. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or any suitable sub-combination. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed in this way, in some cases, one or more features from the claimed combination may be removed from that combination, and the claimed combination may involve sub-combinations or variations thereof.
[0303] Similarly, although operations are depicted in a specific order in the figures, this should not be construed as requiring such operations to be performed in the specific order shown or sequentially, or to perform all the illustrated operations to achieve the desired result. Furthermore, the separation of various system components in the embodiments described in this disclosure should not be construed as requiring such separation in all embodiments.
[0304] Only some implementation methods and examples are described, and other implementation methods, enhancements and variations can be made based on what is described and shown in this disclosure.
[0305] When there are no intermediate components other than a line, trace, or other medium between the first and second components, the first component is directly coupled to the second component. When there are intermediate components other than a line, trace, or other medium between the first and second components, the first component is indirectly coupled to the second component. The term "coupled" and its variations include both direct and indirect coupling. Unless otherwise stated, the term "about" is used to mean a range including ±10% of the subsequent value.
[0306] While several embodiments have been provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of this disclosure. The present examples are intended to be illustrative rather than restrictive and are not intended to be limited to the details given herein. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented.
[0307] Furthermore, without departing from the scope of this disclosure, the technologies, systems, subsystems, and methods described and shown as discrete or independent in the various embodiments may be combined or integrated with other systems, modules, technologies, or methods. Other items shown or discussed as coupled may be directly connected or indirectly coupled or communicated via some interface, device, or intermediate component in an electrical, mechanical, or other manner. Those skilled in the art can identify other examples of changes, substitutions, and modifications, and these changes, substitutions, and modifications may be made without departing from the spirit and scope of this disclosure.
Claims
1. A method for processing media data, comprising: Determine one or more indications for a media unit, wherein the one or more indications include a first indication indicating that one or more Network Abstraction Layer (NAL) units with similar parameter sets required for decoding the bitstream in the associated data are corrupted, the one or more NAL units with similar parameter sets referring to the collective term for Parameter Set NAL units, Decoding Capability Information (DCI) NAL units, and Operation Point Information (OPI) NAL units; and Based on the first instruction, perform the conversion between media data and media data files; The one or more indications include a second indication that indicates that one or more NAL units of a similar set of parameters in the associated data that are not required for decoding the bitstream are corrupted.
2. The method according to claim 1, wherein, The first indication is a first parameter set corruption flag with a value of 0x00000001.
3. The method according to claim 1, wherein, The second indication is a second parameter set corruption flag with a value of 0x00000002.
4. The method according to claim 1, wherein, The one or more indications include a third indication indicating that one or more SEI NAL cells containing Supplemental Enhancement Information (SEI) messages in the associated data are corrupted, the SEI messages affecting the hypothetical reference decoder (HRD) consistency of the bitstream, and The third indication is a consistency SEI corruption flag with a value of 0x00000004.
5. The method according to claim 1, wherein, The one or more indications include a fourth indication indicating that one or more SEI NAL cells containing necessary supplementary enhancement information (SEI) messages in the associated data are corrupted, the necessary SEI messages not affecting the HRD consistency of the bitstream, and The fourth indication is a necessary SEI damage flag with a value of 0x00000008.
6. The method according to claim 1, wherein, The one or more indications include a fifth indication that indicates one or more SEI NAL cells in the associated data containing unnecessary supplementary enhancement information (SEI) messages are corrupted, the unnecessary SEI messages not affecting the HRD consistency of the bitstream, and The fifth indication is a non-essential SEI damage flag with a value of 0x00000010.
7. The method according to claim 1, wherein, The one or more indications include a sixth indication indicating that one or more NAL unit headers, stripe headers, or image headers of the Video Codec Layer (VCL) NAL units in the associated data are corrupted, and The sixth indicator is a VCL head damage flag with a value of 0x00000020.
8. The method according to claim 1, wherein, The one or more indications include a seventh indication indicating that one or more stripes of video codec layer (VCL) data in the associated data are corrupted, and The VCL data refers to the data in the VCL NAL unit other than the NAL unit header, strip header, and image header.
9. The method according to claim 8, wherein, The seventh indication is a VCL data corruption flag with a value of 0x00000040.
10. The method according to claim 1, wherein, The one or more indications include an eighth indication that one or more reference images of the stripes in the associated data are corrupted, and The eighth indication is a reference image damage indicator with a value of 0x00000100.
11. The method according to claim 1, wherein, The one or more indications include a ninth indication that indicates that one or more NAL units of a similar parameter set required for decoding the associated data stripe are corrupted, and The ninth indication is a corrupted reference image decoding parameter set flag with a value of 0x00000200.
12. The method according to claim 1, wherein, The one or more indications include a tenth indication indicating that one or more non-Video Codec Layer (VCL) NAL units in the associated data are corrupted, the non-VCL NAL units being neither NAL units of a similar parameter set nor Supplemental Enhancement Information (SEI) NAL units, and The tenth indication is a non-VCL NAL damage flag with a value of 0x00000080.
13. The method according to claim 1, further comprising: Determine the codec-specific parameter field of the corrupted sample information entry, wherein the codec-specific parameter field indicates codec-specific information about the corruption.
14. The method according to claim 13, wherein, A codec-specific parameter field with a value of 0 indicates that no information is available to describe the corruption.
15. The method according to claim 1, wherein, The conversion includes generating the media data file from the media data.
16. The method according to claim 1, wherein, The conversion includes parsing the media data from the media data file.
17. An apparatus for processing media data, comprising: processor; and a non-transitory memory thereon having instructions, wherein, when executed by the processor, the instructions cause the processor to: Determine one or more indications for a media unit, wherein the one or more indications include a first indication indicating that one or more Network Abstraction Layer (NAL) units with similar parameter sets required for decoding the bitstream in the associated data are corrupted, the one or more NAL units with similar parameter sets referring to the collective term for Parameter Set NAL units, Decoding Capability Information (DCI) NAL units, and Operation Point Information (OPI) NAL units; and Based on the first instruction, perform the conversion between media data and media data files; The one or more indications include a second indication that indicates that one or more NAL units of a similar set of parameters in the associated data that are not required for decoding the bitstream are corrupted.
18. A non-transitory computer-readable storage medium, wherein, The non-transitory computer-readable storage medium stores instructions that cause the processor to: Determine one or more indications for a media unit, wherein the one or more indications include a first indication indicating that one or more Network Abstraction Layer (NAL) units with similar parameter sets required for decoding the bitstream in the associated data are corrupted, the one or more NAL units with similar parameter sets referring to the collective term for Parameter Set NAL units, Decoding Capability Information (DCI) NAL units, and Operation Point Information (OPI) NAL units; and Based on the first instruction, perform the conversion between media data and media data files; The one or more indications include a second indication that indicates that one or more NAL units of a similar set of parameters in the associated data that are not required for decoding the bitstream are corrupted.
19. A non-transitory computer-readable recording medium storing a bitstream of video generated by a video processing apparatus performing a method, wherein, The method includes: Determine one or more indications for a media unit, wherein the one or more indications include a first indication indicating that one or more Network Abstraction Layer (NAL) units with similar parameter sets required for decoding the bitstream in the associated data are corrupted, the one or more NAL units with similar parameter sets referring to the collective term for Parameter Set NAL units, Decoding Capability Information (DCI) NAL units, and Operation Point Information (OPI) NAL units; and The bit stream is generated based on the first instruction; The one or more indications include a second indication that indicates that one or more NAL units of a similar set of parameters in the associated data that are not required for decoding the bitstream are corrupted.
20. A method for storing a video bitstream, comprising: Determine one or more indications for a media unit, wherein the one or more indications include a first indication that one or more network abstraction layer (NAL) units with similar parameter sets required for decoding the bitstream in the associated data are corrupted, the one or more NAL units with similar parameter sets being a collective term for parameter set NAL units, decoding capability information (DCI) NAL units, and operation point information (OPI) NAL units. The bit stream is generated based on the determination; and The bitstream is stored in a non-transitory computer-readable recording medium; The one or more indications include a second indication that indicates that one or more NAL units of a similar set of parameters in the associated data that are not required for decoding the bitstream are corrupted.
Citation Information
Patent Citations
Image decoding device and image decoding method
CN106165422A
Method and apparatus for video coding
US20130272372A1