Associating motion point information characteristics with VVC image items

By signaling recommended transition durations and defining file brands and picture item types, the VVC image file format addresses precision and interoperability issues, enabling efficient and user-friendly image processing with multiple transition effects.

JP7758431B2Active Publication Date: 2025-10-22LEMON CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2021142855
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-09-02
Filing Date
2021-09-02
Publication Date
2025-10-22
Estimated Expiration
2041-09-02

AI Technical Summary

Technical Problem

The current design of the VVC image file format and image transition effect signaling lacks precision in transition duration determination, interoperability points, and support for multiple transition effects, leading to suboptimal user experience and inefficiencies in image processing.

Method used

Implementing methods to signal recommended transition durations, define file brands and picture item types that ensure single access units, allow multiple transition effects, and associate image items with different operating points, while adhering to specific VVC standards.

Benefits of technology

Enhances user experience by ensuring precise transition durations, improves interoperability, and supports multiple transition effects, thereby optimizing image processing and file format compliance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007758431000009
    Figure 0007758431000009
  • Figure 0007758431000010
    Figure 0007758431000010
  • Figure 0007758431000011
    Figure 0007758431000011
Patent Text Reader

Abstract

To provide a system, a method, and a device for processing image data.SOLUTION: A method according to an embodiment includes a step of performing conversion between a visual media file and a bitstream. The visual media file has an image item including a sequence of one or more pictures according to the media file format, and the bitstream includes an access unit consisting of one or more pictures each belonging to a layer according to a video coding format. The media file format specifies that an image item including a picture originating from the bitstream is allowed to be associated with an instance having a different characteristic descriptor indicating the high level characteristics of the bitstream.SELECTED DRAWING: Figure 11
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] Under applicable patent law and / or the rules pursuant to the Paris Convention, this application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 073,829, filed September 2, 2020. For all purposes under law, the entire disclosure of the above application is incorporated by reference as part of the disclosure of this application.

[0002] [Technical field] This patent document relates to image and video encoding and decoding. [Background technology]

[0003] Digital video accounts for the largest bandwidth usage on the Internet and other digital communication networks, and as the number of connected user devices capable of receiving and displaying video increases, the bandwidth demands for digital video usage are expected to continue to grow. Summary of the Invention

[0004] This patent document discloses techniques that can be used by video encoders and decoders to process coded representations of video or images according to a file format.

[0005] In one exemplary aspect, a method for processing image data includes performing a conversion between a visual media file and a bitstream. The visual media file has a sequence of one or more pictures according to a media file format, and the bitstream has one or more access units according to a video coding format. The bitstream is coded according to the video coding format. The media file format specifies that an image item of a particular type value in the visual media file contains a single access unit of the bitstream. The single access unit is either an Intra Random Access Picture (IRAP) access unit according to the video coding format or a Gradual Decoding Refresh (GDR) access unit according to the video coding format. All pictures within the GDR access unit are identified as recovery points in the bitstream.

[0006] In another example aspect, a method for processing image data includes performing a conversion between a visual media file and a bitstream, where the visual media file has a sequence of one or more pictures according to a media file format, and the bitstream has one or more access units according to a video coding format, and the bitstream is coded according to the video coding format, and the media file format specifies that image items of a particular type value in the visual media file exclude layers that do not belong to a target output layer set.

[0007] In another example aspect, a method for processing image data includes performing a conversion between a visual media file and a bitstream, where the visual media file has a sequence of one or more pictures according to a media file format, and the bitstream has one or more access units according to a video coding format. The bitstream is coded according to the video coding format. The media file format specifies that an image item of a particular type value in the visual media file includes at least a portion of an access unit in which a picture has one or more sub-pictures.

[0008] In another exemplary aspect, a method for processing image data includes performing a conversion between a visual media file and a bitstream. The visual media file has image items, each having a sequence of one or more pictures according to a media file format. The bitstream includes access units, each consisting of one or more pictures, each belonging to a layer according to a video coding format. The media file format specifies that image items having pictures originating from the bitstream are allowed to be associated with different instances of characteristic descriptors that indicate high-level characteristics of the bitstream.

[0009] In another exemplary aspect, a method for processing image data includes performing a conversion between a visual media file and a bitstream. The visual media file has image items each including a sequence of one or more pictures according to a media file format, and the bitstream includes access units each consisting of one or more pictures, each belonging to a layer according to a video coding format. The media file format specifies, in response to an operation point record included in an operation point characteristic descriptor indicating high-level characteristics of the bitstream, that at least one of a value of a first syntax element in the record or a value of a second syntax element in the record is constrained to be a predetermined value.

[0010] In another example aspect, a video processing method is disclosed that includes performing a conversion between visual media including a sequence of one or more images and a bitstream representation according to a file format configured to include one or more syntax elements that indicate transition characteristics between one or more images during display of the one or more images.

[0011] In another example aspect, another video processing method is disclosed, the method including performing a conversion between visual media including a sequence of one or more images and a bitstream representation according to a file format, the file format specifying that when the visual media is represented in a file having a particular file brand, the file format is restricted according to rules.

[0012] In another example aspect, another video processing method is disclosed, the method including performing a conversion between visual media including a sequence of one or more images and a bitstream representation according to a file format configured to indicate an image type of the one or more images according to a rule.

[0013] In yet another exemplary aspect, a video encoder apparatus is disclosed, the video encoder having a processor configured to implement the above method.

[0014] In yet another exemplary aspect, a video decoder apparatus is disclosed, the video decoder having a processor configured to implement the above method.

[0015] In yet another exemplary aspect, a computer-readable medium having stored thereon code embodying one of the methods described herein in the form of code executable by a processor is disclosed.

[0016] In yet another exemplary aspect, a computer-readable medium storing a bitstream is disclosed, the bitstream being generated using the methods described herein.

[0017] These and other features are described throughout this specification. [Brief explanation of the drawings]

[0018] [Figure 1] FIG. 1 is a block diagram of an example video processing system. [Figure 2] FIG. 1 is a block diagram of a video processing device. [Figure 3] 1 is a flowchart of an example method of video processing. [Figure 4] 1 is a block diagram illustrating a video coding system in accordance with some embodiments of the present disclosure. [Figure 5] FIG. 2 is a block diagram illustrating an encoder in accordance with some embodiments of the present disclosure. [Figure 6] FIG. 2 is a block diagram illustrating a decoder in accordance with some embodiments of the present disclosure. [Figure 7] 1 shows an example of an encoder block diagram. [Figure 8] 1 is a flowchart representation of a method for processing image data in accordance with one or more embodiments of the present technique. [Figure 9] 1 is a flowchart representation of a method for processing image data in accordance with one or more embodiments of the present technique. [Figure 10] 1 is a flowchart representation of a method for processing image data in accordance with one or more embodiments of the present technique. [Figure 11] 1 is a flowchart representation of a method for processing image data in accordance with one or more embodiments of the present technique. [Figure 12] 1 is a flowchart representation of a method for processing image data in accordance with one or more embodiments of the present technique. DETAILED DESCRIPTION OF THE INVENTION

[0019] Section headings are used herein for ease of understanding and are not intended to limit the applicability of the techniques and embodiments disclosed in each section to that section alone. Furthermore, H.266 terminology is used in some descriptions for ease of understanding only and is not intended to limit the scope of the disclosed techniques. Accordingly, the techniques described herein are applicable to other video codec protocols and designs. Editorial changes herein are indicated in the text by strikethrough of deleted text and highlighting (including bold italics) of added text relative to the current draft of the VVC standard.

[0020] 1. Overview This specification relates to image file formats. Specifically, it relates to the signaling and storage of images and image transitions in media files based on the ISO Basic Media File Format. The ideas may be applied individually or in various combinations to images coded according to any codec, such as the Versatile Video Coding (VVC) standard, and to any image file format, such as the VVC image file format under development.

[0021] 2.Abbreviation AU Access Unit AUD Access Unit Delimiter AVC Advanced Video Coding BP Buffering Period CLVS Coded Layer Video Sequence CLVSS Coded Layer Video Sequence Start CPB Coded Picture Buffer CRA Clean Random Access CTU Coding Tree Unit CVS Coded Video Sequence DCI Decoding Capability Information DPB Decoded Picture Buffer DUI Decoding Unit Information EOB End Of Bitstream EOS End Of Sequence GDR Gradual Decoding Refresh HEVC High Efficiency Video Coding HRD Hypothetical Reference Decoder IDR Instantaneous Decoding Refresh ILP Inter-Layer Prediction ILRP Inter-Layer Reference Picture IRAP Intra Random Access Picture JEM Joint Exploration Model LTRP Long-Term Reference Picture MCTS Motion-Constrained Tile Sets NAL Network Abstraction Layer OLS Output Layer Set PH Picture Header POC Picture Order Count PPS Picture Parameter Set PT Picture Timing PTL Profile, Tier and Level PU Picture Unit RAP Random Access Point RBSP Raw Byte Sequence Payload SEI Supplemental Enhancement Information SLI Subpicture Level Information SPS Sequence Parameter Set STRP Short-Term Reference Picture SVC Scalable Video Coding VCL Video Coding Layer VPS Video Parameter Set VTM VVC Test Model VUI Video Usability Information VVC Versatile Video Coding

[0022] 3. Initial discussion 3.1 Video Coding Standards Video coding standards have evolved primarily through the development of well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, while ISO / IEC produced MPEG-1 and MPEG-4 Visual. The two organizations jointly produced the H.262 / MPEG-2 Video, H264 / MPEG-4 AVC (Advanced Video Coding), and H.265 / HEVC standards. Since H.262, video coding standards have been based on hybrid video coding architectures, utilizing temporal prediction and transform coding. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, many new methods have been introduced by JVET and incorporated into reference software named the Joint Exploration Model (JEM). JVET was later renamed to be the Joint Video Experts Team (JVET) when the Versatile Video Coding (VVC) project was officially launched. VVC is a new coding standard finalized by JVET at its 19th meeting, which ended on July 1, 2020, that aims to reduce the bitrate by 50% compared to HEVC.

[0023] The Versatile Video Coding (VVC) standard (ITU-T H.266 | ISO / IEC 23090-3) and the related Versatile Supplemental Enhancement Information (VSEI) standard (ITU-T H.274 | ISO / IEC 23002-7) are designed for use in the widest possible range of applications, including both traditional uses such as television broadcasting, videoconferencing, or playback from storage media, and newer and more advanced use cases such as adaptive bitrate streaming, video region extraction, compositing and merging content from multiple coded video bitstreams, multiview video, scalable layered coding, and viewport-adaptive 360° immersive media.

[0024] 3.2 File Format Standards Media streaming applications are typically based on IP, TCP, and HTTP transport methods and typically rely on file formats such as the ISO Base Media File Format (ISOBMFF). One such streaming system is Dynamic Adaptive Streaming over HTTP (DASH). To use video formats with ISOBMFF and DASH, video format-specific file format standards, such as the AVC file format and the HEVC file format, became required for encapsulation of video content in ISOBMFF tracks and in DASH representations and segments. Important information about the video bitstream, such as profile, tier, and level, and many others, became required to be exposed as file format-level metadata and / or a DASH Media Presentation Description (MPD) for content selection purposes, e.g., for selection of appropriate media segments for both initialization at the start of a streaming session and stream adaptation during a streaming session.

[0025] Similarly, to use an image format according to ISOBMFF, a file format standard specific to the image format, such as the AVC image file format and the HEVC image file format, may be required.

[0026] 3.3 VVC video file format The VVC video file format is a file format for storage of VVC video content based on ISOBMFF and is currently under development by MPEG.

[0027] 3.4 VVC Image File Format and Image Transitions The VVC image file format is a file format for storage of image content coded using VVC, based on ISOBMFF, and is currently under development by MPEG.

[0028] In some cases, designs for slideshow signaling include support for image transition effects such as wipe, zoom, fade, split, dissolve, etc. Transition effects are signaled in a transition effect characteristics structure, which is associated with the first of two consecutive items involved in the transition and contains the transition type and possibly signals other transition information such as transition direction and transition shape, if applicable.

[0029] 4. Examples of technical problems solved by the disclosed technical solutions The current design of the VVC image file format and image transition effect signaling has the following problems:

[0030] 1) In slideshows or other types of image-based applications with transition effects from one image to another, the time for the transition often does not need to be precise, but for a good user experience, it should not be too long or too short. And the best transition duration depends on the content and transition type. Therefore, from the user experience point of view, it is useful to communicate the recommended transition duration, and the recommended value is determined by the content creator.

[0031] 2) In the latest VVC image file format draft specification, a specific VVC image item type and file brand allow a VVC bitstream of an image item to contain access units containing multiple pictures of multiple layers, where some pictures may be inter-coded, i.e., contain B or P slices predicted using inter-layer prediction as specified in VVC. In other words, interoperability points via either the image item type or file brand are lacking; an image item can only contain one intra-coded picture (i.e., contain only intra-coded I-slices). In the VVC standard itself, such interoperability points are provided by the definition of two still image profiles: the Main10 still image profile and the Main10 4:4:4 still image profile.

[0032] 3) An item of type 'vvc1' is specified as follows: An item of type 'vvc1' consists of a NAL unit of a VVC bitstream that is length-bounded as specified below, where the bitstream contains exactly one access unit. NOTE 2: An item of type 'vvc1' may consist of an IRAP access unit as defined in ISO / IEC 23090-3 and may contain more than one coded picture, including at most one coded picture with any particular value of nuh_layer_id. However, no access unit can be an access unit in such an image item, so the first part of Note 2 above should be moved to the base definition (i.e., the first part quoted above), and the missing part of the GDR access unit should be added.

[0033] 4) The following statement exists: The 'vvc1' image item should contain the layers contained in the layer set identified by the associated TargetOlsProperty, and may also contain other layers. If other layers than those contained in the identified OLS should be allowed, which entity in the application system is supposed to set the correct value of the target OLS index in the associated TargetOlsProperty? In any case, this value needs to be set correctly, for example by the file composer, so it makes sense to not allow unwanted pictures in unwanted layers at all, since discarding unwanted pictures in unwanted layers is also an easy action for the file composer.

[0034] 5) The following constraints exist: Image items that originate from the same bitstream should be associated with the same VvcOperatingPointsInformationProperty. However, a VVC bitstream may contain multiple CVSs that may have different operating points.

[0035] 6) In the following text, the values ​​of other syntax elements of VvcOperatingPointRecord, such as ptl_max_temporal_id[i] (the temporal ID of the highest sublayer representation whose level information is present in the i-th profile_tier_level() syntax structure) and op_max_temporal_id, should also be constrained: When included in the VvcOperatingPointsInformationProperty, the values ​​of the syntax elements of the VvcOperatingPointsRecord are constrained as follows: frame_rate_info_flag must be equal to 0. As a result, avgFrameRate and constantFrameRate do not exist and their semantics are unspecified. bit_rate_info_flag must be equal to 0. As a result, maxBitRate and avgBitRate do not exist and their semantics are unspecified.

[0036] 7) The following text exists: If a VVC sub-picture item is suitable to be decoded by a VVC decoder and consumed without other VVC sub-picture items, it should be stored as an item of type 'vvc1'. Otherwise, it should be stored as an item of type 'vvs1' and formatted as a series of NAL units preceded by a length field as defined in L.2.2.1.2. This has the following problems: a) This condition needs to be clarified because it is not clear enough to be used as a condition for a conformance requirement (e.g., when thinking about how to check whether the requirement is met). b) Here, the use of a picture item of type 'vvc1' does not strictly correspond to the above definition that a bitstream contains exactly one VVC access unit, in that a bitstream of a picture item of type 'vvc1' can contain just a subset of VVC access units. c) It is not clear whether it is permissible to have one VVC image item, for example of type 'vvc1', containing a picture that contains multiple 'extractable' sub-pictures.

[0037] 8) The following statements do not contain OPI NAL units: The VPS, DCI, SPS, PPS, AUD, PH, EOS, and EOB NAL units should not be present in both the item and a sample of the 'vs1' item. However, operation point information (OPI) NAL units should be treated similarly here.

[0038] 9) Only one transition effect (e.g., zoom, rotate) is allowed in a given image or for a region in a given image, although in a real application, multiple effects may be applied to an image or a region in a given image.

[0039] 5. Example Embodiments and Solutions To solve the above problems and others, the methods summarized below are disclosed. The items should be considered as examples to illustrate the general concept and should not be construed in a narrow sense. Furthermore, these items may be applied individually or combined in any way.

[0040] 1) To solve problem 1, a recommended transition duration may be signaled for the transition from one image to another. a. Alternatively, in one example, a forced transition period is signaled for the transition from one image to another. b. In one example, the advertised value, i.e., the recommended or mandatory transition period value, is determined by the content creator. c. In one example, one transition period is signaled per transition characteristic. d. In one example, one transition period is announced for each type of transition. e. In one example, one transition period is advertised with a list of transition characteristics. f. In one example, one transition period is advertised for a list of transition types. g. In one example, one transition period is notified for all transitions.

[0041] 2) To solve problem 2, one or more file brands are defined such that a VVC bitstream contained in a video item conforming to such a brand is required to contain exactly one access unit containing exactly one picture (or portion thereof) that is intra-coded. a. Alternatively, one or more file brands may be defined such that a VVC bitstream contained in a visual item conforming to such a brand is required to contain exactly one access unit containing exactly one picture (or portion thereof) that is Intra / IBC / Palette coded. i. Alternatively, one or more file brands may be defined such that a VVC bitstream contained in a video item conforming to such a brand is required to contain exactly one access unit containing exactly one I-picture (or portion thereof). b. In one example, such file brand values ​​are specified as 'vvic', 'vvi1', 'vvi2'. c. In one example, additionally, the VVC bitstream contained in such image item is required to conform to the Main10 still image profile, the Main10 4:4:4 still image profile, the Main10 profile, the Main10 4:4:4 profile, the multi-layer Main10 profile, or the multi-layer Main10 4:4:4 profile. i. Alternatively, or additionally, the VVC bitstream contained in such image item is required to conform to the Main10 still image profile, the Main10 4:4:4 still image profile, the Main10 profile, or the Main10 4:4:4 profile. ii. Alternatively or additionally, the VVC bitstream contained in such image item is required to conform to the Main10 still image profile or the Main10 4:4:4 still image profile. d. In one example, it may be specified that image items that conform to such a brand should not have any of the following properties: Target Output Layer Set Property (TargetOlsProperty), VV Operating Points Information Property (VvcOperatingPointsInformationProperty).

[0042] 3) To solve problem 2, one or more picture item types are defined such that a VVC bitstream contained in a picture item of such type contains only one access unit that contains only intra-coded pictures. a. Alternatively, one or more picture item types are defined such that a VVC bitstream contained in a picture item of such type contains only one access unit that contains only pictures that are Intra / Palette / IBC coded. i. Alternatively, one or more picture item types are defined such that a VVC bitstream contained in a picture item of such type contains only one access unit that contains only I-pictures. b. In one example, the type value of such an image item type is specified as 'vvc1' or 'vvc2'. c. In one example, additionally, the bitstream in such image item is required to conform to the Main10 still image profile, the Main10 4:4:4 still image profile, the Main10 profile, the Main10 4:4:4 profile, the multi-layer Main10 profile, or the multi-layer Main10 4:4:4 profile. i. Alternatively, or additionally, the bitstream in such image item is required to conform to the Main10 still image profile, the Main10 4:4:4 still image profile, the Main10 profile, or the Main10 4:4:4 profile. ii. Alternatively or additionally, the bitstream in such image items is required to conform to the Main10 still image profile or the Main10 4:4:4 still image profile. d. In one example, it may be specified that such type of image item should not have any of the following properties: target output layer set property (TargetOlsProperty), VVC operating points information property (VvcOperatingPointsInformationProperty).

[0043] 4) To solve problem 3, a VVC picture item, e.g., of type 'vvc1', is defined to consist of a NAL unit of a VVC bitstream containing exactly one access unit which is an IRAP access unit as defined in ISO / IEC 23090-3 or a GDR access unit where all pictures have ph_recovery_poc_cnt equal to 0 as defined in ISO / IEC 23090-3.

[0044] 5) To solve problem 4, a VVC image item, for example of type 'vvc1', is not allowed to contain pictures in layers that do not belong to the target output layer set.

[0045] 6) To solve problem 5, image items that originate from the same bitstream are allowed to be associated with different instances of VvcOperatingPointsInformationProperty.

[0046] 7) To solve problem 6, the syntax elements ptl_max_temporal_id[i] and op_max_temporal_id of a VvcOperatingPointsRecord are constrained to be specific values ​​when the VvcOperatingPointsRecord is included in a VvcOperatingPointsInformationProperty.

[0047] 8) To solve problem 7, either of the following is allowed for a VVC image item, e.g., 'vvc1': a. Contains an entire VVC access unit where each picture may contain multiple "extractable" sub-pictures. b. For each layer present in the bitstream, it contains a subset of VVC access units for which there are one or more "extractable" sub-pictures that collectively form a rectangular region. Here, an "extractable" subpicture refers to a subpicture whose corresponding flag sps_subpic_treated_as_pic_flag[i] specified in VVC is equal to 1.

[0048] 9) To solve problem 8, it may be specified that an OPI NAL unit should not be present in both an item and a sample of a 'vvs1' item.

[0049] 10) To solve problem 9, it is proposed that multiple transition effects from one image (or region thereof) to another image (or region thereof) are allowed in a slideshow. In one example, an indication of multiple transition effects may be signaled, for example, by having multiple transition effect property structures associated with the first of two consecutive image items. b. In one example, an indication of the number of transition effects to be applied between two consecutive image items may be signaled in the file. c. Alternatively or additionally, the method of applying multiple transition effects may be signaled in a file, or predefined, or derived on the fly. i. In one example, the order in which multiple transition effects are applied may be signaled in a file. ii. In one example, the order of applying multiple transition effects may be derived according to the order of the effects' indications in the bitstream.

[0050] 11) To solve problem 9, it is proposed to allow multiple transition effects from one image to another in a slideshow, where each of the multiple transition effects is applied to a specific region in the two image items involved in the transition. In one example, the specific regions in the two image items involved in the transition to which the transition effect is applied are signaled in the transition effect properties.

[0051] 12) To solve problem 9, it is proposed to allow multiple alternative transition effects to be signaled for a pair of consecutive image items, and it is up to the player of the file to select one of the multiple transition effects to be applied. In one example, the priority (or preference order) of multiple transition effects is signaled in a file, or is predefined, or is derived according to the signaling order of the transition characteristics.

[0052] 6. Implementation form Below are some example embodiments of some of the inventive aspects briefly described above in Section 5, which may be applied to the VVC image file format and slideshow support standard. The most relevant parts that have been added or changed are underlined in bold italics, and some parts that have been removed are indicated using [[]].

[0053] 6.1 First embodiment This embodiment relates to at least items 1, 1.b and 1.c. [Table 1] TIFF0007758431000002.tif222161TIFF0007758431000003.tif200121TIFF0007758431000004.tif242161TIFF0007758431000005.tif174161

[0054] 6.2 Second embodiment This embodiment relates to at least items 4 and 5. [Table 2]

[0055] 6.3 Third embodiment This embodiment relates to at least items 6 and 7. [Table 3]

[0056] 6.4 Fourth embodiment This embodiment relates to at least item 9. [Table 4]

[0057] 1 is a block diagram illustrating an example video processing system 1900 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 that receives video content. The video content may be received in raw or uncompressed format, such as 8- or 10-bit multi-component pixel values, or in a compressed or encoded format. Input 1902 may correspond to a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet or a passive optical network (PON), and wireless interfaces such as Wi-Fi or a cellular interface.

[0058] System 1900 may include a coding component 1904 that may implement various coding or encoding methods described herein. Coding component 1904 may reduce the average bitrate of video from input 1902 to an output of coding component 1904 to generate a coded representation of the video. Accordingly, coding techniques are sometimes referred to as video compression or video transcoding techniques. The output of coding component 1904 may be stored or transmitted via a connected communication, as represented by component 1906. The stored or communicated bitstream (or coded) representation of the video received at input 1902 may be used by component 1908 to generate pixel values ​​or displayable video that is sent to display interface 1910. The process of generating user-viewable video from the bitstream is sometimes referred to as video decompression. Furthermore, while certain video processing operations are referred to as “coding” operations or tools, it will be understood that such coding tools or operations are used in an encoder, and that corresponding decoding tools or operations that transpose the results of the coding would be performed by a decoder.

[0059] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB) or High-Definition Multimedia Interface (HDMI®) or Displayport®, etc. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), PCI, IDE interfaces, etc. The techniques described herein may be embodied in various electronic devices such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0060] FIG. 2 is a block diagram of a video processing device 3600. The device 3600 may be used to implement one or more of the methods described herein. The device 3600 may be embodied in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. The device 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processor 3602 may be configured to implement one or more of the methods described herein. The memory(s) 3604 may be used to store data and code used to implement the methods and techniques described herein. The video processing hardware 3606 may be used to implement some of the techniques described herein in a hardware circuit. In some embodiments, the video processing hardware 3606 may be at least partially included in the processor 3602, e.g., a graphics coprocessor.

[0061] FIG. 4 is a block diagram illustrating an example video coding system 100 that may utilize techniques of this disclosure.

[0062] 4, video coding system 100 may include source device 110 and destination device 120. Source device 110 generates encoded video data and may be referred to as a video encoding device. Destination device 120 can decode the encoded video data generated by source device 110 and may be referred to as a video decoding device.

[0063] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface .

[0064] The video source 112 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a coded representation of the video data. The bitstream may include coded pictures and associated data. A coded picture is a coded representation of a picture. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The coded video data may be transmitted directly over the network 130a to the destination device 120 via the I / O interface 116. The coded video data may also be stored on a storage medium / server 130b for access by the transmitter.

[0065] The destination device 120 may include an I / O interface 126 , a video decoder 124 , and a display device 122 .

[0066] I / O interface 126 may include a receiver and / or modem. I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120 or may be external to destination device 120 and configured to interface with an external display device.

[0067] Video encoder 114 and video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or future standards.

[0068] FIG. 5 is a block diagram illustrating an example of a video encoder 200, which may be the video encoder 114 of the system 100 illustrated in FIG.

[0069] Video encoder 200 may be configured to perform any or all of the techniques of this disclosure. In the example of FIG. 5, video encoder 200 includes multiple functional components. The techniques described in this disclosure may be shared among various components of video encoder 200. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.

[0070] Functional components of the video encoder 200 may include a partition unit 201, a prediction unit 202, which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.

[0071] In other examples, video encoder 200 may include more, fewer, or different functional components. In an example, prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.

[0072] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, may be highly integrated, but are depicted separately in the example of FIG. 5 for illustrative purposes.

[0073] Partition unit 201 may partition a picture into one or more video blocks. Video encoder 200 and video decoder 300 may support a variety of video block sizes.

[0074] The mode select unit 203 may select one of intra or inter coding modes, for example, based on an error result, and provide the resulting intra- or inter-coded block to a residual generation unit 207, which generates residual block data, and to a reconstruction unit 212, which reconstructs the coded block for use as a reference picture. In some examples, the mode select unit 203 may select a combination of intra and inter prediction (CIIP) mode, in which prediction is based on an inter prediction signal and an intra prediction signal. The mode select unit 203 may also select a resolution (e.g., sub-pixel or integer pixel precision) for the motion vector of the block in the case of inter prediction.

[0075] To perform inter prediction on a current video block, motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from buffer 213 other than the picture associated with the current video block.

[0076] Motion estimation unit 204 and motion compensation unit 205 may perform different operations for the current video block depending on whether the current video block is an I slice, a P slice, or a B slice, for example.

[0077] In some examples, motion estimation unit 204 may perform unidirectional prediction for the current video block, and motion estimation unit 204 may look up reference pictures in list 0 or list 1 for reference video blocks for the current video block. Motion estimation unit 204 may then generate a reference index indicating a reference picture in list 0 or list 1 that contains the reference video block, and a motion vector indicating a spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0078] In another example, motion estimation unit 204 may perform bidirectional prediction for the current video block, and motion estimation unit 204 may look up reference pictures in list 0 for a reference video block for the current video block and may also look up reference pictures in list 1 for another reference video block for the current video block. Motion estimation unit 204 may then generate reference indexes that indicate the reference pictures in lists 0 and 1 that contain the reference video blocks, and motion vectors that indicate the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 204 may output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0079] In some examples, the motion estimation unit 204 may output a full set of motion information for the decoding process of the decoder.

[0080] In some examples, motion estimation unit 204 may not output a full set of motion information for the current video. Rather, motion estimation unit 204 may signal the motion information of the current video block by reference to motion information of other video blocks. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0081] In one example, motion estimation unit 204 may indicate in a syntax structure associated with the current video block a value that indicates to video decoder 300 that the current video block has the same motion information as other video blocks.

[0082] In other examples, motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the identified video block. Video decoder 300 may use the motion vector and the motion vector difference of the identified video block to determine the motion vector of the current video block.

[0083] As described above, video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by video encoder 200 include Advanced Motion Vector Prediction (AMVP) and merge mode signaling.

[0084] Intra prediction unit 206 may perform intra prediction on the current video block. When intra prediction unit 206 performs intra prediction on the current video block, intra prediction unit 206 may generate predictive data for the current video block based on decoded samples of other video blocks in the same picture. The predictive data for the current video block may include the predicted video block and various syntax elements.

[0085] Residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., as indicated by a minus sign) the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks that correspond to different sample components of the samples in the current video block.

[0086] In other examples, for example, in skip mode, residual data for the current video block may not be present and residual generation unit 207 may not perform the subtraction operation.

[0087] Transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0088] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0089] Inverse quantization unit 210 and inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. Reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more predicted video blocks generated by prediction unit 202 to generate a reconstructed video block related to the current block for storage in buffer 213.

[0090] After reconstruction unit 212 reconstructs the video blocks, a loop filtering operation may be performed to reduce video blocking artifacts in the video blocks.

[0091] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy-encoded data and generate a bitstream that includes the entropy-encoded data.

[0092] FIG. 6 is a block diagram illustrating an example of a video decoder 300, which may be the video decoder 124 of the system 100 illustrated in FIG.

[0093] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. In the example of FIG. 6, video decoder 300 includes multiple functional components. The techniques described in this disclosure may be shared among various components of video decoder 300. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.

[0094] 6, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. Video decoder 300 may, in some examples, perform a decoding pass that is generally the reverse of the encoding pass described with respect to video encoder 200 (FIG. 5).

[0095] The entropy decoding unit 301 may retrieve an encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 may decode the entropy-coded video data, and from the entropy-decoded video data, the motion compensation unit 302 may determine motion information including motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 may determine such information by, for example, implementing AMVP and merge mode.

[0096] The motion compensation unit 302 may optionally perform interpolation based on an interpolation filter to generate the motion-compensated block. Identifiers for the interpolation filters used with sub-pixel precision may be included in syntax elements.

[0097] Motion compensation unit 302 may use the interpolation filters used by video encoder 200 during encoding of the video block to calculate interpolated values ​​for sub-integer pixels of the reference block. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 according to received syntax information and use the interpolation filters to generate the predictive block.

[0098] The motion compensation unit 302 may use some of the syntax information to determine the size of the blocks used to code the frames and / or slices of the coded video sequence, partition information describing how each macroblock of a picture of the coded video sequence is partitioned, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the coded video sequence.

[0099] The intra prediction unit 303 may use, for example, an intra prediction mode received in the bitstream to form a prediction block from spatially neighboring blocks. The inverse quantization unit 304 inverse quantizes, or dequantizes, the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.

[0100] The reconstruction unit 306 may add to the corresponding prediction block and residual block generated by the motion compensation unit 302 or the intra prediction unit 303 to form a decoded block. If desired, a deblocking filter may also be applied to filter the decoded block to remove blockiness artifacts. The decoded video block is then stored in a buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and further generates the decoded video for presentation on a display device.

[0101] A list of solutions desired by some embodiments is given below.

[0102] The following solutions represent exemplary implementations of the techniques discussed in previous sections (eg, items 1, 10, and 11).

[0103] 1. A visual media processing method (e.g., method 700 depicted in FIG. 3), comprising: performing (702) a conversion between visual media comprising a sequence of one or more images and a bitstream representation in accordance with a file format; the file format is configured to include one or more syntax elements that indicate transition characteristics between one or more images during display of the one or more images; method.

[0104] 2. The method of Solution 1, The transition characteristic is a transition time, and the file format includes another syntax element that indicates the type of transition time, the type having mandatory transition time or recommended transition time. method.

[0105] 3. The method of Solution 1, The transition feature has one or more transition effects between one or more images. method.

[0106] 4. The method of Solution 2, The file format includes one or more syntax elements that describe one or more transition effects that can be applied to transitions between successive images or portions of successive images, method.

[0107] 5. The method of Solution 3, The file format includes a syntax structure that specifies multiple transition effects and corresponding portions of images to which the multiple transition effects are applicable during a transition from one image to the next. method.

[0108] The following solutions illustrate exemplary implementations of the techniques discussed in the previous section (eg, item 2).

[0109] 6. A visual media processing method comprising: performing a conversion between visual media comprising a sequence of one or more images and a bitstream representation in accordance with a file format; When visual media is represented by files with a particular file brand, the file format is restricted according to the rules; method.

[0110] 7. The method of Solution 6, The rules provide that a single access unit of a portion of an image being coded using a particular coding tool, method.

[0111] 8. Solutions 6 and 7, Specific coding tools include intra-coding tools; method.

[0112] 9. Solutions 6 and 7, Specific coding tools include an intra-block copy coding tool; method.

[0113] 10. The method of Solution 6, Specific coding tools include palette coding tools; method.

[0114] 11. The method of Solution 6, The regulations provide that a file format is not permitted to store more than one image coded according to a coding characteristic, method.

[0115] 12. The method of Solution 11, The coding characteristics include a target output layer set characteristic. method.

[0116] The following solutions represent exemplary implementations of the techniques discussed in the previous sections (eg, items 3, 4, 5, and 8).

[0117] 13. A visual media processing method comprising: performing a conversion between visual media comprising a sequence of one or more images and a bitstream representation in accordance with a file format; The file format is configured to indicate the image type of one or more images according to a rule; method.

[0118] 14. The method of Solution 13, The rules specify that the file format further provides that, for one picture type, the file format allows the inclusion of only one access unit containing an intra-coded picture. method.

[0119] 15. The method of Solution 13, The rule specifies that a particular picture type is only allowed to contain Network Abstraction Layer units that contain exactly one access unit that is an intra-random access picture unit, method.

[0120] 16. The method of Solution 13, The rule specifies that for a particular image type, the file format does not allow pictures in layers to be contained that are from different target output layer sets. method.

[0121] 17. The method of Solution 13, The rule specifies that the file format allows for a particular image type to contain an entire access unit that contains one or more pictures that contain multiple extractable subpictures. method.

[0122] 18. Any of the methods 1 to 17, The conversion comprises encoding one or more images to generate a bitstream representation according to a file format; method.

[0123] 19. The method of Solution 18, The bitstream representation according to the file format may be stored on a computer-readable medium or transmitted over a communications connection. method.

[0124] 20. Any of the methods 1-17, The conversion comprises decoding and reconstructing one or more images from a bitstream representation. method.

[0125] 21. The method of Solution 20, and further comprising the step of prompting display of the one or more images after decoding and reconstruction. method.

[0126] 22. A video decoding device having a processor configured to implement the methods described in one or more of solutions 1 to 21.

[0127] 23. A video encoding device having a processor configured to implement the methods described in one or more of solutions 1 to 21.

[0128] 24. A computer program product storing computer code, The code, when executed by a processor, causes the processor to implement a method as recited in any of solutions 1 to 21. Computer program products.

[0129] 25. A computer-readable medium having recorded thereon a bitstream representation conforming to a file format produced by any of solutions 1 to 21.

[0130] 26. Any method, apparatus, or system described herein.

[0131] 8 is a flowchart representation of a method for processing image data in accordance with one or more embodiments of the present technology. Method 800 includes, at operation 810, performing a conversion between a visual media file and a bitstream. The visual media file has a sequence of one or more pictures according to a media file format, and the bitstream has one or more access units according to a video coding format. The bitstream is coded according to the video coding format. The media file format specifies that an image item of a particular type value in the visual media file contains a single access unit of the bitstream. The single access unit is either an Intra Random Access Picture (IRAP) access unit according to the video coding format or a Gradual Decoding Refresh (GDR) access unit according to the video coding format. All pictures within the GDR access unit are identified as recovery points in the bitstream.

[0132] 9 is a flowchart representation of a method for processing image data in accordance with one or more embodiments of the present technology. The method 900 includes, at operation 910, performing a conversion between a visual media file and a bitstream. The visual media file has a sequence of one or more pictures according to a media file format, and the bitstream has one or more access units according to a video coding format. The bitstream is coded according to the video coding format. The media file format specifies that image items of a particular type value in the visual media file exclude layers that do not belong to the target output layer set.

[0133] 10 is a flowchart representation of a method for processing image data in accordance with one or more embodiments of the present technology. Method 1000 includes, at operation 1010, performing a conversion between a visual media file and a bitstream. The visual media file has a sequence of one or more pictures according to a media file format, and the bitstream has one or more access units according to a video coding format. The bitstream is coded according to the video coding format. The media file format specifies that an image item of a particular type value in the visual media file includes at least a portion of an access unit in which the picture has one or more sub-pictures.

[0134] The following are examples of the techniques discussed in connection with FIGS.

[0135] 1. An example of a method for processing video data, comprising: performing a conversion between a visual media file and a bitstream; The visual media file comprises a sequence of one or more pictures according to a media file format, and the bitstream comprises one or more access units according to a video coding format; The bitstream is coded according to a video coding format; The media file format specifies that an image item of a particular type value in a visual media file contains a single access unit of a bitstream, the single access unit being either an Intra Random Access Picture (IRAP) access unit according to the video coding format or a Gradual Decoding Refresh (GDR) access unit according to the video coding format, and all pictures within the GDR access unit are identified as recovery points within the bitstream. method.

[0136] 2. The method of Example 1, The video coding format complies with the VVC (Versatile Video Coding) standard in accordance with ISO / IEC 23090-3. method.

[0137] 3. The method of Example 1 or 2, The specific type value is specified as 'vvc1', method.

[0138] 4. The method of any of Examples 1 to 3, Each of all pictures in a GDR access unit includes a picture header field with a value of zero indicating that the corresponding picture is a recovery point. method.

[0139] 5. The method of Example 4, The picture header field corresponds to the ph_recovery_poc_cnt_field. method.

[0140] 6. A method for processing video data, comprising: performing a conversion between a visual media file and a bitstream; The visual media file comprises a sequence of one or more pictures according to a media file format, and the bitstream comprises one or more access units according to a video coding format; The bitstream is coded according to a video coding format; The media file format specifies that image items of a particular type value in a visual media file exclude layers that do not belong to the target output layer set. method.

[0141] 7. The method of Example 6, The video coding format complies with the VVC (Versatile Video Coding) standard in accordance with ISO / IEC 23090-3. method.

[0142] 8. The method of Example 6 or 7, The specific type value is specified as 'vvc1', method.

[0143] 9. The method of any of Examples 6-8, The image item includes layers in the output layer set identified by a characteristic indicating the target output layer set, and does not include other layers; method.

[0144] 10. A method for processing video data, comprising: performing a conversion between a visual media file and a bitstream; The visual media file comprises a sequence of one or more pictures according to a media file format, and the bitstream comprises one or more access units according to a video coding format; The bitstream is coded according to a video coding format; The media file format specifies that an image item of a particular type value in a visual media file includes at least a portion of an access unit in which a picture has one or more sub-pictures. method.

[0145] 11. The method of Example 10, The video coding format corresponds to the VVC (Versatile Video Coding) standard according to ISO / IEC23090-3, and the specific type value is designated as 'vvc1'. method.

[0146] 12. The method of Example 10 or 11, An image item includes an entire access unit. method.

[0147] 13. The method of any of Examples 10-12, further comprising: The image item comprises a portion of the access unit, For each layer present in the bitstream, one or more subpictures form a rectangular region. method.

[0148] 14. A video processing device having a processor, The processor is configured to execute the method of any of Examples 1 to 13. Video processing equipment.

[0149] 15. A non-transitory computer-readable recording medium storing a video bitstream generated by the method of any of Examples 1 to 13 performed by a video processing device.

[0150] 11 is a flowchart representation of a method for processing image data in accordance with one or more embodiments of the present technology. Method 1100 includes, at operation 1110, performing a conversion between a visual media file and a bitstream. The visual media file has image items, each having a sequence of one or more pictures according to a media file format. The bitstream includes access units, each consisting of one or more pictures, each belonging to a layer according to a video coding format. The media file format specifies that image items with pictures originating from the bitstream are allowed to be associated with different instances of characteristic descriptors that indicate high-level characteristics of the bitstream.

[0151] 12 is a flowchart representation of a method for processing image data in accordance with one or more embodiments of the present technology. Method 1200 includes, at operation 1210, performing a conversion between a visual media file and a bitstream. The visual media file has image items, each including a sequence of one or more pictures according to a media file format. The bitstream has access units, each consisting of one or more pictures, each belonging to a layer according to a video coding format. The media file format specifies, in response to an operation point record included in an operation point characteristic descriptor indicating high-level characteristics of the bitstream, that at least one of a value of a first syntax element in the record or a value of a second syntax element in the record be constrained to a predetermined value.

[0152] The following are examples of the techniques discussed in connection with FIGS.

[0153] 1. An exemplary solution for a method of processing image data, comprising: performing a conversion between a visual media file and a bitstream; the visual media files have image items each having a sequence of one or more pictures according to a media file format, and the bitstream includes access units each consisting of one or more pictures each belonging to a layer according to a video coding format; The media file format specifies that a visual item having pictures originating from a bitstream is allowed to be associated with different instances of characteristic descriptors that indicate high-level characteristics of the bitstream, method.

[0154] 2. The method of example solution 1, The video coding format complies with the VVC (Versatile Video Coding) standard in accordance with ISO / IEC23090-3. method.

[0155] 3. Exemplary solution 1 or 2, A characteristic descriptor is represented as a VvcOperatingPointsInformationProperty. method.

[0156] 4. An exemplary solution of a method for processing image data, comprising: performing a conversion between a visual media file and a bitstream; the visual media files have image items each including a sequence of one or more pictures according to a media file format, and the bitstream includes access units each consisting of one or more pictures each belonging to a layer according to a video coding format; the media file format, in response to an operation point record included in an operation point characteristic descriptor indicating high-level characteristics of the bitstream, specifies that at least one of a value of a first syntax element in the record or a value of a second syntax element in the record is constrained to be a predetermined value; method.

[0157] 5. The method of example solution 4, The video coding format complies with the VVC (Versatile Video Coding) standard in accordance with ISO / IEC23090-3. method.

[0158] 6. The method of exemplary solution 4 or 5, The first syntax element specifies the maximum time discrimination associated with the i-th profile tier level syntax structure, where i ranges from 0 to the number of profile tier levels minus 1; method.

[0159] 7. The method of example solution 6, The first syntax element is represented as ptl_max_temporal_id[i], method.

[0160] 8. Any of the exemplary solutions 4 to 7, The second syntax element specifies the maximum time discrimination associated with recording the operating point. method.

[0161] 9. The method of example solution 8, The second syntax element is denoted as max_temporal_id. method.

[0162] 10. Any of the exemplary solutions 4 to 9, The record includes a third syntax element that specifies whether frame rate information is present, and the value of the third syntax element is constrained to be a predetermined value. method.

[0163] 11. Any of the exemplary solutions 4 to 10, The recording includes a fourth syntax element that specifies whether bit rate information is present, and the value of the fourth syntax element is constrained to be a predetermined value. method.

[0164] 12. Any of the exemplary solutions 4 to 10, The given value is equal to 0, method.

[0165] 13. A video processing device having a processor, The processor is configured to perform the method of any of Examples 1 to 12. Video processing equipment.

[0166] 14. A non-transitory computer-readable recording medium storing a video bitstream generated by the method of any of Examples 1 to 12 executed by a video processing device.

[0167] In the solutions described herein, an encoder may conform to the formatting rules by generating a coded representation in accordance with the formatting rules. In the solutions described herein, a decoder may use the formatting rules to parse syntax elements in the coded representation, knowing their presence or absence, in accordance with the formatting rules, to generate decoded video.

[0168] As used herein, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. The bitstream representation of a current video block may correspond to bits that are either co-located or spread across different locations in the bitstream, e.g., as defined by syntax. For example, a macroblock may be encoded with respect to a transformed and coded error residual value, as well as using bits in headers and other fields in the bitstream. Furthermore, during conversion, a decoder may parse the bitstream knowing the possible presence or absence of some fields based on a decision, as described in the above solution. Similarly, an encoder may determine whether a particular syntax field should be included and generate the bitstream representation by including or excluding the syntax field from the bitstream representation accordingly.

[0169] The disclosed and other solutions, examples, embodiments, modules, and functional operations described herein can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in one or more combinations thereof. The disclosed and other embodiments can be implemented as one or more computer program products, e.g., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or to control the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage carrier, a memory device, a composition of matter bearing a machine-readable propagated signal, or one or more combinations thereof. The term "data processing apparatus" encompasses all apparatus, devices, and machines that process data, including, by way of example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, an apparatus can include code that creates an execution environment for the computer program in question, such as code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations thereof. A propagated signal is an artificially generated signal, for example, a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device.

[0170] A computer program (also known as a program, software, software application, script, or code) can be written in any programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a single file dedicated to the program in question, or in multiple coordinated files (e.g., a file storing one or more modules, subprograms, or portions of code), or in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communications network.

[0171] The processes and logic flows described herein may be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be performed by, and apparatus may be implemented as, special purpose logic circuitry, such as a field programmable gate array (FPGA) or application specific integrated circuit (ASIC).

[0172] Processors suitable for executing a computer program include, by way of example, both general-purpose and special-purpose microprocessors, as well as one or more processors of any kind of digital computer. Typically, a processor will read instructions and data from read-only memory or random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices, e.g., magnetic, optical-magnetic, or optical disks, for storing data, or be operatively coupled to receive data from, transfer data to, or both, such one or more mass storage devices. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include, by way of example, semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal or removable disks; optical-magnetic disks; and all forms of non-volatile memory, media, and memory devices, including CD-ROM and DVD-ROM disks. The processor and memory may be enhanced by, or incorporated in, dedicated logic circuitry.

[0173] While this specification contains numerous details, these should not be construed as limitations on the scope of any subject matter or what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular technology. Certain features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, while features may be described above as operating in a particular combination and even initially claimed as such, one or more features from a claimed combination may in some cases be deleted from that combination, and the claimed combination may be directed to a subcombination or a variation of the subcombination.

[0174] Similarly, although operations are depicted in the figures in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown, or in a sequential order, or that all of the depicted operations be performed, to achieve desired results. Further, the separation of various system components in the embodiments described herein should not be understood as requiring such separation in all embodiments.

[0175] Only a few implementations and examples have been described; other implementations, enhancements and variations can be made based on what is described and illustrated in this patent document.

Claims

1. 1. A method for processing image data, comprising: performing a conversion between a visual media file and a bitstream; the visual media file comprises image items, each image item being an image according to a media file format, the bitstream comprising access units each consisting of one or more pictures each belonging to a layer according to a video coding format; the media file format specifies that image items containing images originating from the bitstream are allowed to be associated with different content of characteristic descriptors that indicate characteristics of the bitstream; the media file format further specifies that, in response to inclusion of a record of the operation point in an operation point characteristic descriptor that describes a characteristic of the bitstream, at least one of a value of a first syntax element in the record or a value of a second syntax element in the record is constrained to be a predetermined value; the first syntax element specifies a maximum time discrimination associated with an i-th profile tier level syntax structure, where i ranges from 0 to (number of profile tier levels minus 1), and the second syntax element specifies a maximum time discrimination associated with the recording of the operating point. method.

2. The video coding format corresponds to the VVC standard according to ISO / IEC 23090-3; The method of claim 1.

3. The characteristic descriptor is represented as VvcOperatingPointsInformationProperty.

3. The method according to claim 1 or 2.

4. The first syntax element is represented as ptl_max_temporal_id[i] and the second syntax element is represented as max_temporal_id. The method of claim 1.

5. the predetermined value is equal to 0; The method according to claim 1 or 4.

6. the converting includes encoding the visual media file into the bitstream.

6. The method according to any one of claims 1 to 5.

7. the converting includes decoding the visual media file from the bitstream.

6. The method according to any one of claims 1 to 5.

8. 1. An apparatus for processing visual media file data, comprising: a processor; a non-transitory memory having instructions; and The instructions, when executed by the processor, cause the processor to convert between a visual media file and a bitstream; the visual media file comprises image items, each image item being an image according to a media file format, the bitstream comprising access units each consisting of one or more pictures each belonging to a layer according to a video coding format; the media file format specifies that image items containing images originating from the bitstream are allowed to be associated with different content of characteristic descriptors that indicate characteristics of the bitstream; the media file format further specifies that, in response to inclusion of a record of the operation point in an operation point characteristic descriptor that describes a characteristic of the bitstream, at least one of a value of a first syntax element in the record or a value of a second syntax element in the record is constrained to be a predetermined value; the first syntax element specifies a maximum time discrimination associated with an i-th profile tier level syntax structure, where i ranges from 0 to (number of profile tier levels minus 1), and the second syntax element specifies a maximum time discrimination associated with the recording of the operating point. Device.

9. A non-transitory computer-readable storage medium storing instructions, comprising: The instructions cause the processor to perform a conversion between a visual media file and a bitstream; the visual media file comprises image items, each image item being an image according to a media file format, the bitstream comprising access units each consisting of one or more pictures each belonging to a layer according to a video coding format; the media file format specifies that image items containing images originating from the bitstream are allowed to be associated with different content of characteristic descriptors that indicate characteristics of the bitstream; the media file format further specifies that, in response to inclusion of a record of the operation point in an operation point characteristic descriptor that describes a characteristic of the bitstream, at least one of a value of a first syntax element in the record or a value of a second syntax element in the record is constrained to be a predetermined value; the first syntax element specifies a maximum time discrimination associated with an i-th profile tier level syntax structure, where i ranges from 0 to (number of profile tier levels minus 1), and the second syntax element specifies a maximum time discrimination associated with the recording of the operating point. A non-transitory computer-readable storage medium.

10. 1. A method for generating and storing a video bitstream, comprising: generating the bitstream from a visual media file based on a media file format; storing the bitstream on a non-transitory computer-readable recording medium; and the visual media file comprises image items, each image item being an image according to the media file format, and the bitstream comprises access units each consisting of one or more pictures, each belonging to a layer according to a video coding format; the media file format specifies that image items containing images originating from the bitstream are allowed to be associated with different content of characteristic descriptors that indicate characteristics of the bitstream; the media file format further specifies that, in response to inclusion of a record of the operation point in an operation point characteristic descriptor that describes a characteristic of the bitstream, at least one of a value of a first syntax element in the record or a value of a second syntax element in the record is constrained to be a predetermined value; the first syntax element specifies a maximum time discrimination associated with an i-th profile tier level syntax structure, where i ranges from 0 to (number of profile tier levels minus 1), and the second syntax element specifies a maximum time discrimination associated with the recording of the operating point. method.

Citation Information

Patent Citations

  • Designing a Multilayer Video File Format

    JP2016540416A

  • Method, device and computer program for dynamically setting operational origin descriptors for retrieving media data and metadata from encapsulated bitstreams

    JP2018524877A