Media file processing method and apparatus
The VVC standard configuration file format solves the problem of efficient storage and transmission of high-resolution and high-quality images and videos, simplifies media file generation and processing, and improves data transmission efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- LG ELECTRONICS INC
- Filing Date
- 2021-09-23
- Publication Date
- 2026-05-19
AI Technical Summary
As the demand for high-resolution and high-quality images and videos increases, existing technologies struggle to effectively compress, transmit, or store these media files, leading to increased transmission and storage costs.
By adopting the VVC standard configuration file format, and generating and processing media files, it prevents incomplete slices in sub-picture tracks, ensures that slices are fully contained in sub-pictures, and simplifies the relationship between VVC sub-picture tracks, sub-pictures, and slices.
It enables efficient storage and transmission of video/image data, simplifies the generation and processing of media files, and improves data transmission efficiency.
Smart Images

Figure CN116406505B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to image coding technology, and more specifically, to a method for processing a media file containing encoded image information for use in an image coding system, and an apparatus for processing a media file containing encoded image information for use in an image coding system. Background Technology
[0002] Recently, there has been a growing demand across various fields for high-resolution and high-quality images and videos, such as 4K, 8K, or even higher Ultra High Definition (UHD) images and videos. As the resolution and quality of image and video data increase, the amount of information or bits transmitted increases relatively compared to existing image and video data. Therefore, using media such as existing wired or wireless broadband lines to transmit image data, or using existing storage media to store image and video data, increases both transmission and storage costs.
[0003] Furthermore, interest in and demand for immersive media, such as virtual reality (VR), artificial reality (AR) content, or 360-degree holographic video, is growing. Broadcasting of images and videos with characteristics distinct from real images, such as game graphics, is also increasing.
[0004] Therefore, there is a need for an efficient image and video compression technology to effectively compress and send or store and play information with such high resolution and high quality images and videos.
[0005] However, even if the video information is compressed, as mentioned above, with the increasing demand for high-quality images, the size (or amount) of the image information being processed increases, and the size of the compressed image information is expected to be even larger than that of existing technologies.
[0006] Therefore, in order to process high-resolution, high-capacity video information, the following problem exists: how to define the format of video information files and the file containing information for processing such video information files so that the media can be processed effectively. Summary of the Invention
[0007] Technical solution
[0008] According to one embodiment of this disclosure, a method and apparatus for efficiently storing and transmitting video / audio data are provided herein.
[0009] According to one embodiment of this disclosure, this document provides a method and apparatus for configuring (or forming) a file format that can be used with support for VVC.
[0010] According to one embodiment of this disclosure, this document provides a method and apparatus for forming tracks for storing and transmitting video / image data and for generating media files.
[0011] According to one embodiment of this disclosure, this document provides a method and apparatus for preventing a situation where a sub-image track is allowed to include one or more complete slices but not all slices are included in a single sub-image.
[0012] According to one embodiment of this disclosure, this document provides a method and apparatus for preventing a situation where a sub-image track includes slices but not all slices are included in a single sub-image.
[0013] According to one embodiment of this disclosure, a method for generating media files, performed by a media file generating device, is provided herein.
[0014] According to one embodiment of this disclosure, a media file generation device for generating media files is provided herein.
[0015] According to one embodiment of this disclosure, a method for processing media files, performed by a media file processing device, is provided herein.
[0016] According to one embodiment of this disclosure, a media file processing device for processing media files is provided herein.
[0017] Beneficial effects
[0018] According to one embodiment of this disclosure, video / image data can be stored and transmitted efficiently.
[0019] According to one embodiment of this disclosure, a file format that can be used by supporting VVC can be configured (or formed).
[0020] According to one embodiment of this disclosure, a track for storing and transmitting video / image data can be formed, and a media file can be generated.
[0021] According to one embodiment of this disclosure, it is possible to prevent situations where a sub-picture track is allowed to include one or more complete slices but not all slices are included in a single sub-picture.
[0022] According to one embodiment of this disclosure, it is possible to prevent a situation where a sub-image track includes slices but not all slices are included in a single sub-image.
[0023] According to one embodiment of this disclosure, the relationship between VVC sub-image tracks, VVC sub-images, and slices within VVC sub-images can become simple and clear. Attached Figure Description
[0024] Figure 1 and Figure 2 This is a diagram illustrating an example of a media file structure.
[0025] Figure 3 This is a diagram illustrating the overall operation of a DASH-based adaptive flow model according to one embodiment of the present disclosure.
[0026] Figure 4 Examples of video / image coding systems to which the exemplary implementations of this document are applicable are illustrated schematically.
[0027] Figure 5 This is a diagram illustrating the configuration of a video / image encoding device to which the exemplary implementation of this document applies.
[0028] Figure 6 This is a diagram illustrating the configuration of a video / image decoding device to which the exemplary embodiments of this document apply.
[0029] Figure 7 An exemplary layered structure for encoded video / images is shown.
[0030] Figure 8 An example of a method for generating media files using embodiments of the present disclosure is shown.
[0031] Figure 9 An example of a method for processing a media file generated by applying the embodiments proposed in this disclosure is shown.
[0032] Figure 10 An overall view of a method for generating media files performed by a device for generating media files, according to this disclosure, is shown.
[0033] Figure 11 An overall view of a device for generating media files is shown, illustrating a method for generating media files according to this disclosure.
[0034] Figure 12 An overall view of a method for generating media files performed by a device for processing media files, according to this disclosure, is shown.
[0035] Figure 13 An overall view of a device for processing media files is shown, illustrating a method for processing media files according to this disclosure.
[0036] Figure 14 An exemplary structural diagram is shown for a content streaming system to which the implementation methods disclosed in this document are applied. Detailed Implementation
[0037] This document may be modified in various forms, and specific embodiments thereof will be described and illustrated in the accompanying drawings. However, these embodiments are not intended to limit this document. The terminology used in the following description is for the purpose of describing particular embodiments only and is not intended to limit this document. Singular expressions include plural expressions, provided that different interpretations are clear. Terms such as “comprising” and “having” are intended to indicate the presence of the features, quantities, steps, operations, elements, components, or combinations thereof used in the following description, and therefore it should be understood that the possibility of having or adding one or more different features, quantities, steps, operations, elements, components, or combinations thereof is not excluded.
[0038] On the other hand, for ease of explanation of different specific functions, the elements in the accompanying drawings described in this document are drawn independently and do not imply that these elements are implemented by independent hardware or independent software. For example, two or more elements may be combined to form a single element, or a single element may be divided into multiple elements. The implementation of combining and / or dividing elements is part of this document without departing from its concepts.
[0039] In the following, preferred embodiments of this document will be described in more detail with reference to the accompanying drawings. Throughout this specification, the same reference numerals will be used to refer to the same components, and redundant descriptions of identical parts may be omitted.
[0040] Figure 1 and Figure 2 This is a diagram illustrating an example of the structure of a media file.
[0041] A media file according to an implementation may include at least one frame. Here, a frame may be a data block or object that includes media data or metadata associated with the media data. The frames may be in a hierarchical structure, and thus the data can be categorized, and the media file may have a format suitable for storing and / or transmitting large volumes of media data. Furthermore, the media file may have a structure that allows users to easily access media information such as moving to a specific media content point.
[0042] A media file according to one implementation may include an ftyp box, a moov box, and / or a mdat box.
[0043] The ftyp box (file type box) provides information about the file type or compatibility of the corresponding media file. The ftyp box can include configuration version information about the media data of the corresponding media file. Decoders can refer to the ftyp box to identify the corresponding media file.
[0044] A moov frame (movie frame) can be a frame that includes metadata about the media data of the corresponding media file. A moov frame can serve as a container for all metadata. The moov frame can be the highest-level frame associated with metadata. According to one implementation, only one moov frame may exist in a media file.
[0045] A mdat frame (media data frame) can be a frame containing the actual media data of a corresponding media file. Media data can include audio samples and / or video samples. The mdat frame can be used as a container to hold such media samples.
[0046] According to the implementation method, the above-mentioned moov box may also include an mvhd box, a trak box and / or an mvex box as a lower box.
[0047] The MVHD header frame can include information related to the media presentation of the media data included in the corresponding media file. In other words, the MVHD frame can include information such as the media generation time, change time, time standard, and time period of the corresponding media presentation.
[0048] A trak frame provides information about the corresponding media track. A trak frame can include information such as streaming-related information, rendering-related information, and access-related information about audio or video tracks. Multiple trak frames can exist, depending on the number of tracks.
[0049] The `trak` box can also include a `tkhd` box (track header box) as its lower box. The `tkhd` box can include information about the track indicated by the `trak` box. The `tkhd` box can include information such as the creation time, change time, and track identifier of the corresponding track.
[0050] The mvex box (movie extension box) can indicate that the corresponding media file may have a moof box, which will be described later. Scanning the moof boxes may be necessary to identify all media samples for a specific track.
[0051] According to one implementation (200), a media file can be divided into multiple segments. Therefore, the media file can be segmented and stored or transmitted. The media data (mdat frames) of the media file can be divided into multiple segments, each segment including a moof frame and the divided mdat frame. According to one implementation, information from the ftyp frame and / or moov frame may be required to use the segment.
[0052] A moof (movie clip frame) can provide metadata about the media data of the corresponding clip. The moof can be the highest-level frame among the frames associated with the metadata of the corresponding clip.
[0053] The mdat frame (media data frame) can include the actual media data as described above. The mdat frame can include media samples corresponding to the media data of each segment it corresponds to.
[0054] According to one implementation, the above-mentioned moof box may further include an mfhd box and / or a traf box as a lower box.
[0055] The MFHD (Movie Clip Header) box can include information about the correlation between the segmented clips. The MFHD box can indicate the order of the segmented media data by including sequence numbers. Furthermore, the MFHD box can be used to check for missing data within the segmented data.
[0056] A TRAFF box (track segment box) can contain information about the corresponding track segment. A TRAFF box can provide metadata about the segmented track segments included in the corresponding segment. The metadata provided by the TRAFF box allows the media samples in the corresponding track segment to be decoded / reproduced. Depending on the number of track segments, multiple TRAFF boxes can exist.
[0057] According to one implementation, the above-mentioned traf box may also include a tfhd box and / or a trun box as the lower box.
[0058] The TFHD frame (track segment header frame) can include header information for the corresponding track segment. The TFHD frame can provide information such as the basic sample size, time period, offset, and identifier of the media sample for the track segment indicated by the TRAFF frame mentioned above.
[0059] The trun box can include information related to the corresponding track segment. The trun box can include information such as the duration, size, and playback time of each media sample.
[0060] The aforementioned media files and their fragments can be processed into segments and sent. Segmentation may include initialization segments and / or media segments.
[0061] In addition to media data, the file in embodiment 210 may also include information related to media decoder initialization. For example, the file may correspond to the aforementioned initialization segment. The initialization segment may include the aforementioned ftyp box and / or moov box.
[0062] The file in embodiment 220 shown may include the aforementioned segments. For example, the file may correspond to the aforementioned media segments. Media segments may also include styp frames and / or sidx frames.
[0063] A styp box (segment type box) can provide information for identifying media data that is divided into segments. A styp box can be used as the aforementioned ftyp box for segmenting segments. In some implementations, the styp box can have the same format as the ftyp box.
[0064] The sidx box (segment index box) can provide information indicating the index of the segment. Therefore, it can indicate the order of segmentation.
[0065] According to embodiment 230, an ssix box may also be included. The ssix box (subsegment index box) can provide information indicating the index of the subsegment when a segment is divided into subsegments.
[0066] For example, boxes in a media file may include further extended information based on boxes or FullBoxes, as shown in Embodiment 250. In this embodiment, the size and largesize fields may represent the length of the corresponding box in bytes. The version field may indicate the version of the corresponding box format. The type field may indicate the type or identifier of the corresponding box. The flag field may indicate a flag associated with the corresponding box.
[0067] Furthermore, the video / image fields (characteristics) in this document can be transmitted by being included in the DASH-based adaptive streaming model.
[0068] Figure 3 This is a diagram illustrating the overall operation of a DASH-based adaptive streaming model according to an embodiment of this disclosure. The DASH-based adaptive streaming model according to the embodiment shown in (400) describes the operation between an HTTP server and a DASH client. Here, Dynamic Adaptive Streaming over HTTP (DASH) (which is a protocol for supporting HTTP-based adaptive streaming) can dynamically support streaming based on network conditions. As a result, AV content can be reproduced without interruption.
[0069] First, the DASH client can obtain the MPD. The MPD can be delivered from a service provider such as an HTTP server. The DASH client can use information about access to the segment described in the MPD to request that segment from the server. Here, this request can be performed taking network conditions into account.
[0070] After acquiring the segments, the DASH client can use the media engine to process them and display them on the screen. The DASH client can request and acquire the necessary segments in real time, taking into account playback time and / or network conditions (adaptive streaming). As a result, content can be played back without interruption.
[0071] A Media Presentation Description (MPD) is a file that includes detailed information about segments that enables DASH clients to dynamically retrieve them, and can be expressed in XML format.
[0072] The DASH client controller can generate commands for requesting MPD and / or segments, taking network conditions into account. Additionally, the controller can perform controls to make the acquired information usable in internal blocks, such as those of the media engine.
[0073] The MPD parser can parse the acquired MPD in real time. While doing so, the DASH client controller can generate commands to retrieve the necessary segments.
[0074] The segment parser can parse the acquired segments in real time. Internal blocks (such as media engines) can perform specific operations based on the information included in the segments.
[0075] An HTTP client can request the necessary MPD and / or necessary segments from an HTTP server. Alternatively, the HTTP client can pass the MPD and / or segments obtained from the server to an MPD parser or a segment parser.
[0076] Media engines can use media data included in segments to display content. In this case, information from the MPD can be used.
[0077] The DASH data model can have a hierarchical structure (410). Media presentation can be described by an MPD. An MPD can describe the time series of multiple time periods in which media presentation takes place. A time period can indicate a part of the media content.
[0078] Within a given time period, data can be included in an adapter set. An adapter set can be a collection of media content components that can be exchanged with each other. An adapter can include a set of representations. One representation can correspond to one media content component. Within a representation, content can be divided into multiple segments over time. This can be used for appropriate access and delivery. A URL for each segment can be provided to facilitate access to each segment.
[0079] MPD can provide information related to media presentation. Time period elements, adaptation set elements, and representation elements can describe the corresponding time period, adaptation set, and representation, respectively. A representation can be divided into sub-representations. Sub-representation elements can describe the corresponding sub-representations.
[0080] Here, public properties / elements can be defined. Public properties / elements can be applied to (or included in) adapter sets, representations, and sub-representations. EssentialProperty and / or SupplementalProperty can be included in public properties / elements.
[0081] EssentialProperty can be information that includes elements deemed necessary for processing data related to media presentation. SupplementalProperty can be information that includes elements that can be used to process data related to media presentation. In some implementations, when signaling information (described later) is transmitted via MPD, the signaling information can be transmitted simultaneously in both EssentialProperty and / or SupplementalProperty.
[0082] Figure 4 Examples of video / image coding systems to which the implementation methods described in this document can be applied are illustrated schematically.
[0083] Reference Figure 4 A video / image encoding system may include a first device (source device) and a second device (receiving device). The source device may send the encoded video / image information or data to the receiving device in the form of a file or stream via a digital storage medium or network.
[0084] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0085] Video sources can acquire video / images through processes that capture, synthesize, or generate video / images. Video sources may include video / image capture devices and / or video / image generation devices. For example, a video / image capture device may include one or more cameras, a video / image archive containing previously captured video / images, etc. For example, a video / image generation device may include a computer, tablet computer, and smartphone, and can generate video / images (electronically). For example, virtual video / images can be generated via a computer, etc. In this case, the video / image capture process can be replaced by a process that generates related data.
[0086] Encoding devices can encode input video / images. For compression and encoding efficiency, encoding devices can perform a series of processes such as prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output as a bitstream.
[0087] A transmitter can send encoded images / image information or data, output as a bitstream, to a receiver in a receiving device via a digital storage medium or network, either as a file or a stream. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include elements for generating media files according to a predetermined file format and may include elements for transmission over a broadcast / communication network. The receiver can receive / extract the bitstream and send the received bitstream to a decoding device.
[0088] Decoding devices can decode video / images by performing a series of processes such as dequantization, inverse transform, and prediction, which correspond to the operations of encoding devices.
[0089] The renderer can render the decoded video / image. The rendered video / image can be displayed on a monitor.
[0090] This document relates to video / image coding. For example, the methods / implementations disclosed in this document can be applied to methods disclosed in the Universal Video Coding (VVC) standard, the Basic Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the Audio Video Coding Standard 2 (AVS2) or other next-generation video / image coding standards (e.g., H.267, H.268, etc.).
[0091] This disclosure provides various implementations related to video / image encoding, and unless otherwise specified, these implementations may also be combined with each other.
[0092] In this document, video can refer to a series of images over time. An image typically refers to a unit representing an image within a specific time period, and a sub-image / slice / tile refers to a unit that constitutes part of an image during encoding. A sub-image / slice / tile may include one or more Code Tree Units (CTUs). An image can be composed of one or more sub-images / slices / tiles. An image can be composed of one or more sub-images / slices / tiles. An image can be configured from one or more tile groups. A tile group may include one or more tiles. A tile can represent a rectangular area of CTU rows within a tile in an image. A tile can be divided into multiple tiles, and each tile can consist of one or more CTU rows within the tile. A tile that is not divided into multiple tiles can also be called a tile. A tile scan can represent a specific order of CTUs that segment an image. In this paper, CTUs can be aligned by CTU raster scans within tiles, tiles within tiles can be aligned continuously (or coherently) by raster scans of tiles within tiles, and tiles within an image can be aligned continuously by raster scans of tiles within the image. A sub-image can represent a rectangular region of one or more slices within an image. That is, a sub-image can include one or more slices that collectively cover a rectangular region of the image. A tile is a rectangular region of CTUs within a specific tile column and a specific tile column. A tile column is a rectangular region of CTUs with the same height as the image, and the width of the rectangular region can be specified by syntax elements in the image parameter set. A tile row is a rectangular region of CTUs with the same width as the image height, specified by syntax elements in the image parameter set. A tile scan can represent a specific order in which CTUs are segmented within an image; CTUs can be aligned continuously by CTU raster scans within tiles, and tiles within an image can be aligned continuously by raster scans of tiles within the image. A slice can include an integer number of tiles from an image, where the integer number of tiles can belong to NAL cells. A slice can consist of multiple complete tiles, or it can be a continuous sequence of complete tiles. In this disclosure, tile groups and slices can be used interchangeably. For example, in this disclosure, a tile group / tile group header can also be referred to as a slice / slice header.
[0093] A pixel, or image unit, can refer to the smallest unit that makes up a picture (or image). Additionally, the term "sample" can be used as the counterpart to a pixel. A sample can typically represent a pixel or a pixel value, and can represent pixel / pixel values for either the luminance component or the chrominance component.
[0094] A unit can represent the basic unit of image processing. A unit may include a specific region of an image and at least one of the information associated with that region. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "region". In general, an M×N block may include a set (or array) of samples (or sample arrays) or transform coefficients in M columns and N rows.
[0095] In this document, the term "A or B" may mean "A only", "B only", or "both A and B". In other words, in this document, the term "A or B" may be interpreted as indicating "A and / or B". For example, in this document, the term "A, B, or C" may mean "A only", "B only", "C only", or "any combination of A, B, and C".
[0096] The forward slash ( / ) or comma used in this document can mean "and / or". For example, "A / B" can mean "A and / or B". Therefore, "A / B" can mean "A only", "B only", or "both A and B". For example, "A, B, C" can mean "A, B, or C".
[0097] In this document, "at least one of A and B" may mean "A only", "B only", or "both A and B". Additionally, in this document, the expression "at least one of A or B" or "at least one of A and / or B" may be interpreted as the same as "at least one of A and B".
[0098] Furthermore, in this document, "at least one of A, B, and C" may mean "A only", "B only", "C only" or "any combination of A, B, and C". Additionally, "at least one of A, B, or C" or "at least one of A, B, and / or C" may mean "at least one of A, B, and C".
[0099] Furthermore, the parentheses used in this document may mean "for example". Specifically, when expressing "prediction (intra-prediction)", it may indicate that "intra-prediction" is proposed as an example of "prediction". In other words, the term "prediction" in this document is not limited to "intra-prediction", and may indicate that "intra-prediction" is proposed as an example of "prediction". Moreover, even when expressing "prediction (i.e., intra-prediction)", it may indicate that "intra-prediction" is proposed as an example of "prediction".
[0100] The technical features described individually in a diagram in this document can be implemented individually or simultaneously.
[0101] The following figures are provided to illustrate specific examples in this document. Since the names of specific devices or signals / messages / fields described in the figures are provided as examples, the technical features of this document are not limited to the specific names used in the following figures.
[0102] Figure 5 This is a schematic diagram illustrating the configuration of a video / image encoding apparatus to which the embodiments described in this document can be applied. Hereinafter, the encoding apparatus may include a video encoding apparatus and / or an image encoding apparatus.
[0103] Reference Figure 5 The encoding device 200 includes an image segmenter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstruction block generator. According to one embodiment, the image segmenter 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 may be configured by at least one hardware component (e.g., an encoder chipset or processor). Additionally, the memory 270 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware component may also include the memory 270 as an internal / external component.
[0104] Image segmenter 210 can segment an input image (or picture or frame) input to encoding device 200 into one or more processors. For example, a processor may be called a coding unit (CU). In this case, the coding unit can be recursively segmented from a coding tree unit (CTU) or a maximum coding unit (LCU) according to a quadtree-binary-trinary tree (QTBTTT) structure. For example, a coding unit can be segmented into multiple deeper coding units based on a quadtree structure, a binary tree structure, and / or a ternary structure. In this case, for example, a quadtree structure can be applied first, followed by a binary tree structure and / or a ternary structure. Alternatively, a binary tree structure can be applied first. The encoding process according to this document can be performed based on the final coding unit that is no longer segmented. In this case, the maximum coding unit can be used as the final coding unit based on image characteristics, coding efficiency, etc., or, if necessary, the coding unit can be recursively segmented into deeper coding units, and the coding unit with the optimal size can be used as the final coding unit. Here, the encoding process may include prediction, transformation, and reconstruction processes (described later). As another example, the processor may also include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit can be split or divided from the final encoding unit described above. The prediction unit may be a unit for predicting samples, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving residual signals from transform coefficients.
[0105] In some cases, a unit can be used interchangeably with terms such as block or region. Generally, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. Samples can typically represent pixels or pixel values, and can represent pixel / pixel values of only the luminance component or only the chrominance component. A sample can be used as a term corresponding to a picture (or image) of pixels or cells.
[0106] In the encoding device 200, a residual signal (residual block, residual sample array) is generated by subtracting the prediction signal (prediction block, prediction sample array) output from the inter-frame predictor 221 or the intra-frame predictor 222 from the input image signal (original block, original sample array), and the generated residual signal is sent to the converter 232. In this case, as shown, the portion in the encoder 200 that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) can be referred to as the subtractor 231. The predictor can perform prediction on the block to be processed (hereinafter referred to as the current block) and generate a prediction block that includes the prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. As described later in the description of each prediction mode, the predictor can generate various prediction-related information such as prediction mode information and send the generated information to the entropy encoder 240. The information about the prediction can be encoded in the entropy encoder 240 and output as a bitstream.
[0107] Intra-predictor 222 can refer to samples in the current image to predict the current block. Depending on the prediction mode, the referenced samples may be located near or separated from the current block. In intra-prediction, the prediction mode may include multiple non-directional modes and multiple directional modes. For example, non-directional modes may include DC mode and planar mode. For example, depending on the level of detail in the prediction direction, the directional modes may include 33 or 65 directional prediction modes. However, this is just an example, and more or fewer directional prediction modes may be used depending on the settings. Intra-predictor 222 can use the prediction modes applied to neighboring blocks to determine the prediction mode applied to the current block.
[0108] Inter-frame predictor 221 can deduce the predicted block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference image. Here, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. The reference image including the reference block and the reference image including the temporally neighboring block may be the same or different. The temporally neighboring block may be referred to as a juxtaposed reference block, juxtaposed CU (colCU), etc., and the reference image including the temporally neighboring block may be referred to as a juxtaposed image (colPic). For example, inter-frame predictor 221 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to deduce the motion vector and / or reference image index of the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame predictor 221 can use motion information from neighboring blocks as motion information for the current block. In skip mode, unlike merge mode, residual signals may not be sent. In motion vector prediction (MVP) mode, motion vectors from neighboring blocks can be used as motion vector predictors, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0109] Predictor 220 can generate prediction signals based on various prediction methods described below. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but also apply both intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as Intra-Frame and Inter-Frame Prediction Combination (CIIP). Alternatively, the predictor can predict blocks based on Intra-Block Copy (IBC) prediction mode or Palette mode. IBC prediction mode or Palette mode can be used for content image / video coding such as games, for example, Screen Content Coding (SCC). IBC essentially performs prediction in the current frame, but can be performed similarly to inter-frame prediction in terms of deriving reference blocks in the current frame. That is, IBC can use at least one of the inter-frame prediction techniques described in this document. Palette mode can be considered as an example of intra-frame coding or intra-frame prediction. When applying Palette mode, sample values within the frame can be signaled based on information about the palette table and palette index.
[0110] The predicted signal generated by the predictor (including inter-frame predictor 221 and / or intra-frame predictor 222) can be used to generate a reconstructed signal or a residual signal. Transformer 232 can generate transform coefficients by applying transform techniques to the residual signal. For example, the transform technique can include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-loève Transform (KLT), Graph-Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is represented as a graph. CNT refers to a transform generated based on the predicted signal generated using all previously reconstructed pixels. Furthermore, the transform processing can be applied to square pixel blocks of the same size, or to blocks of variable size that are not square.
[0111] Quantizer 233 can quantize the transform coefficients and send them to entropy encoder 240, which can encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients can be referred to as residual information. Quantizer 233 can rearrange the block-type quantized transform coefficients into a one-dimensional vector based on the coefficient scan order, and generate information about the quantized transform coefficients based on this one-dimensional vector form. Entropy encoder 240 can perform various encoding methods such as (for example) exponential Golomb, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). Entropy encoder 240 can encode information necessary for video / image reconstruction (e.g., values of syntax elements) other than the quantized transform coefficients, either together or separately. The encoded information (e.g., encoded video / image information) can be sent or stored in NAL (Network Abstraction Layer) units as a bitstream. The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). Additionally, the video / image information may include general constraint information. In this document, information and / or syntax elements sent / signed from the encoding device to the decoding device may be included in the video / image information. The video / image information can be encoded using the encoding process described above and included in a bitstream. The bitstream may be transmitted over a network or stored in a digital storage medium. The network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that sends signals output from the entropy encoder 240 or a storage device (not shown) that stores signals may be included as an internal / external element of the encoding device 200, and alternatively, the transmitter may be included within the entropy encoder 240.
[0112] The quantized transform coefficients output from quantizer 233 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients using dequantizer 234 and inverse transformer 235. Adder 250 adds the reconstructed residual signal to the prediction signal output from inter-frame predictor 221 or intra-frame predictor 222 to generate a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample array). If the target block to be processed has no residual (e.g., in the case of applying a skip mode), the prediction block can be used as a reconstructed block. Adder 250 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current image, and can be filtered for inter-frame prediction of the next image, as described below.
[0113] In addition, a luminance mapping with chroma scaling (LMCS) can be applied during image encoding and / or reconstruction processing.
[0114] Filter 260 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 260 can generate a corrected reconstructed image by applying various filtering methods to the reconstructed image and store the corrected reconstructed image in memory 270, specifically in the DPB of memory 270. Various filtering methods can include, for example, deblocking filtering, sample adaptive offset (SAO), adaptive loop filtering, bilateral filtering, etc. Filter 260 can generate various types of filtering-related information and transmit the generated information to entropy encoder 290, as described in the subsequent descriptions of each filtering method. The filtering-related information can be encoded by entropy encoder 290 and output as a bitstream.
[0115] The corrected reconstructed image sent to memory 270 can be used as a reference image in inter-frame predictor 221. When inter-frame prediction is applied by the encoding device, prediction mismatch between the encoding device 200 and the decoding device can be avoided, and encoding efficiency can be improved.
[0116] The DPB of memory 270 can store the corrected reconstructed image for use as a reference image in inter-frame predictor 221. Memory 270 can store motion information of blocks from which motion information in the current image is derived (or encoded) and / or motion information of blocks in already reconstructed images. The stored motion information can be transmitted to inter-frame predictor 221 to be used as motion information for spatially or temporally neighboring blocks. Memory 270 can store reconstructed samples of reconstructed blocks in the current image and can transmit these reconstructed samples to intra-frame predictor 222.
[0117] Figure 6This is a schematic diagram illustrating the configuration of a video / image decoding device to which the embodiments described in this document can be applied. Hereinafter, the decoding device may include a video decoding device and / or an image decoding device.
[0118] Reference Figure 6 The decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 332 and an intra-frame predictor 331. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321. According to one embodiment, the entropy decoder 310, residual processor 320, predictor 330, adder 340, and filter 350 may be configured by hardware components (e.g., a decoder chipset or processor). Additionally, the memory 360 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.
[0119] When the input includes a bitstream containing video / image information, the decoding device 300 can reconstruct and... Figure 5 The encoding device processes video / image information corresponding to the image. For example, the decoding device 300 can deduce units / blocks based on block segmentation information obtained from the bitstream. The decoding device 300 can use a processor applied in the encoding device to perform decoding. Therefore, for example, the processor for decoding can be an encoding unit, and the encoding unit can be segmented from encoding tree units or maximum encoding units according to a quadtree structure, binary tree structure, and / or ternary tree structure. One or more transform units can be derived from the encoding units. The reconstructed image signal decoded and output by the decoding device 300 can be reproduced by a reproduction device.
[0120] Decoding device 300 can receive data from... in the form of a bitstream. Figure 5The signal output by the encoding device can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive the information necessary for image reconstruction (or picture reconstruction) (e.g., video / image information). The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). In addition, the video / image information may also include general constraint information. The decoding device can also decode the picture based on the information about the parameter sets and / or general constraint information. The signaling / receiving information and / or syntax elements described subsequently in this document can be decoded by the decoding process and obtained from the bitstream. For example, the entropy decoder 310 can decode the information in the bitstream based on encoding methods such as Exponential Golomb coding, CABAC, or CAVLC, and output the syntax elements required for image reconstruction and the quantized values of the transform coefficients for the residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine the context model using information about the target syntax element, the decoding information of the target block, or information about the symbols / bins decoded in the previous stage, and perform arithmetic decoding on the bins by predicting the probability of bin occurrence based on the determined context model, generating symbols corresponding to the value of each syntax element. In this case, the CABAC entropy decoding method can update the context model by using the information of the decoded symbols / bins for the context model of the next symbol / bin after determining the context model. Prediction-related information from the information decoded by the entropy decoder 310 can be provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values (i.e., quantized transform coefficients and related parameter information) from the entropy decoding performed in the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive the residual signals (residual blocks, residual samples, residual sample arrays). In addition, filtering information from the information decoded by the entropy decoder 310 can be provided to the filter 350. Furthermore, a receiver (not shown) for receiving signals output from the encoding device may be additionally configured as an internal / external element of the decoding device 300, or the receiver may be a component of the entropy decoder 310. Additionally, the decoding device according to this document may be referred to as a video / image / picture decoding device, and the decoding device may be classified as an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0121] Dequantizer 321 can dequantize the quantized transform coefficients and output the transform coefficients. Dequantizer 321 can rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement can be performed based on the coefficient scan order performed in the encoding device. Dequantizer 321 can use quantization parameters (e.g., quantization step size information) to perform dequantization on the quantized transform coefficients and obtain the transform coefficients.
[0122] The inverse transformer 322 performs inverse transformation on the transformation coefficients to obtain the residual signal (residual block, residual sample array).
[0123] The predictor can perform prediction on the current block and generate a prediction block that includes prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on information about the prediction output from the entropy decoder 310, and can determine a specific intra-frame / inter-frame prediction mode.
[0124] Predictor 330 can generate prediction signals based on various prediction methods described below. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but also apply both intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as Intra-Frame and Inter-Frame Prediction Combination (CIIP). Alternatively, the predictor can predict blocks based on Intra-Block Copy (IBC) prediction mode or Palette mode. IBC prediction mode or Palette mode can be used for content image / video coding such as games, for example, Screen Content Coding (SCC). IBC essentially performs prediction in the current frame, but can be performed similarly to inter-frame prediction in terms of deriving reference blocks in the current frame. That is, IBC can use at least one of the inter-frame prediction techniques described in this document. Palette mode can be considered as an example of intra-frame coding or intra-frame prediction. When applying Palette mode, sample values within the frame can be signaled based on information about the palette table and palette index.
[0125] Intra-predictor 331 can predict the current block by referencing samples in the current image. Depending on the prediction mode, the referenced samples may be located among the neighbors of the current block, or their location may be separate from the current block. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. Intra-predictor 331 can determine the prediction mode to be applied to the current block by using the prediction modes applied to neighboring blocks.
[0126] Inter-frame predictor 332 can deduce the predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference image. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the motion information correlation between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. For example, inter-frame predictor 332 can construct a motion information candidate list based on neighboring blocks and deduce the motion vector and / or reference image index for the current block based on received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information about the prediction may include information indicating the mode of inter-frame prediction for the current block.
[0127] Adder 340 can generate a reconstruction signal (reconstructed image, reconstruction block, or reconstruction sample array) by adding the obtained residual signal to the prediction signal (prediction block or prediction sample array) output from predictor 330. If there is no residual to process the target block (e.g., in the case of applying a jump mode), the prediction block can be used as the reconstruction block.
[0128] Adder 340 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current image, and as described later, it can also be output by filtering or used for inter-frame prediction of the next image.
[0129] In addition, Luminance Mapping with Chroma Scaling (LMCS) can also be applied to image decoding processing.
[0130] Filter 350 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 350 can generate a corrected reconstructed image by applying various filtering methods to the reconstructed image and store the corrected reconstructed image in memory 360, specifically in the DPB of memory 360. Various filtering methods may include, for example, deblocking filtering, adaptive sample shifting, adaptive loop filtering, bilateral filtering, etc.
[0131] The (corrected) reconstructed image stored in the DPB of memory 360 can be used as a reference image in inter-frame predictor 332. Memory 360 can store motion information of blocks from which motion information in the current image is derived (or decoded) and / or motion information of blocks in already reconstructed images. The stored motion information can be transmitted to inter-frame predictor 332 to be used as motion information for spatially or temporally neighboring blocks. Memory 360 can store reconstructed samples of reconstructed blocks in the current image and transmit the reconstructed samples to intra-frame predictor 331.
[0132] In this disclosure, the embodiments described in the filter 260, inter-frame predictor 221, and intra-frame predictor 222 of the encoding device 200 can be the same as or respectively applied to the filter 350, inter-frame predictor 332, and intra-frame predictor 331 of the decoding device 300. The same applies to the inter-frame predictor 332 and the intra-frame predictor 331.
[0133] Furthermore, the encoded image / video information described above can be configured based on a media file format to generate a media file. For example, the encoded image / video information can be used to form a media file (segment) based on one or more NAL units / sample entries on the encoded image / video information. Media fields can include sample entries and tracks. For example, a media file (segment) can include various records, and each record can include information related to the image / video information or information related to the media file format. Additionally, for example, one or more NAL units can be stored in a configuration record (or decoder configuration record or VVC decoder configuration record) field of the media file. In this document, this field can also be referred to as a syntax element.
[0134] For example, the ISO Basic Media File Format (ISOBMFF) can be used as a media file format for which the methods / implementations disclosed in this disclosure can be applied. ISOBMFF can be used as the basis for various codec encapsulation formats (e.g., AVC, HEVC, and / or VVC formats, etc.) and various multimedia container formats (e.g., MPEG-4, 3GPP (3GP), and / or DVB formats, etc.). In addition to continuous media (e.g., audio and video), static media (e.g., images) and metadata can also be stored in files according to ISOBMFF. Files structured according to ISOBMFF can be used for various purposes such as local media file playback, progressive download of remote files, segmentation for Dynamic Adaptive Streaming (DASH) over HTTP, containerization and packetization instructions for content to be streamed, recording of received real-time media streams, etc.
[0135] The 'box' described below can serve as a basic syntax element of ISOBMFF. An ISOBMFF file can be configured with a series of boxes, and another box can be included within a box. For example, a movie box (a box with grouping type 'moov') can include metadata for a continuous media stream belonging to a media file, and each stream can be indicated as a track in the file. Metadata on a track can belong to a track box (a box with grouping type 'trak'), and the media content of a track can be included in a media data box (a box with grouping type 'mdat'), or it can belong directly to a separate file. The media content of a track can be configured with a series of samples (e.g., audio or video access units). For example, ISOBMFF can specify various types of tracks such as media tracks that include basic media streams, cue tracks that include media transmission instructions or indications of received packet streams, and timing metadata tracks that include timing metadata tracks.
[0136] Additionally, although ISOBMFF is designed for storage purposes, it is also very useful when performing progressive downloads or streaming such as DASH. For streaming purposes, movie clips defined in ISOBMFF can be used. Fragmented ISOBMFF files can, for example, be designated as two separate files associated with video and audio. For instance, when random access is included after receiving 'moov', all movie clips 'moof' can be decoded along with the associated media data.
[0137] Additionally, the metadata for each track can include a list of sample description entries, which provides the track required for processing the corresponding format and the encoding or encapsulation format used in the initialization data. Furthermore, each sample can be associated with one of the sample description entries for the track.
[0138] When using ISOBMFF, sample-specific metadata can be specified by various mechanisms. Specific boxes within a sample table frame (a box with grouping type 'stbl') can be standardized to respond to general requirements. For example, a Sync sample frame (a box with grouping type 'stss') can be used to list random access samples for a track. By using sample grouping mechanisms, samples can be mapped to specified sample groups that share the same characteristics (specified as sample group description entries in a file) based on a four-character grouping type. Various grouping types can be specified to ISOBMFF.
[0139] Figure 7 An exemplary layered structure for encoded video / images is shown.
[0140] Reference Figure 7The encoded image / video is divided into a Video Coding Layer (VCL), subsystems, and a Network Abstraction Layer (NAL). The VCL processes the image / video and its own decoding process, the subsystems send and store the encoded information, and the NAL is responsible for the functions and exists between the VCL and the subsystems.
[0141] In VCL, VCL data including compressed image data (slice data) is generated, or parameter sets including Picture Parameter Set (PSP), Sequence Parameter Set (SPS), and Video Parameter Set (VPS), or additional supplemental enhancement information (SEI) messages required by the image decoding process, can be generated.
[0142] In NAL, NAL cells can be generated by adding header information (NAL cell header) to the raw byte sequence payload (RBSP) generated in VCL. In this case, RBSP refers to slice data, parameter sets, SEI messages, etc., generated in VCL. The NAL cell header can include NAL cell type information specified based on the RBSP data included in the corresponding NAL cell.
[0143] like Figure 7 As shown, based on the RBSP generated in the VCL, NAL units can be classified into VCL NAL units and non-VCL NAL units. A VCL NAL unit can refer to a NAL unit that includes information about the image (slice data), while a non-VCL NAL unit can refer to a NAL unit that includes information (parameter set or SEI message) required for decoding the image.
[0144] The aforementioned VCL NAL units and non-VCL NAL units can be transmitted over a network by attaching header information according to the subsystem's data standard. For example, NAL units can be converted into predetermined standard data formats such as H.266 / VVC file format, Real-time Transport Protocol (RTP), Transport Stream (TS), etc., and transmitted over various networks.
[0145] As described above, a NAL cell can be specified by the NAL cell type according to the RBSP data structure included in the corresponding NAL cell, and information about the NAL cell type can be stored in the NAL cell header for concurrent signaling notification.
[0146] For example, NAL units can be classified into VCLNAL unit types and non-VCL NAL unit types based on whether they include information about the image (slice data). VCL NAL unit types can be classified according to the nature and type of the image included in the VCL NAL unit, while non-VCL NAL unit types can be classified according to the type of parameter set.
[0147] The following is an example of a NAL cell type specified based on the type of the parameter set included in a non-VCL NAL cell type.
[0148] - Adaptive Parameter Set (APS) NAL Unit: The type of NAL unit that includes the APS.
[0149] - Decoding Parameter Set (DPS) NAL Unit: Includes the type of NAL unit for the DPS.
[0150] - Video Parameter Set (VPS) NAL Unit: Includes the type of NAL unit for the VPS.
[0151] - Sequence Parameter Set (SPS) NAL Unit: The type of NAL unit that includes the SPS.
[0152] - Picture Parameter Set (PPS) NAL Unit: Includes the type of NAL unit for PPS.
[0153] - Picture Header (PH) NAL Unit: Includes the type of NAL unit for the PH.
[0154] The NAL unit types mentioned above can have syntax information specific to the NAL unit type, and this syntax information can be stored in the NAL unit header for concurrent signaling. For example, the syntax information can be `nal_unit_type`, and the NAL unit type can be specified through the `nal_unit_type` value.
[0155] Furthermore, as mentioned above, an image can include multiple slices, and a slice can include a slice header and slice data. In this case, an image header can also be added to multiple slices (slice headers and slice data sets) in an image. The image header (image header syntax) can include information / parameters typically applicable to the image. For example, an image can be configured with different types of slices (e.g., intra-coded slices (i.e., I-slices) and / or inter-coded slices (i.e., P-slices and B-slices)). In this case, the image header can include information / parameters applied to both intra-coded and inter-coded slices. Alternatively, an image can also be configured with one type of slice.
[0156] The slice header (slice header syntax) can include information / parameters typically applicable to a slice. APS (APS syntax) or PPS (PPS syntax) can include information / parameters typically applicable to one or more slices or images. SPS (SPS syntax) can include information / parameters typically applicable to one or more sequences. VPS (VPS syntax) can include information / parameters typically applicable to multiple layers. DPS (DPS syntax) can include information / parameters typically applicable to the entire video. DPS can also include information / parameters related to the concatenation of encoded video sequences (CVS).
[0157] In this specification (or document), the video / image information encoded and signaled from the encoding device to the decoding device in the form of a bitstream may include not only information related to intra-frame segmentation, intra / inter-frame prediction information, information related to inter-layer prediction, residual information, and intra-loop filtering information, but also information included in the slice header, the image header, the APS, the PPS, the SPS, the VPS, and / or the DPS. Additionally, the video / image information may also include information from the NAL unit header.
[0158] In this disclosure, the embodiments described in the filter 260, inter-frame predictor 221 and intra-frame predictor 222 of the encoding device 200 can be the same as or correspond to the filter 350, inter-frame predictor 332 and intra-frame predictor 331 of the decoding device 300.
[0159] Furthermore, the encoded image / video information described above can be configured based on a media file format to facilitate the generation of a media file. For example, the encoded image / video information can be used to form a media file (segment) based on one or more NAL units / sample entries for the encoded image / video information. The media file may include sample entries and tracks. For example, the media file (segment) may include various records, and each record may include information related to the image / video or information related to the media file format. Additionally, for example, one or more NAL units may be stored in a configuration record (or decoder configuration record or VVC decoder configuration record) field of the media file. In this document, fields may also be referred to as syntax elements.
[0160] Furthermore, the 'sample' described below can refer to a single element of one of the three sample arrays (Y, Cb, Cr) representing an image, or all data associated with a single time. For example, when the term 'sample' is used in the context of a track (in a media file format), 'sample' can refer to all data associated with a single time of the corresponding track. In this document, time can be decoding time or synthesis time. Additionally, for example, when the term 'sample' is used in the context of an image, i.e., when the term is used in a phrase such as "luminance sample," a sample can indicate a single element belonging to one of the three sample arrays representing an image.
[0161] In addition, the following three types of basic streams can be defined to store VVC content:
[0162] - Excluding the video base stream containing parameter sets, where all parameter sets can be stored in one sample entry or multiple sample entries.
[0163] - May include parameter sets and video and parameter set basic streams storing one or more sample entries of the parameter sets.
[0164] - Includes non-VCL NAL units and non-VCL elementary streams synchronized with the elementary streams included in the video track; in this paper, VVC non-VCL tracks do not include the parameter set in the sample entries.
[0165] Furthermore, operation point information for the ISO-based media file format (ISOBMFF) for VVC can be signaled to samples from group boxes of group type 'vopi' or entity groups of group type 'opeg'. In this paper, operation points can be temporal subsets of the Output Layer Set (OLS) identified by the OLS index and the highest TemporalId value. Each operation point can be associated with a profile, hierarchy, and level (i.e., PTL) defining the consistency point for each operation point. Operation point information may be needed to identify samples and sample entries from each operation point.
[0166] Applications can provide information about the composition of operation points by using various operation points and operation point information sample groups 'vopi' provided from a given VVC bitstream. Each operation point can be associated with OL, the highest TemporalId value, and profile, level, and hierarchy signaling. All of the above information can be captured by the 'vopi' sample group. In addition to the above information, the sample group can also provide information on dependencies between layers.
[0167] Furthermore, when there are one or more VVC tracks for a VVC bitstream and no operation point entity group exists for the VVC bitstream, all of the following details can be applied:
[0168] - In the VVC tracks for VVC bitstreams, there should only be one track group that transmits 'vopi' samples.
[0169] - All other VVC tracks for the VVC bitstream should have an 'oref' type track reference for the track used to pass 'vopi' samples.
[0170] Additionally, for a specific sample on a given track, a time-juxtaposed sample on another track can be defined as a sample with the same decoding time as the specific sample. For each sample SN on track TN that has an 'oref' track reference to track Tk that delivers a 'vopi' sample group, the following can be applied:
[0171] - When a time-juxtaposed sample Sk exists in orbital Tk, sample SN can be associated with the same 'vopi' sample group entity as sample Sk.
[0172] - Otherwise, sample SN can be associated with the same 'vopi' sample group entity as sample Sk.
[0173] When referencing multiple VPSs in a VVC bitstream, it may be necessary to include multiple entities in the sample group description box to which grouping_type 'vopi' belongs. In the more general case where a single VPS exists, it is recommended to include the operation point information sample group in the sample table box, rather than including it in each track segment, by using the default sample group mechanism defined in ISO / IEC 14496-12.
[0174] Alternatively, you don't need to define the grouping_type_parameter for SampleToGroupBox with grouping type 'vopi'.
[0175] The syntax of the 'vopi' sample group, which includes the above-mentioned operation point information (i.e., the operation point information sample group), can be shown in the following table.
[0176] [Table 1]
[0177]
[0178] In addition, the semantics of the syntax of the operation point information sample group can be shown in the table below.
[0179] [Table 2]
[0180]
[0181]
[0182]
[0183] Additionally, for example, an operation point entity group can be defined as one that can provide track mapping and profile-level information for operation points.
[0184] When the aggregated samples of tracks mapped to the aforementioned operands in the operand entity group are aggregated, the implicit reconstruction process no longer requires the removal of any additional NAL units to obtain a conforming VVC bitstream. Tracks belonging to the operand entity group should have a track reference of type 'oref' for the group_id indicated in the operand entity group.
[0185] Additionally, all entity_id values included in the operation point entity group should belong to the same VVC bitstream. If present (or existing), the OperatingPointGroupBox is included in the GroupsListBox of the movie-level MetaBox and not in the file-level or track-level MetaBox. In this document, the OperatingPointGroupBox can indicate the operation point entity group.
[0186] The syntax for the above operation point entity group can be shown in the table below.
[0187] [Table 3]
[0188]
[0189]
[0190] In addition, the semantics of the syntax for operation point entity groups can be seen in the following table.
[0191] [Table 4]
[0192]
[0193]
[0194] Additionally, for example, media files may include decoder configuration information for image / video content. That is, media files may include a VVC decoder configuration record that contains decoder configuration information.
[0195] When a VVC decoder configuration record is stored in a sample entry, the VVC decoder configuration record may include not only a parameter set but also the size of a length field for each sample, indicating the length of the NAL unit included in the VVC decoder configuration record. The VVC decoder configuration record may be formed (or configured) by an external source (frame) (in this document, the size of the VVC decoder configuration record is provided from the structure that includes the VVC decoder configuration record).
[0196] Additionally, the VVC decoder configuration record may include a version field. For example, the version in this disclosure may define version 1 of the VVC decoder configuration record. Changes incompatible with the VVC decoder configuration record may be indicated as changes to the version number. In the case of an unrecognized version number, the reader should not decode the VVC decoder configuration record or the stream to which the corresponding record is applied.
[0197] Compatible extensions to the VVC decoder configuration record can be made without changing the configuration version code. The reader should be prepared to ignore (or disregard) unrecognized data that exceeds the data definition understood by the reader.
[0198] When a track essentially comprises the VVC bitstream, or when the track resolves this issue via a 'subp' track reference, the VvcPtlRecord should be present in the decoder configuration record. Additionally, when the ptl_present_flag in the decoder configuration record of a track is equal to 0, the track should include an 'oref' track record.
[0199] When decoding a stream as described in the VVC decoder configuration record, the values of the syntax elements VvcPTrecord, chroma_format_idc, and bit_depth_minus8 can be valid for all active parameter sets. More specifically, the following restrictions can be applied:
[0200] - The profile indicator general_profile_idc indicates that a profile is followed by a stream associated with the current configuration record.
[0201] - The tier indicator general_tier_flag can indicate the same or higher tier as the highest tier indicated in all parameter sets.
[0202] Each bit of -general_constraint_info can be configured only when the corresponding bit is configured in all parameter sets.
[0203] - The level indicator general_level_idc can indicate a capacity level that is equal to or higher than the highest level indicated for the highest level in the parameter set.
[0204] Alternatively, the restrictions can be applied to chroma_format_idc as follows:
[0205] - When the value of sps_chroma_format_idc as defined in ISO / IEC 23090-3 is the same in all SPS referenced by the NAL element of the track, chroma_format_idc should be the same as sps_chroma_format_idc.
[0206] - Conversely, when ptl_present_flag equals 1, chroma_format_idc should equal vps_ols_dpb_chroma_format[output_layer_set_idx], which is defined in ISO / IEC 23090-3.
[0207] Otherwise (i.e., without applying the above conditions), chroma_format_idc does not exist.
[0208] In addition to other important format information used in the VVC video primary stream, explicit indications of chroma format and bit depth can be provided from the VVC decoder configuration record. If the color space information differs in the VUI information of two sequences, two different VVC sample entries may be required.
[0209] Additionally, for example, the array set of initial NAL units passed can be included in the VVC decoder configuration record. The NAL unit type can be limited to those indicating only DCI, VPS, SPS, PPS, prefix APS, and prefix SEI NAL units. NAL unit types reserved by ISO / IEC 23090-3 and this disclosure may be defined in the future, and readers may need to ignore (or disregard) arrays with reserved NAL unit types or unauthorized values.
[0210] Furthermore, the array can exist in the order of DCI, VPS, SPS, PPS, prefix APS, and prefix SEI.
[0211] The syntax for the VVC decoder configuration record described above can be shown in the table below.
[0212] [Table 5]
[0213]
[0214] In addition, the semantics of the syntax of the VVC decoder configuration record can be seen in the following table.
[0215] [Table 6]
[0216]
[0217]
[0218]
[0219] For example, referring to Table 6 above, for the stream to which the VVC decoder configuration record is applied, as defined in ISO / IEC 23090-3, the syntax elements general_profile_idc, general_tier_flag, general_sub_profile_idc, general_constraint_info, general_level_idc, ptl_frame_only_constraint_flag, ptl_multilayer_enabled_flag, sublayer_level_present, and sublayer_level_idc[i] may include the matching values of the fields general_profile_idc, general_tier_flag, and general_sub_profile_idc, as well as the bits in general_constraint_info(), general_level_idc, ptl_multilayer_enabled_flag, ptl_frame_only_constraint_flag, sublayer_level_present, and sublayer_level_idc[i]. In this paper, avgFrameRate can provide the average frame rate in frames / (256 seconds) for the stream to which the VVC decoder configuration record is applied. A value of 0 can indicate an unspecified (or undefined) average frame rate.
[0220] Additionally, for example, referring to Table 6, the syntax element `constantFrameRate` can indicate the constant frame rate of the VVC decoder configuration record. For example, a `constantFrameRate` value of 1 can indicate that the VVC decoder configuration record applied to the stream has a constant frame rate. A `constantFrameRate` value of 2 can indicate that the representation of each time layer of the stream has a constant frame rate. A `constantFrameRate` value of 0 can indicate that the stream may or may not have a constant frame rate.
[0221] Additionally, for example, referring to Table 6, the syntax element `numTemporalLayers` can indicate the number of time layers included in the track to which the VVC decoder configuration record is applied. For instance, if `numTemporalLayers` is greater than 1, this indicates that the track to which the VVC decoder configuration record is applied is time-scalable, and the number of time layers (also known as time sublayers or sub-layers in ISO / IEC 23090-3) included in the track is equal to `numTemporalLayers`. A value of 1 for `numTemporalLayers` indicates that the track to which the VVC decoder configuration record is applied is not time-scalable. A value of 0 for `numTemporalLayers` indicates that it is unknown whether the track to which the VVC decoder configuration record is applied is time-scalable.
[0222] Additionally, for example, referring to Table 6, the syntax element `lengthSizeMinusOne` can indicate the length in bytes of the `NALUnitLength` field included in the VVC video stream sample to which this configuration record is applied. For example, the size of one byte can be indicated by the value 0. The value of `lengthSizeMinusOne` can be one of 0, 1, or 3, corresponding to a length encoded with 1, 2, or 4 bytes respectively.
[0223] Additionally, for example, referring to Table 6, the syntax element `ptl_present_flag` can indicate that a track includes a VVC bitstream corresponding to a specific (or specified) set of output layers, and accordingly indicates whether PTL information is included. For example, a value of 1 for `ptl_present_flag` can indicate that a track includes a VVC bitstream corresponding to a specific set of output layers (a specific OLS). Furthermore, a value of 0 for `ptl_present_flag` can indicate that a track may not include a VVC bitstream corresponding to a specific OLS, but may instead include one or more individual layers that do not form an OLS, or individual sublayers other than those with TemporalId equal to 0.
[0224] Additionally, for example, referring to Table 6, the syntax element num_sub_profiles can define the number of subprofiles marked in the VVC decoder configuration record.
[0225] Additionally, for example, referring to Table 6, the syntax element track_ptl can indicate the profile, hierarchy, and level of the OLS as indicated by the VVC bitstream included in the track.
[0226] Additionally, for example, referring to Table 6, the syntax element `output_layer_set_idx` can indicate the output layer set index, which is indicated by the set of output layers included in the VVC bitstream in the track. The value of `output_layer_set_idx` can be used as the value of the `TargetOlsIdx` parameter, as given in ISO / IEC 23090-3, for decoding the bitstream included in the track, provided by an external device to the VVC decoder.
[0227] Additionally, for example, referring to Table 6, the syntax element `chroma_format_present_flag` can indicate whether `chroma_format_idc` exists. For instance, a value of 0 for `chroma_format_present_flag` indicates that `chroma_format_idc` does not exist, while a value of 1 indicates that `chroma_format_idc` exists.
[0228] Additionally, for example, referring to Table 6, the syntax element `chroma_format_idc` can indicate the chroma format applied to the track. For example, the following constraints can be applied to `chroma_format_idc`:
[0229] - When the value of sps_chroma_format_idc is the same in all SPS referenced by the NAL unit of the track as defined in ISO / IEC 23090-3, chroma_format_idc should be the same as sps_chroma_format_idc.
[0230] - Alternatively, if ptl_present_flag equals 1, then chroma_format_idc should be the same as vps_ols_dpb_chroma_format[output_layer_set_idx], as defined in ISO / IEC 23090-3.
[0231] - Otherwise (i.e., if none of the above restrictions apply), chroma_format_idc does not exist.
[0232] Additionally, for example, referring to Table 6, the syntax element `bit_depth_present_flag` can indicate whether `bit_depth_minus8` exists. For instance, `bit_depth_present_flag` can indicate that `bit_depth_minus8` does not exist. A value of 1 for `bit_depth_present_flag` can indicate that `bit_depth_minus8` exists.
[0233] Additionally, for example, referring to Table 6, the syntax element bit_depth_minus8 can indicate the bit depth applied to a track. For example, the following constraints can be applied to bit_depth_minus8:
[0234] - When the value of sps_bitdepth_minus8 as defined in ISO / IEC 23090-3 is the same in all SPSs of the NAL element reference of the track, bit_depth_minus8 should be the same as sps_bitdepth_minus8.
[0235] - Alternatively, if ptl_present_flag equals 1, then bit_depth_minus8 should be the same as vps_ols_dpb_bitdepth_minus8[output_layer_set_idx], as defined in ISO / IEC 23090-3.
[0236] - Otherwise (i.e., if none of the above restrictions apply), bit_depth_minus8 does not exist.
[0237] Additionally, for example, referring to Table 6, the syntax element numArrays can indicate the number of NAL cell arrays of the indicated type.
[0238] Additionally, for example, referring to Table 6, the syntax element `array_completeness` can indicate whether additional NAL units exist in the stream. For example, if `array_completeness` equals 1, this can indicate that all NAL units of a given type are included in the subsequent (or next) array and not included in the stream. Alternatively, for example, if `array_completeness` equals 0, this can indicate that additional NAL units of the indicated type can be included in the stream. Default values and allowed (or authorized) values can be restricted (constrained) by sample entry names.
[0239] Additionally, for example, referring to Table 6, the syntax element NAL_unit_type can indicate the type of NAL units (which should include all types) in the subsequent array. NAL_unit_type can have values as defined in ISO / IEC 23090-2. Furthermore, NAL_unit_type can be restricted to having one of the values indicating DCI, VPS, SPS, PPS, APS, prefix SEI, or suffix SEI NAL units respectively.
[0240] Additionally, for example, referring to Table 6, the syntax element `numNalus` can indicate the number of NAL units of the indicated type included in the VVC decoder configuration record of the stream to which the VVC decoder configuration record is applied. The SEI array can include only SEI messages of a 'declarative' nature, i.e., messages that provide information about the entire (or the whole) stream. An example of such an SEI could be a user data SEI.
[0241] Additionally, for example, referring to Table 6, the syntax element nalUnitLength can indicate the byte length of a NAL unit.
[0242] Alternatively, for example, nalUnit can include DCI, VPS, SPS, PPS, APS, or declarative SEI NAL units, as defined in ISO / IEC 23090-3.
[0243] Furthermore, the operation point can be determined first to facilitate the reconstruction of the access unit from samples of multiple tracks passing a multi-layer VVC bitstream. For example, when a VVC bitstream is represented as multiple VVC tracks, the file parser can identify the track required for the selected operation point as described below.
[0244] For example, a file parser can find all tracks with VVC sample entries. If a track includes an 'oref' track reference for the same ID, the corresponding ID can be verified as either a VVC track or an 'opeg' entity group. The operation point can be selected from either the 'opeg' entity group or the 'vopi' entity group (whichever suits the decoding capacity and application purpose).
[0245] With the presence of 'opeg' entity groups, a set of tracks that precisely indicates the selected operation point can be indicated. Therefore, the VVC bitstream can be reconstructed from the track set and can be decoded.
[0246] Additionally, in the absence of the 'opeg' entity group (i.e., in the presence of the 'vopi' sample group), the set of tracks required to decode the selected operation point can be searched from the 'vopi' sample group and the 'linf' sample group.
[0247] To reconstruct the bitstream from multiple VVC tracks that transmit the VVC bitstream, it may be necessary to determine the target highest (or maximum) value of TemporalId. In cases where multiple tracks include data from access units, alignment (or arraying) of each sample in the track can be performed based on the sample decoding time. That is, a time-sample table can be used without considering edit lists.
[0248] When representing a VVC bitstream using multiple VVC tracks, if the tracks are combined into a single stream by extending the decoding time, the decoding time for a sample may require the precise access unit order given in ISO / IEC 23090-3. Furthermore, the access unit sequence can be reconstructed from each sample of the tracks required according to the implicit reconstruction process, which will be described below. For example, the implicit reconstruction process for a VVC bitstream can be described as follows.
[0249] For example, in the presence of a sample set of operation point information, the desired (or required) orbit can be selected based on the layers and reference layers to be passed, as indicated in the sample sets of operation point information and layer information.
[0250] Additionally, for example, in the presence of an operating point entity group, the desired (or required) track can be selected based on the information in the OperatingPointGroupBox.
[0251] Additionally, for example, in the case of reconstructing a bitstream that includes sub-layers with TemporalId greater than 0 using VCL NAL units, all lower-level sub-layers within the same layer (i.e., sub-layers with VCL NAL units that include lower TemporalId values) can be included in the resulting bitstream, and the desired tracks can be selected accordingly.
[0252] Additionally, for example, in the case of a reconfigured access unit, image units (adjusted in ISO / IEC 23090-3) can be assigned to access units according to the ascending order of nuh_layer_id values from samples with the same decoding time.
[0253] Additionally, for example, if the access unit is reconstructed using a dependency layer and max_tid_il_ref_pics_plus1 is greater than 0, the sub-layers of the VCL NAL units within the same layer whose TemporalId value is less than or equal to max_tid_il_ref_pics_plus1-1 (indicated in the operation point information sample group) can also be included in the resulting bitstream, and the desired tracks can be selected accordingly.
[0254] Additionally, for example, in the case where the VVC track includes a 'subp' track reference, each picture cell can be reconstructed as given in Section 11.7.3 of ISO / IEC 23090-3, with additional constraints (or limitations) on the EOS and EOB NAL cells (which will be given below). The process described in Section 11.7.3 of ISO / IEC 23090-3 can be repeated for each layer at the target operation point according to the increasing order of nuh_layer_id. Otherwise, each picture cell can be reconstructed as described below.
[0255] The reconstructed access units can be assigned to the VVC bitstream in ascending order of decoding time, and as described separately below, copies of the End of Bitstream (EOB) and End of Sequence (EOS) NAL units can be removed from the VVC bitstream.
[0256] Additionally, for example, for access units existing in the same coded video sequence within a VVC bitstream and belonging to different sublayers stored in multiple tracks, one or more tracks may exist, wherein one or more tracks include EOS NAL units with a specific nuh_layer_id in each of the samples. In this case, only one of the EOSNAL units can be maintained in the last access unit (the unit with the maximum decoding time) among the access units in the final reconstructed bitstream, and only one EOS NAL unit can be assigned after all NAL units except for the EOB NAL unit (if present) of the last access unit in the access unit, and other EOS NAL units can be deleted. Similarly, one or more tracks including EOB NAL units may exist in each sample. In this case, among the EOB NAL units, only one EOB NAL unit can be maintained in the final reconstructed bitstream and can be assigned to the end of such access units, and other EOB NAL units can be deleted.
[0257] Additionally, for example, since a particular layer or sublayer can be represented by one or more tracks, when searching for the desired track for an operation point, it may be necessary to select a track from the set of all tracks that pass through the particular layer or sublayer.
[0258] Additionally, for example, in the absence of an operation point entity group, after selecting a track from those passing the same layer or sublayer, the required final track can be shared by a portion of the layer or sublayer that still does not belong to the target operation point. Although the bitstream reconstructed for the target operation point is passed by the final required track, it may exclude layers or sublayers that do not belong to the target operation point.
[0259] Furthermore, according to this disclosure, for efficient processing, region-based independent processing can be supported. To this end, specific regions can be extracted and / or processed to configure independent bitstreams, and file formats can be formed (or configured) for specific regions through extraction and / or processing. In this case, the circular coordinate information of the extracted regions can be signaled to support efficient image region decoding and rendering at the receiving end. In the following, the regions supported by independent processing of the input image can be referred to as sub-pictures. For example, sub-pictures can be used for 360-degree video / image content such as VR or AR. However, sub-pictures are not limited to this. For example, in 360-degree video / image content, the portion seen (or viewed) by the user may not be the entirety (or whole) of the 360-degree video / image content, but may only be a portion of it. Sub-pictures can be used to independently process portions of the 360-degree video / image content viewed by the user that are separate from the rest of the content. The input image can be split into a sequence of sub-pictures before encoding, and each sub-picture sequence can cover a subset of the spatial regions of the 360-degree video / image content. Each sub-picture sequence can be independently encoded and output as a single-layer bitstream. Each sub-picture bitstream can be encapsulated within a file based on a single track, or it can be streamed. In this case, the receiving device can decode and render tracks covering the entire area, or it can select tracks associated with a specific sub-picture and then decode and render the selected sub-picture. The concept and characteristics of sub-pictures used in the VVC standard will be described in detail below.
[0260] The subpicks used in VVC can be called VVC subpicks. A VVC subpick can be a rectangular region that includes one or more slices of a picture. That is, a subpick can include one or more slices that cover a rectangular region of a picture. A VVC subpick can include one or more complete tiles or a portion of a tile. The encoder can treat the boundaries of the subpicks as the same as the boundaries of the picture and can avoid using loop filtering across the subpick boundaries. Therefore, selected subpicks can be extracted from the VVC bitstream, or subpicks can be encoded so that they can be integrated with the destination VVC bitstream. Alternatively, this VVC bitstream extraction or integration process can be performed without modifying (or correcting) the VCL NAL units. Subpick identifiers (IDs) of subpicks present in the bitstream can be marked by SPS or PPS. In the following text, for simplicity, 'VVC' can be omitted from VVC subpicks, VVC tracks, VVC subpick tracks, etc.
[0261] Furthermore, the aforementioned media files can be included in tracks. That is, bitstreams including video / image data can be stored in the aforementioned tracks, thereby forming (or configuring) media files. The types of tracks, more specifically, those used for transmitting VVC elementary streams, are shown in the table below.
[0262] [Table 7]
[0263]
[0264]
[0265] a) VVC track:
[0266] A VVC track can be represented by including NAL units in the samples and / or sample entries of the VVC track, by referencing a VVC track that includes a sublayer of another VVC bitstream, or by referencing a VVC subpicture track. In the case where a VVC track references a VVC subpicture track, the VVC track can be called a VVC base track.
[0267] b) VVC non-VCL track
[0268] APS and other non-VCL NAL units that transmit ALF, LMCS, or scaling list parameters can be stored and transmitted via a different track than the track containing VCLNAL units. This track is the VVC non-VCL track.
[0269] c) VVC sub-image track
[0270] 1) Sub-image tracks include one of the following:
[0271] 1-1) A sequence of one or more VVC sub-images
[0272] 1-2) Forming one or more complete slice sequences of a rectangular region
[0273] 2) Samples of sub-image tracks include one of the following:
[0274] 2-1) One or more complete sub-pictures specified in ISO / IEC 23090-3 and having a consecutive decoding order.
[0275] 2-2) One or more complete slices that have a sequential decoding order and form a rectangular region, as specified in ISO / IEC 23090-3.
[0276] VVC sub-picture tracks, or slices included in random samples within VVC sub-picture tracks, can be consecutive in the decoding order.
[0277] VVC non-VCL tracks and VVC subpicture tracks can preferably transmit VVC video within a streaming application as described below. Each track can be transmitted (or carried) with its own DASH representation, wherein the DASH representation of a subset of VVC subpictures, including those used for decoding and rendering tracks, and the DASH representation of the non-VCL tracks can be requested by the client for each segment. By using this method, redundant transmission of APS and other non-VCLNAL units can be prevented.
[0278] The method for reconstructing image units from samples in the VVC track of the reference VVC sub-image track can be described as follows.
[0279] The samples of the VVC track can be decomposed into access units including the following NAL units according to the order shown in the table below.
[0280] [Table 8]
[0281]
[0282]
[0283] - AUD NAL cells when they exist (or are located) in the sample (and when they are the first NAL cells)
[0284] - When the sample is the first sample in a sample sequence associated with the same sample entry, include the parameter set and SEI NAL unit (if any) in the sample entry.
[0285] - NAL units (if any) present in samples containing up to PH NAL units.
[0286] - The contents of the time-aligned decomposed samples (time-aligned according to decoding time) from each reference VVC subpicture track, in the order given by the 'spor' sample group description entry mapped to the sample (if any), excluding all VPS, DCI, SPS, PPS, AUD, PH, EOS, and EOB NAL units. Track references can be decomposed as follows. When a positively referenced VVC subpicture track is associated with a VVC non-VCL track, the decomposed samples of the VVC subpicture track include time-aligned non-VCL units (if any) within the VVC non-VCL track.
[0287] - NAL units following the PH NAL unit in the sample. NAL units following the PH NAL unit in the sample may include suffix SEI NAL units, suffix APS NAL units, EOS NAL units, EOB NAL units, or reserved NAL units authorized after the last VCLNAL unit.
[0288] Furthermore, the 'subp' orbital reference index of the 'spor' sample group description entries can be decomposed as shown in the table below.
[0289] [Table 9]
[0290]
[0291]
[0292] - In the case where the track reference indicates the track ID of the VVC sub-image track, the track reference can be decomposed into VVC sub-image tracks.
[0293] - In other cases, i.e., when the orbital reference indicates an 'alte' orbital group, the orbital reference can be decomposed into random orbits within the 'alte' orbital group. When a particular orbital reference index value is decomposed into a specific sample of a previous sample, the corresponding value should be decomposed into one of the following in the current sample:
[0294] -The same specific orbit, and
[0295] -Includes another random orbit in the 'alte' orbital group of sync samples that are aligned with the current sample time.
[0296] To avoid decoding mismatches, VVC subpicture tracks within the 'alte' track should be strictly independent of all other VVC subpicture tracks referenced by the same VVC base track, and may be subject to the following restrictions:
[0297] - All VVC sub-picture tracks should include VVC sub-pictures.
[0298] - The boundaries (or multiple boundaries) of the sub-image should be the same as the boundaries (or multiple boundaries) of the image.
[0299] - Loop filtering should be turned off across the boundaries (or multiple boundaries) of the sub-image. In other words, loop filtering can be omitted from the vicinity of the sub-image boundaries (or multiple boundaries).
[0300] In addition, when the reader makes an initial selection, or when the user selects a VVC subpicture track that includes a set of VVC subpictures with subpicture ID values different from the previously selected ones, the following steps can be performed as shown in the table below.
[0301] [Table 10]
[0302]
[0303]
[0304] - To determine whether changes to the PPS or SPS NAL units are necessary, the 'spor' sample group description entries can be examined. SPS changes may only be possible at the starting point of the CLVS.
[0305] - If the 'spor' sample group description entry indicates that the start code emulation prevention byte exists within or before the sub-picture ID of the NAL cell containing that byte, the RBSP can be derived from the NAL cell (i.e., the start code emulation prevention byte can be removed). After performing coverage in the next stage, start code emulation prevention can be performed again.
[0306] - The reader can use the 'spor' sample group to describe the bit positions and sub-picture ID length information within the entry, in order to determine the bits that will be used to overwrite and update the sub-picture ID to the selected sub-picture ID.
[0307] - If the sub-image ID value of PPS or SPS is initially selected, the reader may need to rewrite PPS or SPS with the sub-image ID value that has been selected in the reconstructed access unit.
[0308] - When the sub-image ID value of a PPS or SPS changes after performing a comparison between the same PPS ID value or SPS ID value and (each) a previous PPS ID value or SPS ID value, the reader should include copies of the previous PPS and SPS, and (each) PPS and SPS may need to be rewritten to the updated sub-image ID value in the reconstructed access unit.
[0309] In addition, as described below, the following issues (or problems) may occur related to the aforementioned tracks, sub-images, and slices.
[0310] In ISOBMFF, the current specification for transmitting VVC allows (or permits) subpicture tracks to include one or more complete slices as specified in ISO / IEC 23090-3 (or VVC), which are consecutive in decoding order and configured (or form) rectangular regions. In this document, although subpicture tracks include one or more complete slices, the usefulness of allowing cases where not all slices are included in a subpicture is unclear. That is, although a subpicture track includes slices, if not all slices are included in a subpicture, the track becomes useless unless those slices are always referenced by the VVC base track (which also references other subpicture tracks that include the remaining slices of the same subpicture).
[0311] Therefore, this disclosure proposes the following solutions to the above-mentioned problems. The proposed implementations can be applied individually or in combination.
[0312] 1. A sub-picture may be limited to all slices comprising one or more sub-pictures, as specified in ISO / IEC 23090-3.
[0313] 2. As an alternative, if a sub-image does not contain all slices, then all of the following conditions can be applied even if the sub-image track contains one or more slices:
[0314] a) All slices within a sub-image track can belong to the same sub-image.
[0315] b) All VVC base tracks that reference a sub-picture track also reference sub-picture tracks that include the remaining slices of the same sub-picture.
[0316] The above solution will be described in more detail below.
[0317] According to one embodiment of this disclosure, a VVC sub-image track may be limited to one or more VVC sub-images. That is, according to Table 1 above, a VVC sub-image track is designed to include a sequence of one or more VVC sub-images forming a rectangular region or a sequence of one or more complete slices. According to this embodiment, the case where a VVC sub-image track includes a sequence of one or more complete slices forming a rectangular region can be excluded.
[0318] Furthermore, the samples of a VVC sub-picture track can be limited to include only one or more complete sub-pictures with a sequential decoding order as specified in ISO / IEC 23090-3. That is, according to Table 1 above, a VVC sub-picture track is designed to include one or more complete sub-pictures with a sequential decoding order as specified in ISO / IEC 23090-3, or one or more complete slices with a sequential decoding order forming a rectangular region as specified in ISO / IEC 23090-3. According to this embodiment, the case where a VVC sub-picture track includes a sequence of one or more complete slices with a sequential decoding order forming a rectangular region as specified in ISO / IEC 23090-3 can be excluded.
[0319] When this embodiment is applied, the sub-image track can be configured to always include only one or more complete sub-images. Therefore, this has the effect of preventing the aforementioned problem, namely, allowing the sub-image track to include one or more complete slices but not all slices within a single sub-image. Additionally, when this embodiment is applied, it may be advantageous that the relationship between the VVC sub-image track, the VVC sub-image, and the slices within the VVC sub-image becomes simple and clear.
[0320] In summary, according to one embodiment of this disclosure, the VVC sub-image track can be as shown in the table below, for example.
[0321] [Table 11]
[0322]
[0323] According to another embodiment of this disclosure, where a VVC sub-picture track is allowed to include one or more complete slices but not all slices in a sub-picture, all of the following can be applied:
[0324] All slices within a sub-image track should belong to the same sub-image.
[0325] - All VVC base tracks that reference a sub-picture track should also reference sub-picture tracks that include the remaining slices of the same sub-picture.
[0326] In applying this embodiment, when a sub-image track includes one or more complete slices but not all slices within a single image, a restrictive configuration ensures that all slices within a sub-image belong to the same sub-image, and that all VVC base tracks referencing a sub-image track also refer to sub-image tracks containing remaining slices of the same sub-image. This has the advantageous effect of solving the aforementioned problems. In other words, even if a sub-image track includes slices, but the slices are not all contained within a single sub-image, this can prevent the following situation: slices are not always referenced by the same VVC base track, but also reference other sub-image tracks containing remaining slices of the same sub-image, which then renders the track useless. Furthermore, in applying this embodiment, it may be advantageous that the relationship between VVC sub-image tracks, VVC sub-images, and slices within VVC sub-images can become simple and clear.
[0327] In summary, according to one embodiment of this disclosure, the VVC sub-image track can be as shown in the table below, for example.
[0328] [Table 12]
[0329]
[0330]
[0331] Figure 8 An example of a method for generating media files using embodiments of the present disclosure is shown.
[0332] Figure 8 The process can be performed by a first device. The first device may include, for example, a sending end, an encoding end, a media file generating end, etc. However, the first device is not limited to these.
[0333] Reference Figure 8 The first device can form (or configure) a sub-image track (S800). As described above, the sub-image track can include a sequence of one or more VVC sub-images forming a rectangular region or a sequence of one or more complete slices. In other words, the sequence of one or more VVC sub-images forming a rectangular region or the sequence of one or more complete slices can be stored in the sub-image track and then transmitted.
[0334] After the sub-image track is formed, the first device can generate a media file based on the sub-image track (S810).
[0335] Figure 9 An example of a method for processing a media file generated by applying the embodiments proposed in this disclosure is shown.
[0336] Figure 9 The process can be performed by a second device. This second device may include, for example, a receiver, a decoder, or a renderer. However, the second device is not limited to these.
[0337] Reference Figure 9 The second device can acquire / receive a media file including sub-picture tracks (S900). The media file can be a media file generated by the first device. The media file can include the aforementioned frames and tracks, and a bitstream including video / image data can be stored in both the frames and tracks within the media file. The tracks can include the aforementioned VVC tracks, VVC non-VCL tracks, or VVC sub-picture tracks, etc.
[0338] The second device can parse / obtain sub-picture tracks (S910). The second device can parse / obtain sub-picture tracks included in the media file. For example, the second device can reproduce a video / image by using the video / image data stored in the sub-picture tracks.
[0339] Media files generated by the first device and acquired / received by the second device may include the aforementioned (VVC) subpicture tracks. Based on the subpicture tracks, the second device can deduce one or more subpictures or one or more slices within a subpicture. Video / image decoding can be performed based on the slices or subpictures.
[0340] Figure 10 An overall view of a method for generating a media file performed by a device for generating media files, according to this disclosure, is shown. Figure 10 The method disclosed in the document can be used by [the party in question]. Figure 11 The device (or media file generating device) disclosed herein for generating media files shall be used to perform the operation. The media file generating device may represent the first device described above. More specifically, for example, Figure 10 S1000 to S1010 can be executed by the image processor of the media file generation device, and S1020 can be executed by the media file generator of the media file generation device. Additionally, although not shown in the figure, the process of encoding the bitstream including image information can be executed by the encoder of the media file generation device.
[0341] The media file generation device can form (or configure) sub-picture tracks (S1000). For example, the sub-picture track formed in S1000 may include a VVC sub-picture track. For example, the sub-picture track may include a sequence of one or more (VVC) sub-pictures or a sequence of one or more complete slices. In other words, a sequence of one or more (VVC) sub-pictures or a sequence of one or more complete slices can be stored in the sub-picture track and then transmitted. The sequence of one or more complete slices can form a rectangular area.
[0342] For example, a sample of a sub-picture track may include one or more complete sub-pictures as specified in ISO / IEC 23090-3, or one or more complete slices as specified in ISO / IEC 23090-3. One or more complete sub-pictures may have a sequential decoding order. One or more complete slices may form a rectangular region. Alternatively, one or more complete slices may also have a sequential decoding order.
[0343] For example, a media file generating device can obtain encoded image information via a network or (digital) storage medium. In this document, the network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. Alternatively, for example, the media file generating device may include an encoder and may derive the encoded image information.
[0344] The media file generation device can form (or configure) a base track (S1010). The base track formed in step S1010 may, for example, include a VVC base track. The VVC base track can be a VVC track that references a VVC sub-image. That is, when a VVC track references a VVC sub-image track, the VVC track can be called a VVC base track.
[0345] According to one embodiment, the sub-image track formed in step S1000 may include multiple sub-image tracks. For example, a sub-image track may include a first sub-image track and one or more second sub-image tracks. The first sub-image track may include one or more slices, wherein, if one or more slices are not all slices of a sub-image, all slices of the first sub-image track are included in the same sub-image as the corresponding sub-image, and the basic track of the first sub-image track may refer to one or more second sub-image tracks that include the remaining slices of the same sub-image.
[0346] More specifically, according to one implementation, a sub-image may include a first slice and a second slice. That is, a sub-image may be formed by a first slice and a second slice corresponding to a remaining slice after excluding the first slice. In other words, the second slice may be a slice remaining after excluding the first slice.
[0347] There may be a situation where the first sub-image track includes the first slice but does not include the second slice. As described above, when the first sub-image track includes the first slice, all slices within the first sub-image can form (or be configured) the same sub-image, and the base track referencing the first sub-image track can refer to one or more second sub-image tracks, wherein one or more second sub-image tracks can include the second slice.
[0348] In other words, even when a particular sub-picture track includes only a portion of a sub-picture's slices instead of all of them, the above problem can be solved by setting a limitation (or constraint) that includes the remaining slices of the sub-picture in another sub-picture track that is also referenced by the base track that references the particular sub-picture track.
[0349] According to one embodiment, each sub-image track in a sub-image track may include one or more sub-images or one or more slices. Herein, one or more slices may form a rectangular region. Therefore, when each sub-image track in a sub-image track includes one or more slices, the one or more slices included in each sub-image track in the sub-image track can form a rectangular region.
[0350] According to one implementation, each sub-image track in a sub-image track may include one or more sub-images, wherein one or more slices may not be included. In this document, one or more slices may form a portion of a sub-image. The aforementioned problem can be solved by excluding the case where a sub-image track includes only a portion of a sub-image, i.e., the case where a sub-image track includes slices that only form (or configure) a portion of a sub-image. In this case, each sub-image track in the sub-image track may be presented as described above in Table 11.
[0351] According to one implementation, a sample of a sub-image track may include one or more complete sub-images or one or more complete slices. In this document, one or more complete slices may form a rectangular region. Therefore, when a sample of a sub-image track includes one or more complete slices, the one or more complete slices included in the sample of the sub-image track may form a rectangular region. One or more complete sub-images may be consecutive in decoding order. One or more complete slices may be consecutive in decoding order.
[0352] According to one implementation, a sample of a sub-image track may include one or more complete sub-images, wherein the sample of a sub-image track may not include one or more complete slices. In this document, one or more complete slices may form a portion of a sub-image. The aforementioned problem can be solved by excluding the case where the sample of a sub-image track includes only a portion of a complete sub-image, i.e., the case where the sample of a sub-image track includes only complete slices that form (or configure) a portion of a complete sub-image. In this case, the sample of a sub-image track can be presented as described above in Table 11.
[0353] After forming sub-image tracks and a base track using the above method, the media file generation device can generate a media file based on the sub-image tracks (S1020).
[0354] Furthermore, although not shown in the figure, the media file generating device can store the generated media file on a (digital) storage medium, or can transfer the generated media file to a media file processing device via a network or (digital) storage medium. In this document, the network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.
[0355] Figure 11 An overall view of a device for generating media files is shown, illustrating a method for generating media files according to this disclosure. Figure 10 The methods disclosed herein can be used by [the party] in [the context of] [the process]. Figure 11 The process is performed using a publicly disclosed device for generating media files (or a media file generation device). More specifically, for example, Figure 11 The image processor of the media file generation device can perform Figure 10 S1000 to S1010, and Figure 11 The media file generator of the media file generation device can execute S1020. Furthermore, although not shown in the figure, the process of encoding the bitstream including image information can be performed by the encoder of the media file generation device.
[0356] Figure 12 An overall view of a method for processing media files performed by a device for processing media files, according to this disclosure, is shown. Figure 12 The method disclosed in the document can be used by [the party in question]. Figure 13 The device (or media file processing device) disclosed herein for processing media files shall be used to perform the operation. The media file processing device may refer to the second device described above. More specifically, for example, Figure 12 S1200 can be executed by the receiver of the media file processing device, and S1210 and S1220 can be executed by the media file processor of the media file processing device. Additionally, although not shown in the figure, the process of decoding the bitstream based on the decoder configuration record can be executed by the encoder of the media file generation device.
[0357] The media file processing device acquires a media file including sub-picture tracks and a base track (S1200). For example, the media file processing device can acquire a media file including sub-picture tracks and a base track via a network or a (digital) storage medium. In this document, the network may include a broadcast network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.
[0358] The media file processing device parses the sub-picture track (S1210). For example, the media file processing device can parse / derive the sub-picture track.
[0359] The sub-picture track parsed in S1210 may, for example, include a VVC sub-picture track. For instance, a sub-picture track may include a sequence of one or more VVC sub-pictures or a sequence of one or more complete slices. In other words, a sequence of one or more VVC sub-pictures or a sequence of one or more complete slices can be transmitted through a sub-picture track. A sequence of one or more complete slices can form a rectangular region.
[0360] For example, a sample of a sub-picture track may include one or more complete sub-pictures as specified in ISO / IEC 23090-3, or one or more complete slices as specified in ISO / IEC 23090-3. One or more complete sub-pictures may have a sequential decoding order. One or more complete slices may form a rectangular region. Alternatively, one or more complete slices may also have a sequential decoding order.
[0361] The media file processing device parses the basic tracks (S1220). For example, the media file processing device can parse / derive the basic tracks. The basic tracks parsed in S1220 may include, for example, VVC basic tracks. A VVC basic track can be a VVC track in the case of a VVC track referencing a VVC sub-picture. That is, when a VVC track refers to a VVC sub-picture track, the VVC track can be called a VVC basic track.
[0362] According to one implementation, the sub-image track parsed in S1210 may include multiple sub-image tracks. For example, a sub-image track may include a first sub-image track and one or more second sub-image tracks. The first sub-image track may include one or more slices, wherein, if one or more slices are not all slices of a sub-image, all slices of the first sub-image track are included in the same sub-image as the corresponding sub-image, and the base track of the first sub-image track may refer to one or more second sub-image tracks that include the remaining slices of the same sub-image.
[0363] More specifically, according to one implementation, a sub-image may include a first slice and a second slice. That is, a sub-image may be formed from a first slice and a second slice. More specifically, a sub-image may be formed from a second slice corresponding to a remaining slice after excluding the first and second slices. In other words, the second slice may be a slice remaining after excluding the first slice.
[0364] There may be a situation where the first sub-image track includes the first slice but does not include the second slice. As described above, when the first sub-image track includes the first slice, all slices within the first sub-image can form (or be configured) the same sub-image, and the base track referencing the first sub-image track can refer to one or more second sub-image tracks, wherein one or more second sub-image tracks can include the second slice.
[0365] In other words, even when a particular sub-picture track includes only a portion of a slice of a sub-picture rather than all of it, the above problem can be solved by setting a limitation (or constraint) such that the remaining slices of the sub-picture are included in another sub-picture track that is also referenced by the base track that references the particular sub-picture track.
[0366] According to one implementation, each sub-image track in a sub-image track may include one or more sub-images or one or more slices. Herein, one or more slices may form a rectangular region. Therefore, when each sub-image track in a sub-image track includes one or more slices, the one or more slices included in each sub-image track in a sub-image track can form a rectangular region.
[0367] According to one implementation, each sub-image track in a sub-image track may include one or more sub-images, wherein one or more slices may not be included. In this document, one or more slices may form a portion of a sub-image. The aforementioned problem can be solved by excluding the case where a sub-image track includes only a portion of a sub-image, i.e., the case where a sub-image track includes slices that only form (or configure) a portion of a sub-image. In this case, each sub-image track in the sub-image track may be presented as described above in Table 11.
[0368] According to one implementation, a sample of a sub-image track may include one or more complete sub-images or one or more complete slices. In this document, one or more complete slices may form a rectangular region. Therefore, when a sample of a sub-image track includes one or more complete slices, the one or more complete slices included in the sample of the sub-image track may form a rectangular region. One or more complete sub-images may be consecutive in decoding order. One or more complete slices may be consecutive in decoding order.
[0369] According to one implementation, a sample of a sub-image track may include one or more complete sub-images, wherein the sample of a sub-image track may not include one or more complete slices. In this document, one or more complete slices may form a portion of a sub-image. The aforementioned problem can be solved by excluding the case where the sample of a sub-image track includes only a portion of a complete sub-image, i.e., the case where the sample of a sub-image track includes only complete slices that form (or configure) a portion of a complete sub-image. In this case, the sample of a sub-image track can be presented as described above in Table 11.
[0370] Furthermore, although not shown in the accompanying drawings, the media file processing device can decode the bitstream based on sub-picture tracks and base tracks. For example, the media file processing device can decode image information in the bitstream of a sub-picture based on the sub-picture tracks and base tracks, and can generate a reconstructed image based on the image information.
[0371] Figure 13 An overall view of a device for processing media files is shown, illustrating a method for processing media files according to this disclosure. Figure 12 The method disclosed in the document can be used by [the party in question]. Figure 13 The device (or media file processing device) disclosed in the document is used to execute the operation. More specifically, for example, Figure 13 The receiver of the media file processing device can perform Figure 12 The S1200, and Figure 13 The media file processor of the media file processing device can perform Figure 12 S1210 and S1220. Furthermore, although not shown in the figure, the media file processing device may include a decoder that can decode the bitstream based on sub-picture tracks and a base track.
[0372] According to the embodiments described above in this disclosure, a sub-image track can always include one or more complete sub-images. Therefore, this has the effect of preventing the aforementioned problems, namely, allowing a sub-image track to include one or more complete slices but not all slices within a single sub-image. Furthermore, when applying this embodiment, it may be advantageous that the relationship between the VVC sub-image track, the VVC sub-image, and the slices within the VVC sub-image can become simple and clear.
[0373] According to other embodiments of this disclosure, when a sub-image track includes one or more complete slices but not all slices in a single image, a restrictive configuration ensures that all slices in the sub-image track belong to the same sub-image, and that all VVC base tracks referencing the sub-image track also refer to sub-image tracks including the remaining slices of the same sub-image. This has the advantageous effect of solving the aforementioned problems. In other words, even if a sub-image track includes slices, but the slices are not all included in a single sub-image, this can prevent the following situation: slices are not always referenced by the same VVC base track, but also reference other sub-image tracks including the remaining slices of the same sub-image, which then renders the track useless. Furthermore, when applying this embodiment, it may be advantageous that the relationship between VVC sub-image tracks, VVC sub-images, and slices within VVC sub-images can become simple and clear.
[0374] In the above embodiments, the method is explained based on the flowchart by means of a series of steps or blocks. However, this disclosure is not limited to the order of the steps, and certain steps may be performed in a different order or sequence than those described above, or may be performed simultaneously with another step. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, and may be incorporated into another step, or one or more steps of the flowchart may be removed without affecting the scope of this disclosure.
[0375] The implementations described in this document can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional components shown in each figure can be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, the information used for implementation (e.g., information about instructions) or algorithms can be stored in a digital storage medium.
[0376] Furthermore, the devices to which this document applies may include multimedia broadcasting transmitting and receiving devices, mobile communication terminals, home theater video devices, digital cinema video devices, surveillance cameras, video chat devices, real-time communication devices such as video communication, mobile streaming devices, storage media, portable cameras, video-on-demand (VoD) service providers, over-the-top (OTT) video devices, internet streaming service providers, three-dimensional (3D) video devices, video telephony devices, transportation terminals (e.g., vehicle terminals, aircraft terminals, or ship terminals), and medical video devices; and may be used to process video or data signals. For example, over-the-top (OTT) video devices may include game consoles, Blu-ray players, networked TVs, home theater systems, smartphones, tablet PCs, and digital video recorders (DVRs).
[0377] Furthermore, the processing methods to which this document applies can be generated in the form of computer-executable programs and can be stored in computer-readable recording media. Multimedia data having the data structure according to this document can also be stored in computer-readable recording media. Computer-readable recording media include all types of storage devices storing computer-readable data. For example, computer-readable recording media can include Blu-ray discs (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disks, and optical data storage devices. Additionally, computer-readable recording media include media implemented in carrier wave form (e.g., transmission over the Internet). Furthermore, bitstreams generated using encoding methods can be stored in computer-readable recording media or transmitted via wired and wireless communication networks.
[0378] Furthermore, the embodiments described in this document can be implemented as a computer program product using program code. The program code can be executed by a computer according to the embodiments described in this document. The program code can be stored on a computer-readable medium.
[0379] Figure 14 Examples of content streaming systems to which the implementation methods disclosed in this document can be applied are illustrated.
[0380] A content streaming system using embodiments of this disclosure may substantially include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.
[0381] An encoding server compresses content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and then sends the bitstream to a streaming server. As another example, if the bitstream is generated directly by a multimedia input device such as a smartphone, camera, or camcorder, the encoding server can be omitted.
[0382] Bitstreams can be generated using the encoding methods or bitstream generation methods described in this document. Furthermore, the streaming server can temporarily store the bitstream during transmission or reception.
[0383] A streaming server sends multimedia data to a user's device via a web server based on a user's request. The web server acts as a medium for informing the user of available services. When a user requests a desired service from the web server, the web server forwards the request to the streaming server, which then sends the multimedia data to the user. In this scenario, the content streaming system may include a separate control server. In this case, the control server manages the commands / responses between the various devices within the content streaming system.
[0384] A streaming server can receive content from media storage and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a predetermined period of time.
[0385] Examples of user equipment can include mobile phones, smartphones, laptops, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation systems, board PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays), digital TVs, desktop computers, digital signage, etc. Each server in the content streaming system can operate as a distributed server, and in this case, data received from each server can be distributed.
[0386] The claims described herein can be combined in various ways. For example, the technical features in the method claims of this specification can be combined and implemented as a device, and the technical features in the device claims of this specification can be combined and implemented as a method. Furthermore, the technical features in the method claims and the device claims of this specification can be combined to implement a device, and the technical features in the method claims and the device claims of this specification can be combined to implement a method.
Claims
1. A method for generating media files, the method comprising the following steps: The sub-image tracks are configured, wherein each sub-image track includes one or more sub-images or one or more slices; Configure a base track that references at least one sub-image track; and Generate a media file that includes the sub-image tracks and the base track. The sub-image track includes a first sub-image track and one or more second sub-image tracks. Specifically, this includes the case where the first sub-image track comprises one or more slices that are not equal to all slices of the sub-image. All slices in the first sub-image track belong to the same sub-image, and The basic track, referring to the first sub-image track, also refers to the one or more second sub-image tracks containing the remaining slices of the same sub-image.
2. The method according to claim 1, wherein, The sub-image includes one or more first slices and one or more second slices. Wherein, the one or more second slices are the remaining slices of the sub-image besides the one or more first slices. Specifically, this applies to cases where the first sub-image track includes one or more first slices. All slices in the first sub-image track constitute the same sub-image. The basic track, referring to the first sub-image track, also refers to the one or more second sub-image tracks, wherein the one or more second sub-image tracks include the one or more second slices.
3. The method according to claim 1, wherein, Based on the case where each of the sub-image tracks includes one or more slices, the one or more slices form a rectangular region.
4. The method according to claim 1, wherein, Each sub-image track in the sub-image track includes one or more sub-images in addition to the one or more slices. The one or more slices constitute a part of the sub-image.
5. The method according to claim 1, wherein, The sample of the sub-image track includes one or more complete sub-images or one or more complete slices.
6. The method according to claim 5, wherein, The sample based on the sub-image track includes one or more complete slices, where the one or more complete slices form a rectangular region.
7. The method according to claim 5, wherein, The sample of the sub-image track includes the one or more complete sub-images in addition to the one or more complete slices. The one or more complete slices constitute a part of the sub-image.
8. A method for processing media files, the method comprising the following steps: Obtain a media file, the media file including sub-picture tracks and a base track referencing at least one sub-picture track, wherein each sub-picture track includes one or more sub-pictures or one or more slices; The sub-image track is parsed; and The basic orbit is analyzed. The sub-image track includes a first sub-image track and one or more second sub-image tracks. Specifically, this includes the case where the first sub-image track comprises one or more slices that are not equal to all slices of the sub-image. All slices in the first sub-image track belong to the same sub-image, and The basic track, referring to the first sub-image track, also refers to the one or more second sub-image tracks containing the remaining slices of the same sub-image.
9. The method according to claim 8, wherein, The sub-image includes one or more first slices and one or more second slices. Wherein, the one or more second slices are the remaining slices of the sub-image besides the one or more first slices. Specifically, this applies to cases where the first sub-image track includes one or more first slices. All slices in the first sub-image track constitute the same sub-image. The basic track, referring to the first sub-image track, also refers to the one or more second sub-image tracks, wherein the one or more second sub-image tracks include the one or more second slices.
10. The method according to claim 8, wherein, Based on the case where each of the sub-image tracks includes one or more slices, the one or more slices form a rectangular region.
11. The method according to claim 8, wherein, Each sub-image track in the sub-image track includes one or more sub-images in addition to the one or more slices. The one or more slices constitute a part of the sub-image.
12. The method according to claim 8, wherein, The sample of the sub-image track includes one or more complete sub-images or one or more complete slices.
13. The method according to claim 12, wherein, The sample based on the sub-image track includes one or more complete slices, where the one or more complete slices form a rectangular region.
14. The method according to claim 12, wherein, The sample of the sub-image track includes the one or more complete sub-images in addition to the one or more complete slices. The one or more complete slices constitute a part of the sub-image.