Bitstream Merge
By using merge identifiers to determine the appropriate complexity of the merge procedure within the video codec, the challenges of merging multiple video bitstreams are addressed, resulting in efficient and scalable video encoding and decoding.
Patent Information
- Application Number
- JP2024102170
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-09-13
- Filing Date
- 2024-06-25
- Publication Date
- 2025-06-12
- Estimated Expiration
- 2039-09-12
AI Technical Summary
Current video encoding and decoding technologies face challenges in efficiently merging multiple video bitstreams due to computational complexity and scalability issues, particularly in systems with limited resources.
Incorporating merge identifiers into the video codec to reduce computational load and accelerate the merge process by determining the appropriate complexity of the merge procedure based on the encoding parameters.
The proposed solution enables efficient merging of video bitstreams with reduced computational requirements, improving scalability and maintaining rate distortion performance.
Smart Images

Figure 0007692090000002 
Figure 0007692090000003 
Figure 0007692090000004
Abstract
Description
Technical Field
[0001] The present invention relates to video encoding / decoding.
Background Art
[0002] Video composition is used in a number of applications where the composition of multiple video sources is presented to a user. Common examples are, for example, picture-in-picture (PiP) composition and fusion with overlay video content for advertising and user interfaces. Generating such a composition in the pixel domain requires complex calculations and may even be infeasible on a device with a single hardware decoder or other limited resources, necessitating parallel decoding of the input video bitstream. For example, in current IPTV system designs, the corresponding set-top boxes perform the composition, which is a major service cost factor due to their complexity, distribution, and limited lifespan. Reduction of these cost factors has motivated the ongoing effort to virtualize the set-top box functionality, for example, the shift to cloud resources for user interface generation. In such an approach, only the video decoder, a so-called zero client, is the only hardware remaining in the customer's home. The state-of-the-art technology in such system designs is composition in its simplest form based on transcoding, i.e., decoding, composition in the pixel domain, and re-encoding before or during transport. To reduce the workload from the full decoding and encoding cycle, operations in the transform coefficient domain rather than the pixel domain were first proposed for PiP composition. Since then, numerous techniques have been proposed to fuse or shorten the individual composition steps and apply them to current video codecs. However, the transcoding-based approach for general composition is still computationally complex and impairs the scalability of the system. Depending on the transcoding approach, such composition may also affect the rate distortion (RD) performance.
[0003] In addition, there are a wide range of applications built on tiles, where a tile is a spatial subset of an encoded video plane that is independent of adjacent tiles. A tile-based streaming system for 360° video works by splitting the 360° video into tiles that are encoded into sub-bitstreams of various resolutions and merged on the client side into a single bitstream according to the current user's viewing direction. Another application involving the merging of sub-bitstreams is, for example, the combination of conventional video content and banner ads. Further, the recombination of sub-bitstreams can be an important part in video conferencing systems where multiple users send individual video streams to a receiver where the streams are ultimately merged. Additionally, a tile-based cloud encoding system where the video is split into tiles and tile encoding is distributed to separate independent instances relies on fail-safe checks before the resulting tile bitstreams are merged back into a single bitstream. In all these applications, the streams are merged to enable decoding with a single video decoder using known compliance points. Here, merging means a lightweight bitstream rewrite that does not require a complete decoding operation and encoding operation of the entropy-encoded data for the reconstruction of the bitstream or pixel values. However, the techniques used to ensure the success of forming the bitstream merge, i.e., the compliance of the resulting merged bitstream, have been derived from very different application scenarios.
[0004] For example, a legacy codec supports a technique called Motion Constrained Tile Set (MCTS), where the encoder constrains inter-prediction between pictures to be limited within a tile or the image boundary, i.e., without using sample values or syntax element values that do not belong to the same tile or are located outside the image boundary. The origin of this technology is Region of Interest (ROI) decoding, where the decoder encounters irresolvable dependencies and can only decode specific subsections of the bitstream and the encoded image without avoiding drift. Another state-of-the-art technology in this context is the Structure of Pictures (SOP) SEI message that gives instructions on the applied bitstream structure, i.e., the order of pictures and the reference structure of inter-prediction. This information is provided by summarizing the Picture Order Count (POC) value, the active Sequence Parameter Set (SPS) identifier, and the Reference Picture Set (RPS) index to the active SPS for each picture between two Random Access Points (RAPs). Based on this information, a transcoder, a middlebox, or a media-aware network entity (MANE) or a media player can identify a bitstream structure that can assist in operating or modifying the bitstream, such as bitrate adjustment, frame dropping, and fast forwarding.
[0005] Both of the above exemplary notification techniques are essential when understanding whether they can perform a lightweight merge of sub-bitstreams without complex syntax changes or even complete transcoding, but they are never sufficient. More specifically, lightweight merge in this context interleaves the NAL units of the source bitstream with only minor rewriting operations, i.e., describes the parameter sets used together with the new image size and tile structure so that all bitstreams to be merged are located in separate tile regions. The next level of merge complexity is ideally composed of only minor rewriting of slice header elements without changing the variable-length codes within the slice header. There are further levels of merge complexity, and these levels are, for example, for re-executing entropy coding on slice data and are beneficial compared to full transcoding involving video decoding and encoding, and rather can be considered lightweight in that they can be changed without requiring the reconstruction of pixel values which would not be considered lightweight.
[0006] In the merged bitstream, all slices need to refer to the same parameter set. If the parameter sets of the original bitstreams use significantly different settings, many of the syntax elements of the parameter sets will have further effects on the syntax of the slice header and slice payload, and since the degree of the effects is different, lightweight merge may not be possible. The further downstream in the decoding process where syntax elements are involved, the more complex the merge / rewriting becomes. Some notable general categories of syntax dependencies (of parameter sets and other structures) can be distinguished as follows. A. Existence indication of syntax B. Dependency of value calculation C. Encoding tool control of slice payload a. Syntax used at the beginning of the decoding process (such as coefficient sign hiding, block splitting limit, etc.) b. Syntax used in the later stage of the decoding process (such as motion compensation (motion comp), loop filter, etc.) or general decoding process control (reference picture, bitstream order) D. Source format parameters
[0007] In Category A, the parameter set includes a number of presence flags for various tools (such as dependent_slice_segments_enabled_flag and output_flag_present_flag). If there are differences in such flags, the flags can be set to be valid in the combined parameter set, and the default values can be explicitly written to the merged slice header of the slices in the bitstream that did not contain syntax elements before merging. That is, in this case, the merge requires changes to the syntax of the parameter set and the slice header.
[0008] In Category B, the notified value of the syntax of the parameter set may be used together with the parameters of other parameter sets or slice headers in calculations. For example, in HEVC, the slice quantization parameter (quantization parameter (QP)) of a slice is used to control the coarseness of quantization of the transform coefficients of the residual signal of the slice. The notification of the slice QP (SliceQpY) in the bitstream depends on the QP notification at the picture parameter set (PPS) level as follows. SliceQpY = 26 + init_qp_minus26 (from PPS) + slice_qp_delta (from slice header)
[0009] Since all slices within the (merged) encoded video picture must reference the same active PPS, differences in init_qp_minus26 of the PPSs of the individual streams to be merged require adjustment of the slice headers such that they reflect the new common value of init_qp_minus26. That is, in this case, merging requires changes to the parameter set and the syntax of the slice header, similar to the case of category 1.
[0010] In category C, additional syntax elements of the parameter set control the encoding tools that affect the bitstream structure of the slice payload. Sub-category C.a and sub-category C.b can be distinguished according to how far downstream in the decoding process the syntax elements are included and (in relation to this) the complexity associated with changes to these syntax elements, i.e., where the syntax elements are included between entropy encoding and pixel-level reconstruction.
[0011] For example, one element of category C.a is the sign_data_hiding_enabled_flag that controls the derivation of the sign data of the encoded transform coefficients. The sign data hiding can be easily deactivated and the corresponding inferred sign data is explicitly written into the slice payload. However, for such changes to the slice payload in this category, it is not necessary to proceed to the pixel domain by complete decoding before re-encoding the per-pixel merged video. Another example is the determination of the inferred block partitioning or any other syntax where the inferred values can be easily written into the bitstream. That is, in this case, merging requires changes to the parameter set, the syntax of the slice header, and the slice payload that require entropy decoding / encoding.
[0012] However, subcategory C.b requires syntax elements related to processes much further downstream in decoding, and thus many of the complex decoding processes need to be executed in a way that is undesirable in terms of implementation and computation to avoid the remaining decoding steps. For example, due to differences in syntax elements connected to various decoding process-related syntaxes such as the motion compensation constraint "temporal motion constraint tile set SEI (Structure of Picture Information) message" or "SEI message", pixel-level transcoding cannot be avoided. In contrast to subcategory C.a, there are many encoder decisions that affect the slice payload in a way that cannot be changed without complete pixel-level transcoding.
[0013] In category D, there are syntax elements of the parameter set (such as chroma subsampling indicated by chroma_format_idc) that make pixel-level transcoding inevitable to merge sub-bitstreams with different values. That is, in this case, the merge requires a process of complete decoding, pixel-level merge, and complete encoding.
[0014] The above list is by no means exhaustive, but it becomes clear that a wide range of parameters affect the merits of merging sub-bitstreams into a common bitstream in various ways, and it is time-consuming to track and analyze these parameters. SUMMARY OF THE INVENTION PROBLEMS TO BE SOLVED BY THE INVENTION
[0015] An object of the present invention is to provide a video codec for efficiently merging video bitstreams. MEANS FOR SOLVING THE PROBLEMS
[0016] This object is achieved by the subject matter of the claims of the present application.
[0017] The basic idea underlying the present invention is that the improvement of the merging of multiple video streams is achieved by including one or more merge identifiers. This approach makes it possible to reduce the load on computing resources and accelerate the merge process.
[0018] According to an embodiment of the present invention, a video coder provides a video stream that includes encoded parameter information describing a plurality of encoding parameters, encoded video content information, i.e., information encoded using the encoding parameters defined by the parameter information, and one or more merge identifiers indicating whether the encoded video representation can be merged with another encoded video representation and / or how, for example, with what complexity, it can be merged. What complexity to use can be determined based on the parameter values defined by the parameter information. The merge identifier can be a concatenation of a plurality of encoding parameters that must be equal in two different encoded video representations in order to be able to merge the two different encoded video representations using a predetermined complexity. Additionally, the merge identifier can be a hash value of a concatenation of a plurality of encoding parameters that must be equal in two different encoded video representations in order to be able to merge the two different encoded video representations using a predetermined complexity.
[0019] According to an embodiment of the present invention, the merge identifier can indicate a merge procedure that merges by rewriting a parameter set, or merges by rewriting a parameter set and a slice header, or merges by rewriting a parameter set, a slice header, and a slice payload, or generally, a merge identifier type representing the complexity of an "appropriate" merge method. The merge identifier is associated with the merge identifier type, and the merge identifier associated with the merge identifier type includes encoded parameters that must be equal in two different encoded video representations so that the two different encoded video representations can be merged using the complexity of the merge procedure represented by the merge identifier type. The value of the merge identifier type can indicate the merge process, and the video coder is configured to switch between at least two of the following values of the merge identifier type: a first value of the merge identifier type representing a merge process by rewriting a parameter set, a second value of the merge identifier type representing a merge process by rewriting a parameter set and a slice header, and a third value of the merge identifier type representing a merge process by rewriting a parameter set, a slice header, and a slice payload.
[0020] According to an embodiment of the present invention, a plurality of merge identifiers are associated with different complexities of merge procedures. For example, each identifier indicates a parameter set / hash and type of the merge procedure. The coder can be configured to check whether the encoded parameters evaluated for providing the merge identifier are the same in all units of the video sequence and provide the merge identifier depending on the check.
[0021] According to an embodiment of the present invention, a plurality of encoding parameters can include merge-related parameters that are the same in different video streams, i.e., encoded video representations, in order to enable a merge that is less complex than a merge by complete pixel decoding. The video coder is configured to provide one or more merge identifiers based on the merge-related parameters, i.e., if there are no common parameters between two encoded video representations, then complete pixel decoding is performed, i.e., there is no possibility of reducing the complexity of the merge process. The merge-related parameters can include one or more or all of the following parameters: motion constraints at tile boundaries, e.g., parameters describing motion constrain tail set supplemental enhancement information (MCTS SEI), group of picture (GOP) structure, i.e., mapping of the encoding order to the display order, random access point indication, e.g., SOP: temporal hierarchicalization using structure of picture, information regarding SEI, parameters describing chroma encoding format and parameters describing luma encoding format, e.g., a set including at least luma / chroma of chroma format and bit depth, parameters describing advanced motion vector prediction, parameters describing sample adaptive offset, parameters describing temporal motion vector prediction, and parameters describing loop filter and other encoding parameters, i.e., the parameter set (including reference picture set, reference quantization parameter, etc.), slice header, and slice payload are rewritten to merge two encoded video representations.
[0022] According to an embodiment of the present invention, a merge identifier associated with a first complexity of a merge procedure, and a merge parameter associated with a second complexity of a merge procedure higher than the first complexity are determined based on a second set of encoding parameters that is a true subset of a first set of encoding parameters, and can be determined based on the first set of encoding parameters. A merge identifier associated with a third complexity of a merge procedure higher than the second complexity can be determined based on a third set of encoding parameters that is a true subset of the second set of encoding parameters. A video coder is configured to determine a merge identifier associated with a first complexity of a merge procedure based on a first set of encoding parameters, for example, a set of parameters applicable to a plurality of slices while keeping the encoding parameter set, such as a slice header and a slice payload, unchanged, i.e., to enable a merge process by rewriting (only) the parameter set, in two different video streams, for example, a first set that must be equal in an encoded video representation. The video coder can be configured to determine a merge identifier associated with a second complexity of a merge procedure based on a second set of encoding parameters, for example, to enable a merge process by rewriting the parameter set and the slice header, where the slice payload is kept unchanged and the parameter set applicable to a plurality of slices is changed, and the slice header is also changed, in two different video streams, for example, a second set that must be equal in an encoded video representation. The video coder can be configured to determine a merge identifier associated with a third complexity of a merge procedure based on a third set of encoding parameters, for example, to enable a merge process by rewriting the parameter set, the slice header, and the slice payload, where the parameter set applicable to a plurality of slices is changed, the slice header and the slice payload are also changed, but full pixel decoding and pixel re-encoding are not performed, in two different video streams, for example, a third set that must be equal in an encoded video representation.
[0023] According to an embodiment of the present invention, there is provided a video merger for providing a merged video representation based on a plurality of provided encoded video representations, such as a video stream. The video merger receives a plurality of video streams including encoded parameter information describing a plurality of encoding parameters, encoded video content information, that is, information encoded using the encoding parameters defined by the parameter information, and one or more merge identifiers indicating whether the encoded video representation can be merged with another encoded video representation and / or how, for example, with what complexity, it can be merged. The video merger is configured to determine depending on the merge identifier, that is, depending on the comparison of the merge identifiers of different video streams, regarding the usage method of the merge method, for example, the merge type, the merge process. The video merger is configured to select a merge method from among a plurality of merge methods depending on the merge identifier. The video merger can be configured to select at least two of the following merge methods: a first merge method which is a merge of a video stream that only changes a parameter set applicable to a plurality of slices while keeping the slice header and the slice payload unchanged; a second merge method which is a merge of a video stream that changes a parameter set applicable to a plurality of slices while keeping the slice payload unchanged and also changes the slice header; and a third merge method which is a merge of a video stream that changes a parameter set applicable to a plurality of slices and also changes the slice header and the slice payload, but does not perform complete pixel decoding and pixel re-encoding.
[0024] According to an embodiment of the present invention, a video merger compares merge identifiers of two or more video streams associated with the same given merge method or associated with the same merge identifier type, and determines whether to perform a merge using the given merge method depending on the result of the comparison. The video merger may be configured to selectively perform a merge using the given merge method when the comparison indicates that the merge identifiers of two or more video streams associated with the given merge method are equal. When the comparison of the merge identifiers indicates that the merge identifiers of two or more video streams associated with the given merge method are different, that is, without further comparing the encoding parameters themselves, the video merger may be configured to use a merge method having a higher complexity than the given merge method associated with the compared merge identifiers. When the comparison of the merge identifiers indicates that the merge identifiers of two or more video streams associated with the given merge method are equal, the video merger may be configured to selectively compare encoding parameters that must be equal in the two or more video streams to enable the merge of the video streams using the given merge method. When the comparison of the encoding parameters, that is, the encoding parameters that must be equal in the two or more video streams to enable the merge of the video streams using the given merge method, indicates that the encoding parameters are equal, the video merger is configured to selectively perform a merge using the given merge method. When the comparison of the encoding parameters indicates that the encoding parameters include a difference, the video merger is configured to perform a merge using a merge method having a higher complexity than the given merge method.
[0025] According to an embodiment of the present invention, the video merger may be configured to compare merge identifiers associated with merge methods having different complexities, i.e., to compare hashes. The video merger is configured to identify the merge method with the lowest complexity in which the associated merge identifiers in two or more video streams to be merged are equal. The video merger is configured to compare, instead of the hash version of the encoding parameters, the encoding parameter sets, i.e., the individual encoding parameters, which must be equal in two or more video streams to be merged in order to enable merging using the identified merge method. Different, usually overlapping, encoding parameter sets are associated with merge methods of different complexities. The video merger is configured to selectively merge two or more video streams using the identified merge method if the comparison indicates that the encoding parameters of the encoding parameter set associated with the identified merge method are equal in the video streams to be merged. The video merger is configured to merge two or more video streams using a merge method having a higher complexity than the identified merge method if the comparison indicates that the encoding parameters of the encoding parameter set associated with the identified merge method contain differences. The video merger is configured to determine, for example, which encoding parameters should be changed in the merge process, i.e., the merge of the video streams, depending on one or more differences between the merge identifiers associated with the same merge method or "merge identifier type" of the different video streams to be merged.
[0026] According to an embodiment of the present invention, a video merger obtains combined encoding parameters or a set of combined encoding parameters, such as a sequence parameter set SPS and a picture parameter set PPS, associated with slices of all video streams to be merged, based on the encoding parameters of the video streams to be merged. For example, when the values of all encoding parameters of one encoded video representation and another encoded video representation are the same, the combined encoding parameters are configured to be included in the merged video stream. The encoding parameters are updated by copying common parameters when there are differences between the encoding parameters of the encoded video representations. The encoding parameters are updated based on the main (i.e., one is main and the other is sub) encoded video representation. For example, some encoding parameters can be adapted according to a combination of video streams, such as the total image size. The video merger is configured to adapt the encoding parameters associated individually with individual video slices, for example, defined in the slice header, or when using a merging method having a complexity higher than the minimum, in order to obtain the modified slices to be included in the merged video stream. The adapted encoding parameters include parameters representing the image size of the merged encoded video representation, and the image size is calculated based on the image sizes of the encoded video representations to be merged, i.e., in each dimension, in the context of their spatial arrangement.
[0027] According to an embodiment of the present invention, there is provided an encoded video representation, i.e., a video coder for providing a video stream. The video coder is configured to provide a coarse granularity capability demand information, e.g., level information such as level 3 or level 4 or level 5, which describes the compatibility of the video stream with a video decoder having a function level among a plurality of predetermined function levels, e.g., the decodability of the video stream by the video decoder. The video coder is configured to describe which fraction of the acceptable function demands associated with one of the predetermined function levels is required for decoding the encoded video representation, i.e., how much decoder functionality is needed, and / or which fraction of the acceptable function demands, i.e., the "level limit of the merged bitstream", contributes to the merged video stream in which the video stream whose function demand matches one of the predetermined function levels, i.e., whose function demand is below the acceptable function demand, is merged. The video coder is configured to provide fine granularity function demand information, e.g., merge level limit information. The "acceptable function demand of the merged bitstream that matches one of the predetermined function levels" corresponds to the "level limit of the merged bitstream". The video coder is configured to provide fine granularity function demand information such that the fine granularity function demand information includes a ratio value or a percentage value that refers to one of the predetermined function levels. The video coder is configured to provide fine granularity function demand information such that the fine granularity function demand information includes reference information and fraction information. The reference information describes which of the predetermined function levels the fraction information refers to so that the fine granularity function demand information describes a fraction of one of the predetermined function levels as a whole.
[0028] According to an embodiment of the present invention, there is provided a video merger for providing a merged video representation based on a plurality of encoded video representations, i.e., video streams. The video merger includes encoded parameter information describing a plurality of encoding parameters, encoded video content information, e.g., information encoded using the encoding parameters defined by the parameter information, and coarse-grained functional requirement information describing the compatibility of the video stream with a video decoder having a functional level among a plurality of predetermined functional levels, i.e., the decodability of the video stream by the video decoder, e.g., level information, level 3 or level 4 or level 5, and fine-grained functional requirement information, e.g., merge level limit information. The video merger may be configured to receive a plurality of video streams. The video merger is configured to merge two or more video streams depending on the coarse-grained functional requirement information and the fine-grained functional requirement information. The video merger is configured to determine, i.e., without violating acceptable functional requirements, i.e., such that the functional requirements of the merged video stream match one of the predetermined functional levels depending on the fine-resolution functional requirement information, which video streams can be included in the merged video stream or whether to include them. The video merger may be configured to determine, for example, whether an effective merged video stream that does not violate acceptable functional requirements, i.e., such that the functional requirements of the merged video stream match one of the predetermined functional levels, can be obtained by merging two or more video streams depending on the fine-resolution functional requirement information. The video merger is configured to summarize the fine-grained functional requirement information of the plurality of video streams to be merged, for example, to determine which video streams can be included or whether an effective merged video stream can be obtained.
[0029] Preferred embodiments of the present invention will be described below with reference to the drawings.
Brief Description of the Drawings
[0030]
Figure 1
Figure 2
Figure 3
Figure 4-1
Figure 4-2
Figure 4-3
Figure 4-4
Figure 5a
Figure 5b
Figure 6a
Figure 6b
Figure 7a
Figure 7b
Figure 7c
Figure 7d
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
[0031] In the following description, specific details such as specific embodiments, procedures, techniques, etc. are described for illustrative purposes rather than limitation. It will be understood by those skilled in the art that other embodiments may be used separately from these specific details. For example, the following description proceeds using non-limiting exemplary applications, but this technology can be applied to any type of video codec. In some cases, detailed descriptions of well-known methods, interfaces, circuits, and devices are omitted so as not to obscure the description with unnecessary details.
[0032] In the following description, equivalent or equivalent elements having equivalent or equivalent functions are represented by equivalent or equivalent reference numerals.
[0033] The invention of this specification enables the identification of sub-bitstreams that can be merged into a legal bitstream together and sub-bitstreams that cannot be merged into a legal bitstream together with a given level of complexity (related merging method) in future video codecs such as VVC (Versatile Video Coding). This is achieved by providing means for inserting into each sub-bitstream an instruction that enables the identification of sub-bitstreams that can be merged into a legal bitstream together and sub-bitstreams that cannot be merged into a legal bitstream together with a given level of complexity (related merging method). This instruction is hereinafter referred to as a "merge identifier" and further provides information regarding an appropriate merging method via an instruction called a "merge identifier type". If two bitstreams have the "merge identifier" and the same value, it is possible to merge the sub-bitstreams into a new combined bitstream with a given level of merge complexity (related merging method).
[0034] FIG. 1 shows a video encoder 2 for providing an encoded video representation, i.e., an encoded video stream 12, based on a provided (input) video stream 10 that includes an encoder core 4 including an encoding parameter determination member 14 and a merge identifier provider 6. The provided video stream 10 and the encoded video stream 12 each have a bitstream structure, such as shown as a simplified configuration in FIG. 2. This bitstream structure is composed of a plurality of Network Abstraction Layer (NAL) units, and each NAL unit includes various parameters and / or data, for example, a Sequence Parameter Set (SPS) 20, a Picture Parameter Set (PPS) 22, an instantaneous decoder refresh (IDR) 24, Supplemental Enhancement Information (SEI) 26, and a plurality of slices 28. The SEI 26 includes various messages, i.e., a Structure of Picture, a Motion Constraint Tile Set, etc. The slice 28 includes a header 30 and a payload 32. The encoding parameter determination member 14 determines encoding parameters based on the SPS 20, PPS 22, SEI 26, and the slice header 30. The IDR 24 is not necessarily a factor for determining the encoding parameters according to the present invention, but the IDR 24 can optionally be included for determining the encoding parameters. The merge identifier provider 6 provides a merge identifier indicating whether the encoded video stream can be merged with another encoded video stream and / or how (with what complexity) it can be merged.The merge identifier determines a merge identifier that indicates, for example, the complexity of the merge procedure that merges by rewriting the parameter set, or by rewriting the parameter set and the slice header, or by rewriting the parameter set, the slice header, and the slice payload, or, generally, the merge identifier type associated with the "appropriate" merge method. The merge identifier is associated with the merge identifier type, and the merge identifier associated with the merge identifier type must contain encoding parameters that are equal in two different encoded video representations so that the two different encoded video representations can be merged using the complexity of the merge procedure represented by the merge identifier type.
[0035] As a result, the encoded video stream 12 includes a plurality of encoding parameters, encoded video content information, and encoded parameter information that describes one or more merge identifiers. In this embodiment, the merge identifier is determined based on the encoding parameters. However, the merge identifier value can be set according to the wishes of the encoder operator, and there may be little guarantee of sufficient collision avoidance in a closed system. Third-party entities such as DVB (Digital Video Broadcasting) and ATSC (Advanced Television Systems Committee) can also define the values of the merge identifiers to be used within their systems.
[0036] The merge identifier type will be described as follows in consideration of another embodiment according to the present invention using FIGS. 3 to 9.
[0037] Figure 3 shows an encoder 2(2a) including an encoder core 4 and a merge identifier 6a, and shows the data flow in the encoder 2a. The merge identifier 6a includes a first hash member 16a and a second hash member 16b that provide a hash value as a merge identifier for indicating a merge type. That is, the merge value can be formed from a concatenation of encoded values of a defined set of syntax elements (encoding parameters) of a bitstream, hereinafter referred to as a hash set. The merge identifier value can also be formed by applying the above concatenation of encoded values to a well-known hash function such as MD5, SHA-3 or any other suitable function. As shown in Figure 3, the input video stream 10 includes input video information, and the input video stream 10 is processed by the encoder core 4. The encoder core 4 encodes the input video content, and the encoded video content information is stored in the payload 32. The encoding parameter determination member 14 receives parameter information including SPS, PPS, slice headers and SEI messages. The parameter information is stored in each corresponding unit and received by the merge identifier provider 6a, that is, received by the first hash member 16a and the second hash member 16b respectively. The first hash member 16a generates, for example, a hash value indicating a merge identifier type 2 based on the encoding parameters, and the second hash member 16b generates, for example, a hash value indicating a merge identifier type 1 based on the encoding parameters.
[0038] The quality of the merge possibility indication for the above syntax category is determined by the content of the hash set, that is, the value of the syntax element (i.e., merge parameter) concatenated to form the merge identifier value.
[0039] For example, the merge identifier type indicates different levels of merge possibility corresponding to appropriate merge methods for the merge identifier, that is, syntax elements incorporated in the hash set, as follows. Type 0: Merge identifier for merge by rewriting parameter set Type 1: Merge identifier for merge by rewriting parameter set and slice header Type 2: Merge identifier for merge by rewriting parameter set, slice header, and slice payload
[0040] For example, if two input sub-bitstreams to be merged in the context of an application are provided, the device can compare the values of the merge identifier and the merge identifier type and conclude on the possibility of using the method associated with the merge identifier type for the two sub-bitstreams.
[0041] The following table shows the mapping between syntax element categories and merge methods associated therewith. [Table 1]
[0042] As already mentioned above, the syntax categories are never exhaustive, but it becomes clear that a wide range of parameters affect the merits of merging sub-bitstreams into a common bitstream in various ways and that it is time-consuming to track and analyze these parameters. In addition, the syntax categories and merge methods (types) do not exactly correspond, so, for example, some parameters are required for category B, but the same parameters are not required for merge method 1.
[0043] To enable the device to easily identify applicable merge methods, a merge identifier value is generated by two or more of the above merge identifier type values and written into the bitstream. Merge method (type) 0 indicates the first value of the merge identifier type, merge method (type) 1 indicates the second value of the merge identifier type, and merge method (type) 2 indicates the third value of the merge identifier type.
[0044] The following exemplary syntax needs to be incorporated into a hash set. <Temporal Motion Constraint Tile Set SEI Message (Merge Methods 0, 1, 2, Syntax Category C.b) that Indicates Motion Constraints at Tiles and Image Borders> <Structure of Picture Information SEI Message (Merge Methods 0, 1, 2, Syntax Category C.b) that Defines GOP Structure (i.e., Mapping from Encoding Order to Display Order, Random Access Point Indication, Temporal Hierarchy, Reference Structure)> <Syntax Element Values of Parameter Set> Reference Picture Set (Merge Method 0, Syntax Category B) Chroma Format (Merge Methods 0, 1, 2, Syntax Category D) Reference QP, Chroma QP Offset (Merge Method 0, Syntax Category B) Bit Depth Luma / Chroma (Merge Methods 0, 1, 2, Syntax Category D) <hrd Parameters> Initial Arrival Delay (Merge Method 0, Syntax Category B) Initial Removal Delay (Merge Method 0, Syntax Category B) <Coding Tools> Coding Block Structure (Maximum / Minimum Block Size, Presumed Partition) (Merge Methods 0, 1, Syntax Category C.a) Transformation Size (Minimum / Maximum) (Merge Methods 0, 1, Syntax Category C.a) PCM Block Usage (Merge Methods 0, 1, Syntax Category C.a) High Motion Vector Prediction (Merge Methods 0, 1, 2, Syntax Category C.b) Sample Adaptive Offset (Merge Method 0, Syntax Category C.b) Temporal Motion Vector Prediction (Merge Methods 0, 1, 2, Syntax Category C.b) Intra Smoothing (Merge Methods 0, 1, Syntax Category C.a) Dependent Slices (Merge Method 0, Syntax Category A) Sign Concealment (Merge Methods 0, 1, Syntax Category C.a) Weighted Prediction (Merge Method 0, Syntax Category A) Transquant Bypass (Merge Methods 0, 1, Syntax Category C.a) Entropy Coding Synchronization (Merge Methods 0, 1, Syntax Category C.a) (Skup4) Loop Filter (Merge Methods 0, 1, 2, Syntax Category C.b) <Slice Header Value> Parameter Set ID (Merge Method 0, Syntax Category C.a) Reference Picture Set (Merge Methods 0, 1, Syntax Category B) <Usage of Implicit CTU Address Signaling cp. (Referenced by European Patent Application EP18153516) (Merge Method 0, Syntax Category A)>
[0045] That is, for the first value of the merge identifier type, i.e., type 0, the syntax elements (parameters), namely the motion constraints at tile and picture boundaries, GOP structure, reference picture set, chroma format, reference quantization parameter and chroma quantization parameter, bit depth luma / chroma, parameters regarding initial arrival delay and parameter initial deletion delay, virtual reference decoder parameters including those for coded block structure, transform minimum and / or maximum size, usage of pulse code modulation blocks, advanced motion vector prediction, sample adaptive offset, temporal motion vector prediction, intra smoothing, dependent slices, sign concealment, weighted prediction, transquantization bypass, entropy coding synchronization, loop filter, slice header values including parameter set ID, slice header values including reference picture set, and the usage of implicit coding transform unit address signaling need to be incorporated into the hash set.
[0046] For the second value of the merge identifier type, i.e., type 1, the syntax elements (parameters), namely the motion constraints at tile and picture boundaries, GOP structure, chroma format, bit depth luma / chroma, coded block structure, transform minimum and / or maximum size, usage of pulse code modulation blocks, advanced motion vector prediction, sample adaptive offset, temporal motion vector prediction, intra smoothing, sign concealment, transquantization bypass, entropy coding synchronization, loop filter, and slice header values including reference picture set.
[0047] The third value of the merge identifier type, i.e., in type 2, is the syntax element (parameter), i.e., the motion constraints at the tile and picture boundaries, GOP structure, chroma format, bit depth luma / chroma, advanced motion vector prediction, sample adaptive offset, temporal motion vector prediction, loop filter.
[0048] FIG. 4 shows a schematic diagram showing an example of encoding parameters according to an embodiment of the present invention. In FIG. 4, reference numeral 40 represents type 0, and the syntax elements belonging to type 0 are indicated by dotted lines. Reference numeral 42 represents type 1, and the syntax elements belonging to type 1 are indicated by normal lines. Reference numeral 44 represents type 2, and the syntax elements belonging to type 2 are indicated by broken lines.
[0049] FIGS. 5A and 5B are an example of a sequence parameter set (SPS) 20, and the syntax elements necessary for type 0 are indicated by reference numeral 40. Similarly, the syntax elements necessary for type 1 are indicated by reference numeral 42, and the syntax elements necessary for type 2 are indicated by reference numeral 44.
[0050] FIGS. 6A and 6B are an example of a picture parameter set (PPS) 22, and the syntax elements necessary for type 0 are indicated by reference numeral 40, and the syntax elements necessary for type 1 are indicated by reference numeral 42.
[0051] FIGS. 7A to 7D are an example of a slice header 30. As indicated by reference numeral 40 in FIG. 7C, only one syntax element of the slice header is necessary for type 0.
[0052] FIG. 8 is an example of a structure of picture (SOP) 26a, and it is necessary for type 2 that all syntax elements belong to the SOP.
[0053] FIG. 9 is an example of a motion constraint tile set (MCTS) 26b, and it is necessary for type 2 that all syntax elements belong to the MCTS.
[0054] As described above, the merge identifier 6a in FIG. 3 generates a merge identifier value using the hash functions in the hash members 16a and 16b. When two or more merge identifier values are generated using the hash function, the hash is connected in the sense that the input to the hash function of the second merge identifier that includes additional elements in the hash set with respect to the first merge identifier uses the first merge identifier (hash result) instead of the respective syntax element values for the concatenation of the inputs to the hash function.
[0055] In addition, the presence of the merge identifier also provides a guarantee that the syntax elements incorporated in the hash set have the same value in all access units (AUs) of the encoded video sequence (CVS) and / or the bitstream. Further, this guarantee has the form of a constraint flag in the profile / level syntax of the parameter set.
[0056] The merge process will be described as follows in view of another embodiment according to the present invention using FIGS. 10 to 13.
[0057] FIG. 10 shows a video merger for providing a merged video stream based on a plurality of encoded video representations. The video merger 50 includes a receiver 52 that receives input video streams 12 (12a and 12b indicated in FIG. 12), a merge method identifier 54, and a merge processor 56. The merged video stream 60 is transmitted to a decoder. When the video merger 50 is included in the decoder, the merged video stream is transmitted to a user device or any other device for displaying the merged video stream.
[0058] The merge process is driven by the above merge identifier and merge identifier type. The merge process can only perform the generation of parameter sets and the interleaving of NAL units, which is the lightest form of merge associated with merge identifier type value 0, i.e., the first complexity. FIG. 12 shows an example of the merge method of the first complexity. As indicated in FIG. 12, the parameter sets, i.e., SPS1 of video stream 12a and SPS2 of video stream 12b are merged (a merged SPS is generated based on SPS1 and SPS2), and PPS1 of video stream 12a and PPS2 of video stream 12b are merged (a merged PPS is generated based on PPS1 and PPS2). IDR is optional data and thus is omitted from the description. In addition, slice 1,1 and slice 1,2 of video stream 12a and slice 2,1 and slice 2,2 of video stream 12b are interleaved as shown as the merged video stream 60. In FIGS. 10 and 12, two video streams 12a and video stream 12b are input as an example. However, more video streams may be input and merged in the same way.
[0059] Optionally, the merge process can also include the rewriting of slice headers in the bitstream during NAL unit interleaving associated with merge identifier type value 1, i.e., the second complexity. Finally, there may be a need to adjust the syntax elements in the slice payload that require entropy decoding and encoding during NAL unit interleaving associated with merge identifier type value 2, i.e., the third complexity. The merge identifier and merge identifier type drive the selection of one of the merge processes to be performed and the determination of its details.
[0060] The input to the merge process is a list of input sub-bitstreams that also represent a spatial arrangement. The output of the process is the merged bitstream. In general, in all bitstream merge processes, a new set of output bitstream parameters needs to be generated, which can be based on the set of parameters of the input sub-bitstreams, for example, the first input sub-bitstream. The necessary updates to the parameter set include the image size. For example, the image size of the output bitstream is calculated as the sum of the image sizes of the input sub-bitstreams in the context of their spatial arrangement for each dimension.
[0061] It is a requirement of the bitstream merge process that all input sub-bitstreams have the same value for at least one instance of a merge identifier and a merge identifier type. In one embodiment, a merge process is performed that associates all sub-bitstreams with the lowest value of a merge identifier type that has the same value for the merge identifier.
[0062] For example, the difference in merge identifier values having a particular merge identifier type value is used in the merge process to determine the details of the merge process according to the difference value of the merge identifier type for which the merge identifier values match. For example, as shown in FIG. 11, a first merge identifier 70 for which the merge identifier type value is equal to 0 does not match between two input sub-bitstreams 70a, 70b, but a second merge identifier 80 for which the merge identifier type value is equal to 1 matches between two input sub-bitstreams 80a, 80b. In this case, the difference in the bit positions of the first merge identifier 70 indicates the syntax elements (related to the slice header) that require adjustment in all slices.
[0063] FIG. 13 shows a video merger 50(50a) including a receiver (not shown), a merge method identifier including a merge identifier comparator 54a and an encoding parameter comparator 54b, and a merge processor 56. When the input video stream 12 includes a hash value as a merge identifier value, the values of each input video stream are compared by the merge identifier comparator 54a. For example, when both input video streams have the same merge identifier value, the individual encoding parameters of each input video stream are compared by the encoding parameter comparator 54b. Based on the encoding parameter comparison result, a merge method is determined, and the merge processor 56 merges the input video stream 12 using the determined merge method. When the merge identifier value (hash value) also indicates the merge method, the comparison of the individual encoding parameters is unnecessary.
[0064] In the above, three merge methods, a first complexity, a second complexity, and a third complexity merge method are described. The fourth merge method is the merge of video streams using complete pixel decoding and pixel re-encoding. The fourth merge method is applied when all three merge methods are not applicable.
[0065] Considering another embodiment according to the present invention using FIGS. 14 and 15, a process for identifying the level of the merge result will be described below. Identifying the level of the merge result means placing information in a sub-bitstream and how much the contribution of the sub-bitstream to the level limit of the merged bitstream incorporating the sub-bitstream is.
[0066] FIG. 14 shows an encoder core including an encoding parameter determination 14, a merge identifier provider (not shown), and an encoder 2(2b) including a granularity function provider 8.
[0067] Generally, when a sub-bitstream is to be merged into a combined bitstream, instructions on how each individual sub-bitstream contributes to codec-system-level specific restrictions that the potential merged bitstream must comply with are essential to ensure the creation of a legal combined bitstream. Conventionally, the codec-level granularity has been rather coarse, e.g., at identifying key resolutions such as 720p, 1080p, or 4K, but instructions on merge-level restrictions require much finer granularity. This conventional level of granularity of instructions is insufficient to represent the contribution of each individual sub-bitstream to the merged bitstream. Assuming the number of tiles to be merged is unknown in advance, a reasonable trade-off between flexibility and bitrate cost needs to be found, which generally far exceeds the granularity of conventional level restrictions. One exemplary use case is a 360-degree video stream where a service provider needs the freedom to choose from different tile structures such as 12 tiles, 24 tiles, or 96 tiles per 360-degree video, and each tile stream, assuming an equal rate distribution, would contribute 1 / 12, 1 / 24, or 1 / 96 of the overall level restriction such as 8K. Further, assuming a non-uniform rate distribution between tiles to achieve uniform quality across the video plane, any fine granularity may be required.
[0068] Such signaling can be, for example, a notified ratio and / or an additionally notified percentage of a level. For example, in a four-participant meeting scenario, each participant would transmit a legal level 3 bitstream that includes an instruction, i.e., level information included in coarse-grained function information, and the instruction indicates that the transmitted bitstream complies with 1 / 3 and / or 33% of the level 5 restriction, i.e., this information can be included in fine-grained function information. A receiver of multiple such streams, i.e., a video merger 50(50b) as shown in FIG. 15, can thus know, for example, that three such bitstreams can be merged into a single combined bitstream that complies with level 5.
[0069] The granularity function information may have an indication of a ratio and / or percentage as a vector of values, where each dimension relates to a different aspect of codec level limitations, such as the maximum allowable number of luma samples per second, the maximum image size, the bitrate, the buffer occupancy, the number of tiles, etc. Additionally, the ratio and / or percentage refers to the general codec level of the video bitstream.
[0070] Although some aspects are described in the context of an apparatus or system, it is clear that these aspects also represent a description of a corresponding method where a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of an apparatus and / or system. Some or all of the method steps may be performed by (or using) a hardware device such as a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such a device.
[0071] The data stream of the present invention can be stored in a digital storage medium or transmitted via a transmission medium such as a wireless transmission medium or a wired transmission medium like the Internet.
[0072] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software. The embodiments can be executed using a digital storage medium storing electronically readable control signals that cooperate (or can cooperate) with a programmable computer system such that each method is executed, for example, a floppy disk, a DVD, a Blu - Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM, or a flash memory. Thus, the digital storage medium can be computer - readable.
[0073] Some embodiments according to the present invention include a data carrier having electronically readable control signals that can cooperate with a programmable computer system so that one of the methods described herein is executed.
[0074] In general, embodiments of the present invention can be implemented as a computer program product having program code, and the program code operates to execute one of the methods when the computer program product operates on a computer. The program code can be stored, for example, in a machine-readable carrier.
[0075] Other embodiments include a computer program for executing one of the methods described herein, stored in a machine-readable carrier.
[0076] In other words, one embodiment of the method of the present invention is thus a computer program having program code for executing one of the methods described herein when the computer program operates on a computer.
[0077] Another embodiment of the method of the present invention is thus a data carrier (or digital storage medium, or computer-readable medium) on which is recorded a computer program for executing one of the methods described herein. The data carrier, digital storage medium, or recorded medium is typically tangible and / or non-transitory.
[0078] Another embodiment of the method of the present invention is thus a sequence of data streams or signals representing a computer program for executing one of the methods described herein. The sequence of data streams or signals can be configured to be transferred, for example, via a data communication connection, such as via the Internet.
[0079] Another embodiment includes processing means, such as a computer or a programmable logic device, configured or adapted to perform one of the methods described herein.
[0080] Another embodiment includes a computer having installed thereon a computer program for performing one of the methods described herein.
[0081] Another embodiment according to the present invention includes an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver can be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system can include, for example, a file server for transferring the computer program to the receiver.
[0082] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Generally, these methods are preferably performed by any hardware device.
[0083] The apparatuses described herein can be implemented using a hardware device, or using a computer, or using a combination of a hardware device and a computer.
[0084] The apparatuses described herein, or any component of the apparatuses described herein, can be implemented at least partially as hardware and / or as software.
[0085] The methods described herein may be performed using a hardware device, or using a computer, or using a combination of a hardware device and a computer.
[0086] The methods described herein, or any component of the methods described herein, may be implemented at least in part by hardware and / or by software.
[0087] The above-described embodiments are merely illustrative of the principles of the present invention. It should be understood that modifications and variations of the configurations and details described herein will be apparent to those skilled in the art. Accordingly, it is intended to be limited only by the appended claims, and not by the specific details presented herein as descriptions and explanations of embodiments.
Claims
1. 1. A video merger for providing a merged video representation based on a plurality of encoded video representations, comprising: the video merger is configured to receive a plurality of video streams including encoded parameter information describing a plurality of encoding parameters, encoded video content information, level information indicating compatibility of the video stream with a video decoder having a functionality level among a plurality of predefined functionality levels, and fraction level information indicating a fraction of one of the plurality of predefined functionality levels; A video merger, wherein the video merger is configured to merge two or more video streams depending on the level information and the fractional level information.
2. The video merger of claim 1 , wherein the fractional level information comprises a ratio or percentage value.
3. The video merger of claim 1 , wherein the fraction level information indicates a fraction of a maximum allowed amount of luma samples per second.
4. The video merger of claim 1 , wherein said fraction level information indicates a fraction of a maximum image size.
5. 2. The video merger of claim 1, wherein said fraction level information indicates a fraction of a maximum bit rate.
6. The video merger of claim 1 , wherein the fractional level information indicates a fractional amount of buffer fullness.
7. The video merger of claim 1 , wherein the fraction level information indicates a fraction of a maximum number of tiles.
Citation Information
Patent Citations
Image encoder and image encoding method, image recorder and image transmitter
JP2003299089A
Encoder, encoding method and program
JP2012231295A
Merging encoded bitstreams
JP2013513999A
Merging encoded bitstreams
WO2011081643A2