Object mask information for additional extended information messages
By incorporating Object Mask Information (OMI) Enhanced Information (SEI) messages, the challenges of efficient video encoding and decoding are addressed, enhancing compression efficiency and quality in video coding standards.
Patent Information
- Authority / Receiving Office
- JP Β· JP
- Patent Type
- Applications
- Current Assignee / Owner
- ALIBABA (CHINA) CO LTD
- Filing Date
- 2024-04-11
- Publication Date
- 2026-05-19
AI Technical Summary
Existing video coding standards, such as HEVC and VVC, face challenges in efficiently encoding and decoding video data to achieve high compression efficiency while maintaining quality, particularly in handling object masks and additional enhancement information.
The implementation of Object Mask Information (OMI) Enhanced Information (SEI) messages in video coding processes, which include encoding and decoding methods to represent object masks using auxiliary pictures and additional enhancement information, enhancing the encoding and decoding of video data.
Improves the encoding and decoding efficiency by providing detailed object mask information, allowing for better compression and quality maintenance in video data processing.
Smart Images

Figure 2026515776000001_ABST
Abstract
Description
Technical Field
[0001] (Cross - reference to Related Applications) This disclosure claims priority to U.S. Provisional Application No. 63 / 495,546, filed on April 11, 2023; U.S. Provisional Application No. 63 / 587,750, filed on October 4, 2023; U.S. Provisional Application No. 63 / 615,294, filed on December 28, 2023; and U.S. Application No. 18 / 624,636, filed on April 2, 2024, the entire contents of which are hereby incorporated by reference in their entirety.
[0002] This disclosure generally relates to video processing, and more specifically, to methods and apparatuses for signaling an object mask information (OMI) - added enhancement information (SEI) message.
Background Art
[0003] Video is a collection of still pictures (or "frames") that capture visual information. To reduce memory storage and transmission bandwidth, video can be compressed before being stored or transmitted and decompressed before being displayed. The compression process is usually called encoding, and the decompression process is usually called decoding. There are various video coding formats that use standardized video coding techniques, and the most common ones are based on prediction, transformation, quantization, entropy coding, and in - loop filtering. Video coding standards that define specific video coding formats, such as High - Efficiency Video Coding (HEVC / H.265), Versatile Video Coding (VVC / H.266), and AVS standards, are developed by standardization organizations. As the advanced video coding techniques adopted in video standards are increasing, the coding efficiency of new video coding standards is also increasing.
Summary of the Invention
[0004] Embodiments of this disclosure provide a method and apparatus for signaling Object Mask Information (OMI) Enhanced Information (SEI) messages.
[0005] According to some exemplary embodiments, a method for detecting an object is provided. This method includes receiving a bitstream, decoding the encoded information of the bitstream to obtain a primary picture and an auxiliary picture, wherein the auxiliary picture shows the mask of an object in the primary picture and the mask of the object is represented by the sample values ββof the auxiliary picture, and decoding the encoded information of the bitstream to obtain an Additional Enhancement Information (SEI) message, wherein the SEI message shows the attributes of the mask of the object.
[0006] According to some exemplary embodiments, an encoding method is provided. This encoding method includes the steps of receiving a video sequence and encoding one or more pictures of the video sequence to generate a bitstream, which includes encoding an auxiliary picture showing the mask of an object in the primary picture, wherein the mask of the object is represented by the sample values ββof the auxiliary picture, and generating an additional extended information (SEI) message showing the attributes of the mask of the object.
[0007] According to some exemplary embodiments, a non-temporary, computer-readable storage medium for storing a bitstream of video is provided. The bitstream includes a primary picture having an object, an auxiliary picture showing the mask of the object, wherein the mask of the object is represented by a sample value of the auxiliary picture, and an additional extended information (SEI) message indicating the attributes of the mask of the object. [Brief explanation of the drawing]
[0008] Embodiments and various aspects of this disclosure are shown in the following detailed description and accompanying drawings. Various features shown in the drawings are not depicted to scale.
[0009] [Figure 1] This is a schematic diagram illustrating an exemplary system for image data preprocessing and coding according to some embodiments of the present disclosure.
[0010] [Figure 2A] This is a schematic diagram illustrating an exemplary encoding process of a hybrid video coding system according to an embodiment of the present disclosure.
[0011] [Figure 2B] This is a schematic diagram showing another exemplary encoding process for a hybrid video coding system according to an embodiment of the present disclosure.
[0012] [Figure 3A] This is a schematic diagram illustrating an exemplary decoding process for a hybrid video coding system according to an embodiment of the present disclosure.
[0013] [Figure 3B] This is a schematic diagram illustrating another exemplary decoding process for a hybrid video coding system according to an embodiment of the present disclosure.
[0014] [Figure 4] This is a block diagram of an exemplary apparatus for image data preprocessing or coding according to some embodiments of the present disclosure.
[0015] [Figure 5] This is a syntax chart of an exemplary Object Mask Information (OMI)SEI message according to some embodiments of the present disclosure.
[0016] [Figure 6]Schematic diagram showing an exemplary method for encoding a video sequence into a bitstream according to an embodiment of the present disclosure.
[0017] [Figure 7] Schematic diagram showing an exemplary primary picture and auxiliary picture according to an embodiment of the present disclosure.
[0018] [Figure 8] Schematic diagram showing sub-steps of an exemplary method for encoding a video sequence into a bitstream according to an embodiment of the present disclosure.
[0019] [Figure 9] Exemplary binary representation of sample value p[x][y] according to some embodiments of the present disclosure.
[0020] [Figure 10] Schematic diagram showing an exemplary method for detecting an object according to an embodiment of the present disclosure.
[0021] [Figure 11] Schematic diagram showing the content of an exemplary bitstream.
Mode for Carrying Out the Invention
[0022] Here, reference is made in detail to the exemplary embodiments illustrated in the accompanying drawings. The following description refers to the accompanying drawings, and the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following description of the exemplary embodiments do not represent all implementations of the present invention. Rather, they are merely examples of devices and methods according to aspects related to the present invention described in the appended claims. Specific aspects of the present disclosure will be described in more detail below. The terms and definitions provided in this specification shall prevail in case of conflict with the terms and / or definitions incorporated by reference.
[0023] The Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (ITU-TVCEG) and the ISO / IEC Video Experts Group (ISO / IECMPEG) is currently developing the Multipurpose Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 using half the bandwidth.
[0024] To achieve this goal, JVET has been working since 2015 to develop a technology that surpasses HEVC using the Joint Search Model (JEM) reference software. As coding techniques were incorporated into JEM, JEM achieved significantly higher coding performance than HEVC. In October 2017, VCEG and MPEG jointly issued a Call for Proposals (CfP) to formally begin development of a next-generation video compression standard that surpasses HEVC. Responses to the CfP were evaluated at the JVET meeting held in San Diego in April 2018, and the formal development process for the VVC standard began in April 2018.
[0025] The VVC standard has been steadily progressing since April 2018, with further coding technologies being added to provide even better compression performance. VVC is based on the same hybrid video coding system used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, and H.263.
[0026] Figure 1 is a block diagram showing a system 100 for image data preprocessing and coding according to some disclosed embodiments. Image data may include images (also called βpicturesβ or βframesβ), multiple images, or video. An image is a still picture. Multiple images may or may not be spatially or temporally related. A video is a collection of images arranged in a temporal sequence.
[0027] As shown in Figure 1, the system 100 includes a source device 120 that provides encoded video data to be later decoded by a destination device 140. According to the disclosed embodiments, each of the source device 120 and the destination device 140 may include any of a wide range of devices, including desktop computers, notebook (e.g., laptop) computers, servers, tablet computers, set-top boxes, mobile phones, vehicles, cameras, image sensors, robots, televisions, wearable devices (e.g., smartwatches or wearable cameras), display devices, digital media players, video game consoles, video streaming devices, and the like. The source device 120 and the destination device 140 may be equipped for wireless or wired communication.
[0028] As shown in Figure 1, the source device 120 may include an image / video preprocessor 122, an image / video encoder 124, and an output interface 126. The destination device 140 may include an input interface 142, an image / video decoder 144, and one or more machine vision applications 146. The image / video preprocessor 122 preprocesses image data, i.e., images or videos, to generate an input bitstream for the image / video encoder 124. The image / video encoder 124 encodes the input bitstream and outputs the encoded bitstream 162 via the output interface 126. The encoded bitstream 162 is transmitted via the communication medium 160 and received by the input interface 142. The image / video decoder 144 then decodes the encoded bitstream 162 to generate decoded data available for use by the machine vision application 146.
[0029] More specifically, the source device 120 may further include various devices (not shown) for providing source image data to be preprocessed by the image / video preprocessor 122. Devices for providing source image data may include image / video capture devices such as cameras, image / video archives or storage devices containing previously captured images / videos, or image / video feed interfaces for receiving images / videos from image / video content providers.
[0030] The image / video encoder 124 and the image / video decoder 144 may each be implemented as one or more suitable encoder or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. If encoding or decoding is partially implemented in software, the image / video encoder 124 or the image / video decoder 144 may implement the techniques of this disclosure by storing instructions for the software in a suitable non-temporary computer-readable medium and executing the instructions in hardware using one or more processors. Each of the image / video encoder 124 or the image / video decoder 144 may be included in one or more encoders or decoders, which may be integrated as part of a combined encoder / decoder (CODEC) in separate devices.
[0031] The image / video encoder 124 and image / video decoder 144 may operate according to any video coding standard such as Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), Multipurpose Video Coding (VVC), AO Media Video 1 (AV1), Joint Photographic Professional (JPEG), or Video Professional (MPEG). Alternatively, the image / video encoder 124 and image / video decoder 144 may be customized devices that do not conform to existing standards. Although not shown in Figure 1, in some embodiments, the image / video encoder 124 and image / video decoder 144 may be integrated with an audio encoder and decoder, respectively, to handle the encoding of both audio and video within a common data stream or separate data streams, or may include an appropriate MUX-DEMUX unit or other hardware and software.
[0032] The output interface 126 may include any type of medium or device capable of transmitting the encoded bitstream 162 from the source device 120 to the destination device 140. An example of the output interface 126 may be a transmitter or transceiver configured to transmit the encoded bitstream 162 directly from the source device 120 to the destination device 140 in real time. The encoded bitstream 162 may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to the destination device 140.
[0033] The communication medium 160 may include transient media such as wireless broadcast or wired network transmission. Examples of the communication medium 160 may include a radio frequency (RF) spectrum or one or more physical transmission lines (e.g., cables). The communication medium 160 may form part of a packet-based network such as a local area network, a wide area network, or a global network (such as the Internet). In some embodiments, the communication medium 160 may include routers, switches, base stations, or any other equipment useful for facilitating communication from the source device 120 to the destination device 140. For example, a network server (not shown) may receive the encoded bitstream 162 from the source device 120 and provide the encoded bitstream 162 to the destination device 140, for example, by network transmission.
[0034] The communication medium 160 may be in the form of a storage medium (e.g., a non-temporary storage medium) such as a hard disk, flash drive, compact disc, digital video disc, Blu-ray disc, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded image data. In some embodiments, a computing device of a media manufacturing facility, such as a disc stamping machine, may receive encoded image data from the source device 120 and produce a disc containing the encoded video data.
[0035] The input interface 142 may include any type of medium or device capable of receiving information from the communication medium 160. The received information includes the encoded bitstream 162. An example of the input interface 142 may be a receiver or transceiver configured to receive the encoded bitstream 162 in real time.
[0036] The machine vision application 146 includes various hardware and / or software for utilizing the decoded image data generated by the image / video decoder 144. Examples of the machine vision application 146 include a display device for displaying the decoded image data to a user, and one of various display devices such as a cathode ray tube (CRT), liquid crystal display (LCD), plasma display, organic light-emitting diode (OLED) display, or any other type of display device. As another example, the machine vision application 146 may include one or more processors configured to perform various machine vision applications using the decoded image data, such as object recognition / tracking, face recognition, image matching, image / video search, augmented reality, robot vision / navigation, autonomous driving, 3D structure construction, stereo-enabled, and motion tracking.
[0037] Next, with reference to Figures 2A-2B and 3A-3B, exemplary image data encoding and decoding techniques (such as those implemented by encoder 124 and decoder 144 in Figure 1) will be described.
[0038] Figure 2A shows a schematic diagram of an exemplary encoding process 200A according to an embodiment of the present disclosure. For example, the encoding process 200A may be performed by an encoder, such as the image / video encoder 124 in Figure 1. As shown in Figure 2A, the encoder may encode a video sequence 202 into a video bitstream 228 according to process 200A. The video sequence 202 may include a set of images (called βoriginal picturesβ) arranged in chronological order. Each original picture in the video sequence 202 may be divided by the encoder into a basic processing unit, a basic processing subunit, or a region for processing. In some embodiments, the encoder may perform process 200A at the level of a basic processing unit for each original picture in the video sequence 202. For example, the encoder may perform process 200A in an iterative manner, in which case the encoder may encode a basic processing unit in one iteration of process 200A. In some embodiments, the encoder may perform process 200A in parallel for the region of each original picture in the video sequence 202.
[0039] As shown in Figure 2A, the encoder can supply the basic processing unit (called the "original BPU") of the original picture of the video sequence 202 to the prediction stage 204 to generate prediction data 206 and prediction BPU 208. The encoder can subtract the prediction BPU 208 from the original BPU to generate residual BPU 210. The encoder can supply the residual BPU 210 to the conversion stage 212 and the quantization stage 214 to generate quantization conversion coefficients 216. The encoder can supply the prediction data 206 and quantization conversion coefficients 216 to the binary coding stage 226 to generate video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be called the "forward path". During process 200A, after the quantization stage 214, the encoder can supply the quantization conversion coefficients 216 to the inverse quantization stage 218 and the inverse conversion stage 220 to generate the reconstructed residual BPU 222. The encoder can add the reconstructed residual BPU 222 to the prediction BPU 208 to generate the prediction criterion 224, which will be used in the prediction stage 204 for the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A may be referred to as the βreconstruction pathβ. The reconstruction path may be used to ensure that both the encoder and decoder use the same reference data for prediction.
[0040] The encoder can iteratively perform process 200A to encode each original BPU of the original picture (in the forward path) and generate a predictive criterion 224 for encoding the next original BPU of the original picture (in the reconstruction path). After encoding all original BPUs of the original picture, the encoder can proceed to encode the next picture in the video sequence 202.
[0041] Referring to process 200A, the encoder can receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term βreceiveβ can mean any action by any means for receiving, inputting, acquiring, retrieving, obtaining, reading, accessing, or inputting data.
[0042] In prediction stage 204, in the current iteration, the encoder receives the original BPU and prediction criterion 224 and can perform prediction calculations to generate prediction data 206 and prediction BPU 208. The prediction criterion 224 may be generated from the reconstruction path of the previous iteration in process 200A. The objective of prediction stage 204 is to reduce information redundancy by extracting prediction data 206 that can be used to reconstruct the original BPU as prediction BPU 208 from the prediction data 206 and prediction criterion 224.
[0043] Ideally, the predicted BPU 208 can be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 is generally slightly different from the original BPU. To record such differences, the encoder can generate the predicted BPU 208 and then subtract it from the original BPU to generate the residual BPU 210. For example, the encoder can subtract the pixel values ββ(e.g., grayscale values ββor RGB values) of the predicted BPU 208 from the corresponding pixel values ββof the original BPU. Each pixel of the residual BPU 210 can have a residual value as a result of such subtraction between the original BPU and the corresponding pixels of the predicted BPU 208. Although the predicted data 206 and residual BPU 210 may have fewer bits compared to the original BPU, they can be used to reconstruct the original BPU without significant quality degradation. Thus, the original BPU is compressed.
[0044] To further compress the residual BPU210, in the transformation stage 212, the encoder can reduce the spatial redundancy of the residual BPU210 by decomposing it into a set of two-dimensional "basis patterns," each associated with a "transformation coefficient." The basis patterns can have the same size (e.g., the size of the residual BPU210). Each basis pattern can represent the fluctuating frequency components (e.g., the frequency of the luminance fluctuations) of the residual BPU210. No basis pattern can be reconstructed from any combination of any other basis patterns (e.g., a linear combination). In other words, this decomposition allows the fluctuations of the residual BPU210 to be decomposed into the frequency domain. Such a decomposition is analogous to the discrete Fourier transform of a function, where the basis patterns are analogous to the base functions of the discrete Fourier transform (e.g., trigonometric functions), and the transformation coefficients are analogous to the coefficients associated with the base functions.
[0045] Different transformation algorithms can use different basis patterns. Various transformation algorithms can be used in transformation stage 212, such as discrete cosine transform and discrete sine transform. The transformation in transformation stage 212 is reversible; that is, the encoder can reconstruct the residual BPU210 by performing the inverse operation of the transformation (called the "inverse transform"). For example, to reconstruct the pixels of the residual BPU210, the inverse transform can be used to multiply the values ββof the corresponding pixels in the basis pattern by the coefficients associated with each pixel, and then add the products to generate a weighted sum. In video coding standards, both the encoder and decoder can use the same transformation algorithm (i.e., the same basis pattern). Therefore, the encoder can record only the transformation coefficients that can reconstruct the residual BPU210 without receiving the basis pattern from the encoder. While the transformation coefficients have fewer bits compared to the residual BPU210, they can be used to reconstruct the residual BPU210 without significant quality degradation. Thus, the residual BPU210 is further compressed.
[0046] The encoder can further compress the conversion coefficients in the quantization stage 214. In the conversion process, different basis patterns may represent different fluctuation frequencies (e.g., luminance fluctuation frequencies). Because the human eye is generally more perceptible to low-frequency fluctuations, the encoder can ignore information about high-frequency fluctuations without causing significant quality degradation during decoding. For example, in the quantization stage 214, the encoder can generate quantization conversion coefficients 216 by dividing each conversion coefficient by an integer value (called a "quantization parameter") and rounding the quotient to the nearest integer. After such an operation, some conversion coefficients for high-frequency basis patterns are converted to zero, and conversion coefficients for low-frequency basis patterns are converted to smaller integers. The conversion coefficients are further compressed because the encoder can ignore the zero-value quantization conversion coefficients 216. The quantization process is reversible, and in this quantization process, the quantization conversion coefficients 216 can be reconstructed into conversion coefficients by the inverse operation of quantization (called "inverse quantization").
[0047] Because the encoder ignores the remainder of such division in rounding operations, the quantization stage 214 can be irreversible. Typically, the quantization stage 214 can result in the greatest loss of information in process 200A. The greater the loss of information, the fewer bits are required for the quantization conversion coefficient 216. To obtain different levels of loss of information, the encoder can use different values ββfor the quantization parameter or any other parameter of the quantization process.
[0048] In the binary coding stage 226, the encoder can encode the prediction data 206 and quantization conversion coefficients 216 using binary coding techniques such as entropy coding, variable-length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other reversible or reversible compression algorithm. In some embodiments, in addition to the prediction data 206 and quantization conversion coefficients 216, the encoder can encode other information in the binary coding stage 226, such as the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the conversion type in the conversion stage 212, the parameters of the quantization process (e.g., quantization parameters), and encoder control parameters (e.g., bitrate control parameters). The encoder can generate a video bitstream 228 using the output data from the binary coding stage 226. In some embodiments, the video bitstream 228 may be further packetized for network transmission.
[0049] Referring to the reconstruction path of process 200A, in the inverse quantization stage 218, the encoder can perform inverse quantization on the quantization transformation coefficients 216 to generate reconstruction transformation coefficients. In the inverse transformation stage 220, the encoder can generate reconstruction residual BPU 222 based on the reconstruction transformation coefficients. The encoder can add the reconstruction residual BPU 222 to the prediction BPU 208 to generate a prediction criterion 224 to be used in the next iteration of process 200A.
[0050] It should be noted that the video sequence 202 can be encoded using other variations of process 200A. In some embodiments, the stages of process 200A may be performed in a different order by the encoder. In some embodiments, one or more stages of process 200A may be combined into a single stage. In some embodiments, a single stage of process 200A may be divided into multiple stages. For example, the conversion stage 212 and the quantization stage 214 may be combined into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, one or more stages in Figure 2A may be omitted from process 200A.
[0051] Figure 2B shows a schematic diagram of another exemplary coding process 200B according to an embodiment of the present disclosure. For example, coding process 200B may be performed by an encoder such as the image / video encoder 124 in Figure 1. Process 200B may be modified from process 200A. For example, process 200B may be used by an encoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared to process 200A, the forward path of process 200B further includes a mode determination stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B further includes a loop filter stage 232 and a buffer 234.
[0052] Generally, prediction techniques can be classified into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra-prediction") can predict the current BPU using pixels from one or more already coded neighboring BPUs within the same picture. That is, the prediction criterion 224 in spatial prediction may include neighboring BPUs. Spatial prediction can reduce spatial redundancy inherent to a picture. Temporal prediction (e.g., inter-picture prediction or "inter-prediction") can predict the current BPU using regions from one or more already coded pictures. That is, the prediction criterion 224 in temporal prediction may include coded pictures. Temporal prediction can reduce temporal redundancy inherent to a picture.
[0053] Referring to process 200B, in the forward path, the encoder performs prediction calculations in the spatial prediction stage 2042 and the time prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder can perform intra-prediction. For the original BPU of the encoded picture, the prediction criterion 224 may include one or more neighboring BPUs encoded (in the forward path) and reconstructed (in the reconstruction path) within the same picture. The encoder can generate a prediction BPU 208 by extrapolating neighboring BPUs. Examples of extrapolation techniques include linear extrapolation or interpolation, polynomial extrapolation or interpolation, etc. In some embodiments, the encoder can perform extrapolation at the pixel level, for example, by extrapolating the corresponding pixel value for each pixel of the prediction BPU 208. The adjacent BPU used for extrapolation may be located in various directions relative to the original BPU, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., below left, below right, above left, or above right of the original BPU), or in any direction defined by the video coding standard used. For intra-prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the adjacent BPU used, the size of the adjacent BPU used, the extrapolation parameters, and the orientation of the adjacent BPU used relative to the original BPU.
[0054] As another example, in the time prediction stage 2044, the encoder can perform interpretation. For the original BPU of the current picture, the prediction criterion 224 may include one or more pictures (referred to as "reference pictures") that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference pictures may be encoded and reconstructed for each BPU. For example, the encoder may generate a reconstructed BPU by adding the reconstructed residual BPU 222 to the prediction BPU 208. Once all reconstructed BPUs for the same picture have been generated, the encoder can generate the reconstructed picture as a reference picture. The encoder can perform a "motion estimation" operation to search for a matching region within the range of the reference picture (referred to as a "search window"). The position of the search window in the reference picture may be determined based on the position of the original BPU in the current picture. For example, the search window may be centered in the reference picture at a position having the same coordinates as the original BPU in the current picture and extended outward by a predetermined distance. When the encoder identifies a region similar to the original BPU within the search window (e.g., using a Pell recursive algorithm, a block matching algorithm, etc.), the encoder can determine such a region as a matching region. The matching region may have different dimensions from the original BPU (e.g., smaller than, equal to, larger than, or different in shape from the original BPU). Because the reference picture and the current picture are temporally separated in the timeline, the matching region can be considered to "move" to the position of the original BPU over time. The encoder may record the direction and distance of such movement as a "motion vector". If multiple reference pictures are used, the encoder can search for matching regions and determine the associated motion vector for each reference picture. In some embodiments, the encoder can assign weights to the pixel values ββof the matching regions of separate matching reference pictures.
[0055] Motion estimation can be used to identify various types of motion, such as translation, rotation, and zoom. In the case of interpretation, the prediction data 206 may include, for example, the location of the matching region (e.g., coordinates), the motion vector associated with the matching region, the number of reference pictures, and the weights associated with the reference pictures.
[0056] To generate a predicted BPU 208, the encoder can perform a βmotion compensationβ operation. Motion compensation can be used to reconstruct the predicted BPU 208 based on prediction data 206 (e.g., motion vectors) and prediction criteria 224. For example, the encoder can move the matching region of a reference picture according to the motion vectors. In the matching region, the encoder can now predict the original BPU of the picture. If multiple reference pictures are used, the encoder can move the matching region of the reference pictures according to separate motion vectors and the average pixel values ββof the matching regions. In some embodiments, if the encoder has assigned weights to the pixel values ββof the matching regions of separate matching reference pictures, the encoder can add the weighted sum of the pixel values ββof the moved matching regions.
[0057] In some embodiments, interpretation can be unidirectional or bidirectional. In unidirectional interpretation, one or more reference pictures in the same time direction as the current picture can be used. In unidirectional interpretation, a reference picture prior to the current picture can be used. In bidirectional interpretation, one or more reference pictures in both time directions can be used as references to the current picture.
[0058] Further reference to the forward path of process 200B, after the spatial prediction stage 2042 and the time prediction stage 2044, in the mode determination stage 230, the encoder can select a prediction mode for the current iteration of process 200B (e.g., either intra-prediction or inter-prediction). For example, the encoder can perform a rate-distortion optimization technique that selects a prediction mode that minimizes the value of the cost function, depending on the bitrate of the candidate prediction mode and the distortion of the reconstructed reference picture under the candidate prediction mode. Depending on the selected prediction mode, the encoder can generate the corresponding prediction BPU 208 and prediction data 206.
[0059] In the reconstruction path of process 200B, if intra-prediction mode is selected in the forward path, after generating the prediction criterion 224 (e.g., the current BPU encoded and reconstructed within the current picture), the encoder can directly supply the prediction criterion 224 to the spatial prediction stage 2042 for later use (e.g., extrapolation of the next BPU of the current picture). If inter-prediction mode is selected in the forward path, after generating the prediction criterion 224 (e.g., the current picture with all BPUs encoded and reconstructed), the encoder can supply the prediction criterion 224 to the loop filter stage 232, where the encoder can apply loop filters to the prediction criterion 224 to reduce or eliminate distortions (e.g., blocking artifacts) introduced by inter-prediction. In the loop filter stage 232, the encoder can apply various loop filtering techniques, such as deblocking, sample-adaptive offset, and adaptive loop filtering. Loop-filtered reference pictures may be stored in buffer 234 (or βDecoded Picture Bufferβ) for later use (e.g., as inter-predictive reference pictures for future pictures in video sequence 202). The encoder may store one or more reference pictures used in the time prediction stage 2044 in buffer 234. In some embodiments, the encoder may encode the loop filter parameters (e.g., the strength of the loop filter) along with the quantization transformation coefficients 216, the prediction data 206, and other information in the binary coding stage 226.
[0060] In some embodiments, the input video sequence 202 is processed block by block according to the encoding process 200B. In VVC, the coding tree unit (CTU) is the largest block unit and can be up to 128 Γ 128 chroma samples (and corresponding chroma samples depending on the chroma format). The CTU can be further partitioned into coding units (CUs) using a quadtree, binary tree, or ternary tree. At the leaf nodes of the partition structure, coding information is transmitted, such as the coding mode (intra-mode or inter-mode), motion information if intercoded (reference index, motion vector difference, etc.), and quantization conversion coefficients 216. When intra-prediction (also called spatial prediction) is used, spatially adjacent samples are used to predict the current block. When inter-prediction (also called temporal prediction or motion-compensated prediction) is used, samples from an already coded picture called a reference picture are used to predict the current block. Inter-prediction may use single or bi-prediction. In single prediction, only one motion vector pointing to one reference picture is used to generate the prediction signal for the current block. In dual prediction, two motion vectors, each pointing to its own reference picture, are used to generate a prediction signal for the current block. The motion vectors and reference indices are sent to the decoder to identify where the prediction signal for the current block came from. After intra or interpretation, in the mode determination stage 230, the optimal prediction mode for the current block is selected, for example, based on a rate-distortion optimization method. Based on the optimal prediction mode, a prediction BPU 208 is generated and subtracted from the input video block.
[0061] Referring further to Figure 2B, the residual BPU 210 is sent to the transformation stage 212 and quantization stage 214 to generate quantization transformation coefficients 216. The quantization transformation coefficients 216 are then dequantized in the dequantization stage 218 and inverse transformed in the inverse transformation stage 220 to obtain the reconstructed residual BPU 222. The prediction BPU 208 and the reconstructed residual BPU 222 are added together before loop filtering to form the prediction criterion 224. Loop filtering is used to provide reference samples for intra-prediction. In the loop filtering stage 232, loop filtering such as deblocking, sample adaptive offset (SAO), and adaptive loop filtering (ALF) may be applied to the prediction criterion 224 to form reconstructed blocks stored in buffer 234 and used to provide reference samples for intra-prediction. The coding information generated in the mode determination stage 230 (coding mode (intra or inter predictive), intra predictive mode, motion information, quantization residual coefficients, etc.) is sent to the binary coding stage 226 to further reduce the bitrate before being packed into the output video bitstream 228.
[0062] Figure 3A shows a schematic diagram of an exemplary decoding process 300A according to an embodiment of the present disclosure. For example, the decoding process 300A may be carried out by a decoder such as the image / video decoder 144 in Figure 1. Process 300A may be a decompression process corresponding to the compression process 200A in Figure 2A. In some embodiments, process 300A may be similar to the reconstruction path of process 200A. The decoder (e.g., the image / video decoder 144 in Figure 1) can decode the video bitstream 228 into a video stream 304 according to process 300A. The video stream 304 may be very similar to the video sequence 202. However, due to loss of information in the compression and decompression processes (e.g., the quantization stage 214 in Figures 2A-2B), the video stream 304 is generally not identical to the video sequence 202. Similar to processes 200A and 200B in Figures 2A and 2B, the decoder can perform process 300A at the level of the basic processing unit (BPU) for each pixel encoded in the video bitstream 228. For example, the decoder can perform process 300A in an iterative manner, in which case the decoder can decode the basic processing unit in one iteration of process 300A. In some embodiments, the decoder can perform process 300A in parallel for each region of the picture encoded in the video bitstream 228.
[0063] As shown in Figure 3A, the decoder can supply a portion of the video bitstream 228 associated with the basic processing unit of the encoded picture (called the "encoded BPU") to the binary decoding stage 302. In the binary decoding stage 302, the decoder can decode that portion into prediction data 206 and quantization conversion coefficients 216. The decoder can supply the quantization conversion coefficients 216 to the inverse quantization stage 218 and the inverse conversion stage 220 to generate the reconstructed residual BPU 222. The decoder can supply the prediction data 206 to the prediction stage 204 to generate the prediction BPU 208. The decoder can add the reconstructed residual BPU 222 to the prediction BPU 208 to generate the prediction criterion 224. In some embodiments, the prediction criterion 224 may be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder can supply the prediction criterion 224 to the prediction stage 204 for performing the prediction calculation in the next iteration of process 300A.
[0064] The decoder can iteratively perform process 300A to decode each encoding BPU of the encoded picture and generate a prediction criterion 224 for encoding the next encoding BPU of the encoded picture. After decoding all encoding BPUs of the encoded picture, the decoder can output the picture to the video stream 304 for display and proceed to decode the next encoded picture in the video bitstream 228.
[0065] In the binary decoding stage 302, the decoder can perform the inverse operation of the binary coding technique used by the encoder (e.g., entropy coding, variable-length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or other lossless compression algorithms). In some embodiments, in addition to the predicted data 206 and quantization conversion coefficients 216, the decoder can decode other information in the binary decoding stage 302, such as the prediction mode, parameters of the prediction operation, conversion type, parameters of the quantization process (e.g., quantization parameters), and encoder control parameters (e.g., bitrate control parameters). In some embodiments, if the video bitstream 228 is transmitted in packets over the network, the decoder can depacketize the video bitstream 228 before supplying it to the binary decoding stage 302.
[0066] Figure 3B shows a schematic diagram of another exemplary decoding process 300B according to an embodiment of the present disclosure. For example, decoding process 300B may be carried out by a decoder such as the image / video decoder 144 in Figure 1. Process 300B may be modified from process 300A. For example, process 300B may be used by a decoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared to process 300A, process 300B further divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and further includes a loop filter stage 232 and a buffer 234.
[0067] In process 300B, the prediction data 206 decoded by the decoder from the binary decoding stage 302 for the encoding base processing unit ("current BPU") of the encoded picture being decoded ("current picture") may contain various types of data depending on the prediction mode used by the encoder to encode the current BPU. For example, if intra-prediction is used by the encoder to encode the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra-prediction, parameters of the intra-prediction operation, etc. Parameters of the intra-prediction operation may include, for example, the location (e.g., coordinates) of one or more adjacent BPUs used as reference, the size of the adjacent BPUs, extrapolation parameters, the orientation of the adjacent BPUs relative to the original BPU, etc. As another example, if inter-prediction is used by the encoder to encode the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter-prediction, parameters of the inter-prediction operation, etc. Parameters for the interpretation calculation may include, for example, the number of reference pictures currently associated with the BPU, the weights associated with each reference picture, the locations (e.g., coordinates) of one or more matching regions within separate reference pictures, and one or more motion vectors associated with each matching region.
[0068] Based on the prediction mode indicator, the decoder can decide whether to perform spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. Details of performing such spatial or temporal predictions are shown in Figure 2B and will not be repeated below. After performing such spatial or temporal predictions, the decoder can generate a prediction BPU 208. The decoder can then generate a prediction criterion 224 by adding the prediction BPU 208 and the reconstructed residual BPU 222, as shown in Figure 3A.
[0069] In process 300B, the decoder can supply the prediction criterion 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing prediction calculations in the next iteration of process 300B. For example, if the current BPU is decoded using intra-prediction in the spatial prediction stage 2042, after generating the prediction criterion 224 (e.g., decoded current BPU), the decoder can supply the prediction criterion 224 directly to the spatial prediction stage 2042 for later use (e.g., extrapolation of the next BPU of the current picture). If the current BPU is decoded using inter-prediction in the temporal prediction stage 2044, after generating the prediction criterion 224 (e.g., reference picture with all BPUs decoded), the encoder can supply the prediction criterion 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder can apply a loop filter to the prediction criterion 224 as shown in Figure 2B. Loop-filtered reference pictures may be stored in a buffer 234 (e.g., a "decoded picture buffer") in preparation for later use (e.g., use as an inter-prediction reference picture for future encoded pictures of the video bitstream 228). The decoder may store one or more reference pictures used in the time prediction stage 2044 in buffer 234. In some embodiments, if the prediction mode indicator of the prediction data 206 indicates that inter-prediction was used to encode the BPU now, the prediction data may further include loop filter parameters (e.g., loop filter strength).
[0070] Referring back to Figure 1, the image / video preprocessor 122, image / video encoder 124, and image / video decoder 144 may each be implemented as any suitable hardware, software, or combination thereof. Figure 4 is a block diagram of an exemplary apparatus 400 for processing image data according to an embodiment of the present disclosure. For example, apparatus 400 may be a preprocessor, an encoder, or a decoder. As shown in Figure 4, apparatus 400 may include a processor 402. When the processor 402 executes instructions described herein, apparatus 400 may become a dedicated machine for preprocessing, encoding, and / or decoding image data. The processor 402 may be any type of circuit capable of manipulating or processing information. For example, the processor 402 may include any combination of any number of central processing units (or "CPUs"), graphics processing units (or "GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), generic array logic (GALs), composite programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), systems on a chip (SoCs), application-specific integrated circuits (ASICs), and so on. In some embodiments, the processor 402 may also be a set of processors grouped as a single logical component. For example, as shown in Figure 4, the processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.
[0071] The device 400 may also include a memory 404 configured to store data (e.g., a collection of instructions, computer code, intermediate data, etc.). For example, as shown in Figure 4, the stored data may include program instructions (e.g., program instructions for carrying out each stage in processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 can access the program instructions and data for processing (e.g., via the bus 410) and execute the program instructions to perform arithmetic or operations on the data for processing. The memory 404 may include a high-speed random-access storage device or a non-volatile storage device. In some embodiments, the memory 404 may include any number of random-access memories (RAM), read-only memories (ROM), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, security digital (SD) cards, memory sticks, compact flash (CF) cards, and any combination thereof. Memory 404 may also be a memory group (not shown in Figure 4) that is grouped as a single logical component.
[0072] Bus 410 may be a communication device that transfers data between components within the device 400, such as an internal bus (e.g., a CPU-memory bus) or an external bus (e.g., a universal serial bus port, a peripheral component interconnection express port).
[0073] To facilitate explanation without creating ambiguity, the processor 402 and other data processing circuits are collectively referred to as the βdata processing circuitsβ in this disclosure. The data processing circuits may be implemented as hardware as a whole, or as a combination of software, hardware, or firmware. The data processing circuits may also be a single, independent module, or may be incorporated whole or in part into any other component of the device 400.
[0074] The device 400 may further include a network interface 406 for providing wired or wireless communication to a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, etc.). In some embodiments, the network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near-field communication (NFC) adapters, cellular network chips, and the like.
[0075] In some embodiments, the apparatus 400 may further include a peripheral interface 408 for providing connectivity to one or more peripheral devices. As shown in Figure 4, peripheral devices may include, but are not limited to, cursor control devices (e.g., mouse, touchpad, or touchscreen), keyboards, displays (e.g., cathode ray tube displays, liquid crystal displays, or light-emitting diode displays), video input devices (e.g., cameras or input interfaces coupled to a video archive), and the like.
[0076] It should be noted that the video codec (for example, the codec that performs processes 200A, 200B, 300A, or 300B) may be implemented as any combination of any software modules or hardware modules within the device 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more software modules of the device 400, such as program instructions that can be loaded into memory 404. As another example, some or all stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more hardware modules of the device 400, such as dedicated data processing circuits (e.g., FPGA, ASIC, NPU, etc.).
[0077] The video bitstream used in VVC or HEVC is a bit sequence in the form of network abstraction layer (NAL) units or byte streams, forming one or more coded video sequences (CVS), each CVS consisting of one or more coded layer video sequences (CLVS). Between these layers, inter-layer prediction may be applied to achieve high compression performance. Here, a layer is a set of video coding layer (VCL) NAL units, all of which have a specific NAL layer ID value, and associated non-VCL NAL units. A VCL NAL unit is a collective term for coded slice NAL units and a subset of NAL units that have reserved values ββof the NAL unit type classified as a VCL NAL unit. Inter-layer prediction may be applied between different layers.
[0078] Supplemental Enhancement Information (SEI) messages are intended to be transmitted within the coded video bitstream in the manner specified by the video coding specification, or by other means defined by the specifications of the system using such coded video bitstream. SEI messages may contain various types of data indicating the timing of video pictures or describing various characteristics of the coded video and how they may be used or enhanced. SEI messages are also defined to include arbitrary user-defined data. SEI messages do not affect the core decoding process, but they can indicate how the video should be post-processed or displayed.
[0079] To specify SEI messages, the JVET Working Group also developed the H.274 standard, which defines the syntax and semantics of Video Usability Information (VUI) parameters and Additional Extension Information (SEI) messages, particularly intended for use in coded video bitstreams as specified in the VVC standard. However, since neither VUI parameters nor SEI messages affect the decoding process, H.274 SEI messages can also be used in other types of coded video bitstreams such as H.265 / HEVC and H.264 / AVC.
[0080] For object detection and tracking purposes, the latest versions of the HEVC and VSEI standards employ Annotated Area (AR)SEI messages that carry parameters describing the bounding boxes of objects detected or tracked within the compressed video bitstream. Therefore, if the encoder, transcoder, or network node has already performed video analysis to recognize objects, the decoder-side device does not need to perform that analysis. This is beneficial for applications where the decoder device has limited computing resources or power. Conversely, if the encoder performs object detection and tracking and sends the information to the decoder, it can potentially improve the quality of detection and tracking, as the encoder can perform detection and tracking tasks using the original video, which is of much higher quality than the reconstructed video recovered by the decoder.
[0081] In HEVC AR SEI messages, in addition to the bounding box of a detected or tracked object, an object label and confidence level associated with the object may also be provided. The object label provides information about the type of object, and the confidence level indicates the fidelity of the detected or tracked object within its bounding box. Furthermore, a flag is provided indicating whether the bounding box currently in the SEI message represents the location of an object that may be occluded or partially occluded by other objects, or only the location of the visible portion of the object. A flag indicating whether the object represented by the current bounding box is only partially visible may also be optionally signaled for each bounding box.
[0082] The AR SEI message syntax uses parameter persistence to avoid the need to re-signal information already available in previous SEI messages within the same duration. For example, if the first detected object is currently stationary within a picture compared to a previous coded picture, and the second detected object moves from one picture to another, then only the bounding box information of the second object needs to be signaled, and the position / bounding box information of the first object can be copied from the previous SEI message.
[0083] Major video coding standards such as H.264 / AVC, H.265 / HEVC, and H.266 / VVC all support the coding of a special type of picture called an auxiliary picture, which provides supplementary information to a regular picture called a primary picture. The auxiliary picture has no normative effect on the decoding process of the primary picture. The bitstreams of the auxiliary picture and the primary picture are packed into a single coded video sequence (CVS). The information necessary to interpret the auxiliary picture is transmitted via SEI messages.
[0084] In HEVC, auxiliary pictures are coded as one or more auxiliary picture layers separate from the primary picture layer. Instructions for auxiliary picture coding are signaled by the Video Parameter Set (VPS) extension, as shown in Table 1 below. [Table 1]
[0085] If `splitting_flag` is equal to 1, it indicates that the `dimension_id[i][j]` syntax element does not exist, and the binary representation of the `nuh_layer_id` value in the NAL unit header is split into NumScalabilityTypes segments of bit length depending on the value of `dimension_id_len_minus1[j]`, and the value of `dimension_id[LayerIdxInVps[nuh_layer_id]][j]` is inferred from the NumScalabilityTypes segments. If `splitting_flag` is equal to 0, it indicates that the syntax element `dimension_id[i][j]` exists.
[0086] If splitting_flag is equal to 1, the scalability identifiers for existing scalability dimensions can be derived from a bitmasked copy of the nuh_layer_id syntax element in the NAL unit header. A separate bitmask for the i-th existing scalability dimension is defined by the dimension_id_len_minus1[i] syntax element and the value of dimBitOffset[i], as specified by the semantics of dimension_id_len_minus1[j].
[0087] If scalability_mask_flag[i] is equal to 1, it indicates that a dimension_id syntax element exists corresponding to the i-th scalability dimension in Table 2 below. If scalability_mask_flag[i] is equal to 0, it indicates that a dimension_id syntax element does not exist corresponding to the i-th scalability dimension. [Table 2]
[0088] dimension_id_len_minus1[j]+1 specifies the bitwise length of the dimension_id[i][j] syntax element.
[0089] If splitting_flag is equal to 1, the following applies: -If the variable dimBitOffset[0] is set to equal to 0, and j is in the range of 1 or greater and NumScalabilityTypes?1 or less, then dimBitOffset[j] is derived as follows:
number
[0090] The value of -dimension_id_len_minus1[NumScalabilityTypes?1] is presumed to be equal to 5?dimBitOffset[NumScalabilityTypes?1]. The value of -dimBitOffset[NumScalabilityTypes] is set to equal 6. The bitstream compliance requirement is that if NumScalabilityTypes is greater than 0, then dimBitOffset[NumScalabilityTypes?1] is less than 6.
[0091] If vps_nuh_layer_id_present_flag is equal to 1, it specifies that there exists a layer_id_in_nuh[i] where i is between 1 and MaxLayersMinus1 (inclusive). If vps_nuh_layer_id_present_flag is equal to 0, it specifies that there does not exist a layer_id_in_nuh[i] where i is between 1 and MaxLayersMinus1 (inclusive).
[0092] `layer_id_in_nuh[i]` specifies the value of the `nuh_layer_id` syntax element in the VCL NAL unit of the i-th layer. If i is greater than 0, `layer_id_in_nuh[i]` is greater than `layer_id_in_nuh[i-1]`. For any value of i between 0 and MaxLayersMinus1, if `layer_id_in_nuh[i]` does not exist, its value is presumed to be equal to i.
[0093] If i is within the range of 0 or greater and MaxLayersMinus1 or less, the variable LayerIdxInVps[layer_id_in_nuh[i]] will be set to equal to i.
[0094] `dimension_id[i][j]` specifies the identifier of the j-th scalability dimension type in the i-th layer. The number of bits used to represent `dimension_id[i][j]` is `dimension_id_len_minus1[j]+1` bits.
[0095] Depending on the splitting_flag, the following applies: -If splitting_flag is equal to 1, and i is greater than or equal to 0 and MaxLayersMinus1, and j is greater than or equal to 0 and NumScalabilityTypes?1, then dimension_id[i][j] is ((layer_id_in_nuh[i]&((1<<dimBitOffset[j+1])?1))> It is presumed to be equal to >dimBitOffset[j]). -If not (splitting_flag is equal to 0), then if j is greater than or equal to 0 and less than or equal to NumScalabilityTypes?1, then dimension_id[0][j] is presumed to be equal to 0.
[0096] The variable ScalabilityId[i][smIdx], which specifies the identifier of the (smIdx) scalability dimension type for the i-th layer, and the variables DepthLayerFlag[lId], ViewOrderIdx[lId], DependencyId[lId], and AuxId[lId], which specify the depth flag, view order index, spatial / quality scalability identifier, and auxiliary identifier for the layer where nuh_layer_id is equal to lId, are derived as follows:
number
[0097] If AuxId[lId] is equal to 0, it specifies that no auxiliary pictures are included in the layer where nuh_layer_id is equal to lId. If AuxId[lId] is greater than 0, it specifies the type of auxiliary picture in the layer where nuh_layer_id is equal to lId, as specified in Table 3 below. [Table 3]
[0098] The interpretation of auxiliary pictures associated with AuxIds between 128 and 159 is specified by means other than the AuxId value.
[0099] If the bitstream conforms to this version of this specification, AuxId[lId] is in the range of 0 to 2 or 128 to 159. However, in this version of this specification, the decoder may set the value of AuxId[lId] to the range of 0 to 255.
[0100] If AuxId[lId] is equal to AUX_ALPHA or AUX_DEPTH, one of the following bitstream conformance requirements applies: -For layers where nuh_layer_id is equal to lId, chroma_format_idc is equal to 0 in the active SPS. -For all pictures where nuh_layer_id is equal to lId and this VPS RBSP is the active VPS RBSP, the value of all decoded chroma samples is equal to 1 << (BitDepthC ? 1).
[0101] SEI messages may describe the interpretation of auxiliary pictures, including possible associations with one or more primary pictures.
[0102] Unless constrained by the semantics of an SEI message specifying the interpretation of auxiliary pictures, two layers have nuh_layer_id values ββlayerIdA and layerIdB such that AuxId[layerIdA] is equal to AuxId[layerIdB] and both are greater than 0, and for each value of i between 0 and 15, all values ββof ScalabilityId[LayerIdxInVps[layerIdB]][i] are allowed to be equal to ScalabilityId[LayerIdxInVps[layerIdB]][i]. An SEI message specifying the interpretation of auxiliary pictures may specify that both a picture with nuh_layer_id equal to layerIdA and a picture with nuh_layer_id equal to layerIdB may be associated with the same primary picture within the same access unit.
[0103] In VVC, auxiliary pictures are coded as one or more auxiliary picture layers distinct from the primary picture layer. Instructions for the auxiliary picture are signaled by scalability dimension information SEI messages, as shown in Table 4 below. [Table 4]
[0104] Scalability Dimensional Information (SDI) SEI messages provide SDI information to each layer in the current CVS (i.e., the CVS contains SD ISEI messages including, for example, 1) the view ID of each layer if multiple views exist, and 2) the auxiliary ID of each layer if auxiliary information (such as depth or alpha) is carried by one or more layers).
[0105] If an SDI SEI message exists in any AU of the CVS, then an SDI SEI message also exists in the first AU of the CVS. All SDI SEI messages within the CVS are identical in content.
[0106] sdi_max_layers_minus1+1 indicates the current maximum number of layers in CVS.
[0107] If sdi_multiview_info_flag is equal to 1, it indicates that CVS may currently contain multiple views and that the sdi_view_id_val[] syntax element exists in the SDI SEI message. If sdi_multiview_info_flag is equal to 0, it indicates that CVS does not currently contain multiple views and that the sdi_view_id_val[] syntax element does not exist in the SDI SEI message.
[0108] If sdi_auxiliary_info_flag is equal to 1, it indicates that one or more layers in CVS are currently auxiliary layers and may carry auxiliary information, and that the sdi_aux_id[] syntax element exists in the SDI SEI message. If sdi_auxiliary_info_flag is equal to 0, it indicates that CVS currently does not contain any auxiliary layers and that the sdi_aux_id[] syntax element does not exist in the SDI SEI message.
[0109] sdi_view_id_len_minus1+1 specifies the bitwise length of the sdi_view_id_val[i] syntax element.
[0110] sdi_layer_id[i] specifies the layer identifier of the i-th layer that may currently exist in CVS.
[0111] sdi_view_id_val[i] specifies the view identifier of the i-th layer in CVS. The length of the sdi_view_id_val[i] syntax element is sdi_view_id_len_minus1+1 bits.
[0112] The variable NumViews, which specifies the number of views currently in CVS, and the list ViewId, which specifies the view identifiers of the views currently in CVS, are derived as follows:
number
[0113] If sdi_aux_id[i] is equal to 0, it indicates that the i-th layer in the CVS currently does not contain an auxiliary picture. If sdi_aux_id[i] is greater than 0, it indicates the type of auxiliary picture currently in the i-th layer in the CVS, as specified in Table 5 below. If sdi_auxiliary_info_flag is equal to 0, it is inferred that the value of sdi_aux_id[i] is equal to 0. [Table 5]
[0114] The interpretation of auxiliary pictures associated with sdi_aux_id[i] in the range of 128 to 159 is specified by means other than the sdi_aux_id[i] value.
[0115] If the bitstream conforms to this version of this specification, sdi_aux_id[i] is in the range of 0 to 2 or 128 to 159. However, in this version of this specification, the decoder may also make other values ββof sdi_aux_id[i] in the range of 0 to 2 or 128 to 159.
[0116] If sdi_aux_id[i] is equal to 0, the i-th layer is called the primary layer. Otherwise, the i-th layer is called the auxiliary layer. If sdi_aux_id[i] is equal to 1, the i-th layer is also called the alpha auxiliary layer. If sdi_aux_id[i] is equal to 2, the i-th layer is also called the depth auxiliary layer.
[0117] sdi_num_associated_primary_layers_minus1[i]+1 specifies the number of primary layers associated with the i-th auxiliary layer. The value of sdi_num_associated_primary_layers_minus1[i] is less than the total number of primary layers.
[0118] sdi_associated_primary_layer_idx[i][j] specifies the layer index of the j-th primary layer associated with the i-th auxiliary layer. The value of sdi_aux_id[sdi_associated_primary_layer_idx[i][j]] is equal to 0.
[0119] The auxiliary layer describes the properties of its associated primary layer and applies them to that associated primary layer.
[0120] Currently, AR SEI messages can be used to annotate and track objects in video. However, their functionality has limitations. For example, AR SEI messages currently cannot adequately support the following two aspects:
[0121] Currently, in AR SEI messages, detected or tracked objects are represented by bounding boxes. While object position information can be described by bounding boxes, object shape information cannot. Applications that use segmentation to facilitate features such as virtual backgrounds need to describe object shape information more accurately. Also, object segmentation consumes power, placing a significant burden on mobile devices. When object segmentation is performed, it may be desirable to carry such information as side information within the video bitstream. The current AR SEI message syntax, as shown in Table 1, cannot carry such information.
[0122] Furthermore, the AR SEI message currently signals a flag indicating whether an object represented by a bounding box is partially or fully visible. However, when an object is partially visible, there is no parameter to tell the decoder which parts are visible and which are occluded. Therefore, the flag itself does not provide the decoder with much information for determining the visible and invisible regions of an object. Instead, depth information of an object may provide a better mechanism for describing the relative positions of various objects in a picture in terms of their distance to the camera. Such information can be used directly to derive which parts of which objects are occluded or not.
[0123] To solve the above problem, instead of signaling the bounding box of the annotated region, a mask representing the shape and position of the annotated or tracked object is signaled. The mask can be implemented as a binary matrix the same size as the picture, where elements with a value of 0 represent the positions covered by the background and elements with a value of 1 represent the positions covered by the object. Thus, any shape of an object can be represented by the mask. To distinguish between different objects, a multi-valued mask can be used, where elements with a value of 0 represent the background and elements with a value of k (where k is not equal to 0) represent the k-th object.
[0124] While masks can represent the precise shape of an object, the signaling overhead is also much greater than that of sending a bounding box. Therefore, this disclosure proposes encoding the mask as an auxiliary picture instead of signaling it within the SEI message, so that low-level video coding techniques supported by video coding standards can be used to compress the mask picture.
[0125] In this disclosure, at some point in time, there are one or more regular pictures called primary pictures, and each primary picture is associated with one or more object mask auxiliary pictures. In H.264 / AVS, auxiliary pictures are indicated by a special NAL unit type. In H.265 / HEVC and H.266 / VVC, auxiliary pictures are coded as an auxiliary picture layer, which is a separate layer from the primary picture layer. Therefore, there may be multiple primary picture layers and multiple auxiliary picture layers.
[0126] Interpreting an auxiliary picture requires some side information. This disclosure proposes signaling side information regarding the mask auxiliary picture within the SEI message.
[0127] Figure 5 is a syntax chart of a proposed Object Mask Information (OMI) SEI message according to some disclosed embodiments. This chart shows the syntax structure and order of syntax elements of an Object Mask Information SEI message. First, a cancellation flag is signaled indicating whether this OMI SEI is used to cancel the duration of a previous SEI message (e.g., the last OMI message). If the cancellation flag indicates that it does not cancel the duration of a previous OMI SEI message, then information about the object mask is signaled to update the object information signaled in the previous OMI SEI message. Among this information, object mask auxiliary (picture) identifier information used to distinguish object mask auxiliary pictures from other auxiliary pictures is signaled first. Next, the number of object mask pictures (e.g., auxiliary picture layers) is signaled. Subsequently, existence flags (e.g., confidence existence flag, depth existence flag, label existence flag) and syntax elements (if any) for confidence, depth, and identifier length (e.g., confidence length, depth length, label length) are signaled. The syntax elements signaled above are called common information for the object mask indicated in this OMI SEI message, while individual mask information is signaled later. Finally, for each mask within each object mask picture, the mask identifier, followed by the mask confidence, object depth, and mask label (if any) are signaled.
[0128] Figure 6 is a schematic diagram showing an exemplary method 600 for encoding a video sequence into a bitstream according to an embodiment of the present disclosure. As shown in Figure 6, the method 600 includes steps 602 and 604, which may be implemented by an encoder (e.g., the image / video encoder 124 in Figure 1, or the device 400 in Figure 4).
[0129] In step 602, the encoder can receive the video sequence.
[0130] In step 604, the encoder may encode one or more pictures of the video sequence to generate a bitstream. Specifically, the encoder may encode in the bitstream an auxiliary picture to indicate the mask of an object in the primary picture. The object mask may be represented by sample values ββin the auxiliary picture. As understood, an object in the primary picture may be drawn by pixels masked with sample values. The encoder may also generate Additional Enhancement Information (SEI) messages associated with the primary picture. SEI messages may also be applied to auxiliary pictures and may be used to indicate the attributes of the object mask. In this disclosure, SEI messages used to indicate the attributes of the object mask are also referred to as Object Mask Information (OMI) SEI messages.
[0131] Figure 7 is a schematic diagram showing exemplary primary pictures 701, 703 and auxiliary pictures 702, 704 according to embodiments of the present disclosure. Auxiliary picture 702 corresponds to primary picture 701, while auxiliary picture 704 corresponds to primary 703. As shown in Figure 7, primary picture 701 may include a background 711, a person object 712, and an animal object 713 within the picture. Auxiliary picture 702 corresponding to primary picture 701 may include a person mask 722 corresponding to the person object 712 and an animal mask 723 corresponding to the animal object 713. Masks in auxiliary picture 702 may be used to represent the position and contour of separate objects in primary picture 701. Specifically, auxiliary picture 702 may be the same size as primary picture 701. The person mask 722 represents its own position and contour in primary picture 701 by the position and contour of the person object 712 in auxiliary picture 702. As is understood, a mask in auxiliary picture 702 can be represented by separate sample values ββ(also called pixel values). Similarly, a mask in auxiliary picture 704 can be used to represent the position and contours of separate objects in primary picture 703.
[0132] The OME SEI message (not shown) is applied to auxiliary picture 702 or 704 and may be used to indicate the attributes of at least one mask.
[0133] In some embodiments, each OMI SEI message contains information about all masks. A persistence mechanism is used for OMI SEI messages; therefore, for a primary picture, if the mask does not change at all from one point in time to the next, there is no need to signal an OMI SEI. If any information changes, a new OMI SEI containing the new information about the mask needs to be signaled. Its syntax is shown in Table 6, and its semantics are provided below in Table 6 as an example. [Table 6]
[0134] Object Mask Information (OMI)SEI messages provide information about object mask pictures coded as auxiliary pictures. Object mask auxiliary pictures have a nuh_layer_id equal to nuhLayerIdA, and an AuxId[nuhLayerIdA] in the range of 128 to 159. Each overlay auxiliary picture layer is associated with one or more primary picture layers as specified below.
[0135] In some embodiments, the encoder may, in step 604, determine a cancellation flag to indicate whether the SEI message cancels the persistence of a previous SEI message. For example, if omi_cancel_flag is equal to 1, it indicates that the SEI message cancels the persistence of a previous object mask information SEI message in output order associated with one or more primary picture layers to which this SEI applies. If omi_cancel_flag is equal to 0, it indicates that object mask information follows and that the object mask information signaled within this SEI message is used to update the object mask information in which any previous SEI message exists.
[0136] If a cancellation flag (e.g., omi_cancel_flag) indicates that an SEI message cancels the persistence of any SEI message, only the cancellation flag is included in the SEI message and notified. If it is decided to reuse the mask information, a complete SEI message with all the necessary syntax must be generated and signaled to the decoder.
[0137] An SEI message may indicate attributes of multiple masks. In some embodiments, the mask attributes conveyed by an SEI message may include common features of each mask indicated by the SEI message and individual features of the object's mask. As is understood, common features are shared by these masks indicated by the SEI message, while individual features are specified for the target mask.
[0138] If the encoder indicates that the cancellation flag (e.g., omi_cancel_flag) does not cancel the persistence of information from previous SEI messages, it may further determine common and individual features in step 604.
[0139] As described above with reference to Figure 5, common features may include object mask auxiliary (picture) identifier information, the number of object mask pictures, the presence of flags, and the confidence level, depth, and identifier length (e.g., confidence length, depth length, or label length) (if any).
[0140] The identifier of the auxiliary picture to which the SEI message applies can be determined by one of its common features. For example, omi_aux_id_minus128+128 represents the AuxId value of the object mask auxiliary picture. omi_aux_id_minus128 is in the range of 0 to 31.
[0141] The number of bits used to encode the identifier of any of multiple masks can be determined as one of the common features. For example, omi_num_mask_pic_minus1+1 indicates the number of object mask auxiliary pictures associated with the same one or more primary pictures. The value of omi_num_mask_pic_minus1 is in the range of 0 to 63. The value of omi_num_mask_pic_minus1 is the same in all OMI SEI messages within CVS.
[0142] In some embodiments, the SEI message is applied to multiple auxiliary pictures. The number of auxiliary pictures can be determined as one of the common features. For example, omi_mask_id_length_minus8+8 indicates the number of bits used to encode the omi_mask_id[i][j] syntax element.
[0143] The confidence presence flag, used to indicate whether confidence information for multiple masks is included in an SEI message, can be determined as one of the common features. For example, if omi_mask_confidence_info_present_flag is equal to 1, it indicates that the omi_mask_confidence[i][j] syntax element is present. If omi_mask_confidence_info_present_flag is equal to 0, it indicates that the omi_mask_confidence[i][j] syntax element is not present. As a bitstream conformance requirement, the value of omi_mask_confidence_info_present_flag is the same for all object_mask_info() syntax structures within CLVS.
[0144] In some embodiments, if a confidence presence flag (e.g., omi_mask_confidence_info_present_flag) indicates that the SEI message contains confidence information for multiple masks, the length of the confidence information for the multiple masks may also be determined as one of the common features. For example, omi_mask_confidence_length_minus1+1 specifies the bitwise length of the omi_mask_confidence[i][j] syntax element. As a bitstream conformance requirement, the value of omi_mask_confidence_length_minus1 is the same for all object_mask_info() syntax structures within the CLVS.
[0145] The depth presence flag, used to indicate whether depth information for multiple masks is included in an SEI message, can be determined as one of the common features. For example, if omi_object_depth_info_present_flag is equal to 1, it indicates that the omi_object_depth[i][j] syntax element exists. If omi_object_depth_info_present_flag is equal to 0, it indicates that the omi_object_depth[i][j] syntax element does not exist. As a bitstream conformance requirement, the value of omi_object_depth_info_present_flag is the same for all object_mask_info() syntax structures within CLVS.
[0146] In some embodiments, if a depth presence flag (e.g., omi_object_depth_info_present_flag) indicates that the SEI message contains depth information for multiple masks, the length of the depth information for the multiple masks may be determined as one of the common features. For example, omi_object_depth_length_minus1+1 specifies the bitwise length of the omi_mask_confidence[i][j] syntax element. As a bitstream conformance requirement, the value of omi_object_depth_length_minus1 is the same for all object_mask_info() syntax structures within the CLVS.
[0147] The label presence flag, used to indicate whether label language presence information and label information for multiple masks are included in the SEI message, can be determined as one of the common features. For example, if omi_mask_label_info_present_flag is equal to 1, it indicates that omi_mask_label_language_present_flag and omi_mask_label[i][j] exist. If omi_mask_label_info_present_flag is equal to 0, it indicates that omi_mask_label_language_present_flag and omi_mask_label[i][j] do not exist.
[0148] In some embodiments, if a label presence flag (e.g., omi_mask_label_info_present_flag) indicates that label language presence information and label information for multiple masks are included in the SEI message, the language presence flag used to indicate whether or not label language information for multiple masks is included in the SEI message may be determined as one of the common features. For example, if omi_mask_label_language_present_flag is equal to 1, it indicates that omi_mask_label_language exists. If omi_mask_label_language_present_flag is equal to 0, it indicates that omi_mask_label_language does not exist and that the language of the mask label is not specified.
[0149] In some embodiments, omi_bit_equal_to_zero is equal to 0.
[0150] In some embodiments, when the language presence flag indicates that the label language information of multiple masks is included in the SEI message, the label language information of the multiple masks can be determined as one of the common features. For example, omi_mask_label_language includes a language tag as defined in IET FRFC 5646, followed by a null-terminating byte equal to 0x00. The length of the omi_mask_label_language syntax element is 255 bytes or less excluding the null-terminating byte. If not present, the language of the label is not specified.
[0151] As described above with reference to FIG. 5, the individual features can be the mask identifier, mask confidence, object depth, and mask label (if present) of each mask within each object mask picture. As described above, when the SEI message is applied to multiple auxiliary pictures, the number of multiple auxiliary pictures can be determined as one of the common features. Further, the SEI message may include individual features generated for the masks represented by the multiple auxiliary pictures.
[0152] In some embodiments, the individual feature omi_mask_pic_layer_id[i] indicates the nuh_layer_id value of the i-th auxiliary picture layer. AuxId[omi_mask_pic_layer_id[i]] is equal to omi_aux_id_minus128 + 128 for all values within the range of 0 to omi_num_mask_pic_minus1.
[0153] In some embodiments, the individual feature omi_num_mask_in_pic[i] indicates the number of masks in the i-th auxiliary picture. omi_num_mask_in_pic[i] is within the range of 0 to (1 << BitDepthY)-1, where BitDepthY is the bit depth of the samples of the luma component.
[0154] In some embodiments, the individual feature omi_mask_id[i][j] indicates the identifier of the j-th object mask in the i-th object mask auxiliary picture. The object mask identifier associated with the sample position (x,y) in the i-th object mask auxiliary picture is equal to p[i][x][y], where p[i][x][y] points to the luma sample at position (x,y) in the decoded i-th object mask auxiliary picture.
[0155] The variable maskId[i][j], which specifies the object mask identifier of the j-th object mask of the i-th object mask auxiliary picture within the SEI message, is derived as follows.
number
[0156] In some embodiments, the individual feature omi_mask_confidence[i][j] represents the confidence associated with the j-th object mask in the i-th object mask auxiliary picture, 2 -(omi_mask_confidence_length_minus1+1) Expressed in units of omi_mask_confidence[i][j][k], a higher value of omi_mask_confidence[i][j][k] indicates higher confidence. The length of the omi_mask_confidence[i][j][k] syntax element is omi_mask_confidence_length_minus1+1 bits.
[0157] In some embodiments, the individual feature omi_mask_depth[i][j] indicates the object depth associated with the j-th object mask in the i-th object mask auxiliary picture. Smaller values ββof omi_mask_depth indicate shorter distances to the object. The length of the omi_mask_depth[i][j][k] syntax element is omi_object_depth_length_minus1+1 bits.
[0158] In some embodiments, the individual feature omi_mask_label[i][j] specifies the content of the label associated with the j-th object mask in the i-th object mask auxiliary picture. The length of the omi_mask_label[i][j] syntax element is 255 bytes or less, excluding the null-terminating byte.
[0159] The syntax described in Table 6 incurs a high bit cost because, every time the object mask information changes, all mask information, including the unchanged portion, must be re-signaled in the OMI SEI message. In some embodiments, only the changed information is signaled. Figure 8 is a schematic diagram showing the substeps of method 600. As shown in Figure 8, step 604 may include substeps 802 and 804, which can be implemented by the encoder.
[0160] In substep 802, the encoder can determine whether the object's mask is different from the previous mask of the object represented by the previous auxiliary picture, if the cancel flag indicates that the SEI message cancels the persistence of information from the previous SEI message. In some embodiments, the encoder can determine whether the object's mask is different from the previous mask of the object represented by the previous auxiliary picture, regardless of whether the cancel flag indicates that the SEI message cancels the persistence of information from the previous SEI message.
[0161] In substep 804, the encoder may encode the attributes of the object's mask in the SEI message if the object's mask is different from the object's previous mask. In some embodiments, the encoder may skip encoding the attributes of the object's mask in the SEI message if the object's mask is the same as the object's previous mask.
[0162] In some embodiments, an update flag may be introduced for each auxiliary picture. If there is no change in the auxiliary picture of the object mask, the mask information signaling for this auxiliary picture is skipped. If there is a change, the changed information is signaled. For example, if the label, depth, or confidence level of the object mask changes, or if the number of masks changes, only the changed masks need to be signaled. Thus, the signaling overhead is reduced. Referring back to Figure 7, the person mask 742 in auxiliary picture 704 and the person mask 722 in auxiliary picture 702 are the same, but the animal mask 723 changes to animal mask 743. If the previous SEI is signaled to indicate masks 722 and 723, the current SEI, which is signaled to indicate masks 742 and 743, may skip the information that has not changed (e.g., the person mask 742).
[0163] The syntax is shown in Table 7 (differences from Table 6 are shown in italics in Table 7). The semantics are provided below in Table 7 as an example. The definitions of common features, shown as "high-level information," and individual features, shown as "individual mask information" (some of which are omitted), can be inherited from the above embodiments by shared parameter / function names. [Table 7-1] [Table 7-2]
[0164] Object Mask Information (OMI)SEI messages provide information about object mask pictures coded as auxiliary pictures. Object mask auxiliary pictures have a nuh_layer_id equal to nuhLayerIdA, and an AuxId[nuhLayerIdA] in the range of 128 to 159. Each overlay auxiliary picture layer is associated with one or more primary picture layers as specified below.
[0165] Similar to some of the embodiments described above, the encoder may in step 604 determine a cancellation flag to indicate whether the SEI message cancels the persistence of a previous SEI message. For example, if omi_cancel_flag is equal to 1, it indicates that the SEI message cancels the persistence of a previous object mask information SEI message in output order associated with one or more primary picture layers to which this SEI applies. If omi_cancel_flag is equal to 0, it indicates that object mask information follows and that the object mask information signaled in this SEI message is used to update the object mask information in which any previous SEI message exists.
[0166] Similarly, the encoder may also determine other common features for the mask.
[0167] omi_aux_id_minus128+128 indicates the AuxId value of the object mask auxiliary picture. omi_aux_id_minus128 is within the range of 0 to 31 (inclusive).
[0168] omi_num_mask_pic_minus1+1 indicates the number of object mask auxiliary pictures associated with the same one or more primary pictures. The value of omi_num_mask_pic_minus1 is between 0 and 63. The value of omi_num_mask_pic_minus1 is the same in all OMI SEI messages in CVS.
[0169] omi_mask_id_length_minus8+8 indicates the number of bits used to encode the omi_mask_id[i][j] syntax element.
[0170] If omi_mask_confidence_info_present_flag is equal to 1, it indicates that the omi_mask_confidence[i][j] syntax element exists. If omi_mask_confidence_info_present_flag is equal to 0, it indicates that the omi_mask_confidence[i][j] syntax element does not exist. As a bitstream conformance requirement, the value of omi_mask_confidence_info_present_flag must be the same for all object_mask_info() syntax structures within CLVS.
[0171] omi_mask_confidence_length_minus1+1 specifies the bitwise length of the omi_mask_confidence[i][j] syntax element. As a bitstream conformance requirement, the value of omi_mask_confidence_length_minus1 must be the same for all object_mask_info() syntax structures within CLVS.
[0172] If omi_object_depth_info_present_flag is equal to 1, it indicates that the omi_object_depth[i][j] syntax element exists. If omi_object_depth_info_present_flag is equal to 0, it indicates that the omi_object_depth[i][j] syntax element does not exist. As a bitstream conformance requirement, the value of omi_object_depth_info_present_flag must be the same for all object_mask_info() syntax structures within CLVS.
[0173] omi_object_depth_length_minus1+1 specifies the bitwise length of the omi_mask_confidence[i][j] syntax element. As a bitstream conformance requirement, the value of omi_object_depth_length_minus1 must be the same for all object_mask_info() syntax structures within the CLVS.
[0174] If omi_mask_label_info_present_flag is equal to 1, it indicates that omi_mask_label_language_present_flag and omi_mask_label[i][j] exist. If omi_mask_label_info_present_flag is equal to 0, it indicates that omi_mask_label_language_present_flag and omi_mask_label[i][j] do not exist.
[0175] If omi_mask_label_language_present_flag is equal to 1, it indicates that omi_mask_label_language exists. If omi_mask_label_language_present_flag is equal to 0, it indicates that omi_mask_label_language does not exist and the language of the mask label is not specified.
[0176] omi_bit_equal_to_zero is equal to 0.
[0177] The `omi_mask_label_language` element contains a language tag, as defined in IET FRFC 5646, followed by a null-terminating byte equal to 0x00. The length of the `omi_mask_label_language` syntax element is 255 bytes or less, excluding the null-terminating byte. If it is not present, the label language is not specified.
[0178] In some embodiments, the encoder can determine an update flag for indicating whether the mask of an object is signaled within the auxiliary picture. For example, if omi_mask_pic_update_flag[i] is equal to 1, it indicates that the mask information of the i-th object mask auxiliary picture is signaled. If omi_mask_pic_update_flag[i] is equal to 0, it indicates that the mask information of the i-th object mask auxiliary picture is not signaled. If the mask information of the i-th object mask auxiliary picture does not exist, a persistence mechanism is used. That is, information is inherited from the last OMI SEI message that signaled the mask information of the i-th object mask auxiliary picture.
[0179] omi_mask_pic_layer_id[i] indicates the nuh_layer_id value of the i-th auxiliary picture layer. AuxId[omi_mask_pic_layer_id[i]] is equal to omi_aux_id_minus128 + 128 for all values within the range of 0 to omi_num_mask_pic_minus1.
[0180] In some embodiments, omi_num_mask_in_pic_update[i] indicates the number of masks within the i-th auxiliary picture that are signaled. omi_num_mask_in_pic[i] is within the range of 0 to (1 << BitDepthY) - 1, where BitDepthY is the bit depth of the samples of the luma component.
[0181] In some embodiments, omi_mask_id[i][j] indicates the identifier of the j-th object mask updated within the i-th object mask auxiliary picture.
[0182] The object mask identifier associated with the sample position (x,y) in the i-th object mask auxiliary picture is equal to p[i][x][y], where p[i][x][y] points to the Luma sample at position (x,y) in the decoded i-th object mask auxiliary picture.
[0183] The variable maskId[i][j], which specifies the object mask identifier of the j-th object mask of the i-th object mask auxiliary picture within the SEI message, is derived as follows.
number
[0184] In some embodiments, the encoder may determine a mask cancellation flag to indicate whether an object's mask cancels the persistence of an object's previous mask, if the cancellation flag indicates that the SEI message does not cancel the persistence of information from a previous SEI message. In some embodiments, the encoder may determine a mask cancellation flag to indicate whether an object's mask cancels the persistence of an object's previous mask, regardless of whether the cancellation flag indicates that the SEI message cancels the persistence of information from a previous SEI message. For example, if omi_mask_cancel[i][j] is equal to 1, the persistence range of the object mask whose identifier is equal to omi_mask_id[i][j] is canceled. If omi_mask_cancel[i][j] is equal to 0, it indicates that information from the object mask whose identifier is equal to omi_mask_id[i][j] is signaled.
[0185] If the variable maskIdExist[i][id] is equal to 1, it indicates that an object mask with identifier id exists in the i-th object mask auxiliary picture. If the variable maskIdExist[i][id] is equal to 0, it indicates that an object mask with identifier id does not exist in the i-th object mask auxiliary picture. maskIdExist[i][id] is currently initialized to 0 before decoding the CVS.
[0186] omi_mask_confidence[i][j] is the confidence associated with the j-th object mask, which is updated within the i-th object mask auxiliary picture. -(omi_mask_confidence_length_minus1+1) Expressed in units of omi_mask_confidence[i][j][k], a higher value of omi_mask_confidence[i][j][k] indicates higher confidence. The length of the omi_mask_confidence[i][j][k] syntax element is omi_mask_confidence_length_minus1+1 bits.
[0187] omi_mask_depth[i][j] indicates the object depth associated with the j-th object mask, which is updated within the i-th object mask auxiliary picture. Smaller values ββof omi_mask_depth indicate a shorter distance to the object. The length of the omi_mask_depth[i][j][k] syntax element is omi_object_depth_length_minus1+1 bits.
[0188] omi_mask_label[i][j] specifies the content of the label associated with the j-th object mask, which is updated within the i-th object mask auxiliary picture. The length of the omi_mask_label[i][j] syntax element is 255 bytes or less, excluding the null-terminating byte.
[0189] In embodiments relating to Tables 6 and 7, it is assumed that one or more primary picture layers have already been determined, and only the layer identifier of each object mask picture layer is indicated in the OMI SEI message. Furthermore, since it is assumed that all object mask picture layers are associated with one or more primary picture layers, there is no need to signal the primary picture layer to which the object mask auxiliary picture layer is associated.
[0190] However, VVC supports multiple primary picture layers and multiple auxiliary picture layers. An auxiliary picture layer may be associated with multiple primary picture layers, and one primary picture may be associated with multiple auxiliary picture layers. The NAL unit layer identifiers for the primary and auxiliary picture layers are specified in the SDI SEI message. Also, for each auxiliary picture layer, the primary picture layer to which it is associated is also specified in the SDI SEI message. Referring back to Figure 7, the OMI SEI message may be used to indicate the masks of auxiliary pictures 702 and 704. Thus, the OMI SEI message is associated with primary pictures 701 and 703. In some embodiments, primary picture 701 may correspond to multiple auxiliary pictures, which may also be indicated by the OMI SEI message.
[0191] Since a VVC may have multiple primary picture layers, and an OMI SEI message may apply to only some of these primary picture layers, the number of primary picture layers to which the OMI SEI message applies and the layer identifier of each primary picture layer are signaled within the OMI SEI message. According to the SDI SEI message, after the primary picture layers have been determined, it is not necessary to signal the layer identifier of the object mask auxiliary picture layers within the OMI SEI message. The layer identifier of an auxiliary picture associated with a primary picture layer having the layer identifier layerIdA can be derived based on the SDI SEI message.
[0192] The variable numAuxLayer[i] indicates the number of auxiliary picture layers associated with a primary picture layer whose nuh_layer_id (where nuh_layer_id is the syntax element name for the layer identifier) ββis equal to i. The variable associationAuxLayer[j][i] indicates the value of the nuh_layer_id of the i-th auxiliary picture layer associated with a primary picture layer whose nuh_layer_id is equal to j. numAuxLayer[i] and associationAuxLayer[j][i] are derived from the SDI SEI message as follows:
number
[0193] The OMI SEI messages are shown in Table 8, and their semantics are provided below as an example in Table 8. The definitions of common features, indicated as "high-level information," and individual features, indicated as "individual mask information" (some of which are omitted), can be inherited from the above embodiments by shared parameter / function names. [Table 8]
[0194] The Object Mask Information (OMI) SEI message provides information about the object mask picture, which is coded as an auxiliary picture. The object mask auxiliary picture has nuh_layer_id as nuhLayerIdA, sdi_layer_id[i] as nuhLayerIdA, and sdi_aux_id[i] in the range of 128 to 159 (any value of i is in the range of 0 to sid_max_layers_minus1).
[0195] If an access unit contains auxiliary picture picA in a layer indicated as an object mask auxiliary layer by an OMI SEI message (where nuh_layer_id is equal to nuhLayerIdA), and also contains primary picture picB in a layer indicated as a primary layer by an OMI SEI message (where nuh_layer_id is equal to nuhLayerIdB), the OMI SEI messages persist in output order until one or more of the following conditions are met.
[0196] -CLVS, including auxiliary picture picA, terminates. -CLVS, including primary picture picB, terminates. -CVS will shut down. - The bitstream ends.
[0197] The value of sdi_aux_id[sdi_associated_primary_layer_idx[i][j]] is equal to 0.
[0198] If omi_cancel_flag is equal to 1, it indicates that the SEI message cancels the persistence of previous object mask information SEI messages in output order associated with one or more primary picture layers to which this SEI applies. If omi_cancel_flag is equal to 0, it indicates that object mask information follows, and the object mask information signaled within this SEI message is used to update any existing object mask information from previous SEI messages.
[0199] omi_aux_id_minus128+128 indicates the value of sdi_aux_id of the object mask auxiliary picture. omi_aux_id_minus128 is within the range of 0 to 31.
[0200] If the CVS does not contain an SDI SEI message where sdi_aux_id[i] equals omi_aux_id_minus128+128 for at least one value of i, then no picture in the CVS is associated with an OMI SEI message.
[0201] If an AU contains both SDI SEI messages and OMI SEI messages where sdi_aux_id[i] is equal to omi_aux_id_minus128+128 for at least one value of i, then SDI SEI messages take precedence over OMI SEI messages in the decoding order.
[0202] In some embodiments, an SEI message is associated with multiple primary pictures corresponding to an auxiliary picture to which the SEI message applies. The number of primary pictures and the layer identifiers of the primary pictures can be determined as common features. For example, omi_num_primary_pic_layer_minus1+1 indicates the number of primary picture layers associated with the object mask auxiliary picture layer to which this SEI message applies. The value of omi_num_primary_pic_layer_minus1 is in the range of 0 to sdi_max_layers_minus1.
[0203] Additionally, omi_primary_pic_layer_id[i] specifies the nuh_layer_id value of the i-th primary picture layer to which this OMI SEI message applies. Since the value of sdi_aux_id[j] is equal to 0 regardless of the range of j from 0 to sid_max_layers_minus1, sdi_layer_id[j] is equal to omi_primary_pic_layer_id[i].
[0204] omi_mask_id_length_minus8+8 indicates the number of bits used to encode the omi_mask_id[i][j][k] syntax element.
[0205] If omi_mask_confidence_info_present_flag is equal to 1, it indicates that the omi_mask_confidence[i][j][k] syntax element exists. If omi_mask_confidence_info_present_flag is equal to 0, it indicates that the omi_mask_confidence[i][j][k] syntax element does not exist. As a bitstream conformance requirement, the value of omi_mask_confidence_info_present_flag must be the same for all object_mask_info() syntax structures within CLVS.
[0206] omi_mask_confidence_length_minus1+1 specifies the bitwise length of the omi_mask_confidence[i][j][k] syntax element. As a bitstream conformance requirement, the value of omi_mask_confidence_length_minus1 must be the same for all object_mask_info() syntax structures within CLVS.
[0207] If omi_object_depth_info_present_flag is equal to 1, it indicates that the omi_object_depth[i][j][k] syntax element exists. If omi_object_depth_info_present_flag is equal to 0, it indicates that the omi_object_depth[i][j][k] syntax element does not exist. As a bitstream conformance requirement, the value of omi_object_depth_info_present_flag must be the same for all object_mask_info() syntax structures within CLVS.
[0208] omi_object_depth_length_minus1+1 specifies the bitwise length of the omi_mask_confidence[i][j][k] syntax element. As a bitstream conformance requirement, the value of omi_object_depth_length_minus1 must be the same for all object_mask_info() syntax structures within the CLVS.
[0209] If omi_mask_label_info_present_flag is equal to 1, it indicates that omi_mask_label_language_present_flag and omi_mask_label[i][j][k] exist. If omi_mask_label_info_present_flag is equal to 0, it indicates that omi_mask_label_language_present_flag and omi_mask_label[i][j][k] do not exist.
[0210] If omi_mask_label_language_present_flag is equal to 1, it indicates that omi_mask_label_language exists. If omi_mask_label_language_present_flag is equal to 0, it indicates that omi_mask_label_language does not exist and the language of the mask label is not specified.
[0211] omi_bit_equal_to_zero is equal to 0.
[0212] The `omi_mask_label_language` element contains a language tag, as defined in IET FRFC 5646, followed by a null-terminating byte equal to 0x00. The length of the `omi_mask_label_language` syntax element is 255 bytes or less, excluding the null-terminating byte. If it is not present, the label language is not specified.
[0213] omi_num_mask_in_pic[i][j] indicates the number of masks in the j-th object mask auxiliary picture associated with the i-th primary picture. omi_num_mask_in_pic[i][j] is within the range of 0 or more and (1<<BitDepthY)-1 or less, where BitDepthY is the bit depth of the samples of the luma component.
[0214] In some embodiments, the encoder may determine the number of auxiliary pictures corresponding to each of the plurality of primary pictures. Next, the encoder can determine the individual characteristics of the masks represented by the auxiliary pictures for each of the plurality of primary pictures.
[0215] For example, the individual characteristic omi_mask_id[i][j][k] indicates the identifier of the k-th object mask in the j-th object mask auxiliary picture associated with the i-th primary picture. The object mask identifier associated with the sample position (x, y) in the j-th object mask auxiliary picture is equal to p[j][x][y], and p[j][x][y] refers to the luma sample at the position (x, y) in the decoded j-th object mask auxiliary picture.
[0216] The variable maskId[i][j][k] that specifies the object mask identifier of the k-th object mask in the j-th object mask auxiliary picture associated with the i-th primary picture in the SEI message is derived as follows.
Number
[0217] omi_mask_confidence[i][j][k] is the confidence associated with the k-th object mask in the j-th object mask auxiliary picture associated with the i-th primary picture, 2 -(omi_mask_confidence_length_minus1+1)Expressed in units of omi_mask_confidence[i][j][k], a higher value of omi_mask_confidence[i][j][k] indicates higher confidence. The length of the omi_mask_confidence[i][j][k] syntax element is omi_mask_confidence_length_minus1+1 bits.
[0218] omi_mask_depth[i][j][k] indicates the object depth associated with the k-th object mask in the j-th object mask auxiliary picture associated with the i-th primary picture. Smaller values ββof omi_mask_depth indicate shorter distances to the object. The length of the omi_mask_depth[i][j][k] syntax element is omi_object_depth_length_minus1+1 bits.
[0219] omi_mask_label[i][j][k] specifies the content of the label associated with the k-th object mask in the j-th object mask auxiliary picture associated with the i-th primary picture. The length of the omi_mask_label[i][j][k] syntax element is 255 bytes or less, excluding the null-terminating byte.
[0220] Similar to Table 7, in some embodiments, only the updated mask information may be signaled in the OMI SEI message. Furthermore, signaling can be skipped for information that has not changed between this OMI SEI message and the previous OMI SEI message, in order to save bit overhead. The syntax is shown below as Table 9 (differences from Table 6 are shown in italics in Table 9). The definitions of common features, indicated as "high-level information," and individual features, indicated as "individual mask information" (some of which are omitted), can be inherited from the embodiments described above with shared parameter / function names. [Table 9-1] [Table 9-2]
[0221] The variable numAuxLayer[i] indicates the number of auxiliary picture layers associated with the primary picture layer whose nuh_layer_id is equal to i. The variable associationAuxLayer[j][i] indicates the value of the nuh_layer_id of the i-th auxiliary picture layer associated with the primary picture layer whose nuh_layer_id is equal to j. numAuxLayer[i] and associationAuxLayer[j][i] are derived from the SDI SEI message as follows:
number
[0222] The Object Mask Information (OMI) SEI message provides information about the object mask picture, which is coded as an auxiliary picture. The object mask auxiliary picture has nuh_layer_id as nuhLayerIdA, sdi_layer_id[i] as nuhLayerIdA, and sdi_aux_id[i] in the range of 128 to 159 (any value of i is in the range of 0 to sid_max_layers_minus1).
[0223] If an access unit contains auxiliary picture picA in a layer indicated as an object mask auxiliary layer by an OMI SEI message (where nuh_layer_id is equal to nuhLayerIdA), and also contains primary picture picB in a layer indicated as a primary layer by an OMI SEI message (where nuh_layer_id is equal to nuhLayerIdB), the OMI SEI messages persist in output order until one or more of the following conditions are met. -CLVS, including auxiliary picture picA, terminates. -CLVS, including primary picture picB, terminates. -CVS will shut down. - The bitstream ends.
[0224] The value of sdi_aux_id[sdi_associated_primary_layer_idx[i][j]] is equal to 0.
[0225] If omi_cancel_flag is equal to 1, it indicates that the SEI message cancels the persistence of previous object mask information SEI messages in output order associated with one or more primary picture layers to which this SEI applies. If omi_cancel_flag is equal to 0, it indicates that object mask information follows, and the object mask information signaled within this SEI message is used to update any existing object mask information from previous SEI messages.
[0226] omi_aux_id_minus128+128 indicates the value of sdi_aux_id in the object mask auxiliary picture layer. omi_aux_id_minus128 is within the range of 0 to 31.
[0227] If the CVS does not contain an SDI SEI message where sdi_aux_id[i] equals omi_aux_id_minus128+128 for at least one value of i, then no picture in the CVS is associated with an OMI SEI message.
[0228] If an AU contains both SDI SEI messages and OMI SEI messages where sdi_aux_id[i] is equal to omi_aux_id_minus128+128 for at least one value of i, then SDI SEI messages take precedence over OMI SEI messages in the decoding order.
[0229] omi_num_primary_pic_layer_minus1+1 indicates the number of primary picture layers associated with the object mask auxiliary picture layer to which this SEI message applies. The value of omi_num_primary_pic_layer_minus1 is within the range of 0 to sdi_max_layers_minus1.
[0230] omi_primary_pic_layer_id[i] specifies the nuh_layer_id value of the i-th primary picture layer to which this OMI SEI message applies. Since the value of sdi_aux_id[j] is equal to 0 regardless of whether j is a value between 0 and sid_max_layers_minus1, sdi_layer_id[j] is equal to omi_primary_pic_layer_id[i].
[0231] omi_mask_id_length_minus8+8 indicates the number of bits used to encode the omi_mask_id[i][j][k] syntax element.
[0232] If omi_mask_confidence_info_present_flag is equal to 1, it indicates that the omi_mask_confidence[i][j][k] syntax element exists. If omi_mask_confidence_info_present_flag is equal to 0, it indicates that the omi_mask_confidence[i][j][k] syntax element does not exist. As a bitstream conformance requirement, the value of omi_mask_confidence_info_present_flag must be the same for all object_mask_info() syntax structures within CLVS.
[0233] omi_mask_confidence_length_minus1+1 specifies the bitwise length of the omi_mask_confidence[i][j][k] syntax element. As a bitstream conformance requirement, the value of omi_mask_confidence_length_minus1 must be the same for all object_mask_info() syntax structures within CLVS.
[0234] If omi_object_depth_info_present_flag is equal to 1, it indicates that the omi_object_depth[i][j][k] syntax element exists. If omi_object_depth_info_present_flag is equal to 0, it indicates that the omi_object_depth[i][j][k] syntax element does not exist. As a bitstream conformance requirement, the value of omi_object_depth_info_present_flag must be the same for all object_mask_info() syntax structures within CLVS.
[0235] omi_object_depth_length_minus1+1 specifies the bitwise length of the omi_mask_confidence[i][j][k] syntax element. As a bitstream conformance requirement, the value of omi_object_depth_length_minus1 must be the same for all object_mask_info() syntax structures within the CLVS.
[0236] If omi_mask_label_info_present_flag is equal to 1, it indicates that omi_mask_label_language_present_flag and omi_mask_label[i][j][k] exist. If omi_mask_label_info_present_flag is equal to 0, it indicates that omi_mask_label_language_present_flag and omi_mask_label[i][j][k] do not exist.
[0237] If omi_mask_label_language_present_flag is equal to 1, it indicates that omi_mask_label_language exists. If omi_mask_label_language_present_flag is equal to 0, it indicates that omi_mask_label_language does not exist and the language of the mask label is not specified.
[0238] omi_bit_equal_to_zero is equal to 0.
[0239] The `omi_mask_label_language` element contains a language tag, as defined in IET FRFC 5646, followed by a null-terminating byte equal to 0x00. The length of the `omi_mask_label_language` syntax element is 255 bytes or less, excluding the null-terminating byte. If it is not present, the label language is not specified.
[0240] If omi_mask_pic_update_flag[i][j] is equal to 1, it indicates that mask information for the j-th object mask auxiliary picture associated with the i-th primary picture will be signaled. If omi_mask_pic_update_flag[i][j] is equal to 0, it indicates that mask information for the j-th object mask auxiliary picture associated with the i-th primary picture will not be signaled. If mask information for the j-th object mask auxiliary picture associated with the i-th primary picture does not exist, the persistence mechanism is used. That is, the information is inherited from the last OMI SEI message that signals the mask information for the j-th object mask auxiliary picture associated with the i-th primary picture.
[0241] omi_num_mask_in_pic_update[i][j] indicates the number of object masks in the j-th auxiliary picture associated with the i-th primary picture being signaled. omi_num_mask_in_pic[i][j] is within the range of 0 or more and (1<<BitDepthY)-1 or less, where BitDepthY is the bit depth of the samples of the luma component.
[0242] omi_mask_id[i][j][k] indicates the identifier of the k-th object mask in the j-th object mask auxiliary picture associated with the i-th primary picture. The object mask identifier associated with the sample position (x,y) in the j-th object mask auxiliary picture is equal to p[j][x][y], where p[j][x][y] points to the luma sample at the position (x,y) in the decoded j-th object mask auxiliary picture.
[0243] The variable maskId[i][j][k] that specifies the object mask identifier of the k-th object mask in the j-th object mask auxiliary picture associated with the i-th primary picture in the SEI message is derived as follows.
Number
[0244] When omi_mask_cancel[i][j][k] is equal to 1, the duration range of the object mask whose identifier is equal to omi_mask_id[i][j][k] is canceled. When omi_mask_cancel[i][j][k] is equal to 0, it indicates that the information of the object mask whose identifier is equal to omi_mask_id[i][j][k] is being signaled.
[0245] When the variable maskIdExist[i][j][k] is equal to 1, it indicates that there exists an object mask with identifier k in the j-th object mask auxiliary picture associated with the i-th primary picture. When the variable maskIdExist[i][j][k] is equal to 0, it indicates that there is no object mask with an identifier equal to k in the j-th object mask auxiliary picture associated with the i-th primary picture. maskIdExist[i][j][k] is initialized to 0 before currently decrypting the CVS.
[0246] omi_mask_confidence[i][j][k] represents the confidence associated with the k-th object mask in the j-th object mask auxiliary picture associated with the i-th primary picture, -(omi_mask_confidence_length_minus1+1) shown in units of 2, and a higher value of omi_mask_confidence[i][j][k] indicates a higher confidence. The length of the omi_mask_confidence[i][j][k] syntax element is omi_mask_confidence_length_minus1 + 1 bits.
[0247] omi_mask_depth[i][j][k] indicates the object depth associated with the k-th object mask in the j-th object mask auxiliary picture associated with the i-th primary picture. A smaller value of omi_mask_depth indicates a shorter distance to the object. The length of the omi_mask_depth[i][j][k] syntax element is omi_object_depth_length_minus1 + 1 bits.
[0248] omi_mask_label[i][j][k] specifies the content of the label associated with the k-th object mask in the j-th object mask auxiliary picture associated with the i-th primary picture. The length of the omi_mask_label[i][j][k] syntax element is 255 bytes or less excluding the null-terminating byte.
[0249] In some of the above-described embodiments discussed in relation to Tables 7 and 9, in order to save bit overhead, only the updated object mask information is signaled. omi_mask_cancel[i][j][k], which indicates whether the mask having an ID equal to omi_mask_id[i][j][k] is canceled, is signaled only if a mask with an ID equal to omi_mask_id[i][j][k] already exists. Therefore, the decoder has to maintain an object mask ID list for deriving the variable maskIdExist[i][j][omi_mask_id[i][j][k]] in order to determine whether an object mask with an ID equal to omi_mask_id[i][j][k] exists and analyze the syntax element. This introduces an analysis dependency on the previous SEI message.
[0250] In some embodiments, omi_mask_cancel[i][j][k] is always signaled. Further, if there was no previous mask object with an ID equal to omi_mask_id[i][j][k], the value of omi_mask_cancel[i][j][k] is forced to 0. This indicates that the persistence range of the object mask with an ID equal to omi_mask_cancel[i][j][k] is not canceled. The syntax of this method is shown as Table 10 below (the differences from Table 6 are shown in italics in Table 10). The definitions of the common features shown as "high-level information" and the individual features shown as "individual mask information" (part of which is omitted) can be inherited from the above-described embodiments with shared parameter / function names.
Table 10
[0251] The variable numAuxLayer[i] indicates the number of auxiliary picture layers associated with the primary picture layer whose nuh_layer_id is equal to i. The variable associationAuxLayer[j][i] indicates the value of the nuh_layer_id of the i-th auxiliary picture layer associated with the primary picture layer whose nuh_layer_id is equal to j. numAuxLayer[i] and associationAuxLayer[j][i] are derived from the SDI SEI message as follows:
number
[0252] The Object Mask Information (OMI) SEI message provides information about the object mask picture, which is coded as an auxiliary picture. The object mask auxiliary picture has nuh_layer_id as nuhLayerIdA, sdi_layer_id[i] as nuhLayerIdA, and sdi_aux_id[i] in the range of 128 to 159 (any value of i is in the range of 0 to sid_max_layers_minus1).
[0253] If an access unit contains auxiliary picture picA in a layer indicated as an object mask auxiliary layer by an OMI SEI message (where nuh_layer_id is equal to nuhLayerIdA), and also contains primary picture picB in a layer indicated as a primary layer by an OMI SEI message (where nuh_layer_id is equal to nuhLayerIdB), the OMI SEI messages persist in output order until one or more of the following conditions are met. -CLVS, including auxiliary picture picA, terminates. -CLVS, including primary picture picB, terminates. -CVS will shut down. - The bitstream ends.
[0254] The value of sdi_aux_id[sdi_associated_primary_layer_idx[i][j]] should be equal to 0.
[0255] When omi_cancel_flag is equal to 1, it indicates that the SEI message cancels the persistence of the previous object mask information SEI message in the output order associated with one or more primary picture layers to which this SEI is applied. When omi_cancel_flag is equal to 0, it indicates that the object mask information continues, and the object mask information signaled within this SEI message is used to update the object mask information existing in any previous SEI message.
[0256] When omi_mask_pic_update_flag[i][j] is equal to 1, it indicates that the mask information of the j-th object mask auxiliary picture associated with the i-th primary picture is signaled. When omi_mask_pic_update_flag[i][j] is equal to 0, it indicates that the mask information of the j-th object mask auxiliary picture associated with the i-th primary picture is not signaled. If the mask information of the j-th object mask auxiliary picture associated with the i-th primary picture does not exist, a persistence mechanism is used. That is, information is inherited from the last OMI SEI message that signaled the mask information of the j-th object mask auxiliary picture associated with the i-th primary picture.
[0257] omi_num_mask_in_pic_update[i][j] indicates the number of object masks in the j-th auxiliary picture associated with the i-th primary picture being signaled. omi_num_mask_in_pic[i][j] should be within the range of 0 or more and (1<<BitDepthY)-1 or less, where BitDepthY is the bit depth of the samples of the luma component.
[0258] omi_mask_id[i][j][k] represents the identifier of the k-th object mask in the j-th object mask auxiliary picture associated with the i-th primary picture. The object mask identifier associated with the sample position (x,y) in the j-th object mask auxiliary picture is equal to p[j][x][y], where p[j][x][y] points to the luma sample at position (x,y) in the decoded j-th object mask auxiliary picture.
[0259] The variable maskId[i][j][k], which specifies the object mask identifier of the k-th object mask of the j-th object mask auxiliary picture associated with the i-th primary picture in the SEI message, is derived as follows:
number
[0260] If omi_mask_cancel[i][j][k] is equal to 1, the duration of the object mask whose identifier is equal to omi_mask_id[i][j][k] is canceled. If omi_mask_cancel[i][j][k] is equal to 0, it indicates that information about the object mask whose identifier is equal to omi_mask_id[i][j] is signaled.
[0261] If maskIdExist[i][j][k] is equal to 0, the value of omi_mask_cancel[i][j][k] should be equal to 1. In CLVS, if omi_mask_id[i][j][k] has a specific value equal to omiMaskId for the first time, the value of omi_mask_cancel[i][j][k] should be equal to 0.
[0262] The variables maskIdExist[i][j][k] are derived as follows: maskIdExist[i][j][k] is currently initialized to 0 before decrypting the CVS. [Number]
[0263] omi_mask_confidence[i][j][k] indicates the confidence associated with the k-th object mask in the j-th object mask auxiliary picture associated with the i-th primary picture, in units of 2 -(omi_mask_confidence_length_minus1+1) . A higher value of omi_mask_confidence[i][j][k] indicates a higher confidence. The length of the omi_mask_confidence[i][j][k] syntax element is omi_mask_confidence_length_minus1 + 1 bits.
[0264] omi_mask_depth[i][j][k] indicates the object depth associated with the k-th object mask in the j-th object mask auxiliary picture associated with the i-th primary picture. A smaller value of omi_mask_depth indicates a shorter distance to the object. The length of the omi_mask_depth[i][j][k] syntax element is omi_object_depth_length_minus1 + 1 bits.
[0265] omi_mask_label[i][j][k] specifies the content of the label associated with the k-th object mask in the j-th object mask auxiliary picture associated with the i-th primary picture. The length of the omi_mask_label[i][j][k] syntax element should be 255 bytes or less excluding the null-terminating byte.
[0266] In some of the above embodiments, the pixel values of the object mask auxiliary picture represent the mask ID. The decoder determines the mask based on the decoded sample values of the object mask auxiliary picture. Therefore, the encoder needs to encode the object mask auxiliary picture using reversible coding. Otherwise, the mask will be distorted. However, in some cases, the number of object masks may be much less than the range of sample values. Therefore, in some embodiments, the sample values of the auxiliary picture can be non-reversibly encoded. Therefore, given an object mask ID, for any decoded sample value different from any object mask ID, the decoder can restore the decoded sample value to the closest mask ID value.
[0267] maskID[i] (i = 0 to n - 1) indicates the i-th object mask ID within the picture, and it is assumed that maskID[i] <= maskID[j] when i < j. The tolerance boundary is calculated as follows.
Equation
[0268] p[x][y] represents the decoded value of the sample having the coordinates (x, y). When a mask ID is associated with p[x][y], ID(p[x][y]) is derived as follows.
Equation
[0269] In some embodiments, a bounding box is signaled for each object mask to determine the object's position. For example, if the cancel flag indicates that the SEI message does not cancel the persistence of information from previous SEI messages, the encoder may, in step 604, determine a bounding box surrounding the object mask and encode the bounding box within the SEI message. In some embodiments, the encoder may, in step 604, determine a bounding box surrounding the object mask regardless of whether the cancel flag indicates that the SEI message cancels the persistence of information from previous SEI messages. Thus, on the decoder side, only samples within the bounding box are checked, and samples outside the bounding box are treated as background, regardless of their value. The coordinates of the bounding box of the signaled object mask are defined on the cropped portion of the decoded picture relative to the fitted cropping window specified by the active SPS. Furthermore, the gating flag omi_mask_bounding_box_present_flag is added to make signaling the bounding box parameter optional, allowing the encoder to flexibly decide whether to signal the bounding box to demarcate the mask or not to save bit overhead.
[0270] The syntax for this method is shown below as Table 11 (differences from Table 6 are shown in italics in Table 11). The definitions of common features, shown as "high-level information," and individual features, shown as "individual mask information" (some of which are omitted), can be inherited from the above embodiment using shared parameter / function names. [Table 11]
[0271] The variable numAuxLayer[i] indicates the number of auxiliary picture layers associated with the primary picture layer whose nuh_layer_id is equal to i. The variable associationAuxLayer[j][i] indicates the value of the nuh_layer_id of the i-th auxiliary picture layer associated with the primary picture layer whose nuh_layer_id is equal to j. numAuxLayer[i] and associationAuxLayer[j][i] are derived from the SDI SEI message as follows:
number
[0272] The Object Mask Information (OMI) SEI message provides information about the object mask picture, which is coded as an auxiliary picture. The object mask auxiliary picture has nuh_layer_id as nuhLayerIdA, sdi_layer_id[i] as nuhLayerIdA, and sdi_aux_id[i] in the range of 128 to 159 (any value of i is in the range of 0 to sid_max_layers_minus1).
[0273] To use this SEI message, you need to define the following variables. - The width and height of the cropped picture in units of luma samples (referred to herein as CroppedWidth and CroppedHeight, respectively). - Left offset of the fitting cropping window (ConfWinLeftOffset). - The top offset of the matching cropping window (ConfWinTopOffset). - Chroma format indicator (represented by ChromaFormatId in this specification). -The variables SubWidthC and SubHeightC are derived from ChromaFormatIdc.
[0274] If an access unit contains auxiliary picture picA in a layer indicated as an object mask auxiliary layer by an OMI SEI message (where nuh_layer_id is equal to nuhLayerIdA), and also contains primary picture picB in a layer indicated as a primary layer by an OMI SEI message (where nuh_layer_id is equal to nuhLayerIdB), the OMI SEI messages persist in output order until one or more of the following conditions are met. -CLVS, including auxiliary picture picA, terminates. -CLVS, including primary picture picB, terminates. -CVS will shut down. - The bitstream ends.
[0275] The value of sdi_aux_id[sdi_associated_primary_layer_idx[i][j]] should be equal to 0.
[0276] If omi_cancel_flag is equal to 1, it indicates that the SEI message cancels the persistence of previous object mask information SEI messages in output order associated with one or more primary picture layers to which this SEI applies. If omi_cancel_flag is equal to 0, it indicates that object mask information follows, and the object mask information signaled within this SEI message is used to update any existing object mask information from previous SEI messages.
[0277] omi_mask_size_length_minus1+1 specifies the bitwise length of the omi_mask_top[i][j][k] and omi_mask_left[i][j][k] syntax elements. As a bitstream conformance requirement, the value of omi_mask_size_length_minus1 should be the same for all object_mask_info() syntax structures within CLVS.
[0278] When omi_mask_pic_update_flag[i][j] is equal to 1, it indicates that the mask information of the j-th object mask auxiliary picture associated with the i-th primary picture is signaled. When omi_mask_pic_update_flag[i][j] is equal to 0, it indicates that the mask information of the j-th object mask auxiliary picture associated with the i-th primary picture is not signaled. If there is no mask information for the j-th object mask auxiliary picture associated with the i-th primary picture, a persistence mechanism is used. That is, information is inherited from the last OMI SEI message that signaled the mask information of the j-th object mask auxiliary picture associated with the i-th primary picture.
[0279] omi_num_mask_in_pic_update[i][j] indicates the number of object masks in the j-th auxiliary picture associated with the i-th primary picture that is signaled. omi_num_mask_in_pic_update[i][j] should be within the range of 0 or more and (1<<BitDepthY)-1 or less, where BitDepthY is the bit depth of the samples of the luma component.
[0280] If omi_mask_bounding_box_present_flag[i][j][k] is equal to 1, it indicates that the bounding box parameters associated with the k-th object mask in the j-th object mask auxiliary picture associated with the i-th primary picture, omi_mask_top[i][j][k], omi_mask_left[i][j][k], omi_mask_width[i][j][k], and omi_mask_height[i][j][k], exist. If omi_num_mask_in_pic_update[i][j][k] is equal to 0, it indicates that the bounding box parameters associated with the k-th object mask in the j-th object mask auxiliary picture associated with the i-th primary picture, omi_mask_top[i][j][k], omi_mask_left[i][j][k], omi_mask_width[i][j][k], and omi_mask_height[i][j][k], do not exist.
[0281] omi_mask_id[i][j][k] indicates the identifier of the k-th object mask in the j-th object mask auxiliary picture associated with the i-th primary picture.
[0282] The variable maskId[i][j][k], which specifies the object mask identifier of the k-th object mask of the j-th object mask auxiliary picture associated with the i-th primary picture in the SEI message, is derived as follows:
number
[0283] For example, bounding box information may be generated and signaled within SEI messages. The indicators omi_mask_top[i][j][k], omi_mask_left[i][j][k], omi_mask_width[i][j][k], and omi_mask_height[i][j][k] specify the coordinates, width, and height of the top-left corner of the bounding box of the object identified by identifier omi_mask_id[i][j][k] in the cropped decoded picture, relative to the fitted cropping window specified in the active SPS.
[0284] The value of omi_mask_left[i][j][k] should be within the range of 0 or greater and (CroppedWidth / SubWidthC-1), where CroppedWidth and SubWidthC are associated with the j-th object mask auxiliary picture associated with the i-th primary picture. If omi_mask_left[i][j][k] does not exist, its value is presumed to be 0.
[0285] The value of omi_mask_top[i][j][k] should be within the range of 0 or greater and (CroppedHeight / SubHeightC-1), where CroppedHeight and SubHeightC are associated with the j-th object mask auxiliary picture associated with the i-th primary picture. If omi_mask_top[i][j][k] does not exist, its value is presumed to be 0.
[0286] The value of omi_mask_width[i][j][k] should be within the range of 0 or greater and (CroppedWidth / SubWidthC - omi_mask_left[i][j][k]). If omi_mask_width[i][j][k] does not exist, its value is presumed to be (CroppedWidth / SubWidthC - omi_mask_left[i][j][k]).
[0287] The value of omi_mask_height[i][j][k] should be within the range of 0 or greater and (CroppedHeight / SubHeightC - omi_mask_top[i][j][k]). If omi_mask_height[i][j][k] does not exist, its value is presumed to be (CroppedHeight / SubHeightC - omi_mask_top[i][j][k]).
[0288] The identified object mask is located within a bounding box containing a Luma sample whose horizontal picture coordinates are between SubWidthC*(ConfWinLeftOffset+omi_mask_left[i][j][k]) and SubWidthC*(ConfWinLeftOffset+omi_mask_left[i][j][k]+omi_mask_width[i][j][k])-1, and whose vertical picture coordinates are between SubHeightC*(ConfWinTopOffset+omi_mask_top[i][j][k]) and SubHeightC*(ConfWinTopOffset+omi_mask_top[i][j][k]+omi_mask_height[i][j][k])-1.
[0289] The variable p[i][j][x][y] is the decoded value of the sample at the relative sample position (x,y) in the j-th object mask auxiliary picture associated with the i-th primary picture.
number
[0290] If omi_mask_cancel[i][j][k] is equal to 1, the duration of the object mask whose identifier is equal to omi_mask_id[i][j][k] is canceled. If omi_mask_cancel[i][j][k] is equal to 0, it indicates that information about the object mask whose identifier is equal to omi_mask_id[i][j][k] is signaled.
[0291] If the variable maskIdExist[i][j][k] is equal to 1, it indicates that an object mask with identifier k exists in the j-th object mask auxiliary picture associated with the i-th primary picture. If the variable maskIdExist[i][j][k] is equal to 0, it indicates that no object mask with identifier k exists in the j-th object mask auxiliary picture associated with the i-th primary picture. maskIdExist[i][j][k] is currently initialized to 0 before decoding the CVS.
[0292] omi_mask_confidence[i][j][k] represents the confidence associated with the k-th object mask in the j-th object mask auxiliary picture associated with the i-th primary picture, and 2 -(omi_mask_confidence_length_minus1+1) Expressed in units of omi_mask_confidence[i][j][k], a higher value of omi_mask_confidence[i][j][k] indicates higher confidence. The length of the omi_mask_confidence[i][j][k] syntax element is omi_mask_confidence_length_minus1+1 bits.
[0293] omi_mask_depth[i][j][k] indicates the object depth associated with the k-th object mask in the j-th object mask auxiliary picture associated with the i-th primary picture. Smaller values ββof omi_mask_depth indicate shorter distances to the object. The length of the omi_mask_depth[i][j][k] syntax element is omi_object_depth_length_minus1+1 bits.
[0294] omi_mask_label[i][j][k] specifies the content of the label associated with the k-th object mask in the j-th object mask auxiliary picture associated with the i-th primary picture. The length of the omi_mask_label[i][j][k] syntax element should be 255 bytes or less, excluding the null-terminating byte.
[0295] In some other embodiments, the bounding box parameters omi_mask_top[i][j][k], omi_mask_left[i][j][k], omi_mask_width[i][j][k], and omi_mask_height[i][j][k] are coded with fixed-length codes whose length is preset to 8, 16, or 32, etc. In this case, it is not necessary to signal omi_mask_size_length_minus1. As shown in Table 12 below, 16-bit codes are used to code omi_mask_top[i][j][k], omi_mask_left[i][j][k], omi_mask_width[i][j][k], and omi_mask_height[i][j][k]. The definitions of common features, indicated as "high-level information," and individual features, indicated as "individual mask information" (some of which are omitted), can be inherited from the embodiments described above with shared parameter / function names. [Table 12]
[0296] In some of the above embodiments, the value of the sample at position (x, y) (denoted as p[x][y]) indicates a mask identifier associated with the sample. For a sample with a bit depth equal to bitdepthY, the maximum number of mask identifiers is 1 << bitdepthY. However, if two object masks overlap each other, there are multiple masks associated with the sample, and it cannot be represented by the sample value. Therefore, in some of the above embodiments, multiple object mask auxiliary pictures are used. p[i][x][y] indicates the sample value at position (x, y) of the i-th mask auxiliary picture. When two object masks with identifiers idA and idB are associated with the sample position (x, y), p[0][x][y] can be set to idA, and p[1][x][y] can be set to idB. Therefore, the maximum number of supported overlapping masks is equal to the maximum number of object mask auxiliary pictures.
[0297] In some embodiments, the mask identifier may be represented by bits of the sample value rather than directly by the sample value. That is, each bit of the sample value represents a different identifier of the mask. When a mask with identifier idA is associated with the sample at position (x,y), the sample value p[x][y] at position (x,y) is equal to (1<<idA). Assuming the bit depth of the sample is bitdepthY, the maximum number of mask identifiers is bitdepthY. In this method, the maximum number of mask identifiers supported by the object mask auxiliary picture is less than that in the previous embodiments. However, when masks overlap, they can be easily handled. For example, when two object masks with identifiers idA and idB (since there are two different masks, idA is not equal to idB) are associated with the sample position (x,y), p[x][y] can be set to (1<<idA)+(1<<idB). Also, for the sample value at position (x,y), if the k-th bit is "1", the sample (x,y) is covered by the k-th mask. If the k-th bit is "0", the sample (x,y) is not covered by the k-th mask. [FIG. 9] An exemplary binary representation of the sample value p[x][y] according to some embodiments of the present disclosure. As shown in FIG. 9, this is the binary representation of the sample value p[x][y]. Since the least significant bit is "0", it means that the sample (x,y) is not covered by the 0th mask (or the mask with identifier 0). Since the 1st bit position is also "0", it means that the sample (x,y) is not covered by the 1st mask (or the mask with identifier 1). Since both the 2nd bit position and the most significant bit are "1", it means that the sample (x,y) is covered by both the 2nd mask and the (bitdepthY-1)th mask (i.e., the two masks with identifiers 2 and bitdepthY-1 overlap at the sample position (x,y)).
[0298] To support more mask identifiers, multiple object mask auxiliary pictures can be used. For example, if there are m object mask auxiliary pictures with indices from 0 to m - 1 and a bit depth equal to bitdepthY, the object mask identifier associated with the sample position (x, y) is idA. The sample value of each mask auxiliary picture at position (x, y) is derived as follows.
Number
[0299] p[i][x][y] = 1 << idA (when i is equal to n) p[i][x][y] = 0 (when i is not equal to n) p[i][x][y] is the value of the sample at position (x, y) in the i-th mask auxiliary picture. In some of the embodiments described above, the identifier of an object mask is represented by the sample value within the mask region in an auxiliary picture. Therefore, the encoder cannot change the sample value of the mask region to optimize the coding result, nor can it adjust the sample value of the mask region in real time. In some embodiments, the auxiliary picture may contain a plurality of predetermined sample values, and the sample value used to represent the mask of an object may be selected from those sample values ββdepending on the difference in values ββbetween the plurality of predetermined sample values. For example, if there are three object masks in the first frame, the encoder may set the mask sample values ββof these three object masks to 64, 128, and 192, respectively, because a greater distance between the sample values ββresults in a wider sample reconstruction space and therefore greater error tolerance. (i.e., these three object masks have identifiers equal to 64, 128, and 192, respectively). In the second frame, the two objects with identifiers equal to 128 and 192 disappear from the picture, and only the mask with identifier equal to 64 remains in the picture. Changing the sample value of this mask from 64 to 128 could potentially increase its error tolerance, but the encoder cannot change the sample value because it is the identifier for this mask.
[0300] To address the above issues, in some embodiments, the determination of the mask sample value is independent of the mask identifier, so that the mask sample value of the mask can vary from frame to frame. This allows the encoder to optimize the coding results by adjusting the sample value according to the mask number of different frames.
[0301] The syntax is shown in Table 13 below, and the semantics are listed below the table. The syntax element omi_aux_sample_value[i][j][k] is the mask sample value of the object mask having identifier omi_mask_id[i][j][k], and is signaled only if the syntax element omi_mask_id_equal_to_aux_sample_value_flag is equal to false. This means that the mask sample value is different from the mask identifier. If the mask sample value is different from the mask identifier, the bit lengths of the mask sample value and the mask identifier may also be different. Therefore, the two syntax elements omi_mask_id_length and omi_aux_sample_value_length_minus8 are signaled to indicate the bit lengths of the mask sample value and the mask identifier, respectively. The syntax and semantics of these syntax elements are shown below in italics. The definitions of common features shown as "high-level information" and individual features shown as "individual mask information" (some of which are omitted) can be inherited from the above-described embodiment using shared parameter / function names. [Table 13-1] [Table 13-2]
[0302] The Object Mask Information (OMI) SEI message provides information about an object mask picture coded as an auxiliary picture. An object mask auxiliary picture has a nuh_layer_id equal to sdi_layer_id[i] and an sdi_aux_id[i] in the range of 128 to 159 (any value of i is in the range of 0 to sid_max_layers_minus1). Note 1 - Each object mask auxiliary picture layer is associated with one primary picture layer, and one primary picture layer may be associated with one or more object mask auxiliary picture layers.
[0303] To use this SEI message, you need to define the following variables. - The width and height of the cropped picture in units of luma samples (referred to herein as CroppedWidth and CroppedHeight, respectively). - Left offset of the fitting cropping window (ConfWinLeftOffset). - The top offset of the matching cropping window (ConfWinTopOffset). - Chroma format indicator (represented herein as ChromaFormatIdc). -The variables SubWidthC and SubHeightC are derived from ChromaFormatIdc.
[0304] If an access unit contains auxiliary picture picA in a layer indicated as an object mask auxiliary layer by an OMI SEI message (where nuh_layer_id is equal to nuhLayerIdA), and also contains primary picture picB in a layer indicated as a primary layer by an OMI SEI message (where nuh_layer_id is equal to nuhLayerIdB), the OMI SEI messages persist in output order until one or more of the following conditions are met. -CLVS, including auxiliary picture picA, terminates. -CLVS, including primary picture picB, terminates. -CVS will shut down. - The bitstream ends.
[0305] If omi_cancel_flag is equal to 1, it indicates that the SEI message cancels the persistence of previous object mask information SEI messages in output order associated with one or more primary picture layers to which this SEI applies. If omi_cancel_flag is equal to 0, it indicates that object mask information follows, and the object mask information signaled within this SEI message is used to update any existing object mask information from previous SEI messages.
[0306] omi_aux_id_minus128+128 represents the value of sdi_aux_id in the object mask auxiliary picture layer. omi_aux_id_minus128 should be within the range of 0 to 31.
[0307] If the CVS does not contain an SDI SEI message where sdi_aux_id[i] is equal to omi_aux_id_minus128+128 for at least one value of i, then any picture in the CVS should be associated with an OMI SEI message.
[0308] If an AU contains both SDI SEI messages and OMI SEI messages where sdi_aux_id[i] is equal to omi_aux_id_minus128+128 for at least one value of i, then SDI SEI messages should take precedence over OMI SEI messages in the decoding order.
[0309] omi_num_primary_pic_layer_minus1+1 indicates the number of primary picture layers associated with the object mask auxiliary picture layer to which this SEI message applies. The value of omi_num_primary_pic_layer_minus1 should be within the range of 0 to sdi_max_layers_minus1.
[0310] omi_primary_pic_layer_id[i] specifies the nuh_layer_id value of the i-th primary picture layer to which this OMI SEI message applies. If sdi_layer_id[j] is equal to omi_primary_pic_layer_id[i], then the value of sdi_aux_id[j] should be equal to 0, regardless of whether j is greater than or equal to 0 and less than or equal to sid_max_layers_minus1.
[0311] If omi_mask_id_equal_to_aux_sample_value_flag is equal to 1, it indicates that the object mask identifier is equal to the sample value within the mask. If omi_mask_id_equal_to_aux_sample_value_flag is equal to 0, it indicates that the object mask identifier may be different from the sample value within the mask.
[0312] omi_mask_id_length specifies the length in bits of the omi_mask_id[i][j][k] syntax element, if one exists.
[0313] omi_aux_sample_value_length_minus8+8 specifies the bitwise length of the omi_aux_sample_value[i][j][k] syntax element. As a bitstream conformance requirement, the value of omi_aux_sample_value_length_minus8+8 should be equal to BitDepthY.
[0314] omi_mask_id_length_minus8+8 specifies the bitwise length of the omi_mask_id[i][j][k] syntax element.
[0315] If omi_mask_confidence_info_present_flag is equal to 1, it indicates that the omi_mask_confidence[i][j][k] syntax element exists. If omi_mask_confidence_info_present_flag is equal to 0, it indicates that the omi_mask_confidence[i][j][k] syntax element does not exist. As a bitstream conformance requirement, the value of omi_mask_confidence_info_present_flag should be the same for all object_mask_info() syntax structures within CLVS.
[0316] omi_mask_confidence_length_minus1+1 specifies the bitwise length of the omi_mask_confidence[i][j][k] syntax element. As a bitstream conformance requirement, the value of omi_mask_confidence_length_minus1 should be the same for all object_mask_info() syntax structures within CLVS.
[0317] If omi_object_depth_info_present_flag is equal to 1, it indicates that the omi_object_depth[i][j][k] syntax element exists. If omi_object_depth_info_present_flag is equal to 0, it indicates that the omi_object_depth[i][j][k] syntax element does not exist. As a bitstream conformance requirement, the value of omi_object_depth_info_present_flag should be the same for all object_mask_info() syntax structures within CLVS.
[0318] omi_object_depth_length_minus1+1 specifies the bitwise length of the omi_mask_confidence[i][j][k] syntax element. As a bitstream conformance requirement, the value of omi_object_depth_length_minus1 should be the same for all object_mask_info() syntax structures within CLVS.
[0319] If omi_mask_label_info_present_flag is equal to 1, it indicates that omi_mask_label_language_present_flag and the omi_mask_label[i][j][k] syntax element exist. If omi_mask_label_info_present_flag is equal to 0, it indicates that omi_mask_label_language_present_flag and the omi_mask_label[i][j][k] syntax element do not exist.
[0320] If omi_mask_label_language_present_flag is equal to 1, it indicates that the omi_mask_label_language syntax element exists. If omi_mask_label_language_present_flag is equal to 0, it indicates that the omi_mask_label_language syntax element does not exist.
[0321] omi_bit_equal_to_zero should be equal to 0.
[0322] The `omi_mask_label_language` element contains a language tag, as defined in IET FRFC 5646, followed by a null-terminating byte equal to 0x00. The length of the `omi_mask_label_language` syntax element should be 255 bytes or less, excluding the null-terminating byte. If it is not present, the label language is not specified.
[0323] When omi_mask_pic_update_flag[i][j] is equal to 1, it indicates that the mask information of the j-th object mask auxiliary picture associated with the i-th primary picture is signaled. When omi_mask_pic_update_flag[i][j] is equal to 0, it indicates that the mask information of the j-th object mask auxiliary picture associated with the i-th primary picture is not signaled. When there is no mask information for the j-th object mask auxiliary picture associated with the i-th primary picture, a persistence mechanism is used. That is, information is inherited from the last OMI SEI message that signaled the mask information of the j-th object mask auxiliary picture associated with the i-th primary picture.
[0324] omi_num_mask_in_pic_update[i][j] indicates the number of object masks for which information is signaled within the j-th auxiliary picture associated with the i-th primary picture. omi_num_mask_in_pic_update[i][j] should be in the range of 0 or more and (1<<BitDepthY)-1 or less, where BitDepthY is the bit depth of the samples of the luma component. If the current SEI message is the first OMI SEI message within the current CLVS, the variable omiNumMaskInPic[i][j], which indicates the number of object masks within the j-th auxiliary picture associated with the i-th primary picture, is set to omi_num_mask_in_pic_update[i][j].
[0325] The variable numAuxLayer[primaryLayerId] indicates the number of auxiliary picture layers associated with a primary picture layer whose nuh_layer_id is equal to primaryLayerId. The variable associatedAuxLayerId[primaryLayerId][i] indicates the value of nuh_layer_id of the i-th auxiliary picture layer associated with a primary picture layer whose nuh_layer_id is equal to primaryLayerId. numAuxLayer[primaryLayerId] and associatedAuxLayerId[primaryLayerId][i] are derived as follows:
number
[0326] omi_mask_id[i][j][k] indicates the identifier of the k-th object mask in the j-th object mask auxiliary picture associated with the i-th primary picture.
[0327] omi_aux_sample_value[i][j][k] indicates the sample value within the object mask whose identifier is equal to omi_mask_id[i][j][k].
[0328] The variable maskId[i][j][k], which specifies the object mask identifier of the k-th object mask in the j-th object mask auxiliary picture associated with the i-th primary picture in the SEI message, is derived as follows:
number
[0329] If omi_mask_bounding_box_present_flag[i][j][k] is equal to 1, it indicates that the syntax elements omi_mask_top[i][j][k], omi_mask_left[i][j][k], omi_mask_width[i][j][k], and omi_mask_height[i][j][k] exist. If omi_num_mask_in_pic_update[i][j][k] is equal to 0, it indicates that the syntax elements omi_mask_top[i][j][k], omi_mask_left[i][j][k], omi_mask_width[i][j][k], and omi_mask_height[i][j][k] do not exist.
[0330] omi_mask_top[i][j][k], omi_mask_left[i][j][k], omi_mask_width[i][j][k], and omi_mask_height[i][j][k] indicate the coordinates, width, and height of the top-left corner of the bounding box of an object mask whose identifier in the cropped decoded picture is equal to omi_mask_id[i][j][k], respectively, in relation to the fitted cropping window specified in the active SPS.
[0331] The value of omi_mask_left[i][j][k] should be within the range of 0 or greater and (CroppedWidth / SubWidthC-1), where CroppedWidth and SubWidthC are associated with the j-th object mask auxiliary picture associated with the i-th primary picture. If none exist, the value of omi_mask_left[i][j][k] is assumed to be 0.
[0332] The value of omi_mask_top[i][j][k] should be within the range of 0 or greater and (CroppedHeight / SubHeightC-1), where CroppedHeight and SubHeightC are associated with the j-th object mask auxiliary picture associated with the i-th primary picture. If none exist, the value of omi_mask_top[i][j][k] is presumed to be 0.
[0333] The value of omi_mask_width[i][j][k] should be within the range of 0 or greater and (CroppedWidth / SubWidthC - omi_mask_left[i][j][k]). If it does not exist, the value of omi_mask_width[i][j][k] is presumed to be (CroppedWidth / SubWidthC - omi_mask_left[i][j][k]).
[0334] The value of omi_mask_height[i][j][k] should be between 0 and (CroppedHeight / SubHeightC - omi_mask_top[i][j][k]). If it does not exist, the value of omi_mask_height[i][j][k] is presumed to be (CroppedHeight / SubWidthC - omi_mask_top[i][j][k]).
[0335] The identified object mask is located within a bounding box containing a Luma sample whose horizontal coordinates are between SubWidthC*(ConfWinLeftOffset+omi_mask_left[i][j][k]) and SubWidthC*(ConfWinLeftOffset+omi_mask_left[i][j][k]+omi_mask_width[i][j][k])-1, and whose vertical coordinates are between SubHeightC*(ConfWinTopOffset+omi_mask_top[i][j][k]) and SubHeightC*(ConfWinTopOffset+omi_mask_top[i][j][k]+omi_mask_height[i][j][k])-1.
[0336] The variable I[i][j][x][y] is the decoded value of the sample at the relative sample position (x,y) in the j-th object mask auxiliary picture associated with the i-th primary picture. The following process is used to determine each mask region within each auxiliary picture.
number
[0337] If omi_mask_cancel[i][j][k] is equal to 1, the duration of the object mask whose identifier is equal to omi_mask_id[i][j][k] is canceled. If omi_mask_cancel[i][j][k] is equal to 0, it indicates that information about the object mask whose identifier is equal to omi_mask_id[i][j] is signaled.
[0338] As a bitstream conformance requirement, if an omi_mask_id[i][j][k] with a specific value is parsed for the first time within CLVS, the corresponding value of omi_mask_cancel[i][j][k] should be equal to 0.
[0339] omi_mask_confidence[i][j][k] represents the confidence associated with the k-th object mask in the j-th object mask auxiliary picture associated with the i-th primary picture, and 2 -(omi_mask_confidence_length_minus1+1) Expressed in units of omi_mask_confidence[i][j][k], a higher value of omi_mask_confidence[i][j][k] indicates higher confidence. The length of the omi_mask_confidence[i][j][k] syntax element is omi_mask_confidence_length_minus1+1 bits.
[0340] omi_mask_depth[i][j][k] indicates the object depth associated with the k-th object mask in the j-th object mask auxiliary picture associated with the i-th primary picture. Smaller values ββof omi_mask_depth indicate shorter distances to the object. The length of the omi_mask_depth[i][j][k] syntax element is omi_object_depth_length_minus1+1 bits.
[0341] omi_mask_label[i][j][k] specifies the content of the label associated with the k-th object mask in the j-th object mask auxiliary picture associated with the i-th primary picture. The length of the omi_mask_label[i][j][k] syntax element should be 255 bytes or less, excluding the null-terminating byte.
[0342] In some embodiments, methods for detecting objects are also provided. Figure 10 is a schematic diagram showing an exemplary method 1000 for detecting an object according to an embodiment of the present disclosure. As shown in Figure 10, method 1000 may include steps 1002 to 1006 which can be implemented by a decoder (e.g., the image / video decoder 144 in Figure 1, or the device 400 in Figure 4).
[0343] In step 1002, the decoder can receive the bitstream. The bitstream can be encoded according to any of the encoding methods described above.
[0344] In step 1004, the decoder can decode the encoded information of the bitstream to obtain a primary picture and an auxiliary picture. The auxiliary picture may be used to show a mask of an object in the primary picture. The object mask may be represented by sample values ββin the auxiliary picture.
[0345] In step 1006, the decoder can decode the encoded information of the bitstream to obtain Additional Extended Information (SEI) messages associated with the primary picture and applied to the auxiliary picture. As described above, SEI messages can be used to indicate the attributes of the object's mask.
[0346] In some embodiments, a non-temporary computer-readable storage medium for storing the bitstream is also provided. The bitstream may be encoded / decoded according to the methods described above. Figure 11 is a schematic diagram showing the contents of an exemplary bitstream 1100. As shown in Figure 11, the bitstream 1100 may be used to transmit a primary picture 1101, an auxiliary picture 1102, and an Additional Extended Information (SEI) message 1103 (e.g., Figure 5). The auxiliary picture 1102 shows a mask of an object in the primary picture 1101, and the mask of the object may be represented by sample values ββin the auxiliary picture 1102. The SEI message 1103 may be associated with the primary picture 1101 and applied to the auxiliary picture 1102. The SEI message 1103 may be used to indicate attributes of the object mask, as described above.
[0347] In some embodiments, non-temporary computer-readable storage media containing instructions are also provided, and these instructions may be executed by devices for carrying out the methods described above (such as disclosed encoders and decoders). Examples of common forms of non-temporary media include floppy disks, flexible disks, hard disks, solid-state drives, magnetic tapes, or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media having a pattern of holes, RAM, PROMs, EPROMs, FLASH-EPROMs, or any other flash memory, NVRAMs, caches, registers, any other memory chips or cartridges, and networked versions thereof. The device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.
[0348] Embodiments may be further described using the following clauses. 1. A method for encoding a video sequence into a bitstream, wherein the method is: The steps include receiving a video sequence and A step of encoding one or more pictures of the aforementioned video sequence to generate a bitstream, Encoding an auxiliary picture that shows the mask of an object in a primary picture, wherein the mask of the object is represented by the sample values ββof the auxiliary picture. A method comprising the step of generating an additional extended information (SEI) message indicating the attributes of the mask of the object. 2. Generating the aforementioned SEI message The method according to Clause 1, comprising determining a cancellation flag indicating whether the aforementioned SEI message cancels the persistence of a previous SEI message. 3. The attributes of the mask include the individual features of the mask of the object and the common features of the multiple masks indicated by the SEI message. The aforementioned SEI message is generated The method of Clause 2, further comprising determining the common features and the individual features in response to a determination that the cancellation flag indicates that the SEI message does not cancel the persistence of information from the previous SEI message. 4. The attributes of the mask include the individual features of the mask of the object and the common features of the plurality of masks indicated by the SEI message, and the common features are The identifier of the auxiliary picture to which the SEI message applies, The number of bits for encoding an identifier for any of the aforementioned multiple masks, The bit depth of the sample value of the auxiliary picture, A confidence presence flag indicating whether or not the confidence information of the multiple masks is included in the SEI message, In response to the confidence presence flag indicating that the confidence information of the plurality of masks is included in the SEI message, the length of the confidence information of the plurality of masks, A depth presence flag indicating whether or not the depth information of the aforementioned multiple masks is included in the SEI message, In response to the depth presence flag indicating that the depth information of the plurality of masks is included in the SEI message, the length of the depth information of the plurality of masks, A label presence flag indicating whether the label information of the multiple masks is included in the SEI message, In response to the label presence flag indicating that the label information of the plurality of masks is included in the SEI message, a language presence flag indicating whether or not the label language information of the plurality of masks is included in the SEI message, or The method according to Clause 1, comprising at least one of the label language information of the plurality of masks in response to the language presence flag indicating that the label language information of the plurality of masks is included in the SEI message. 5. The method according to Clause 4, wherein the SEI message is applied to a plurality of auxiliary pictures, and the common feature further includes the number of the plurality of auxiliary pictures. 6. The method according to Clause 5, wherein the SEI message includes the individual features generated for a mask represented by the plurality of auxiliary pictures. 7. The method according to any one of the clauses 4 to 6, wherein the SEI message is associated with a plurality of primary pictures corresponding to the auxiliary picture to which the SEI message applies. 8. The method according to Clause 7, wherein the common feature further includes the number of the plurality of primary pictures and the layer identifiers of the plurality of primary pictures. 9. Generating the aforementioned SEI message Determining the second number of auxiliary pictures corresponding to each of the plurality of primary pictures, The method of Clause 8, comprising determining the individual features of the mask represented by the second number of auxiliary pictures for each of the plurality of primary pictures. 10. The SEI message can be generated. Determining whether the mask of the object is different from the previous mask of the object represented by the previous auxiliary picture, The method according to any one of claims 1 to 9, further comprising encoding the attributes of the mask of the object in the SEI message in response to a determination that the mask of the object is different from the previous mask of the object. 11. The method according to clause 10, further comprising the step of skipping the encoding of the attributes of the mask of the object in the SEI message in response to a determination that the mask of the object is the same as the previous mask of the object. 12. The generation of the aforementioned SEI message is The method according to any one of the clauses 1 to 11, further comprising determining a mask cancellation flag indicating whether the mask of the object cancels the persistence of the previous mask of the object. 13. Generating the aforementioned SEI message Determining the bounding box surrounding the mask of the object, The method according to any one of the clauses 1 to 12, comprising encoding the bounding box within the SEI message. 14. The method according to any one of the clauses 1 to 13, wherein the sample values ββof the auxiliary picture are irreversibly encoded. 15. The method according to any one of the clauses 1 to 13, wherein the mask of the object is indicated by the bits of the sample value of the auxiliary picture. 16. The method according to any one of the clauses 1 to 13, wherein the mask of the object is indicated by the sample values ββof the auxiliary picture. 17. The method according to clause 16, wherein the sample value is included in the SEI message. 18. The method according to any one of the clauses 1 to 13, wherein the auxiliary picture includes a plurality of predetermined sample values, and the sample value used to represent the mask of the object is selected from the plurality of predetermined sample values ββin accordance with the difference between the plurality of predetermined sample values. 19. A method for detecting an object, wherein the method is Receiving a bitstream and Decoding the encoded information of the bitstream to obtain a primary picture and an auxiliary picture, wherein the auxiliary picture shows a mask of an object in the primary picture, and the mask of the object is represented by the sample values ββof the auxiliary picture. A method comprising decoding the encoded information of the bitstream to obtain an additional extended information (SEI) message, wherein the SEI message indicates the attributes of the mask of the object. 20. Decoding the encoded information of the bitstream to obtain the SEI message, The method according to clause 19, comprising determining a cancellation flag indicating whether the aforementioned SEI message cancels the persistence of a previous SEI message. 21. The attributes of the mask include the individual features of the mask of the object and the common features of the plurality of masks indicated by the SEI message, Decoding the encoded information of the bitstream to obtain the SEI message is The method according to clause 20, comprising determining the common features and the individual features in response to a determination that the cancellation flag indicates that the SEI message does not cancel the persistence of information from the previous SEI message. 22. The attributes of the mask include the individual features of the mask of the object and the common features of the plurality of masks indicated by the SEI message, and the common features are The identifier of the auxiliary picture to which the SEI message applies, The number of bits for encoding an identifier for any of the aforementioned multiple masks, The bit depth of the sample value of the auxiliary picture, A confidence presence flag indicating whether or not the confidence information of the multiple masks is included in the SEI message, In response to the confidence presence flag indicating that the confidence information of the plurality of masks is included in the SEI message, the length of the confidence information of the plurality of masks, A depth presence flag indicating whether or not the depth information of the aforementioned multiple masks is included in the SEI message, In response to the depth presence flag indicating that the depth information of the plurality of masks is included in the SEI message, the length of the depth information of the plurality of masks, A label presence flag indicating whether the label information of the multiple masks is included in the SEI message, In response to the label presence flag indicating that the label information of the plurality of masks is included in the SEI message, a language presence flag indicating whether or not the label language information of the plurality of masks is included in the SEI message, or The method according to clause 19, comprising at least one of the label language information of the plurality of masks in response to the language presence flag indicating that the label language information of the plurality of masks is included in the SEI message. 23. The method according to Clause 22, wherein the SEI message is applied to a plurality of auxiliary pictures, and the common feature further includes the number of the plurality of auxiliary pictures. 24. The method according to Clause 23, wherein the SEI message includes the individual features generated for a mask represented by the plurality of auxiliary pictures. 25. The method according to any one of the clauses 22 to 24, wherein the SEI message is associated with a plurality of primary pictures corresponding to the auxiliary picture to which the SEI message applies. 26. The method of Clause 25, wherein the common feature further includes the number of the plurality of primary pictures and the layer identifiers of the plurality of primary pictures. 27. Decoding the encoded information of the bitstream to obtain the SEI message, Determining the second number of auxiliary pictures corresponding to each of the plurality of primary pictures, The method according to Clause 26, comprising determining the individual features of the mask represented by the second number of auxiliary pictures for each of the plurality of primary pictures. 28. Decoding the encoded information of the bitstream to obtain the SEI message, The method according to any one of the clauses 19 to 27, further comprising determining a mask cancellation flag indicating whether the mask of the object cancels the persistence of the previous mask of the object. 29. Decoding the encoded information of the bitstream to obtain the SEI message, The method according to any one of the clauses 19 to 28, further comprising determining a bounding box surrounding the mask of the object based on the SEI message. 30. The method according to any one of the clauses 19 to 29, wherein the mask of the object is indicated by the bits of the sample value of the auxiliary picture. 31. The method according to any one of the clauses 19 to 29, wherein the mask of the object is indicated by the sample values ββof the auxiliary picture. 32. The method according to clause 31, wherein the sample value is included in the SEI message. 33. Decoding the encoded information of the bitstream to obtain the SEI message, The method according to clause 32, further comprising determining that the sample value of the auxiliary picture represents the mask having the same identifier or the closest identifier value. 34. The method according to any one of the clauses 19 to 29, wherein the auxiliary picture includes a plurality of predetermined sample values, and the sample value used to represent the mask of the object is selected from the plurality of predetermined sample values ββin accordance with the difference between the plurality of predetermined sample values. 35. A non-temporary computer-readable storage medium for storing a video bitstream, wherein the bitstream is A primary picture containing an object, An auxiliary picture showing the mask of the object, wherein the mask of the object is represented by the sample values ββof the auxiliary picture, A non-temporary computer-readable storage medium including an additional extended information (SEI) message indicating the attributes of the mask of the object. 36. A non-temporary computer-readable storage medium as described in Clause 35, wherein the SEI message includes a cancellation flag indicating whether the SEI message cancels the persistence of a previous SEI message. 37. A non-temporary computer-readable storage medium according to Clause 36, wherein, in response to the cancellation flag indicating that the SEI message does not cancel the persistence of information from the previous SEI message, the attributes of the mask include the individual characteristics of the mask of the object and the common characteristics of a plurality of masks indicated by the SEI message. 38. The attributes of the mask include the individual features of the mask of the object and the common features of the plurality of masks indicated by the SEI message, and the common features are The identifier of the auxiliary picture to which the SEI message applies, The number of bits for encoding an identifier for any of the aforementioned multiple masks, The bit depth of the sample value of the auxiliary picture, A confidence presence flag indicating whether or not the confidence information of the multiple masks is included in the SEI message, In response to the confidence presence flag indicating that the confidence information of the plurality of masks is included in the SEI message, the length of the confidence information of the plurality of masks, A depth presence flag indicating whether or not the depth information of the aforementioned multiple masks is included in the SEI message, In response to the depth presence flag indicating that the depth information of the plurality of masks is included in the SEI message, the length of the depth information of the plurality of masks, A label presence flag indicating whether the label information of the multiple masks is included in the SEI message, In response to the label presence flag indicating that the label information of the plurality of masks is included in the SEI message, a language presence flag indicating whether or not the label language information of the plurality of masks is included in the SEI message, or A non-temporary computer-readable storage medium according to Clause 35, which includes at least one of the label language information of the plurality of masks in response to the language presence flag indicating that the label language information of the plurality of masks is included in the SEI message. 39. A non-temporary computer-readable storage medium as described in Clause 38, wherein the SEI message is applied to a plurality of auxiliary pictures, and the common feature further includes the number of the plurality of auxiliary pictures. 40. A non-temporary computer-readable storage medium according to Clause 39, wherein the SEI message includes the individual features generated for a mask represented by the plurality of auxiliary pictures. 41. A non-temporary computer-readable storage medium as described in any one of clauses 38 to 40, wherein the SEI message is associated with a plurality of primary pictures corresponding to the auxiliary picture to which the SEI message applies. 42. A non-temporary computer-readable storage medium as described in Clause 41, wherein the common features further include the number of the plurality of primary pictures and layer identifiers for the plurality of primary pictures. 43. The aforementioned SEI message further states An operation to determine the second number of auxiliary pictures corresponding to each of the plurality of primary pictures, A non-temporary computer-readable storage medium according to Clause 42, generated based on an operation that determines the individual characteristics of the mask, which are represented by the second number of auxiliary pictures for each of the plurality of primary pictures. 44. The aforementioned SEI message further states: An operation to determine whether the mask of the object is different from the previous mask of the object represented by the previous auxiliary picture, A non-temporary computer-readable storage medium according to any one of clauses 35 to 43, generated based on an operation to encode the attributes of the mask of the object in the SEI message in response to a determination that the mask of the object is different from the previous mask of the object. 45. The aforementioned SEI message further states A non-temporary computer-readable storage medium according to Clause 44, generated based on an operation that skips encoding the attributes of the mask of the object in the SEI message in response to a determination that the mask of the object is the same as the previous mask of the object. 46. ββThe aforementioned SEI message further states A non-temporary computer-readable storage medium according to any one of the clauses 35 to 45, generated based on an operation that determines a mask cancellation flag indicating whether the mask of the object cancels the persistence of the previous mask of the object. 47. The aforementioned SEI message further states An operation to determine the bounding box surrounding the mask of the object, A non-temporary computer-readable storage medium according to any one of clauses 35 to 46, generated based on the operation of encoding the bounding box within the SEI message. 48. A non-temporary computer-readable storage medium according to any one of the clauses 35 to 47, in which the sample values ββof the auxiliary picture are irreversibly encoded. 49. A non-temporary computer-readable storage medium according to any one of the clauses 35 to 47, wherein the mask of the object is indicated by the bits of the sample value of the auxiliary picture. 50. A non-temporary computer-readable storage medium according to any one of the clauses 35 to 47, wherein the mask of the object is indicated by the sample values ββof the auxiliary picture. 51. A non-temporary computer-readable storage medium as described in Clause 50, in which the sample value is included in the SEI message. 52. A non-temporary computer-readable storage medium according to any one of the clauses 35 to 47, wherein the auxiliary picture includes a plurality of predetermined sample values, and the sample values ββused to represent the mask of the object are selected from the plurality of predetermined sample values ββin accordance with the difference between the plurality of predetermined sample values.
[0349] It should be noted that, in this specification, relational terms such as βfirstβ and βsecondβ are used solely to distinguish one entity or action from another, and do not require or imply any actual relationship or order between these entities or actions. Furthermore, the words βinclude,β βhave,β βcontain,β βinclude,β and other similar forms are semantically equivalent, and the items following any of these words are not exhaustive, nor are they limited to, the items listed.
[0350] As used herein, the term βorβ shall, unless otherwise specified, encompass all possible combinations unless impractical. For example, if it is stated that a database may contain A or B, then unless otherwise specified or impractical, the database may contain A, or B, or A and B. As a second example, if it is stated that a database may contain A, B, or C, then unless otherwise specified or impractical, the database may contain A, or B, or C, or A and B, or A and C, or B and C, or A, B and C.
[0351] It will be understood that the embodiments described above can be implemented by hardware, software (program code), or a combination of hardware and software. When implemented by software, it may be stored in the computer-readable medium described above. When the software is executed by a processor, it can perform the disclosed methods. The computing units and other functional units described in this disclosure can be implemented by hardware, software, or a combination of hardware and software. Those skilled in the art will understand that multiple modules / units described above may be combined into a single module / unit, or each of the modules / units described above may be further divided into multiple submodules / subunits.
[0352] In the aforementioned specification, embodiments have been described with reference to numerous specific details, which may vary from implementation to implementation. Certain modifications and changes can be made to the described embodiments. Other embodiments may become apparent to those skilled in the art by examining the specification and practices of the invention disclosed herein. The specification and examples are for illustrative purposes only, and the true scope and spirit of the invention are intended to be shown by the following claims. Furthermore, the order of steps shown in the drawings is for illustrative purposes only and is not intended to limit the sequence of any particular steps. Thus, those skilled in the art will understand that these steps can be performed in different orders while implementing the same method.
[0353] The drawings and specification disclose exemplary embodiments; however, many variations and modifications can be made to these embodiments. Therefore, whereever specific terms are used, they are used in a general and descriptive sense only and are not intended to be limiting.
Claims
1. A method for encoding a video sequence into a bitstream, wherein the method is The steps include receiving a video sequence and A step of encoding one or more pictures of the aforementioned video sequence to generate a bitstream, Encoding an auxiliary picture that shows the mask of an object in a primary picture, wherein the mask of the object is represented by the sample values ββof the auxiliary picture. A method comprising the step of generating an additional extended information (SEI) message indicating the attributes of the mask of the object.
2. The generation of the aforementioned SEI message is The method according to claim 1, comprising determining a cancellation flag indicating whether the SEI message cancels the persistence of a previous SEI message.
3. The attributes of the mask include the individual features of the mask of the object and the common features of the multiple masks indicated by the SEI message, The generation of the aforementioned SEI message is The method of claim 2, further comprising determining the common features and the individual features in response to a determination that the cancellation flag indicates that the SEI message does not cancel the persistence of information from the previous SEI message.
4. The attributes of the mask include the individual features of the mask of the object and the common features of the plurality of masks indicated by the SEI message, and the common features are The identifier of the auxiliary picture to which the SEI message is applied, The number of bits for encoding an identifier for any of the aforementioned multiple masks, The bit depth of the sample value of the auxiliary picture, A confidence presence flag indicating whether or not the confidence information of the multiple masks is included in the SEI message, In response to the confidence presence flag indicating that the confidence information of the plurality of masks is included in the SEI message, the length of the confidence information of the plurality of masks, A depth presence flag indicating whether or not the depth information of the aforementioned multiple masks is included in the SEI message, In response to the depth presence flag indicating that the depth information of the plurality of masks is included in the SEI message, the length of the depth information of the plurality of masks, A label presence flag indicating whether or not the label information of the aforementioned multiple masks is included in the SEI message. In response to the label presence flag indicating that the label information of the plurality of masks is included in the SEI message, a language presence flag indicating whether or not the label language information of the plurality of masks is included in the SEI message is displayed, or The method according to claim 1, comprising at least one of the label language information of the plurality of masks in response to the language presence flag indicating that the label language information of the plurality of masks is included in the SEI message.
5. The generation of the aforementioned SEI message is Determining whether the mask of the object is different from the previous mask of the object represented by the previous auxiliary picture, The method according to claim 1, further comprising encoding the attributes of the mask of the object in the SEI message in response to a determination that the mask of the object is different from the previous mask of the object.
6. The method of claim 5, further comprising skipping encoding the attributes of the mask of the object in the SEI message in response to a determination that the mask of the object is the same as the previous mask of the object.
7. The generation of the aforementioned SEI message is The method according to claim 1, further comprising determining a mask cancellation flag indicating whether the mask of the object cancels the persistence of the previous mask of the object.
8. The generation of the aforementioned SEI message is Determining the bounding box surrounding the mask of the object, The method according to claim 1, further comprising encoding the bounding box within the SEI message.
9. The method according to claim 1, wherein the sample values ββof the auxiliary picture are lossily encoded.
10. The method according to claim 1, wherein the auxiliary picture includes a plurality of predetermined sample values, and the sample value used to represent the mask of the object is selected from the plurality of predetermined sample values ββaccording to the difference in values ββbetween the plurality of predetermined sample values.
11. A method for detecting an object, wherein the method is Receiving a bitstream and Decoding the encoded information of the bitstream to obtain a primary picture and an auxiliary picture, wherein the auxiliary picture shows a mask of an object in the primary picture, and the mask of the object is represented by the sample values ββof the auxiliary picture. A method comprising decoding the encoded information of the bitstream to obtain an additional extended information (SEI) message, wherein the SEI message indicates the attributes of the mask of the object.
12. Decoding the encoded information of the bitstream to obtain the SEI message is The method according to claim 11, comprising determining a cancellation flag indicating whether the SEI message cancels the persistence of a previous SEI message.
13. The attributes of the mask include the individual features of the mask of the object and the common features of the multiple masks indicated by the SEI message, Decoding the encoded information of the bitstream to obtain the SEI message is The method according to claim 12, comprising determining the common features and the individual features in response to a determination that the cancellation flag indicates that the SEI message does not cancel the persistence of information from the previous SEI message.
14. The attributes of the mask include the individual features of the mask of the object and the common features of the plurality of masks indicated by the SEI message, and the common features are The identifier of the auxiliary picture to which the SEI message is applied, The number of bits for encoding an identifier for any of the aforementioned multiple masks, The bit depth of the sample value of the auxiliary picture, A confidence presence flag indicating whether or not the confidence information of the multiple masks is included in the SEI message, In response to the confidence presence flag indicating that the confidence information of the plurality of masks is included in the SEI message, the length of the confidence information of the plurality of masks, A depth presence flag indicating whether or not the depth information of the aforementioned multiple masks is included in the SEI message, In response to the depth presence flag indicating that the depth information of the plurality of masks is included in the SEI message, the length of the depth information of the plurality of masks, A label presence flag indicating whether or not the label information of the aforementioned multiple masks is included in the SEI message. In response to the label presence flag indicating that the label information of the plurality of masks is included in the SEI message, a language presence flag indicating whether or not the label language information of the plurality of masks is included in the SEI message is displayed, or The method according to claim 11, comprising at least one of the label language information of the plurality of masks in response to the language presence flag indicating that the label language information of the plurality of masks is included in the SEI message.
15. Decoding the encoded information of the bitstream to obtain the SEI message is The method according to claim 11, further comprising determining a mask cancellation flag indicating whether the mask of the object cancels the persistence of the previous mask of the object.
16. Decoding the encoded information of the bitstream to obtain the SEI message is The method according to claim 11, further comprising determining a bounding box surrounding the mask of the object based on the SEI message.
17. The method according to claim 11, wherein the mask of the object is represented by the sample value bits of the auxiliary picture.
18. The method according to claim 11, wherein the mask of the object is represented by the sample values ββof the auxiliary picture.
19. The method according to claim 11, wherein the auxiliary picture includes a plurality of predetermined sample values, and the sample value used to represent the mask of the object is selected from the plurality of predetermined sample values ββaccording to the difference in values ββbetween the plurality of predetermined sample values.
20. A non-temporary computer-readable storage medium for storing a video bitstream, wherein the bitstream is A primary picture containing an object, An auxiliary picture showing the mask of the object, wherein the mask of the object is represented by the sample values ββof the auxiliary picture, A non-temporary computer-readable storage medium including an additional extended information (SEI) message indicating the attributes of the mask of the object.