Object mask information for supplemental enhancement information messages

By generating a Supplementary Enhancement Information (SEI) message of object mask information, the problem of low efficiency in processing object mask information in existing video coding technology is solved, and more efficient video compression and object recognition are achieved.

CN120604515APending Publication Date: 2025-09-05ALIBABA (CHINA) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480009870.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-04-02
Filing Date
2024-04-11
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing video coding technologies are inefficient in processing object mask information and are unable to effectively utilize object mask information for enhanced coding, resulting in insufficient video compression efficiency.

Method used

A supplemental enhancement information (SEI) message of object mask information is generated by receiving a bitstream and decoding a main image and an auxiliary image, using sample values ​​of the auxiliary image to represent the object mask, and generating an SEI message indicating object mask properties during encoding.

Benefits of technology

It improves the compression efficiency of video encoding, enhances the ability of object recognition and tracking, and improves the quality and efficiency of video processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120604515A_ABST
    Figure CN120604515A_ABST
Patent Text Reader

Abstract

Methods and apparatus are provided for processing video data by supplementing enhancement information (SEI) messages using object mask information (OMI). An exemplary encoding method includes receiving a video sequence; and encoding one or more images of the video sequence to generate a bitstream, including: encoding a secondary image indicating a mask of an object in a primary image, the mask of the object being characterized by sample values of the secondary image; and generating a supplemental enhancement information (SEI) message indicating an attribute of the mask of the object.
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This disclosure claims priority to U.S. Provisional Application No. 63 / 495,546, filed April 11, 2023, U.S. Provisional Application No. 63 / 587,750, filed October 4, 2023, U.S. Provisional Application No. 63 / 615,294, filed December 28, 2023, and U.S. Application No. 18 / 624,636, filed April 2, 2024, all of which are incorporated herein by reference in their entirety. Technical Field

[0002] The present disclosure relates generally to video processing, and more particularly, to methods and apparatus for sending Object Mask Information (OMI) Supplemental Enhancement Information (SEI) messages. Background Art

[0003] A video is a set of static images (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, the video can be compressed before storage or transmission, and decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are many video coding formats that use standardized video coding techniques. The most common ones are based on prediction, transform, quantization, entropy coding, and loop filtering. Standardization organizations have developed video coding standards that specify specific video coding formats, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard. As more and more advanced video coding technologies are adopted in video standards, the coding efficiency of new video coding standards is also getting higher and higher. Summary of the Invention

[0004] Embodiments of the present disclosure provide a method and apparatus for transmitting an Object Mask Information (OMI) Supplemental Enhancement Information (SEI) message.

[0005] According to some exemplary embodiments, a method for detecting an object is provided, the method comprising: receiving a bitstream; decoding encoded information of the bitstream to obtain a main image and an auxiliary image, wherein the auxiliary image indicates a mask of an object in the main image, and the mask of the object is represented by sample values ​​of the auxiliary image; and decoding the encoded information of the bitstream to obtain a supplemental enhancement information (SEI) message, the SEI message indicating properties of the mask of the object.

[0006] According to some exemplary embodiments, a coding method is provided, comprising: receiving a video sequence; and encoding one or more images of the video sequence to generate a bitstream, comprising: encoding an auxiliary image indicating a mask of an object in a main image, the mask of the object being represented by sample values ​​of the auxiliary image; and generating a supplemental enhancement information (SEI) message indicating attributes of the mask of the object.

[0007] According to some exemplary embodiments, a non-transitory computer-readable storage medium storing a video bitstream is provided. The bitstream includes: a main image containing an object; an auxiliary image indicating a mask of the object, the mask of the object being represented by sample values ​​of the auxiliary image; and a supplemental enhancement information (SEI) message indicating properties of the mask of the object. BRIEF DESCRIPTION OF THE DRAWINGS

[0008]

[0014] Embodiments and aspects of the present disclosure are illustrated in the following detailed description and accompanying drawings.The various features shown in the drawings are not drawn to scale.

[0009] Figure 1 is a schematic diagram illustrating an exemplary system for preprocessing and encoding image data according to some embodiments of the present disclosure.

[0010] Figure 2A is a schematic diagram illustrating an exemplary encoding process of a hybrid video coding system consistent with an embodiment of the present disclosure

[0011] Figure 2B is a schematic diagram illustrating another exemplary encoding process of a hybrid video coding system consistent with an embodiment of the present disclosure.

[0012] Figure 3A is a schematic diagram illustrating an exemplary decoding process of a hybrid video coding system consistent with an embodiment of the present disclosure.

[0013] Figure 3B is a schematic diagram illustrating another exemplary decoding process of a hybrid video coding system consistent with an embodiment of the present disclosure.

[0014] Figure 4 is a block diagram of an exemplary apparatus for preprocessing or encoding image data according to some embodiments of the present disclosure.

[0015] Figure 5 is a syntax diagram illustrating an exemplary object mask information (OMI) SEI message according to some embodiments of the present disclosure.

[0016] Figure 6 is a diagram illustrating an exemplary method for encoding a video sequence into a bitstream consistent with an embodiment of the present disclosure.

[0017] Figure 7 is a schematic diagram illustrating exemplary primary and auxiliary images consistent with an embodiment of the present disclosure.

[0018] Figure 8 is a schematic diagram illustrating sub-steps of an exemplary method for encoding a video sequence into a bitstream consistent with an embodiment of the present disclosure.

[0019] Figure 9 An exemplary binary representation of a sample value p[x][y] is shown according to some embodiments of the present disclosure.

[0020] Figure 10 is a schematic diagram illustrating an exemplary method for detecting an object consistent with an embodiment of the present disclosure.

[0021] Figure 11 is a schematic diagram showing the contents of an exemplary bitstream. DETAILED DESCRIPTION

[0022] Reference will now be made in detail to exemplary embodiments, examples of which are shown in the accompanying drawings. The following description refers to the accompanying drawings, in which, unless otherwise indicated, the same numbers in different figures represent the same or similar elements. The embodiments set forth in the following description of exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of devices and methods consistent with the relevant aspects of the present invention described in the appended claims. Specific aspects of the present disclosure are described in more detail below. In the event of a conflict with terms and / or definitions incorporated by reference, the terms and definitions provided herein shall prevail.

[0023] The Joint Video Experts Group (JVET) of the ITU-T Video Coding Experts Group (ITU-T VCEG) and the ISO / IEC Moving Picture Experts Group (ISO / IEC MPEG) is currently developing the Versatile Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, VVC aims to achieve the same subjective quality as HEVC / H.265 using half the bandwidth of HEVC / H.265.

[0024] To achieve this goal, JVET has been continuously developing technologies that surpass HEVC since 2015 using the Joint Exploration Model (JEM) reference software. With the incorporation of coding technologies into JEM, JEM has achieved significantly higher coding performance than HEVC. In October 2017, VCEG and MPEG issued a joint call for proposals (CfP), formally launching the development of a next-generation video compression standard that would surpass HEVC. Responses to the CfP were evaluated at the JVET meeting in San Diego in April 2018, and the formal development process for the VVC standard began in April 2018.

[0025] Since April 2018, the VVC standard has been progressing smoothly and continues to incorporate more coding techniques to provide better compression performance. VVC follows the hybrid video coding system used by modern video compression standards such as HEVC, H.264 / AVC, MPEG2, and H.263.

[0026] Figure 1 is a block diagram illustrating a system 100 for preprocessing and encoding image data according to some disclosed embodiments. Image data may include an image (also referred to as an "image" or "frame"), multiple images, or a video. An image is a static image. Multiple images may or may not be spatially or temporally related. A video is a set of images arranged in a temporal sequence.

[0027] like Figure 1 As shown, system 100 includes a source device 120 that provides encoded video data for subsequent decoding by a destination device 140. Consistent with the disclosed embodiments, each of source device 120 and destination device 140 may include any of a variety of devices, including a desktop computer, a notebook (e.g., laptop) computer, a server, a tablet computer, a set-top box, a mobile phone, a vehicle, a camera, an image sensor, a robot, a television, a camera, a wearable device (e.g., a smartwatch or wearable camera), a display device, a digital media player, a video game console, a video streaming device, etc. Source device 120 and destination device 140 may be configured to support wireless or wired communication.

[0028] refer to Figure 1, the source device 120 may include an image / video preprocessor 122, an image / video encoder 124, and an output interface 126. The target device 140 may include an input interface 142, an image / video decoder 144, and one or more machine vision applications 146. The image / video preprocessor 122 preprocesses image data, i.e., one or more images or one or more videos, and generates an input bitstream for the image / video encoder 124. The image / video encoder 124 encodes the input bitstream and outputs an encoded bitstream 162 via the output interface 126. The encoded bitstream 162 is transmitted over the communication medium 160 and is received by the input interface 142. The image / video decoder 144 then decodes the encoded bitstream 162 to generate decoded data, which can be used by the machine vision application 146.

[0029] More specifically, the source device 120 may further include various devices (not shown) for providing source image data to be pre-processed by the image / video pre-processor 122. The devices for providing source image data may include an image / video acquisition device, such as a camera, an image / video archive or storage device containing previously acquired images / videos, or an image / video feed interface for receiving images / videos from an image / video content provider.

[0030] The image / video encoder 124 and the image / video decoder 144 can each be implemented as any of a variety of suitable encoder or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When encoding or decoding is partially implemented in software, the image / video encoder 124 or the image / video decoder 144 can store instructions for the software in a suitable, non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform techniques consistent with the present disclosure. Each of the image / video encoder 124 or the image / video decoder 144 can be included in one or more encoders or decoders, any of which can be integrated as part of a composite encoder / decoder (CODEC) in the corresponding device.

[0031] The image / video encoder 124 and the image / video decoder 144 may operate according to any video coding standard, such as Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), AOMedia Video 1 (AV1), Joint Photographic Experts Group (JPEG), Moving Picture Experts Group (MPEG), etc. Alternatively, the image / video encoder 124 and the image / video decoder 144 may be custom devices that do not conform to existing standards. Although Figure 1 Not shown, but in some embodiments, the image / video encoder 124 and the image / video decoder 144 may each be integrated with an audio encoder and decoder, and may include appropriate MUX-DEMUX units, or other hardware and software, to handle the encoding of both audio and video in a common data stream or in separate data streams.

[0032] Output interface 126 may include any type of medium or device capable of transmitting encoded bitstream 162 from source device 120 to destination device 140. For example, output interface 126 may include a transmitter or transceiver configured to transmit encoded bitstream 162 directly from source device 120 to destination device 140 in real time. Encoded bitstream 162 may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to destination device 140.

[0033] The communication medium 160 may include a transient medium, such as a wireless broadcast or a wired network transmission. For example, the communication medium 160 may include a radio frequency (RF) spectrum or one or more physical transmission lines (e.g., cables). The communication medium 160 may form a portion of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. In some embodiments, the communication medium 160 may include a router, a switch, a base station, or any other device that may be used to facilitate communication from the source device 120 to the target device 140. For example, a network server (not shown) may receive the encoded bit stream 162 from the source device 120 and, for example, provide the encoded bit stream 162 to the target device 140 via a network transmission.

[0034] Communication medium 160 may also be in the form of a storage medium (e.g., a non-transitory storage medium) such as a hard drive, a flash drive, an optical disc, a digital video disc, a Blu-ray disc, a volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded image data. In some embodiments, a computing device at a media production facility, such as an optical disc imprinting facility, may receive the encoded image data from source device 120 and produce an optical disc containing the encoded video data.

[0035] Input interface 142 may include any type of medium or device capable of receiving information from communication medium 160. The received information includes encoded bitstream 162. For example, input interface 142 may include a receiver or transceiver configured to receive encoded bitstream 162 in real time.

[0036] The machine vision application 146 includes various hardware and / or software for using the decoded image data generated by the image / video decoder 144. For example, the machine vision application 146 may include a display device that displays the decoded image data to a user, and may include any of a variety of display devices, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices. As another example, the machine vision application 146 may include one or more processors configured to use the decoded image data to perform various machine vision applications, such as object recognition and tracking, facial recognition, image matching, image / video search, augmented reality, robotic vision and navigation, autonomous driving, three-dimensional structure construction, stereo correspondence, motion tracking, etc.

[0037] Next, combine Figures 2A-2B and Figures 3A-3B Describes exemplary image data encoding and decoding techniques (e.g., by Figure 1 those techniques implemented by the encoder 124 and decoder 144).

[0038] Figure 2A Schematic diagram of an example encoding process 200A consistent with an embodiment of the present disclosure is shown. For example, the encoding process 200A may be performed by an encoder, such as Figure 1 The image / video encoder 124 in FIG. Figure 2A As shown, the encoder can encode a video sequence 202 into a video bitstream 228 according to process 200A. The video sequence 202 may include a set of images (referred to as "original images") arranged in a time sequence. The encoder can divide each original image of the video sequence 202 into a plurality of basic processing units, a plurality of basic processing sub-units, or a plurality of regions for processing. In some embodiments, the encoder can perform process 200A at the level of a basic processing unit for each original image of the video sequence 202. For example, the encoder can perform process 200A in an iterative manner, wherein the encoder can encode a basic processing unit in one iteration of process 200A. In some embodiments, the encoder can perform process 200A in parallel for multiple regions of each original image of the video sequence 202.

[0039] exist Figure 2A2, the encoder may feed the basic processing units of the original images of the video sequence 202 (referred to as "original BPUs") to a prediction stage 204 to generate prediction data 206 and a prediction BPU 208. The encoder may subtract the prediction BPU 208 from the original BPU to generate a residual BPU 210. The encoder may feed the residual BPU 210 to a transform stage 212 and a quantization stage 214 to generate quantized transform coefficients 216. The encoder may feed the prediction data 206 and the quantized transform coefficients 216 to a binary encoding stage 226 to generate a video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as a "forward path." During process 200A, after the quantization stage 214, the encoder may feed the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The encoder may add the reconstruction residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224, which is used in the prediction stage 204 for the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A may be referred to as a "reconstruction path." The reconstruction path may be used to ensure that both the encoder and the decoder use the same reference data for prediction.

[0040] The encoder may iteratively perform process 200A to encode each original BPU of the original image (in the forward path) and generate a prediction reference 224 for encoding the next original BPU of the original image (in the reconstruction path). After encoding all original BPUs of the original image, the encoder may proceed to encode the next image in the video sequence 202.

[0041] Referring to process 200A, the encoder may receive a video sequence generated by a video capture device (e.g., a camera) 202. As used herein, the term "receive" may refer to any action of receiving, inputting, acquiring, retrieving, obtaining, reading, accessing, or inputting data in any manner.

[0042] At the current iteration, at the prediction stage 204, the encoder may receive the original BPU and the prediction reference 224 and perform a prediction operation to generate the prediction data 206 and the predicted BPU 208. The prediction reference 224 may be generated from the reconstruction path of the previous iteration of the process 200A. The purpose of the prediction stage 204 is to reduce information redundancy by extracting the prediction data 206 from the prediction data 206 and the prediction reference 224 that can be used to reconstruct the original BPU into the predicted BPU 208.

[0043] Ideally, the predicted BPU 208 would be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 typically differs slightly from the original BPU. To account for such differences, after generating the predicted BPU 208, the encoder may subtract the predicted BPU 208 from the original BPU to generate a residual BPU 210. For example, the encoder may subtract the values ​​of the pixels of the predicted BPU 208 (e.g., grayscale values ​​or RGB values) from the values ​​of the corresponding pixels of the original BPU. Each pixel of the residual BPU 210 may have a residual value generated by this subtraction between the corresponding pixel values ​​of the original BPU and the predicted BPU 208. Compared to the original BPU, the predicted data 206 and the residual BPU 210 may have fewer bits, but they can be used to reconstruct the original BPU without noticeable quality degradation. Thus, the original BPU is compressed.

[0044] To further compress the residual BPU 210, in the transform stage 212, the encoder can reduce the spatial redundancy of the residual BPU 210 by decomposing it into a set of two-dimensional "basis patterns", each of which is associated with a "transform coefficient". The basis patterns can have the same size (e.g., the size of the residual BPU 210). Each basis pattern can represent a frequency-varying component of the residual BPU 210 (e.g., the frequency of luminance variation). No basis pattern can be reproduced by any combination (e.g., linear combination) of any other basis patterns. In other words, the decomposition can decompose the variation of the residual BPU 210 into the frequency domain. This decomposition is analogous to the discrete Fourier transform of a function, where the basis patterns are analogous to the basis functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transform coefficients are analogous to the coefficients associated with the basis functions.

[0045] Different transform algorithms can use different base patterns. Various transform algorithms can be used in the transform stage 212, such as discrete cosine transform, discrete sine transform, and the like. The transform in the transform stage 212 is reversible. That is, the encoder can restore the residual BPU 210 by performing the inverse operation of the transform (referred to as an "inverse transform"). For example, to restore a pixel of the residual BPU 210, the inverse transform may involve multiplying the value of the corresponding pixel in the base pattern by the corresponding correlation coefficient and adding the products to produce a weighted sum. For video coding standards, both the encoder and decoder can use the same transform algorithm (and therefore the same base pattern). Therefore, the encoder can only record the transform coefficients, and the decoder can reconstruct the residual BPU 210 based on the transform coefficients without receiving the base pattern from the encoder. Compared to the residual BPU 210, the transform coefficients may have fewer bits, but they can be used to reconstruct the residual BPU 210 without significant quality degradation. Consequently, the residual BPU 210 is further compressed.

[0046] The encoder can further compress the transform coefficients during the quantization stage 214. During the transform process, different basis patterns can represent different frequencies of variation (e.g., the frequency of brightness variation). Because the human eye is generally better at discerning low-frequency variations, the encoder can ignore information about high-frequency variations without causing noticeable quality degradation in decoding. For example, during the quantization stage 214, the encoder can generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value (called a "quantization parameter") and rounding the quotient to the nearest integer. After this operation, some transform coefficients of high-frequency basis patterns can be converted to zero, while the transform coefficients of low-frequency basis patterns can be converted to smaller integers. The encoder can ignore quantized transform coefficients 216 with zero values, thereby further compressing the transform coefficients. The quantization process described above is also reversible, where the quantized transform coefficients 216 can be reconstructed into the transform coefficients in the inverse operation of quantization (called "inverse quantization").

[0047] Because the encoder ignores the remainder of this division in the rounding operation, the quantization stage 214 may be lossy. In general, the quantization stage 214 may constitute the largest source of information loss in process 200A. The greater the information loss, the fewer bits may be required to quantize the transform coefficients 216. To achieve different degrees of information loss, the encoder may use different values ​​for the quantization parameter or any other parameter of the quantization process.

[0048] In the binary encoding stage 226, the encoder may encode the prediction data 206 and the quantized transform coefficients 216 using a binary encoding technique, such as, for example, entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder may also encode other information in the binary encoding stage 226, such as, for example, the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the transform type of the transform stage 212, the parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bit rate control parameters), etc. The encoder may use the output data of the binary encoding stage 226 to generate a video bitstream 228. In some embodiments, the video bitstream 228 may be further packaged for network transmission.

[0049] Referring to the reconstruction path of process 200A, the encoder may perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients in the inverse quantization stage 218. In the inverse transform stage 220, the encoder may generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224 to be used in the next iteration of process 200A.

[0050] It should be noted that other variations of process 200A may be used to encode video sequence 202. In some embodiments, the stages of process 200A may be performed by the encoder in a different order. In some embodiments, one or more stages of process 200A may be combined into a single stage. In some embodiments, a single stage of process 200A may be divided into multiple stages. For example, transform stage 212 and quantization stage 214 may be combined into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, process 200A may be omitted. Figure 2A one or more stages in a process.

[0051] Figure 2B Schematic diagram of another example encoding process 200B consistent with an embodiment of the present disclosure is shown. For example, the encoding process 200B may be performed by an encoder, such as Figure 12. Process 200B may be modified from process 200A. For example, process 200B may be used by an encoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared to process 200A, the forward path of process 200B additionally includes a mode decision stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes a loop filter stage 232 and a buffer 234.

[0052] In general, prediction techniques can be divided into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-image prediction or "intra-frame prediction") can use pixels from one or more coded adjacent BPUs in the same image to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include the adjacent BPUs. The spatial prediction can reduce the inherent spatial redundancy of the image. Temporal prediction (e.g., inter-image prediction or "inter-frame prediction") can use regions from one or more coded images to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include the coded image. The temporal prediction can reduce the inherent temporal redundancy of the image.

[0053] Referring to process 200B, in the forward path, the encoder performs prediction operations in a spatial prediction stage 2042 and a temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder may perform intra-frame prediction. For a particular original BPU of a particular picture being encoded, the prediction reference 224 may include one or more neighboring BPUs that have been encoded (in the forward path) and reconstructed (in the reconstruction path) in the same picture. The encoder may generate a predicted BPU 208 by extrapolating the neighboring BPUs. The extrapolation techniques may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, etc. In some embodiments, the encoder may perform the extrapolation at the pixel level, for example, extrapolating the corresponding pixel value of each pixel of the predicted BPU 208. The neighboring BPU used for extrapolation can be positioned in various directions relative to the original BPU, such as vertically (e.g., on top of the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., below left, below right, above left, or above right of the original BPU), or any direction defined in the video coding standard used. For intra-frame prediction, the prediction data 206 may include, for example, the position (e.g., coordinates) of the neighboring BPU used, the size of the neighboring BPU used, parameters of the extrapolation, the direction of the neighboring BPU used relative to the original BPU, etc.

[0054] For another example, in the temporal prediction stage 2044, the encoder may perform the inter-frame prediction. For a certain original BPU of the current image, the prediction reference 224 may include one or more images (referred to as "reference images") that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference images may be encoded and reconstructed on a BPU-by-BPU basis. For example, the encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a reconstructed BPU. After all reconstructed BPUs for the same image are generated, the encoder may generate the reconstructed image as a reference image. The encoder may perform a "motion estimation" operation to search for a matching area within a certain range (referred to as a "search window") of the reference image. The position of the search window in the reference image may be determined based on the position of the original BPU in the current image. For example, the search window may be centered at a location in the reference image that has the same coordinates as the original BPU in the current image and may extend outward by a predetermined distance. When the encoder identifies an area similar to the original BPU in the search window (for example, by using a pixel recursive algorithm, a block matching algorithm, etc.), the encoder can determine such an area as a matching area. The matching area can have different specifications from the original BPU (for example, smaller than, equal to, larger than, or having a different shape). Because the reference image and the current image are temporally separated in the time axis, it can be considered that the matching area "moves" to the position of the original BPU over time. The encoder can record the direction and distance of this movement as a "motion vector". When multiple reference images are used, the encoder can search for matching areas for each reference image and determine the associated motion vector of the matching area. In some embodiments, the encoder can assign weights to the pixel values ​​of the matching areas of each matching reference image.

[0055] The motion estimation may be used to identify various types of motion, such as translation, rotation, scaling, etc. For inter-frame prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the matching region, a motion vector associated with the matching region, the number of reference images, weights associated with the reference images, etc.

[0056] To generate the predicted BPU 208, the encoder may perform a "motion compensation" operation. Motion compensation may be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., motion vector) and the prediction reference 224. For example, the encoder may shift the matching region of the reference image according to the motion vector, thereby predicting the original BPU of the current image. When multiple reference images are used, the encoder may shift the matching region of the reference image according to the respective motion vectors and average pixel values ​​of the matching regions. In some embodiments, if the encoder has assigned weights to the pixel values ​​of the matching regions of the respective matching reference images, the encoder may perform a weighted sum of the pixel values ​​of the shifted matching regions.

[0057] In some embodiments, the inter-frame prediction can be unidirectional or bidirectional. Unidirectional inter-frame prediction can use one or more reference pictures in the same temporal direction relative to the current picture. Unidirectional inter-frame prediction uses a reference picture before the current picture. Bidirectional inter-frame prediction can use one or more reference pictures in both temporal directions relative to the current picture.

[0058] Still referring to the forward path of process 200B, after the spatial prediction stage 2042 and the temporal prediction stage 2044, in the mode decision stage 230, the encoder can select a prediction mode (e.g., one of intra prediction or inter prediction) for the current iteration of process 200B. For example, the encoder can perform a rate-distortion optimization technique, wherein the encoder can select a prediction mode based on the bit rate of a candidate prediction mode and the distortion of a reference image reconstructed under the candidate prediction mode to minimize the value of a cost function. Based on the selected prediction mode, the encoder can generate a corresponding prediction BPU 208 and prediction data 206.

[0059] In the reconstruction path of process 200B, if intra-prediction mode has been selected in the forward path, after generating the prediction reference 224 (e.g., the current BPU in the current picture that has been encoded and reconstructed), the encoder can feed the prediction reference 224 directly to the spatial prediction stage 2042 for subsequent use (e.g., for extrapolation of the next BPU of the current picture). If inter-prediction mode has been selected in the forward path, after generating the prediction reference 224 (e.g., the current picture with all BPUs encoded and reconstructed), the encoder can feed the prediction reference 224 to the loop filtering stage 232, where the encoder can apply loop filtering to the prediction reference 224 to reduce or eliminate distortion (e.g., blocking artifacts) introduced by the inter-prediction. The encoder can apply various loop filtering techniques in the loop filtering stage 232, such as deblocking, sample adaptive offset, adaptive loop filtering, etc. The loop-filtered reference pictures may be stored in a buffer 234 (or "decoded picture buffer") for subsequent use (e.g., as inter-frame prediction reference pictures for future pictures in the video sequence 202). The encoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder may encode loop filtering parameters (e.g., loop filter strength) as well as quantized transform coefficients 216, prediction data 206, and other information in a binary encoding stage 226.

[0060] In some embodiments, the input video sequence 202 is processed block by block according to the encoding process 200B. In VVC, the coding tree unit (CTU) is the largest block unit and can be as large as 128×128 luma samples (plus corresponding chroma samples depending on the chroma format). The CTU can be further partitioned into coding units (CUs) using a quadtree, binary tree, or ternary tree. At the leaf nodes of the partitioning structure, coding information such as the coding mode (intra mode or inter mode), motion information (reference index, motion vector difference, etc.) in the case of inter coding, and quantized transform coefficients 216 are sent. If intra prediction (also known as spatial prediction) is used, the current block is predicted using spatially neighboring samples. If inter prediction (also known as temporal prediction or motion compensated prediction) is used, the current block is predicted using samples from an already coded picture called a reference picture. Inter prediction can use unidirectional prediction or bidirectional prediction. In unidirectional prediction, only one motion vector pointing to one reference image is used to generate the prediction signal for the current block; in bidirectional prediction, two motion vectors, each pointing to its own reference image, are used to generate the prediction signal for the current block. The motion vector and reference index are sent to the decoder to identify where the prediction signal or signals for the current block come from. After intra- or inter-frame prediction, the mode decision stage 230 selects the best prediction mode for the current block, for example, based on a rate-distortion optimization method. Based on the best prediction mode, a prediction BPU 208 is generated and subtracted from the input video block.

[0061] Still refer to Figure 2B , the prediction residual BPU 210 is sent to a transform stage 212 and a quantization stage 214 to generate quantized transform coefficients 216. The quantized transform coefficients 216 are then inverse quantized in an inverse quantization stage 218 and inverse transformed in an inverse transform stage 220 to obtain a reconstructed residual BPU 222. The prediction BPU 208 and the reconstructed residual BPU 222 are added together to form a prediction reference 224 before loop filtering, which is used to provide reference samples for intra prediction. Loop filtering (e.g., deblocking, sample adaptive offset (SAO), and adaptive loop filtering (ALF)) can be applied to the prediction reference 224 in a loop filtering stage 232 to form the reconstructed block, which is stored in a buffer 234 and used to provide reference samples for inter prediction. The coding information generated in the mode decision stage 230, such as the coding mode (intra-frame or inter-frame prediction), intra-frame prediction mode, motion information, quantized residual coefficients, etc., is sent to the binary encoding stage 226 to further reduce the bit rate before being packaged into the output video bitstream 228.

[0062] Figure 3ASchematic diagram of an example decoding process 300A consistent with an embodiment of the present disclosure is shown. For example, the decoding process 300A may be performed by a decoder, such as Figure 1 The image / video decoder 144 in FIG. 300A may correspond to Figure 2A In some embodiments, process 300A may be similar to the reconstruction path of process 200A. A decoder (e.g., Figure 1 The image / video decoder 144 in FIG. 300A may decode the video bitstream 228 into a video stream 304 according to process 300A. The video stream 304 may be very similar to the video sequence 202. However, due to information loss during compression and decompression (e.g., Figures 2A-2B quantization stage 214 in), typically, the video stream 304 is different from the video sequence 202. Figures 2A-2B 200A and 200B in the decoder, the decoder may perform process 300A at the basic processing unit (BPU) level for each picture encoded in the video bitstream 228. For example, the decoder may perform process 300A in an iterative manner, wherein the decoder may decode one basic processing unit in one iteration of process 300A. In some embodiments, the decoder may perform process 300A in parallel for multiple regions of each picture encoded in the video bitstream 228.

[0063] exist Figure 3A In the process 300A, the decoder may feed a portion of the video bitstream 228 associated with the basic processing unit of the encoded picture (referred to as the "encoded BPU") to the binary decoding stage 302. In the binary decoding stage 302, the decoder may decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder may feed the quantized transform coefficients 216 to the inverse quantization stage 218 and the inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may feed the prediction data 206 to the prediction stage 204 to generate the prediction BPU 208. The decoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224. In some embodiments, the prediction reference 224 may be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder may feed the prediction reference 224 to the prediction stage 204 for use in performing a prediction operation in the next iteration of the process 300A.

[0064] The decoder may iteratively perform process 300A to decode each coded BPU of a coded picture and generate a prediction reference 224 for encoding the next coded BPU of the coded picture. After decoding all coded BPUs of the coded picture, the decoder may output the picture to a video stream 304 for display and continue decoding the next coded picture in the video bitstream 228.

[0065] In the binary decoding stage 302, the decoder may perform the inverse of the binary coding technique used by the encoder (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the decoder may also decode other information in the binary decoding stage 302, such as, for example, the prediction mode, parameters of the prediction operation, the transform type, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), etc. In some embodiments, if the video bitstream 228 is transmitted over a network in the form of packets, the decoder may depacketize the video bitstream 228 before feeding it to the binary decoding stage 302.

[0066] Figure 3B Schematic diagram of another example decoding process 300B consistent with an embodiment of the present disclosure is shown. For example, the decoding process 300B can be performed by a decoder, such as Figure 1 2044 . The process 300B may be modified from the process 300A. For example, the process 300B may be used by a decoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared to the process 300A, the process 300B further divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and further includes the loop filter stage 232 and the buffer 234.

[0067] In process 300B, for an encoded basic processing unit (referred to as a "current BPU") of an encoded image being decoded (referred to as a "current image"), the prediction data 206 decoded by the decoder from the binary decoding stage 302 may include various types of data, depending on the prediction mode used by the encoder to encode the current BPU. For example, if the encoder used intra-frame prediction to encode the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra-frame prediction, parameters of the intra-frame prediction operation, etc. The parameters of the intra-frame prediction operation may include, for example, the location (e.g., coordinates) of one or more neighboring BPUs used as a reference, the size of the neighboring BPUs, extrapolation parameters, the orientation of the neighboring BPUs relative to the original BPU, etc. For another example, if the encoder used inter-frame prediction to encode the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter-frame prediction, parameters of the inter-frame prediction operation, etc. The parameters of the inter-frame prediction operation may include, for example, the number of reference images associated with the current BPU, the weights associated with the reference images respectively, the positions (e.g., coordinates) of one or more matching regions in the respective reference images, one or more motion vectors associated with the matching regions respectively, and the like.

[0068] Based on the prediction mode indicator, the decoder may decide whether to perform spatial prediction (eg, intra prediction) in a spatial prediction stage 2042 or temporal prediction (eg, inter prediction) in a temporal prediction stage 2044 . Figure 2B The details of performing such spatial prediction or temporal prediction are described in detail in

[15] and will not be repeated here. After performing such spatial prediction or temporal prediction, the decoder may generate a prediction BPU 208. The decoder may add the prediction BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224, as shown in FIG. Figure 3A described.

[0069] In process 300B, the decoder may feed the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing a prediction operation in the next iteration of process 300B. For example, if the current BPU is decoded using the intra prediction in the spatial prediction stage 2042, then after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder may feed the prediction reference 224 directly to the spatial prediction stage 2042 for subsequent use (e.g., for extrapolation of the next BPU of the current picture). If the current BPU is decoded using the inter prediction in the temporal prediction stage 2044, then after generating the prediction reference 224 (e.g., a reference picture in which all BPUs have been decoded), the encoder may feed the prediction reference 224 to the loop filtering stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder may perform the following operations: Figure 2B In-loop filtering is applied to the prediction reference 224 in the manner described in

[15] . The loop-filtered reference picture may be stored in a buffer 234 (e.g., a decoded picture buffer in a computer memory) for subsequent use (e.g., as an inter-frame prediction reference picture for a future encoded picture in the video bitstream 228). The decoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, when the prediction mode indicator of the prediction data 206 indicates that inter-frame prediction was used to encode the current BPU, the prediction data may further include loop filtering parameters (e.g., loop filtering strength).

[0070] Return Reference Figure 1 , each of the image / video preprocessor 122 , the image / video encoder 124 , and the image / video decoder 144 may be implemented as any suitable hardware, software, or combination thereof. Figure 4 is a block diagram of an example apparatus 400 for processing image data consistent with embodiments of the present disclosure. For example, apparatus 400 may be a preprocessor, an encoder, or a decoder. Figure 4As shown, the device 400 may include a processor 402. When the processor 402 executes the instructions described herein, the device 400 may become a special-purpose machine for pre-processing, encoding and / or decoding image data. The processor 402 may be any type of circuit system capable of manipulating or processing information. For example, the processor 402 may include any number of central processing units (or "CPUs"), graphics processing units (or "GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), general array logic (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), systems on chip (SoCs), application-specific integrated circuits (ASICs), and the like. In some embodiments, the processor 402 may also be a group of processors grouped into a single logical component. For example, as Figure 4 As shown, processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.

[0071] The apparatus 400 may further include a memory 404 configured to store data (eg, instruction sets, computer code, intermediate data, etc.). Figure 4 As shown, the stored data may include program instructions (e.g., program instructions for implementing the stages in process 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). Processor 402 may access the program instructions and data for processing (e.g., via bus 410) and execute the program instructions to perform operations or manipulations on the data for processing. Memory 404 may include a high-speed random access memory device or a non-volatile memory device. In some embodiments, memory 404 may include any combination of any number of random access memories (RAM), read-only memories (ROM), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, secure digital (SD) cards, memory sticks, compact flash (CF) cards, etc. Memory 404 may also be a group of memories grouped into a single logical component (e.g., a memory card). Figure 4 not shown).

[0072] The bus 410 may be a communication device for transmitting data between components within the apparatus 400 , such as an internal bus (eg, a CPU-memory bus), an external bus (eg, a Universal Serial Bus port, a Peripheral Component Interconnect Express port), and the like.

[0073] For ease of explanation and to avoid ambiguity, the processor 402 and other data processing circuitry are collectively referred to as "data processing circuitry" in this disclosure. The data processing circuitry may be implemented entirely in hardware, or as a combination of software, hardware, or firmware. Furthermore, the data processing circuitry may be a single standalone module, or may be fully or partially integrated into any other component of the device 400.

[0074] The device 400 may also include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, etc.). In some embodiments, the network interface 406 may include any number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication ("NFC") adapters, cellular network chips, etc.

[0075] In some embodiments, the apparatus 400 may further include a peripheral interface 408 to provide a connection to one or more peripheral devices. Figure 4 As shown, the peripheral devices may include but are not limited to a cursor control device (e.g., a mouse, touchpad, or touch screen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera or an input interface coupled to a video archive), etc.

[0076] It should be noted that a video codec (e.g., a codec that performs processes 200A, 200B, 300A, or 300B) can be implemented as any combination of software or hardware modules in device 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more software modules of device 400, such as program instructions that can be loaded into memory 404. For another example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of device 400, such as dedicated data processing circuits (e.g., FPGAs, ASICs, NPUs, etc.).

[0077] The video bitstream used in VVC or HEVC is a bit sequence in the form of a network abstraction layer (NAL) unit or a byte stream, which forms one or more coded video sequences (CVSs), and each CVS consists of one or more coded layer video sequences (CLVSs). In these layers, inter-layer prediction can be applied to achieve high compression performance. Here, a layer is a set of video coding layer (VCL) NAL units and associated non-VCL NAL units, and the VCL NAL units all have specific NAL layer ID values. And the VCL NAL unit is a general term for a coded slice NAL unit classified as a VCL NAL unit and a subset of NAL units with reserved values ​​of the NAL unit type. Inter-layer prediction can be applied between different layers.

[0078] Supplemental Enhancement Information (SEI) messages are intended to be transmitted within a coded video bitstream in a manner specified in a video coding specification, or by other methods determined by the specifications of the system using the coded video bitstream. SEI messages can contain various types of data that indicate the timing of the video pictures or describe various properties of the coded video or how the coded video can be used or enhanced. SEI messages are also defined to contain arbitrary user-defined data. SEI messages do not affect the core decoding process, but may indicate recommendations for post-processing or display of the video.

[0079] To specify SEI messages, the JVET working group also developed the H.274 standard, which specifies the syntax and semantics of video usability information (VUI) parameters and supplementary enhancement information (SEI) messages, which are specifically intended to be used with coded video bitstreams specified by the VVC standard. However, since neither the VUI parameters nor the SEI messages affect the decoding process, the SEI messages in H.274 can also be used with other types of coded video bitstreams, such as H.265 / HEVC, H.264 / AVC, etc.

[0080] For the purpose of object detection and tracking, the latest versions of the HEVC standard and the VSEI standard adopt the Annotation Region (AR) SEI message, which carries parameters to describe the bounding box of the detected or tracked object within the compressed video bitstream, so that if the encoder, transcoder or network node has already performed video analysis to identify the object, the decoder side device does not need to perform video analysis to identify the object. This is beneficial for application scenarios where the decoder device has limited computing resources or limited power supply. At the same time, performing object detection and tracking on the encoder side and transmitting the information to the decoder can help improve the quality of the detection and tracking because the encoder can use the original video to perform the detection and tracking tasks, and the original video can have a much higher quality than the reconstructed video restored on the decoder side.

[0081] In the AR SEI message in HEVC, in addition to the bounding box of the detected or tracked object, an object label and a confidence level associated with the object can also be provided. The object label provides information about what kind of object it is, and the confidence level characterizes the restoration fidelity of the detected or tracked object within the bounding box. In addition, a flag is provided to identify whether the bounding box in the current SEI message represents the position of an object that may be occluded or partially occluded by other objects or the position of only the visible part of the object. A flag can also be optionally sent for each bounding box, which identifies whether the object represented by the current bounding box is only partially visible.

[0082] The syntax of the AR SEI message uses parameter persistence to avoid the need to resend information already in a previous SEI message within the same persistence range. For example, if a detected first object remains stationary in the current picture relative to a previously coded picture, and a detected second object moves from one picture to another, only the bounding box information of the second object needs to be sent, and the position / bounding box information of the first object can be copied from the previous SEI message.

[0083] Major video coding standards, such as H.264 / AVC, H.265 / HEVC, and H.266 / VVC, all support encoding a special type of picture, known as an auxiliary picture, to provide auxiliary information to the normal picture (the primary picture). Auxiliary pictures have no regulatory effect on the decoding process of the primary picture. The bitstreams of the auxiliary pictures and the primary picture are packaged into a coded video sequence (CVS). The necessary information for parsing the auxiliary pictures is transmitted via SEI messages.

[0084] In HEVC, auxiliary pictures are coded as one or more auxiliary picture layers that are different from the primary picture layer. The identification of auxiliary pictures is sent in the Video Parameter Set (VPS) extension, as shown in Table 1 below. Table 1: VPS extension syntax

[0085] The value of dimension_id[LayerIdxInVps[nuh_layer_id]][j] is derived from the NumScalabilityTypes segments. The value of dimension_id[LayerIdxInVps[nuh_layer_id]][j] is derived from the NumScalabilityTypes segments. The value of splitting_flag equal to 0 indicates that the syntax element dimension_id[i][j] is present.

[0086] When splitting_flag is equal to 1, the scalability identifier of the current scalability dimension can be derived from the nuh_layer_id syntax element in the NAL unit header via a copy of the bit mask. The corresponding bit mask for the i-th current scalability dimension is defined by the value of the dimension_id_len_minus1[i] syntax element and dimBitOffset[i] as specified in the semantics of dimension_id_len_minus1[j].

[0087] scalability_mask_flag[i] equal to 1 indicates that the dimension_id syntax element corresponding to the i-th scalability dimension in Table 2 below is present. scalability_mask_flag[i] equal to 0 indicates that the dimension_id syntax element corresponding to the i-th scalability dimension is not present. Table 2: Correspondence between ScalabiltyId and scalability dimensions

[0088] dimension_id_len_minus1[j] plus 1 specifies the length in bits of the dimension_id[i][j] syntax element.

[0089] When splitting_flag is equal to 1, the following applies: The variable dimBitOffset[0] is set equal to 0, and for values ​​of j in the range 1 to NumScalabilityTypes-1 (inclusive), dimBitOffset[j] is derived as follows: - The value of dimension_id_len_minus1[NumScalabilityTypes-1] is inferred to be equal to 5-dimBitOffset[NumScalabilityTypes-1]. - Set the value of dimBitOffset[NumScalabilityTypes] equal to 6.

[0090] The bitstream conformance requirement is that when NumScalabilityTypes is greater than 0, dimBitOffset[NumScalabilityTypes-1] is less than 6.

[0091] vps_nuh_layer_id_present_flag equal to 1 specifies that layer_id_in_nuh[i] exists for values ​​of i from 1 to MaxLayersMinus1, inclusive. vps_nuh_layer_id_present_flag equal to 0 specifies that layer_id_in_nuh[i] does not exist for values ​​of i from 1 to MaxLayersMinus1, inclusive.

[0092] layer_id_in_nuh[i] specifies the value of the nuh_layer_id syntax element in the VCL NAL unit of layer i. When i is greater than 0, layer_id_in_nuh[i] is greater than layer_id_in_nuh[i-1]. For any value of i in the range 0 to MaxLayersMinus1 (inclusive), when layer_id_in_nuh[i] is not present, the value of layer_id_in_nuh[i] is inferred to be equal to i.

[0093] For values ​​of i from 0 to MaxLayersMinus1 (inclusive), the variable LayerIdxInVps[layer_id_in_nuh[i]] is set equal to i.

[0094] dimension_id[i][j] specifies the identifier of the jth current scalability dimension type of the i-th layer. The number of bits used to represent dimension_id[i][j] is dimension_id_len_minus1[j]+1 bit.

[0095] Depending on splitting_flag, the following applies: - If splitting_flag is equal to 1, then for i values ​​from 0 to MaxLayersMinus1 (inclusive) and j values ​​from 0 to NumScalabilityTypes-1 (inclusive), dimension_id[i][j] is inferred to be equal to ((layer_id_in_nuh[i] & ((1 <<dimBitOffset[j+1])-1))> >dimBitOffset[j]). Otherwise (splitting_flag is equal to 0), dimension_id[0][j] is inferred to be equal to 0 for j values ​​from 0 to NumScalabilityTypes-1 (inclusive).

[0096] The variable ScalabilityId[i][smIdx] specifies the identifier of the (smIdx)th scalability dimension type of the i-th layer, and the variables DepthLayerFlag[lId], ViewOrderIdx[lId], DependencyId[lId], and AuxId[lId] specify the depth flag, view order index, spatial / quality scalability identifier, and auxiliary identifier, respectively, of the layer with nuh_layer_id equal to lId, derived as follows:

[0097] AuxId[lId] equal to 0 specifies that the layer with nuh_layer_id equal to lId does not contain any auxiliary picture. AuxId[lId] greater than 0 specifies that the auxiliary picture in the layer with nuh_layer_id equal to lId is of the type specified in Table 3 below. Table 3: Correspondence between AuxId and auxiliary image type AuxId AuxId Name Auxiliary image type SEI message describing auxiliary image parsing 1 AUX_ALPHA α plane Alpha channel information 2 AUX_DEPTH Depth Image Deep representation information 3..127 reserve 128..159 not specified 160..255 reserve

[0098] The interpretation of auxiliary pictures associated with AuxId values ​​in the range of 128 to 159, inclusive, is not determined by the AuxId value itself.

[0099] For bitstreams conforming to this version of the specification, AuxId[lId] is in the range of 0 to 2 (inclusive) or 128 to 159 (inclusive). Although in this version of the specification, the value of AuxId[lId] is in the range of 0 to 2 (inclusive) or 128 to 159 (inclusive), the decoder may allow the value of AuxId[lId] to be in the range of 0 to 255 (inclusive).

[0100] A bitstream conformance requirement is that when AuxId[lId] is equal to AUX_ALPHA or AUX_DEPTH, either of the following applies: - For layers with nuh_layer_id equal to ld, chroma_format_idc is equal to 0 in the valid SPS. - In all pictures where nuh_layer_id is equal to lid and this VPS RBSP is the valid VPS RBSP, the values ​​of all decoded chroma samples are equal to 1<<(BitDepthC-1).

[0101] The SEI message may describe the auxiliary picture interpretations, including their possible association with one or more primary pictures.

[0102] Unless subject to the semantic constraints of the SEI message specifying the interpretation of the auxiliary picture, two layers with nuh_layer_id values ​​of layerIdA and layerIdB are allowed such that AuxId[layerIdA] is equal to AuxId[layerIdB], both are greater than 0, and for each value of i in the range of 0 to 15 (inclusive), all values ​​of ScalabilityId[LayerIdxInVps[layerIdA]][i] are allowed to be equal to ScalabilityId[LayerIdxInVps[layerIdB]][i]. The SEI message specifying the interpretation of the auxiliary picture may specify that a picture with nuh_layer_id equal to layerIdA and a picture with nuh_layer_id equal to layerIdB in the same access unit may both be associated with the same primary picture.

[0103] In VVC, auxiliary pictures are coded as one or more auxiliary picture layers different from the primary picture layer. The auxiliary indication is sent in the scalability dimension information SEI message, as shown in Table 4 below. Table 4: Syntax of the Scalability Dimension Information SEI message

[0104] The Scalability Dimension Information (SDI) SEI message provides the SDI for each layer in the current CVS (i.e., the CVS containing the SDI SEI message), for example, 1) providing the view ID of each layer when multiple views may exist; and 2) providing the auxiliary ID of each layer when there may be auxiliary information (such as depth or alpha) carried by one or more layers.

[0105] When an SDI SEI message exists in any AU of a CVS, the SDI SEI message exists for the first AU of the CVS. All SDI SEI messages in a CVS have the same content.

[0106] sdi_max_layers_minus1 plus 1 indicates the maximum number of layers in the current CV.

[0107] sdi_multiview_info_flag equal to 1 indicates that the current CVS can have multiple views and the sdi_view_id_val[] syntax element is present in the SDI SEI message. sdi_multiview_info_flag equal to 0 indicates that the current CVS does not have multiple views and the sdi_view_id_val[] syntax element is not present in the SDI SEI message.

[0108] sdi_auxiliary_info_flag is equal to 1, indicating that one or more layers in the current CVS may be auxiliary layers carrying auxiliary information, and the sdi_aux_id[] syntax element is present in the SDI SEI message. sdi_auxiliary_info_flag is equal to 0, indicating that the current CVS has no auxiliary layers, and the sdi_aux_id[] syntax element is not present in the SDI SEI message.

[0109] sdi_view_id_len_minus1 plus 1 specifies the length in bits of the sdi_view_id_val[i] syntax element.

[0110] sdi_layer_id[i] specifies the layer identifier of the i-th layer that may exist in the current CVS.

[0111] sdi_view_id_val[i] specifies the view identifier of layer i in the current CVS. The length of the sdi_view_id_val[i] syntax element is sdi_view_id_len_minus1+1 bits.

[0112] The variable NumViews specifying the number of views in the current CVS and the list ViewId specifying the view identifiers of the views in the current CVS are derived as follows:

[0113] sdi_aux_id[i] equal to 0 indicates that the i-th layer in the current CVS does not contain auxiliary pictures. sdi_aux_id[i] greater than 0 indicates the type of auxiliary pictures in the i-th layer in the current CVS, as specified in Table 5 below. When sdi_auxiliary_info_flag is equal to 0, the value of sdi_aux_id[i] is inferred to be equal to 0. Table 5: Correspondence between sdi_aux_id[i] and auxiliary image types sdi_aux_id[i] name Auxiliary image type 1 AUX_ALPHA α plane 2 AUX_DEPTH Depth Image 3..127 reserve 128..159 not specified 160..255 reserve

[0114] When the value of sdi_aux_id[i] is in the range of 128 to 159 (inclusive), the interpretation of the auxiliary picture associated with sdi_aux_id[i] is not specified by the sdi_aux_id[i] value itself.

[0115] For bitstreams conforming to this version of the specification, sdi_aux_id[i] is in the range of 0 to 2 (inclusive) or 128 to 159 (inclusive). Although in this version of the specification, the value of sdi_aux_id[i] is in the range of 0 to 2 (inclusive) or 128 to 159 (inclusive), the decoder also allows other values ​​of sdi_aux_id[i] in the range of 0 to 255 (inclusive).

[0116] If sdi_aux_id[i] is equal to 0, the i-th layer is called the main layer. Otherwise, the i-th layer is called the auxiliary layer. When sdi_aux_id[i] is equal to 1, the i-th layer is also called the alpha auxiliary layer. When sdi_aux_id[i] is equal to 2, the i-th layer is also called the depth auxiliary layer.

[0117] sdi_num_associated_primary_layers_minus1[i] plus 1 specifies the number of associated primary layers for layer i, which is an auxiliary layer. The value of sdi_num_associated_primary_layers_minus1[i] is less than the total number of primary layers.

[0118] sdi_associated_primary_layer_idx[i][j] specifies the layer index of the j-th associated primary layer of the i-th layer, where the i-th layer is an auxiliary layer. The value of sdi_aux_id[sdi_associated_primary_layer_idx[i][j]] is equal to 0.

[0119] Auxiliary layers describe the properties of, and apply to, their associated primary layers.

[0120] The current AR SEI message can be used to annotate or track the object in the video. However, its functionality is limited. For example, the current AR SEI message cannot fully support the following two aspects.

[0121] In the current AR SEI message, the detected or tracked object is represented by a bounding box. The position information of the object can be described by the bounding box, while the shape information of the object cannot be represented by the bounding box. For application scenarios where segmentation is used to facilitate functions such as virtual backgrounds, a more accurate description of the object shape information is required. In addition, the power consumption of performing object segmentation is large, which is a heavy burden for mobile devices. Once object segmentation is performed, it may be necessary to carry such information in the video bitstream as auxiliary information. The syntax of the current AR SEI message as shown in Table 1 cannot carry such information.

[0122] In addition, a flag is sent in the current AR SEI message to identify whether the object represented by the bounding box is partially visible or fully visible. However, in the case where the object is partially visible, there is no parameter to tell the decoder which part is visible and which part is occluded. Therefore, the flag itself does not provide the decoder with much information for determining the visible and invisible areas of the object. Instead, object depth information can provide a better mechanism to describe the relative position of different objects in the image in terms of their distance from the camera. Such information can be directly used to derive which parts of which objects are occluded or not.

[0123] To address this issue, instead of sending a bounding box signal to the annotated area, a mask is sent to represent the shape and position of the annotated or tracked object. The mask can be implemented as a binary matrix of the same size as the image, where an element with a value of 0 indicates that the location is covered by the background, while an element with a value of 1 indicates that the location is covered by the object. Therefore, any shape of the object can be represented by the mask. To distinguish different objects, a multi-valued mask can be used, where an element with a value of 0 represents the background, and an element with a value of k (k is not equal to 0) represents the kth object.

[0124] The mask can accurately represent the shape of the object, but the signaling overhead is much greater than sending a bounding box. Therefore, in this disclosure, it is proposed to encode the mask as the auxiliary image instead of sending the mask in the SEI message, so that the mask image can be compressed using low-level video coding techniques supported by video coding standards.

[0125] In this disclosure, at a time instance, there are one or more normal images, referred to as primary images, and each primary image is associated with one or more object mask auxiliary images. In H.264 / AVS, auxiliary images are indicated by a special NAL unit type. In H.265 / HEVC and H.266 / VVC, auxiliary images are coded as an auxiliary image layer, i.e., another layer in addition to the primary image layer. Therefore, there can be multiple primary image layers and multiple auxiliary image layers.

[0126] In order to parse the auxiliary pictures, some auxiliary information is needed. In this disclosure, it is proposed to send the auxiliary information about the masked auxiliary pictures in a SEI message.

[0127] Figure 5 The present invention is a syntax diagram of the proposed object mask information (OMI) SEI message according to some disclosed embodiments. The diagram shows the syntax structure and syntax element order of the object mask information SEI message. First, a cancellation flag is sent to identify whether this OMI SEI is used to cancel the continued validity range of the previous SEI message (e.g., the last OMI message). If the cancellation flag indicates that the continued validity range of the previous OMI SEI message is not canceled, information about the object mask is sent to update the object information sent in the previous OMI SEI message, wherein the object mask auxiliary (image) identifier information used to distinguish the object mask auxiliary image from other auxiliary images is first sent. Then, the number of object mask images (e.g., auxiliary image layers) is sent. Thereafter, the presence flag (e.g., confidence presence flag, depth presence flag, or label presence flag) and the syntax elements of confidence, depth, and identifier length (e.g., confidence length, depth length, or label length) (if any) are sent. The syntax elements sent above are called common information for the object mask indicated by this OMISEI message, and individual mask information is sent subsequently. Finally, for each mask in each object mask image, a mask identifier is sent, followed by mask confidence, object depth, and mask label (if present).

[0128] Figure 6 FIG. 6 is a diagram illustrating an exemplary method 600 for encoding a video sequence into a bitstream consistent with an embodiment of the present disclosure. Figure 6 As shown, the method 600 includes steps 602 and 604, which can be performed by an encoder (e.g., Figure 1 Image / video encoder 124 or Figure 4 400 in the device).

[0129] In step 602, the encoder may receive a video sequence.

[0130] In step 604, the encoder may encode one or more images of the video sequence to generate a bitstream. Specifically, the encoder may encode an auxiliary image in the bitstream to indicate a mask of an object in a main image. The mask of the object may be represented by sample values ​​of the auxiliary image. As understood, the object in the main image may be drawn by mask-filled pixels having the sample values. In addition, the encoder may generate a supplemental enhancement information (SEI) message associated with the main image. The SEI message is also applicable to the auxiliary image and may be used to indicate the properties of the mask of the object. In the present disclosure, the SEI message for indicating the properties of the mask of a certain object is also referred to as an object mask information (OMI) SEI message.

[0131] Figure 7 is a schematic diagram showing exemplary main images 701 and 703 and auxiliary images 702 and 704 consistent with an embodiment of the present disclosure. The auxiliary image 702 corresponds to the main image 701, and the auxiliary image 704 corresponds to the main image 703. Figure 7 As shown, the main image 701 may include a background 711, a human object 712, and an animal object 713 in the image. The auxiliary image 702 corresponding to the main image 701 may include a human mask 722 corresponding to the human object 712 and an animal mask 723 corresponding to the animal object 713. The mask in the auxiliary image 702 can be used to represent the position and outline of each object in the main image 701. Specifically, the auxiliary image 702 can have the same size as the main image 701. The human mask 722 depicts the position and outline of the human object 712 in the main image 701 through its own position and outline in the auxiliary image 702. As is known, the mask in the auxiliary image 702 can be represented by corresponding sample values ​​(also called pixel values). Similarly, the mask in the auxiliary image 704 can be used to represent the position and outline of each object in the main image 703.

[0132] An OME SEI message (not shown) may be applied to the auxiliary picture 702 or 704 and may be used to indicate one or more attributes of at least one of the masks.

[0133] In some embodiments, each OMISEI message contains information about all of the masks. Since a persistent scheme is used for OMISEI messages, for a primary image, if the mask does not change at all from one time instance to the next, then there is no need to send an OMISEI. If any information changes, a new OMISEI message containing the new information about the mask needs to be sent. The syntax is shown in Table 6, with the semantics provided as an example in Table 6 below. Table 6: Example syntax of OMISEI message

[0134] The Object Mask Information (OMI) SEI message provides information about an object mask picture coded as an auxiliary picture. The object mask auxiliary picture has a nuh_layer_id equal to nuhLayerIdA and an AuxId[nuhLayerIdA] in the range of 128 to 159 (inclusive). Each overlay auxiliary picture layer is associated with one or more primary picture layers, as described below.

[0135] In some embodiments, the encoder may determine a cancel flag in step 604 to identify whether the SEI message cancels the continuing effect of a previous SEI message. For example, omi_cancel_flag equal to 1 indicates that the SEI message cancels the continuing effect of any previous object mask information SEI message in the output order associated with one or more primary picture layers to which this SEI message is applied. omi_cancel_flag equal to 0 indicates subsequent object mask information, and the object mask information sent in this SEI message will be used to update the existing object mask information of any previous SEI message.

[0136] It should be understood that if the cancellation flag (e.g., omi_cancel_flag) indicates that the SEI message cancels the continued validity of any SEI message, then the cancellation flag is only included and sent in the SEI message. When it is decided to reuse the mask information, a complete SEI message with all required syntax needs to be generated and sent to the decoder.

[0137] The SEI message may indicate properties of multiple masks. In some embodiments, the properties of the mask transmitted by the SEI message may include common characteristics of each mask indicated by the SEI message and individual characteristics of the mask of the object. As will be appreciated, the common characteristics are shared by the masks indicated by the SEI message, while individual characteristics are specific to the target mask.

[0138] If the cancellation flag (eg, omi_cancel_flag) indicates that the SEI message does not cancel the continuing validity of the information of the previous SEI message, the encoder may further determine the common characteristics and the individual characteristics in step 604 .

[0139] As mentioned above, reference Figure 5, the common features may be object mask auxiliary (image) identifier information, the number of object mask images, presence flags, and confidence, depth, and length of the identifier (e.g., confidence length, depth length, or label length), if any.

[0140] The identifier of the auxiliary image to which the SEI message is applied may be determined as one of the common features. For example, omi_aux_id_minus128 plus 128 indicates the value of AuxId of the object mask auxiliary image. omi_aux_id_minus128 is in the range of 0 to 31 (inclusive).

[0141] The number of bits used to encode the identifier of any mask in the plurality of masks may be determined as one of the common characteristics. For example, omi_num_mask_pic_minus1 plus 1 indicates the number of object mask auxiliary pictures associated with the same one or more primary pictures. The value of omi_num_mask_pic_minus1 is in the range of 0 to 63 (inclusive). The value of omi_num_mask_pic_minus1 is the same in all OMI SEI messages within a CVS.

[0142] In some embodiments, the SEI message applies to multiple auxiliary pictures. The number of the multiple auxiliary pictures can be determined as one of the common characteristics. For example, omi_mask_id_length_minus8 plus 8 indicates the number of bits used to encode the omi_mask_id[i][j] syntax element.

[0143] A confidence present flag in the SEI message indicating whether the confidence information for the plurality of masks includes confidence information can be determined as one of the common features. For example, omi_mask_confidence_info_present_flag equal to 1 indicates the presence of the omi_mask_confidence[i][j] syntax element. omi_mask_confidence_info_present_flag equal to 0 indicates the absence of the omi_mask_confidence[i][j] syntax element. A bitstream conformance requirement is that the value of omi_mask_confidence_info_present_flag is the same for all object_mask_info() syntax structures within a CLVS.

[0144] In some embodiments, if a confidence present flag (e.g., omi_mask_confidence_info_present_flag) indicates that the confidence information for the multiple masks is included in the SEI message, then the length of the confidence information for the multiple masks may also be determined as one of the common characteristics. For example, omi_mask_confidence_length_minus1 plus 1 specifies the length in bits of the omi_mask_confidence[i][j] syntax element. A bitstream consistency requirement is that the value of omi_mask_confidence_length_minus1 is the same for all object_mask_info() syntax structures within a CLVS.

[0145] A depth present flag indicating whether the depth information of the plurality of masks is included in the SEI message can be determined as one of the common features. For example, omi_object_depth_info_present_flag equal to 1 indicates the presence of the omi_object_depth[i][j] syntax element. omi_object_depth_info_present_flag equal to 0 indicates the absence of the omi_object_depth[i][j] syntax element. A bitstream conformance requirement is that the value of omi_object_depth_info_present_flag is the same for all object_mask_info() syntax structures within a CLVS.

[0146] In some embodiments, if the depth present flag (e.g., omi_object_depth_info_present_flag) indicates that depth information for the multiple masks is included in the SEI message, the length of the depth information for the multiple masks may be determined as one of the common characteristics. For example, omi_object_depth_length_minus1 plus 1 specifies the length in bits of the omi_object_depth[i][j] syntax element. A bitstream conformance requirement is that the value of omi_object_depth_length_minus1 is the same for all object_mask_info() syntax structures within a CLVS.

[0147] A label presence flag indicating the label language presence information of the plurality of masks and whether the label information is included in the SEI message may be determined as one of the common features. For example, omi_mask_label_info_present_flag is equal to 1 to indicate the presence of omi_mask_label_language_present_flag and omi_mask_label[i][j]. omi_mask_label_info_present_flag is equal to 0 to indicate the absence of omi_mask_label_language_present_flag and omi_mask_label[i][j].

[0148] In some embodiments, if the label presence flag (e.g., omi_mask_label_info_present_flag) indicates that the label language information of the plurality of masks is present and the label information is included in the SEI message, then the language presence flag for indicating whether the label language information of the plurality of masks is included in the SEI message can be determined as one of the common features. For example, omi_mask_label_language_present_flag is equal to 1 to indicate that omi_mask_label_language is present. Omi_mask_label_language_present_flag is equal to 0 to indicate that omi_mask_label_language is not present and the language of the mask label is not specified.

[0149] In some embodiments, omi_bit_equal_to_zero is equal to 0.

[0150] In some embodiments, if the language presence flag indicates that the label language information for the plurality of masks is included in the SEI message, the label language information for the plurality of masks may be determined to be one of the common characteristics. For example, omi_mask_label_language contains a language tag specified by IETF RFC 5646 followed by a null termination byte equal to 0x00. The length of the omi_mask_label_language syntax element is less than or equal to 255 bytes, excluding the null termination byte. When not present, the language of the label is unspecified.

[0151] As mentioned above, reference Figure 5, the individual features may be the mask identifier, mask confidence, object depth, and mask label (if any) of each mask in each object mask image. As described above, when the SEI message is applied to multiple auxiliary images, the number of the multiple auxiliary images may be determined as one of the common features. In addition, the SEI message may include the individual features generated for the masks represented by the multiple auxiliary images.

[0152] In some embodiments, the individual feature omi_mask_pic_layer_id[i] indicates the nuh_layer_id value of the i-th auxiliary image layer. For all values in the range from 0 to omi_num_mask_pic_minus1 (including the end values), AuxId[omi_mask_pic_layer_id[i]] is equal to omi_aux_id_minus128 + 128.

[0153] In some embodiments, the individual feature omi_num_mask_in_pic[i] indicates the number of masks in the i-th auxiliary image. omi_num_mask_in_pic[i] is in the range from 0 to (1 << BitDepthY) - 1 (including the end values), where BitDepthY is the bit depth of the samples of the luminance component.

[0154] In some embodiments, the individual feature omi_mask_id[i][j] indicates the identifier of the j-th object mask in the i-th object mask auxiliary image. The object mask identifier associated with the sample position (x, y) in the i-th object mask auxiliary image is equal to p[i][x][y], where p[i][x][y] refers to the luminance sample at the position (x, y) in the decoded i-th object mask auxiliary image.

[0155] The variable maskId[i][j] specifies the object mask identifier of the j-th object mask in the i-th object mask auxiliary image in the SEI message, and is derived as follows,

[0156] In some embodiments, the individual feature omi_mask_confidence[i][j] indicates the confidence associated with the j-th object mask in the i-th object mask auxiliary image in base 2 -(omi_mask_confidence_length_minus1+1)The confidence level is in units of omi_mask_confidence[i][j][k], so that higher omi_mask_confidence[i][j][k] values ​​indicate higher confidence levels. The length of the omi_mask_confidence[i][j][k] syntax element is omi_mask_confidence_length_minus1+1 bits.

[0157] In some embodiments, the individual feature omi_mask_depth[i][j] indicates the object depth associated with the j-th object mask in the i-th object mask auxiliary image. Smaller values ​​of omi_mask_depth indicate a shorter distance to the object. The length of the omi_mask_depth[i][j][k] syntax element is omi_object_depth_length_minus1+1 bits.

[0158] In some embodiments, the individual feature omi_mask_label[i][j] specifies the content of the label associated with the j-th object mask in the i-th object mask auxiliary image. The length of the omi_mask_label[i][j] syntax element is less than or equal to 255 bytes, excluding the null termination byte.

[0159] In the syntax described in Table 6, every time the object mask information changes, all mask information including the unchanged part needs to be resent in the OMISEI message, which costs a lot of bits. In some embodiments, only the changed information is sent. Figure 8 is a schematic diagram illustrating the sub-steps of method 600. Figure 8 As shown, step 604 may include sub-steps 802 and 804 that may be performed by the encoder.

[0160] In sub-step 802, when the cancellation flag indicates that the SEI message cancels the continuing validity of information of the previous SEI message, the encoder may determine whether the mask of the object is different from the previous mask of the object represented by the previous auxiliary picture. In some embodiments, the encoder may determine whether the mask of the object is different from the previous mask of the object represented by the previous auxiliary picture regardless of whether the cancellation flag indicates that the SEI message cancels the continuing validity of information of the previous SEI message.

[0161] In sub-step 804, when the mask of the object is different from the previous mask of the object, the encoder may encode the attributes of the mask of the object in the SEI message. In some embodiments, when the mask of the object is the same as the previous mask of the object, the encoder may skip encoding the attributes of the mask of the object in the SEI message.

[0162] In some embodiments, an update flag can be introduced for each auxiliary image. If there is no change in the object mask auxiliary image, the mask information of this auxiliary image is skipped. If there is a change, the changed information is sent. For example, if the label, depth or confidence of the object mask changes, or the number of masks changes, only the changed mask needs to be sent. Therefore, the signaling overhead is reduced. Return to reference Figure 7 , the person mask 742 in the auxiliary 704 and the person mask 722 in the auxiliary 702 are the same, while the animal mask 723 becomes the animal mask 743. If the previous SEI is sent to indicate the masks 722 and 723, the current SEI to be sent to indicate the masks 742 and 743 may skip the unchanged information (e.g., the person mask 742).

[0163] The syntax is shown in Table 7 (differences from Table 6 are italicized in Table 7), wherein the semantics are provided as an example in Table 7 below. The definitions of the common features denoted as "high-level information" and the individual features denoted as "individual mask information" (some of which are omitted) can be inherited from the above-mentioned embodiment with shared parameter / function names. Table 7: Example syntax of OMISEI message

[0164] The Object Mask Information (OMI) SEI message provides information about an object mask picture coded as an auxiliary picture. The object mask auxiliary picture has a nuh_layer_id equal to nuhLayerIdA and an AuxId[nuhLayerIdA] in the range of 128 to 159 (inclusive). Each overlay auxiliary picture layer is associated with one or more primary picture layers, as described below.

[0165] Similar to some embodiments described above, the encoder may determine a cancellation flag in step 604 to identify whether the SEI message cancels the continuing effect of a previous SEI message. For example, omi_cancel_flag equal to 1 indicates that the SEI message cancels the continuing effect of any previous object mask information SEI message in the output order associated with one or more primary picture layers to which this SEI message is applied. omi_cancel_flag equal to 0 indicates subsequent object mask information, and the object mask information sent in this SEI message will be used to update the current object mask information of any previous SEI message.

[0166] Similarly, the encoder may also determine other common features of the masks.

[0167] omi_aux_id_minus128 plus 128 indicates the value of AuxId of the object mask auxiliary image. omi_aux_id_minus128 is in the range of 0 to 31 (inclusive).

[0168] omi_num_mask_pic_minus1 plus 1 indicates the number of object mask auxiliary pictures associated with the same primary picture or pictures. The value of omi_num_mask_pic_minus1 is in the range of 0 to 63 (inclusive). The value of omi_num_mask_pic_minus1 is the same in all OMI SEI messages within a CVS.

[0169] omi_mask_id_length_minus8 plus 8 indicates the number of bits used to encode the omi_mask_id[i][j] syntax element.

[0170] omi_msak_confidence_info_present_flag equal to 1 indicates the presence of the omi_mask_confidence[i][j] syntax element. omi_mask_confidence_info_present_flag equal to 0 indicates the absence of the omi_mask_confidence[i][j] syntax element. Bitstream conformance requirement is that the value of omi_mask_confidence_info_present_flag is the same for all object_mask_info() syntax structures within a CLVS.

[0171] omi_mask_confidence_length_minus1 plus 1 specifies the length in bits of the omi_mask_confidence[i][j] syntax element. A bitstream conformance requirement is that the value of omi_mask_confidence_length_minus1 is the same for all object_mask_info() syntax structures within a CLVS.

[0172] omi_object_depth_info_present_flag equal to 1 indicates the presence of the omi_object_depth[i][j] syntax element. omi_object_depth_info_present_flag equal to 0 indicates the absence of the omi_object_depth[i][j] syntax element. A bitstream conformance requirement is that the value of omi_object_depth_info_present_flag is the same for all object_mask_info() syntax structures within a CLVS.

[0173] omi_object_depth_length_minus1 plus 1 specifies the length in bits of the omi_object_depth[i][j] syntax element. A bitstream conformance requirement is that the value of omi_object_depth_length_minus1 is the same for all object_mask_info() syntax structures within a CLVS.

[0174] omi_mask_label_info_present_flag is equal to 1 to indicate that omi_mask_label_language_present_flag and omi_mask_label[i][j] are present. omi_mask_label_info_present_flag is equal to 0 to indicate that omi_mask_label_language_present_flag and omi_mask_label[i][j] are not present.

[0175] omi_mask_label_language_present_flag equal to 1 indicates that omi_mask_label_language is present. omi_mask_label_language_present_flag equal to 0 indicates that omi_mask_label_language is not present and the language of the mask label is not specified.

[0176] omi_bit_equal_to_zero is equal to 0.

[0177] omi_mask_label_language contains a language tag specified by IETF RFC5646 that is followed by a null-terminated byte with a value equal to 0x00. The length of the omi_mask_label_language syntax element is less than or equal to 255 bytes, excluding the null-terminated byte. When not present, the language of the label is unspecified.

[0178] In some embodiments, the encoder may determine an update flag for indicating whether the mask for identifying the object is sent in an auxiliary image. For example, omi_mask_pic_update_flag[i] being equal to 1 indicates that the mask information of the i-th object mask auxiliary image is sent. omi_mask_pic_update_flag[i] being equal to 0 indicates that the mask information of the i-th object mask auxiliary image is not sent. When the mask information of the i-th object mask auxiliary image is not present, the persistence mechanism is used, that is, the information is inherited from the previous OMI SEI message that transmitted the mask information of the i-th object mask auxiliary image.

[0179] omi_mask_pic_layer_id[i] indicates the nuh_layer_id value of the i-th auxiliary image layer. For all values within the range from 0 to omi_num_mask_pic_minus1 (including the end values), AuxId[omi_mask_pic_layer_id[i]] is equal to omi_aux_id_minus128 + 128.

[0180] In some embodiments, omi_num_mask_in_pic_update[i] indicates the number of masks in the i-th auxiliary image to be sent. omi_num_mask_in_pic[i] is within the range from 0 to (1 << BitDepthY) - 1 (including the end values), where BitDepthY is the bit depth of the samples of the luminance component.

[0181] In some embodiments, omi_mask_id[i][j] indicates the identifier of the j-th object mask to be updated in the i-th object mask auxiliary image.

[0182] The object mask identifier associated with the sample position (x, y) in the i-th object mask auxiliary image is equal to p[i][x][y], where p[i][x][y] refers to the luminance sample at position (x, y) in the decoded i-th object mask auxiliary image.

[0183] The variable maskId[i][j] specifies the object mask identifier of the j-th object mask of the i-th object mask auxiliary image in the SEI message, and is derived as follows,

[0184] In some embodiments, when the cancellation flag indicates that the SEI message does not cancel the persistent validity of the information of the previous SEI message, the encoder may determine a mask cancellation flag, the mask cancellation flag being used to identify whether the mask of the object cancels the persistent validity of the previous mask of the object. In some embodiments, the encoder may determine a mask cancellation flag for identifying whether the mask of the object cancels the persistent validity of the previous mask of the object, regardless of whether the cancellation flag indicates that the SEI message cancels the persistent validity of the information of the previous SEI message. For example, omi_mask_cancel[i][j] being equal to 1 cancels the persistent validity range of the object mask whose identifier is equal to omi_mask_id[i][j]. omi_mask_cancel[i][j] being equal to 0 indicates sending the information of the object mask whose identifier is equal to omi_mask_id[i][j].

[0185] The variable maskIdExist[i][id] equal to 1 indicates that the object mask with identifier id exists in the i-th object mask auxiliary image. The variable maskIdExist[i][id] equal to 0 indicates that the object mask with identifier id does not exist in the i-th object mask auxiliary image. maskIdExist[i][id] is initialized to 0 before decoding the current CVS.

[0186] omi_mask_confidence[i][j] indicates the confidence associated with the j-th object mask to be updated in the i-th object mask auxiliary image. -(omi_mask_confidence_length_minus1+1) The confidence level is in units of omi_mask_confidence[i][j][k], so that higher omi_mask_confidence[i][j][k] values ​​indicate higher confidence levels. The length of the omi_mask_confidence[i][j][k] syntax element is omi_mask_confidence_length_minus1+1 bits.

[0187] omi_mask_depth[i][j] indicates the object depth associated with the j-th object mask to be updated in the i-th object mask auxiliary image. Smaller values ​​of omi_mask_depth indicate a shorter distance to the object. The length of the omi_mask_depth[i][j][k] syntax element is omi_object_depth_length_minus1+1 bits.

[0188] omi_mask_label[i][j] specifies the content of the label associated with the j-th object mask to be updated in the i-th object mask auxiliary image. The length of the omi_mask_label[i][j] syntax element is less than or equal to 255 bytes, excluding the null termination byte.

[0189] In the embodiments associated with Tables 6 and 7, it is assumed that one or more primary image layers have been determined, and only the layer identifier of each object mask image layer is indicated in the OMISEI message. It is also assumed that all object mask image layers are associated with one or more primary image layers, and therefore, there is no need to send the primary image layer associated with the object mask auxiliary image layer.

[0190] However, VVC supports multiple primary image layers and multiple auxiliary image layers. An auxiliary image layer can be associated with more than one primary image layer, and one primary image can be associated with more than one auxiliary image layer. The NAL unit layer identifiers of the primary image layer and the auxiliary image layer are specified in the SDI SEI message. And for each auxiliary image layer, the primary image layer associated with the auxiliary image layer is also specified in the SDI SEI message. Return to reference Figure 7 , an OMI SEI message may be used to indicate the masks of the auxiliary pictures 702 and 704. Thus, the OMI SEI message is associated with the primary pictures 701 and 703. In some embodiments, the primary picture 701 may correspond to more than one auxiliary picture, which may also be indicated by the OMI SEI message.

[0191] Since there are multiple primary image layers in VVC and the OMI SEI message can only be applied to some primary image layers, the number of primary image layers and the layer identifier of each primary image layer to which the OMI SEI message is applied are sent in the OMI SEI message. According to the SDI SEI message, after the primary image layer is determined, there is no need to send the layer identifier of the object mask auxiliary image layer in the OMI SEI message. The layer identifier of the auxiliary image associated with the primary image layer with the layer identifier layerIdA can be derived based on the SDI SEI message.

[0192] The variable numAuxLayer[i] indicates the number of auxiliary picture layers associated with the primary picture layer with nuh_layer_id (nuh_layer_id is the syntax element name of the layer identifier) ​​equal to i. The variable associatedAuxLayer[j][i] indicates the value of nuh_layer_id of the i-th auxiliary picture layer associated with the primary picture layer with nuh_layer_id equal to j. numAuxLayer[i] and associatedAuxLayer[j][i] are derived from the SDI SEI message as follows.

[0193] The OMISEI message is shown in Table 8, and the semantics are provided as an example in Table 8 below. The definitions of the common features denoted as "high-level information" and the individual features denoted as "individual mask information" (some of which are omitted) can be inherited from the above embodiments with shared parameter / function names. Table 8: Example syntax of OMISEI message

[0194] The Object Mask Information (OMI) SEI message provides information about an object mask picture coded as an auxiliary picture. For any value of i in the range of 0 to sid_max_layers_minus1 (inclusive), the object mask auxiliary picture has nuh_layer_id equal to nuhLayerIdA, sdi_layer_id[i] equal to nuhLayerIdA, and sdi_aux_id[i] in the range of 128 to 159 (inclusive).

[0195] When an access unit contains an auxiliary picture picA in a layer with nuh_layer_id equal to nuhLayerIdA (indicated by the OMISEI message as an object mask auxiliary layer) and a primary picture picB in a layer with nuh_layer_id equal to nuhLayerIdB (indicated by the OMISEI message as a primary layer), the OMISEI message remains valid in output order until one or more of the following conditions are true:

[0196] - End of CLVS containing auxiliary image picA. - End of CLVS containing the main image picB. -CVS ends. - The bitstream ends.

[0197] The value of sdi_aux_id[sdi_associated_primary_layer_idx[i][j]] is equal to 0.

[0198] omi_cancel_flag equal to 1 indicates that the SEI message cancels the continuing effect of any previous object mask information SEI message in the output order associated with one or more primary picture layers to which this SEI is applied. omi_cancel_flag equal to 0 indicates that the subsequent object mask information, and the object mask information sent in this SEI message, will be used to update the existing object mask information of any previous SEI message.

[0199] omi_aux_id_minus128 plus 128 indicates the value of sdi_aux_id of the object mask auxiliary image. omi_aux_id_minus128 is in the range of 0 to 31 (inclusive).

[0200] When a CVS does not contain an SDI SEI message with sdi_aux_id[i] equal to omi_aux_id_minus128+128 (for any value of i), no picture in the CVS is associated with an OMI SEI message.

[0201] When an AU contains both an SDISEI message and an OMISEI message with sdi_aux_id[i] equal to omi_aux_id_minus128+128 (for any value of i), the SDISEI message precedes the OMISEI message in decoding order.

[0202] In some embodiments, the SEI message is associated with a plurality of primary pictures corresponding to the auxiliary picture to which the SEI message applies. The number of the plurality of primary pictures and the layer identifiers of the plurality of primary pictures may be determined as the common characteristics. For example, omi_num_primary_pic_layer_minus1 plus 1 indicates the number of primary picture layers associated with the object mask auxiliary picture layer to which this SEI message applies. The value of omi_num_primary_pic_layer_minus1 is in the range of 0 to sdi_max_layers_minus1.

[0203] Furthermore, omi_primary_pic_layer_id[i] specifies the nuh_layer_id value of the i-th primary picture layer to which this OMI SEI message applies. For any value of j in the range 0 to sid_max_layers_minus1, the value of sdi_aux_id[j] is equal to 0, making sdi_layer_id[j] equal to omi_primary_pic_layer_id[i].

[0204] omi_mask_id_length_minus8 plus 8 indicates the number of bits used to encode the omi_mask_id[i][j][k] syntax element.

[0205] omi_mask_confidence_info_present_flag equal to 1 indicates the presence of the omi_mask_confidence[i][j][k] syntax element. omi_mask_confidence_info_present_flag equal to 0 indicates the absence of the omi_mask_confidence[i][j][k] syntax element. Bitstream conformance requirement is that the value of omi_mask_confidence_info_present_flag is the same for all object_mask_info() syntax structures within a CLVS.

[0206] omi_mask_confidence_length_minus1 plus 1 specifies the length in bits of the omi_mask_confidence[i][j][k] syntax element. A bitstream conformance requirement is that the value of omi_mask_confidence_length_minus1 is the same for all object_mask_info() syntax structures within a CLVS.

[0207] omi_object_depth_info_present_flag equal to 1 indicates the presence of the omi_object_depth[i][j][k] syntax element. omi_object_depth_info_present_flag equal to 0 indicates the absence of the omi_object_depth[i][j][k] syntax element. A bitstream conformance requirement is that the value of omi_object_depth_info_present_flag is the same for all object_mask_info() syntax structures within a CLVS.

[0208] omi_object_depth_length_minus1 plus 1 specifies the length in bits of the omi_object_depth[i][j][k] syntax element. The requirement for bitstream conformance is that the value of omi_object_depth_length_minus1 is the same for all object_mask_info() syntax structures within a CLVS.

[0209] omi_mask_label_info_present_flag being equal to 1 indicates the presence of omi_mask_label_language_present_flag and omi_mask_label[i][j][k]. omi_mask_label_info_present_flag being equal to 0 indicates the absence of omi_mask_label_language_present_flag and omi_mask_label[i][j][k].

[0210] omi_mask_label_language_present_flag being equal to 1 indicates the presence of omi_mask_label_language. omi_mask_label_language_present_flag being equal to 0 indicates the absence of omi_mask_label_language and the language of the mask label is not specified.

[0211] omi_bit_equal_to_zero is equal to 0.

[0212] omi_mask_label_language contains a language tag specified by IETF RFC5646 followed by a null-terminating byte with a value equal to 0x00. The length of the omi_mask_label_language syntax element is less than or equal to 255 bytes, excluding the null-terminating byte. When absent, the language of the label is not specified.

[0213] omi_num_mask_in_pic[i][j] indicates the number of masks in the j-th object mask auxiliary image associated with the j-th primary image. omi_num_mask_in_pic[i][j] is in the range of 0 to (1<<BitDepthY)-1 (including the end values), where BitDepthY is the bit depth of the samples of the luminance component.

[0214] In some embodiments, the encoder may determine a number of auxiliary images corresponding to each of a plurality of primary images.The encoder may then determine the individual features for the mask represented by the auxiliary images for each of the plurality of primary images.

[0215] For example, the individual feature omi_mask_id[i][j][k] indicates the identifier of the kth object mask in the jth object mask auxiliary image associated with the i-th primary image. The object mask identifier associated with the sample position (x, y) in the jth object mask auxiliary image is equal to p[j][x][y], where p[j][x][y] refers to the luminance sample at position (x, y) in the decoded jth object mask auxiliary image.

[0216] The variable maskId[i][j] defines the object mask identifier of the kth object mask of the jth object mask auxiliary picture associated with the ith primary picture in the SEI message, derived as follows,

[0217] omi_mask_confidence[i][j][k] indicates the confidence of the kth object mask in the jth object mask auxiliary image associated with the i-th primary image in 2 -(omi_mask_confidence_length_minus1+1) The confidence level is in units of omi_mask_confidence[i][j][k], so that higher omi_mask_confidence[i][j][k] values ​​indicate higher confidence levels. The length of the omi_mask_confidence[i][j][k] syntax element is omi_mask_confidence_length_minus1+1 bits.

[0218] omi_mask_confidence[i][j][k] indicates the object depth associated with the kth object mask in the jth object mask auxiliary image associated with the i-th primary image. Smaller values ​​of omi_mask_depth indicate a shorter distance to the object. The length of the omi_mask_depth[i][j][k] syntax element is omi_object_depth_length_minus1+1 bits.

[0219] omi_mask_label[i][j][k] indicates the content of the label associated with the kth object mask in the jth object mask auxiliary image associated with the i-th primary image. The length of the omi_mask_label[i][j][k] syntax element is less than or equal to 255 bytes, excluding the null termination byte.

[0220] In some embodiments, similar to Table 7, the OMI SEI message may only transmit the mask information to be updated. Furthermore, for information that remains unchanged between this OMI SEI message and the previous OMI SEI message, the transmission of that information may be skipped to save bit overhead. The syntax is shown in Table 9 below (differences from Table 6 are italicized in Table 9). The common features denoted as "high-level information" and the definitions of the individual features denoted as "individual mask information" (some of which are omitted) may be inherited from the above-described embodiment with shared parameter / function names. Table 9: Example syntax of OMISEI message

[0221] The variable numAuxLayer[i] indicates the number of auxiliary picture layers associated with the primary picture layer with nuh_layer_id equal to i. The variable associatedAuxLayer[j][i] indicates the value of nuh_layer_id of the i-th auxiliary picture layer associated with the primary picture layer with nuh_layer_id equal to j. numAuxLayer[i] and associatedAuxLayer[j][i] are derived from the SDI SEI message as follows.

[0222] The Object Mask Information (OMI) SEI message provides information about an object mask picture coded as an auxiliary picture. For any value of i in the range of 0 to sid_max_layers_minus1 (inclusive), the object mask auxiliary picture has nuh_layer_id equal to nuhLayerIdA, sdi_layer_id[i] equal to nuhLayerIdA, and sdi_aux_id[i] in the range of 128 to 159 (inclusive).

[0223] When an access unit contains an auxiliary picture picA located in a layer whose nuh_layer_id is equal to nuhLayerIdA, and the layer is identified as an object mask auxiliary layer by the OMI SEI message; and contains a main picture picB located in a layer whose nuh_layer_id is equal to nuhLayerIdB, and the layer is identified as the main layer by the OMI SEI message, the OMI SEI message will remain valid in the output order until one or more of the following conditions are met: - End of CLVS containing the auxiliary image picA. - End of CLVS containing the main image picB. -CVS ends. - The bitstream ends.

[0224] The value of sdi_aux_id[sdi_associated_primary_layer_idx[i][j]] is equal to 0.

[0225] omi_cancel_flag equal to 1 indicates that the SEI message cancels the continuing effect of any previous object mask information SEI message in the output order associated with one or more primary picture layers to which this SEI is applied. omi_cancel_flag equal to 0 indicates that the subsequent object mask information, and the object mask information sent in this SEI message, will be used to update the existing object mask information of any previous SEI message.

[0226] omi_aux_id_minus128 plus 128 indicates the value of sdi_aux_id of the object mask auxiliary picture layer. omi_aux_id_minus128 is in the range of 0 to 31 (inclusive).

[0227] When a CVS does not contain an SDI SEI message with sdi_aux_id[i] equal to omi_aux_id_minus128+128 (for any value of i), no picture in the CVS is associated with an OMI SEI message.

[0228] When an AU contains both an SDI SEI message and an OMI SEI message with sdi_aux_id[i] equal to omi_aux_id_minus128+128 (for any value of i), the SDI SEI message precedes the OMI SEI message in decoding order.

[0229] omi_num_primary_pic_layer_minus1 plus 1 indicates the number of primary picture layers associated with the object mask auxiliary picture layer to which this SEI message applies. The value of omi_num_primary_pic_layer_minus1 is in the range of 0 to sdi_max_layers_minus1.

[0230] omi_primary_pic_layer_id[i] specifies the nuh_layer_id value of the i-th primary picture layer to which this OMI SEI message applies. For any value of j in the range 0 to sid_max_layers_minus1, the value of sdi_aux_id[j] is equal to 0, making sdi_layer_id[j] equal to omi_primary_pic_layer_id[i].

[0231] omi_mask_id_length_minus8 plus 8 indicates the number of bits used to encode the omi_mask_id[i][j][k] syntax element.

[0232] omi_mask_confidence_info_present_flag equal to 1 indicates the presence of the omi_mask_confidence[i][j][k] syntax element. omi_mask_confidence_info_present_flag equal to 0 indicates the absence of the omi_mask_confidence[i][j][k] syntax element. Bitstream conformance requirement is that the value of omi_mask_confidence_info_present_flag is the same for all object_mask_info() syntax structures within a CLVS.

[0233] omi_mask_confidence_length_minus1 plus 1 specifies the length in bits of the omi_mask_confidence[i][j][k] syntax element. A bitstream conformance requirement is that the value of omi_mask_confidence_length_minus1 is the same for all object_mask_info() syntax structures within a CLVS.

[0234] omi_object_depth_info_present_flag equal to 1 indicates the presence of the omi_object_depth[i][j][k] syntax element. omi_object_depth_info_present_flag equal to 0 indicates the absence of the omi_object_depth[i][j][k] syntax element. A bitstream conformance requirement is that the value of omi_object_depth_info_present_flag is the same for all object_mask_info() syntax structures within a CLVS.

[0235] omi_object_depth_length_minus1 plus 1 specifies the length in bits of the omi_object_depth[i][j][k] syntax element. A bitstream conformance requirement is that the value of omi_object_depth_length_minus1 is the same for all object_mask_info() syntax structures within a CLVS.

[0236] omi_mask_label_info_present_flag is equal to 1 to indicate that omi_mask_label_language_present_flag and omi_mask_label[i][j][k] are present. omi_mask_label_info_present_flag is equal to 0 to indicate that omi_mask_label_language_present_flag and omi_mask_label[i][j][k] are not present.

[0237] omi_mask_label_language_present_flag equal to 1 indicates that omi_mask_label_language is present. omi_mask_label_language_present_flag equal to 0 indicates that omi_mask_label_language is not present and the language of the mask label is not specified.

[0238] omi_bit_equal_to_zero is equal to 0.

[0239] The omi_mask_label_language syntax element contains the language tag specified by IETF RFC 5646 followed by a null termination byte equal to 0x00. The length of the omi_mask_label_language syntax element is less than or equal to 255 bytes, excluding the null termination byte. When not present, the language of the label is unspecified.

[0240] omi_mask_pic_update_flag[i][j] being equal to 1 indicates that the mask information of the j-th object mask auxiliary image associated with the i-th main image is to be sent. omi_mask_pic_update_flag[i][j] being equal to 0 indicates that the mask information of the j-th object mask auxiliary image associated with the i-th main image is not to be sent. When there is no such mask information of the j-th object mask auxiliary image associated with the i-th main image, the persistence mechanism is used, i.e., the information is inherited from the last OMI SEI message that sent the mask information of the j-th object mask auxiliary image associated with the i-th main image..

[0241] omi_num_mask_in_pic_update[i][j] indicates the number of object masks in the j-th auxiliary image associated with the i-th main image to be sent. omi_num_mask_in_pic[i][j] is in the range of 0 to (1<<BitDepthY)-1 (including the end values), where BitDepthY is the bit depth of the samples of the luminance component.

[0242] omi_mask_id[i][j][k] indicates the identifier of the k-th object mask in the j-th object mask auxiliary image associated with the i-th main image. The object mask identifier associated with the sample position (x, y) in the j-th object mask auxiliary image is equal to p[j][x][y], where p[j][x][y] refers to the luminance sample at position (x, y) in the decoded j-th object mask auxiliary image.

[0243] The variable maskId[i][j][k] specifies the object mask identifier of the k-th object mask in the j-th object mask auxiliary image associated with the i-th main image in the SEI message, derived as follows,

[0244] omi_mask_cancel[i][j][k] being equal to 1 indicates canceling the persistence scope of the object mask with the identifier equal to omi_mask_id[i][j][k]. omi_mask_cancel[i][j][k] being equal to 0 indicates sending the information of the object mask with the identifier equal to omi_mask_id[i][j].

[0245] The variable maskIdExist[i][j][k] being equal to 1 indicates that the object mask with identifier k is present in the j-th object mask auxiliary image associated with the i-th primary image. The variable maskIdExist[i][j][k] being equal to 0 indicates that the object mask with identifier k is not present in the j-th object mask auxiliary image associated with the i-th primary image. maskIdExist[i][j][k] is initialized with 0 before decoding the current CVS.

[0246] omi_mask_confidence[i][j][k] indicates the confidence of the kth object mask in the jth object mask auxiliary image associated with the i-th primary image in 2 -(omi_mask_confidence_length_minus1+1) The confidence level is in units of omi_mask_confidence[i][j][k], so that higher omi_mask_confidence[i][j][k] values ​​indicate higher confidence levels. The length of the omi_mask_confidence[i][j][k] syntax element is omi_mask_confidence_length_minus1+1 bits.

[0247] omi_mask_depth[i][j][k] indicates the object depth associated with the kth object mask in the jth object mask auxiliary image associated with the i-th primary image. Smaller values ​​of omi_mask_depth indicate a shorter distance to the object. The length of the omi_mask_depth[i][j][k] syntax element is omi_object_depth_length_minus1+1 bits.

[0248] omi_mask_label[i][j][k] indicates the content of the label associated with the kth object mask in the jth object mask auxiliary image associated with the i-th primary image. The length of the omi_mask_label[i][j][k] syntax element is less than or equal to 255 bytes, excluding the null termination byte.

[0249] In some of the above embodiments discussed in conjunction with Tables 7 and 9, to save bit overhead, only updated object mask information is sent. omi_mask_cancel[i][j][k], indicating whether the mask with ID equal to omi_mask_id[i][j][k] is canceled, is sent only when the mask with ID equal to omi_mask_id[i][j][k] already exists. Therefore, the decoder must maintain a list of object mask IDs to derive the variable maskIdExist[i][j][omi_mask_id[i][j][k]] to determine whether the object mask with ID equal to omi_mask_id[i][j][k] exists to parse syntax elements. This introduces a parsing dependency on the previous SEI message.

[0250] In some embodiments, omi_mask_cancel[i][j][k] is always sent. In addition, if there is no mask object with ID equal to omi_mask_id[i][j][k] before, the value of omi_mask_cancel[i][j][k] is forced to be 0, which indicates that the persistent scope of the object mask with ID equal to omi_mask_cancel[i][j][k] is not canceled. The syntax of this method is shown in Table 10 below (the differences from Table 6 are italicized in Table 10). The definitions of the common features denoted as "high-level information" and the individual features denoted as "individual mask information" (some of which are omitted) can be inherited from the above-mentioned embodiments with shared parameter / function names. Table 10: Syntax of OMI SEI message

[0251] The variable numAuxLayer[i] indicates the number of auxiliary picture layers associated with the primary picture layer with nuh_layer_id equal to i. The variable associatedAuxLayer[j][i] indicates the value of nuh_layer_id of the i-th auxiliary picture layer associated with the primary picture layer with nuh_layer_id equal to j. numAuxLayer[i] and associatedAuxLayer[j][i] are derived from the SDI SEI message as follows.

[0252] The Object Mask Information (OMI) SEI message provides information about an object mask picture coded as an auxiliary picture. For any value of i in the range of 0 to sid_max_layers_minus1 (inclusive), the object mask auxiliary picture has nuh_layer_id equal to nuhLayerIdA, sdi_layer_id[i] equal to nuhLayerIdA, and sdi_aux_id[i] in the range of 128 to 159 (inclusive).

[0253] When an access unit contains an auxiliary picture picA in a layer whose nuh_layer_id is equal to nuhLayerIdA (which is indicated by the OMI SEI message as an object mask auxiliary layer) and a primary picture picB in a layer whose nuh_layer_id is equal to nuhLayerIdB (which is indicated by the OMI SEI message as a primary layer), the OMI SEI message remains valid in the output order until one or more of the following conditions are true: When an access unit contains an auxiliary picture picA located in a layer whose nuh_layer_id is equal to nuhLayerIdA and the layer is identified as an object mask auxiliary layer by the OMI SEI message; and also contains a primary picture picB located in a layer whose nuh_layer_id is equal to nuhLayerIdB and the layer is identified as a primary layer by the OMI SEI message, the OMI SEI message remains valid in the output order until one or more of the following conditions are true: - End of CLVS containing the auxiliary image picA. - End of CLVS containing the main image picB. -CVS ends. - The bitstream ends.

[0254] The value of sdi_aux_id[sdi_associated_primary_layer_idx[i][j]] shall be equal to 0.

[0255] omi_cancel_flag equal to 1 indicates that the SEI message cancels the continuing effect of any previous object mask information SEI message in the output order associated with one or more primary picture layers to which this SEI is applied. omi_cancel_flag equal to 0 indicates that the subsequent object mask information, and the object mask information sent in this SEI message, will be used to update the existing object mask information of any previous SEI message.

[0256] omi_mask_pic_update_flag[i][j] being equal to 1 indicates that the mask information of the j-th object mask auxiliary picture associated with the i-th main picture is to be sent. omi_mask_pic_update_flag[i][j] being equal to 0 indicates that the mask information of the j-th object mask auxiliary picture associated with the i-th main picture is not to be sent. When there is no such mask information of the j-th object mask auxiliary picture associated with the i-th main picture, a persistent validity mechanism is used, i.e., the information is inherited from the last OMI SEI message that sent the mask information of the j-th object mask auxiliary picture associated with the i-th main picture.

[0257] omi_num_mask_in_pic_update[i][j] indicates the number of object masks in the j-th auxiliary picture associated with the i-th main picture to be sent. omi_num_mask_in_pic[i][j] shall be in the range of 0 to (1<<BitDepthY)-1 (including the end values), where BitDepthY is the bit depth of the samples of the luminance component.

[0258] omi_mask_id[i][j][k] indicates the identifier of the k-th object mask in the j-th object mask auxiliary picture associated with the i-th main picture. The object mask identifier associated with the sample position (x,y) in the j-th object mask auxiliary picture is equal to p[j][x][y], where p[j][x][y] refers to the luminance sample at position (x,y) in the decoded j-th object mask auxiliary picture.

[0259] The variable maskId[i][j][k] specifies the object mask identifier of the k-th object mask in the j-th object mask auxiliary picture associated with the i-th main picture in the SEI message, which is derived as follows.

[0260] omi_mask_cancel[i][j][k] being equal to 1 indicates canceling the persistent validity range of the object mask whose identifier is equal to omi_mask_id[i][j][k]. omi_mask_cancel[i][j][k] being equal to 0 indicates sending the information of the object mask whose identifier is equal to omi_mask_id[i][j].

[0261] When maskIdExist[i][j][k] is equal to 0, the value of omi_mask_cancel[i][j][k] shall be equal to 1. The value of omi_mask_cancel[i][j][k] shall be equal to 0 when omi_mask_id[i][j][k] has the same value as omiMaskId for the first time in the CLVS.

[0262] The variable maskIdExist[i][j][k] is derived as: maskIdExist[i][j][k] is initialized with 0 before decoding the current CVS. maskIdExist[i][j][omi_mask_id[i][j][k]]=! omi_mask_cancel[i][j][k].

[0263] omi_mask_confidence[i][j][k] indicates the confidence of the kth object mask in the jth object mask auxiliary image associated with the i-th primary image in 2 -(omi_mask_confidence_length_minus1+1) The confidence level is in units of omi_mask_confidence[i][j][k], so that higher omi_mask_confidence[i][j][k] values ​​indicate higher confidence levels. The length of the omi_mask_confidence[i][j][k] syntax element is omi_mask_confidence_length_minus1+1 bits.

[0264] omi_mask_depth[i][j][k] indicates the object depth associated with the kth object mask in the jth object mask auxiliary image associated with the i-th primary image. Smaller values ​​of omi_mask_depth indicate a shorter distance to the object. The length of the omi_mask_depth[i][j][k] syntax element is omi_object_depth_length_minus1+1 bits.

[0265] omi_mask_label[i][j][k] indicates the content of the label associated with the kth object mask in the jth object mask auxiliary picture associated with the i-th primary picture. The length of the omi_mask_label[i][j][k] syntax element shall be less than or equal to 255 bytes, excluding the null termination byte.

[0266] In some of the above embodiments, the pixel values of the object mask-assisted image represent mask IDs. The decoder determines the mask based on the decoded sample values of the object mask-assisted image. Therefore, the encoder must use lossless coding to encode the object mask-assisted image, otherwise the mask will be distorted. However, in some cases, the number of object masks is much smaller than the range of sample values. Therefore, in some embodiments, the sample values of the auxiliary image can be encoded in a lossy manner. Thus, given the object mask ID, for the decoded sample values that are different from any object mask ID, the decoder can restore the decoded sample values to the closest mask ID value.

[0267] maskID[i] (where i = 0 to n - 1) indicates the i-th object mask ID in the image, assuming that maskID[i] <= maskID[j] when i < j. The tolerance boundary is calculated as: th[i] = (maskID[i] + maskID[i + 1]) / 2, i = 0…n - 2

[0268] p[x][y] represents the decoded value of the sample with coordinates (x, y), and the mask ID associated with p[x][y], ID(p[x][y]) is derived as:

[0269] In some embodiments, a bounding box is sent for each object mask to locate the object. For example, when the cancel flag indicates that the SEI message does not cancel the continued effect of the information in the previous SEI message, the encoder can determine the bounding box of the mask surrounding the object in step 604 and encode the bounding box in the SEI message. In some embodiments, the encoder can determine the bounding box of the mask surrounding the object in step 604 regardless of whether the cancel flag indicates that the SEI message cancels the continued effect of the information in the previous SEI message. Therefore, on the decoder side, only the samples within the bounding box are checked, and for the samples outside the bounding box, regardless of their values, the samples are considered as the background. The coordinates of the bounding box of the sent object mask are defined on the cropped part of the decoded image, relative to the consistency cropping window specified by the active SPS. Additionally, to enable the encoder to flexibly select whether to send the bounding box to define the mask (range) or not send the bounding box to save bit overhead, a gating flag omi_mask_bounding_box_present_flag is added to make the sending of the bounding box parameters optional.

[0270] The syntax of this method is shown in Table 11 below (differences from Table 6 are italicized in Table 11). The definitions of the common features denoted as "high-level information" and the individual features denoted as "individual mask information" (some of which are omitted) can be inherited from the above embodiments with shared parameter / function names. Table 11: Syntax of OMI SEI message

[0271] The variable numAuxLayer[i] indicates the number of auxiliary picture layers associated with the primary picture layer with nuh_layer_id equal to i. The variable associatedAuxLayer[j][i] indicates the value of nuh_layer_id of the i-th auxiliary picture layer associated with the primary picture layer with nuh_layer_id equal to j. numAuxLayer[i] and associatedAuxLayer[j][i] are derived from the SDI SEI message as follows.

[0272] The Object Mask Information (OMI) SEI message provides information about an object mask picture coded as an auxiliary picture. The object mask auxiliary picture has nuh_layer_id equal to nuhLayerIdA, sdi_layer_id[i] equal to nuhLayerIdA, and sdi_aux_id[i] in the range of 128 to 159 (inclusive), where any value of i is in the range of 0 to sid_max_layers_minus1 (inclusive).

[0273] Using this SEI message requires the definition of the following variables: - The cropped image width and image height in units of luma samples, denoted in this paper by CroppedWidth and CroppedHeight respectively. -Consistent cropping window left offset, ConfWinLeftOffset -Consistent cropping window top offset, ConfWinTopOffset - Chroma format indicator, represented by ChromaFormatIdc in this document - Derive the variables SubWidthC and SubHeightC from ChromaFormatIdc

[0274] When an access unit contains an auxiliary picture picA in a layer with nuh_layer_id equal to nuhLayerIdA (indicated by the OMI SEI message as an object mask auxiliary layer) and a primary picture picB in a layer with nuh_layer_id equal to nuhLayerIdB (indicated by the OMI SEI message as a primary layer), the OMI SEI message remains valid in output order until one or more of the following conditions are true: - End of CLVS containing the auxiliary image picA. - End of CLVS containing the main image picB. -CVS ends. - The bitstream ends.

[0275] The value of sdi_aux_id[sdi_associated_primary_layer_idx[i][j]] shall be equal to 0.

[0276] omi_cancel_flag equal to 1 indicates that the SEI message cancels the continuing effect of any previous object mask information SEI message in the output order associated with one or more primary picture layers to which this SEI is applied. omi_cancel_flag equal to 0 indicates that the subsequent object mask information, and the object mask information sent in this SEI message, will be used to update the existing object mask information of any previous SEI message.

[0277] omi_mask_size_length_minus1 plus 1 specifies the length in bits of the omi_mask_top[i][j][k] and omi_mask_left[i][j][k] syntax elements. A bitstream conformance requirement is that the value of omi_mask_size_length_minus1 shall be the same for all object_mask_info() syntax structures within a CLVS.

[0278] omi_mask_pic_update_flag[i][j] equal to 1 indicates that the mask information of the j-th object mask auxiliary picture associated with the i-th primary picture is sent. omi_mask_pic_update_flag[i][j] equal to 0 indicates that the mask information of the j-th object mask auxiliary picture associated with the i-th primary picture is not sent. The persistence mechanism is used when the mask information of the j-th object mask auxiliary picture associated with the i-th primary picture does not exist, that is, the information is inherited from the last OMISEI message in which the mask information of the j-th object mask auxiliary picture associated with the i-th primary picture was sent.

[0279] omi_num_mask_in_pic_update[i][j] indicates the number of object masks in the j-th auxiliary image associated with the i-th primary image to be sent. omi_num_mask_in_pic_update[i][j] shall be in the range of 0 to (1<<BitDepthY)-1 (including the end values), where BitDepthY is the bit depth of the samples of the luminance component.

[0280] omi_mask_bounding_box_present_flag[i][j][k] being equal to 1 indicates the existence of the bounding box parameters omi_mask_top[i][j][k], omi_mask_left[i][j][k], omi_mask_width[i][j][k], and omi_mask_height[i][j][k] related to the k-th object mask in the j-th object mask auxiliary image associated with the i-th primary image. omi_num_mask_in_pic_update[i][j][k] being equal to 0 indicates the non-existence of the bounding box parameters omi_mask_top[i][j][k], omi_mask_left[i][j][k], omi_mask_width[i][j][k], and omi_mask_height[i][j][k] related to the k-th object mask in the j-th object mask auxiliary image associated with the i-th primary image.

[0281] omi_mask_id[i][j][k] indicates the identifier of the k-th object mask in the j-th object mask auxiliary image associated with the i-th primary image.

[0282] The variable maskId[i][j][k] specifies the object mask identifier of the k-th object mask in the j-th object mask auxiliary image associated with the i-th primary image in the SEI message, and is derived as follows:

[0283] For example, information about the bounding box may be generated and sent in the SEI message. The indicators omi_mask_top[i][j][k], omi_mask_left[i][j][k], omi_mask_width[i][j][k], and omi_mask_height[i][j][k] specify the coordinates of the top left corner and the width and height of the bounding box of the object identified by the identifier omi_mask_id[i][j][k] in the cropped decoded image, respectively, which are relative to the consistent cropping window specified by the activated SPS.

[0284] The value of omi_mask_left[i][j][k] should be in the range of 0 to (CroppedWidth / SubWidthC-1), inclusive, where CroppedWidth and SubWidthC are relative to the j-th object mask auxiliary image associated with the i-th primary image. When omi_mask_left[i][j][k] is not present, the value of omi_mask_left[i][j][k] is inferred to be 0.

[0285] The value of omi_mask_top[i][j][k] shall be in the range of 0 to (CroppedWidth / SubWidthC-1), inclusive, where CroppedHeight and SubHeightC are relative to the j-th object mask auxiliary image associated with the i-th primary image. When omi_mask_top[i][j][k] is not present, the value of omi_mask_top[i][j][k] is inferred to be 0.

[0286] The value of omi_mask_width[i][j][k] shall be in the range of 0 to (CroppedWidth / SubWidthC - omi_mask_left[i][j][k]), inclusive. When omi_mask_width[i][j][k] is not present, the value of omi_mask_width[i][j][k] is inferred to be (CroppedWidth / SubWidthC - omi_mask_left[i][j][k]).

[0287] The value of omi_mask_height[i][j][k] shall be in the range of 0 to (CroppedHeight / SubHeightC - omi_mask_top[i][j][k]), inclusive. When omi_mask_height[i][j][k] is not present, the value of omi_mask_height[i][j][k] is inferred to be (CroppedHeight / SubHeightC - omi_mask_top[i][j][k]).

[0288] The identified object mask is within the bounding box containing the luma samples with horizontal image coordinates from SubWidthC*(ConfWinLeftOffset+omi_mask_left[i][j][k]) to SubWidthC*(ConfWinLeftOffset+omi_mask_left[i][j][k]+omi_mask_width[i][j][k])-1 (inclusive), and vertical image coordinates from SubHeightC*(ConfWinTopOffset+omi_mask_top[i][j][k]) to SubHeightC*(ConfWinTopOffset+omi_mask_top[i][j][k]+omi_mask_height[i][j][k])-1 (inclusive).

[0289] The variable p[i][j][x][y] is the decoded value of the sample at the relative sample position (x, y) in the j-th object mask auxiliary image associated with the i-th primary image.

[0290] omi_mask_cancel[i][j][k] equal to 1 indicates canceling the persistent validity range of the object mask with the identifier equal to omi_mask_id[i][j][k]. omi_mask_cancel[i][j][k] equal to 0 indicates sending the information of the object mask with the identifier equal to omi_mask_id[i][j][k].

[0291] The variable maskIdExist[i][j][k] being equal to 1 indicates that the object mask with identifier k is present in the j-th object mask auxiliary image associated with the i-th primary image. The variable maskIdExist[i][j][k] being equal to 0 indicates that the object mask with identifier k is not present in the j-th object mask auxiliary image associated with the i-th primary image. maskIdExist[i][j][k] is initialized with 0 before decoding the current CVS.

[0292] omi_mask_confidence[i][j][k] indicates the confidence of the kth object mask in the jth object mask auxiliary image associated with the i-th primary image in 2 -(omi_mask_confidence_length_minus1+1) The confidence level is in units of omi_mask_confidence[i][j][k], so that higher omi_mask_confidence[i][j][k] values ​​indicate higher confidence levels. The length of the omi_mask_confidence[i][j][k] syntax element is omi_mask_confidence_length_minus1+1 bits.

[0293] omi_mask_depth[i][j][k] indicates the object depth associated with the kth object mask in the jth object mask auxiliary image associated with the i-th primary image. Smaller values ​​of omi_mask_depth indicate a shorter distance to the object. The length of the omi_mask_depth[i][j][k] syntax element is omi_object_depth_length_minus1+1 bits.

[0294] omi_mask_label[i][j][k] specifies the content of the label associated with the kth object mask in the jth object mask auxiliary picture associated with the i-th primary picture. The length of the omi_mask_label[i][j][k] syntax element shall be less than or equal to 255 bytes, excluding the null termination byte.

[0295] In some other embodiments, the bounding box parameters omi_mask_top[i][j][k], omi_mask_left[i][j][k], omi_mask_width[i][j][k], and omi_mask_height[i][j][k] are encoded with a fixed-length code, and the length is preset, such as 8, 16, or 32. In this case, there is no need to send omi_mask_size_length_minus1. As shown in Table 12 below, 16-bit encoding values are used to encode omi_mask_top[i][j][k], omi_mask_left[i][j][k], omi_mask_width[i][j][k], and omi_mask_height[i][j][k]. The definitions of the common features represented as "advanced information" and the individual features represented as "individual mask information" (some of which are omitted) can be inherited from the above embodiments with shared parameter / function names. Table 12: Syntax of OMI SEI Message

[0296] In some of the above embodiments, the sample value at position (x, y) is represented as p[x][y], indicating the mask identifier associated with the sample. For the samples with a bit depth equal to bitdepthY, the maximum number of mask identifiers is 1 << bitdepthY. However, when two object masks overlap with each other, the sample value cannot represent the mask identifier because there are multiple masks associated with the sample. Therefore, in some of the above embodiments, multiple object masks are used to assist the image. p[i][x][y] represents the sample value at position (x, y) of the i-th mask-assisted image. If there are two object masks with identifiers idA and idB associated with the sample position (x, y), then p[0][x][y] can be set to idA, and p[1][x][y] can be set to idB. Therefore, the maximum number of overlapping masks that can be supported is equal to the maximum number of object mask-assisted images.

[0297] In some embodiments, the mask identifier is not directly represented by the sample value, but can be represented by one bit of the sample value. That is, each bit of the sample value represents a different identifier of the mask. If a mask with identifier idA is associated with the sample at position (x,y), then the sample value p[x][y] at position (x,y) is equal to (1<<idA). Assuming the sample bit depth is bitdepthY, the maximum number of mask identifiers is bitdepthY. Using this method, the maximum number of mask identifiers supported by the object mask-assisted image is less than that in the previous embodiments. However, the case of mask overlap can be easily handled. For example, if there are two object masks with identifiers idA and idB associated with the sample position (x,y) (idA is not equal to idB because there are two different masks), then p[x][y] can be set to (1<<idA)+(1<<idB). And for the sample value at position (x,y), if the k-th bit is "1", then the sample (x,y) is covered by the k-th mask; if the k-th bit is "0", then the sample (x,y) is not covered by the k-th mask. Figure 9 shows an exemplary binary representation of the sample value p[x][y] according to some embodiments of the present disclosure. As Figure 9 shown, it is the binary representation of the sample value p[x][y]. The least significant bit is "0", so this means that the sample (x,y) is not covered by the 0-th mask (or the mask with identifier 0); the 1st bit position is also "0", which means that the sample (x,y) is not covered by the 1st mask (or the mask with identifier 1); the 2nd bit position and the most significant bit are both "1", so this means that the sample (x,y) is covered by the 2nd mask and the (bitdepthY-1)-th mask (i.e., the two masks with identifiers 2 and bitdepthY-1 overlapping at the sample position (x,y)).

[0298] To support more mask identifiers, multiple object mask-assisted images can be used. For example, there are m object mask-assisted images with indices from 0 to m-1 and bit depth equal to bitdepthY, and the object mask identifier associated with the sample position (x,y) is idA. The sample value of each mask-assisted image at position (x,y) can be derived as where p[i][x][y] is the sample value at position (x,y) in the i-th mask-assisted image.

[0299] In some of the above-described embodiments, the identifier of the object mask is represented by the sample values ​​within the mask region in the auxiliary image. Therefore, the encoder cannot change the sample values ​​of the mask region to optimize the encoding result, and cannot adjust the sample values ​​of the mask region in real time. In some embodiments, the auxiliary image may include multiple predetermined sample values, and the sample values ​​representing the mask of the object may be selected from the multiple predetermined sample values ​​based on the value difference between the multiple predetermined sample values. For example, if there are three object masks in the first frame, the encoder may set the mask sample values ​​of these three object masks to 64, 128, and 192, respectively (i.e., the identifiers of these three object masks are equal to 64, 128, and 192, respectively), because a longer sample value distance provides more sample recovery space and thus has more error tolerance. In the second frame, the two objects with identifiers equal to 128 and 192, respectively, leave the image, and only the mask with identifier equal to 64 remains in the image. Although changing the sample value of the mask from 64 to 128 may give more error resilience, the encoder cannot change the sample value because the sample value is the identifier of the mask.

[0300] To address the above issues, in some embodiments, the determination of the mask sample value is separated from the mask identifier, so that the mask sample value of a mask can change from frame to frame. This enables the encoder to optimize the encoding result by adjusting the sample value based on the number of masks in different frames.

[0301] The syntax is shown in Table 13 below, and the semantics are given in the table below. The syntax element omi_aux_sample_value[i][j][k] is the mask sample value for the object mask with identifier omi_mask_id[i][j][k] and is only sent when the syntax element omi_mask_id_equal_to_aux_sample_value_flag is equal to false, meaning that the mask sample value is different from the mask identifier. In the case where the mask sample value is different from the mask identifier, the bit lengths of the mask sample value and the mask identifier may be different. Therefore, two syntax elements, omi_mask_id_length and omi_aux_sample_value_length_minus8, are sent to indicate the bit lengths of the mask sample value and the mask identifier, respectively. The syntax and semantics of these syntax elements are italicized below. The definitions of the common features denoted as "high-level information" and the individual features denoted as "individual mask information" (some of which are omitted) can be inherited from the above-mentioned embodiments with shared parameter / function names. Table 13: Syntax of OMI SEI message

[0302] The Object Mask Information (OMI) SEI message provides information about an object mask picture coded as an auxiliary picture. The object mask auxiliary picture has nuh_layer_id equal to sdi_layer_id[i], and sdi_layer_id[i] is in the range of 128 to 159 (inclusive) for any value of i in the range of 0 to sid_max_layers_minus1 (inclusive). NOTE 1 - Each object mask auxiliary image layer is associated with one primary image layer, and one primary image layer may be associated with one or more object mask auxiliary image layers.

[0303] Using this SEI message requires the definition of the following variables: - The cropped image width and image height in units of luma samples, denoted in this paper by CroppedWidth and CroppedHeight respectively. -Consistent cropping window left offset, ConfWinLeftOffset -Consistent cropping window top offset, ConfWinTopOffset - Chroma format indicator, represented herein by ChromaFormatIdc.

[0304] The variables SubWidthC and SubHeightC are derived from ChromaFormatIdc.

[0305] When an access unit contains an auxiliary picture picA in a layer with nuh_layer_id equal to nuhLayerIdA (which is indicated by the OMI SEI message as an object mask auxiliary layer) and a primary picture picB in a layer with nuh_layer_id equal to nuhLayerIdB (which is indicated by the OMI SEI message as a primary layer), the OMI SEI message remains valid in output order until one or more of the following conditions are true: - End of CLVS containing the auxiliary image picA. - End of CLVS containing the main image picB. -CVS ends. - The bitstream ends.

[0306] omi_cancel_flag equal to 1 indicates that the SEI message cancels the continuing effect of any previous object mask information SEI message in the output order associated with one or more primary picture layers to which this SEI is applied. omi_cancel_flag equal to 0 indicates that the subsequent object mask information, and the object mask information sent in this SEI message, will be used to update the existing object mask information of any previous SEI message.

[0307] omi_aux_id_minus128 indicates the value of sdi_aux_id of the object mask auxiliary picture layer plus 128. om_aux_id_minus128 should be in the range of 0 to 31 (inclusive).

[0308] When a CVS does not contain an SDI SEI message with sdi_aux_id[i] equal to omi_aux_id_minus128+128 (for at least one value of i), no picture in the CVS shall be associated with an OMI SEI message.

[0309] When an AU contains both an SDI SEI message and an OMI SEI message with sdi_aux_id[i] equal to omi_aux_id_minus128+128 (for at least one i value), the SDI SEI message should precede the OMI SEI message in decoding order.

[0310] omi_num_primary_pic_layer_minus1 plus 1 indicates the number of primary picture layers associated with the object mask auxiliary picture layer to which this SEI message applies. The value of omi_num_primary_pic_layer_minus1 shall be in the range of 0 to sdi_max_layers_minus1, inclusive.

[0311] omi_primary_pic_layer_id[i] specifies the nuh_layer_id value of the i-th primary picture layer to which this OMI SEI message applies. If sdi_layer_id[j] is equal to omi_primary_pic_layer_id[i], then the value of sdi_aux_id[j] shall be equal to 0 for any value of j in the range of 0 to sid_max_layers_minus1 (inclusive).

[0312] omi_mask_id_equal_to_aux_sample_value_flag equal to 1 indicates that the identifier of the object mask is equal to the sample value within the mask. omi_mask_id_equal_to_aux_sample_value_flag equal to 0 indicates that the identifier of the object mask may be different from the sample value within the mask.

[0313] omi_mask_id_length specifies the length in bits of the omi_mask_id[i][j][k] syntax element (when present).

[0314] omi_aux_sample_value_length_minus8 plus 8 specifies the length in bits of the omi_aux_sample_value[i][j][k] syntax element. A bitstream conformance requirement is that the value of omi_aux_sample_value_length_minus8 plus 8 shall be equal to BitDepth Y .

[0315] omi_mask_id_length_minus8 plus 8 specifies the length in bits of the omi_mask_id[i][j][k] syntax element.

[0316] omi_mask_confidence_info_present_flag equal to 1 indicates the presence of the omi_mask_confidence[i][j][k] syntax element. omi_mask_confidence_info_present_flag equal to 0 indicates the absence of the omi_mask_confidence[i][j][k] syntax element. A bitstream conformance requirement is that the value of omi_mask_confidence_info_present_flag shall be the same for all object_mask_info() syntax structures within a CLVS.

[0317] omi_mask_confidence_length_minus1 plus 1 specifies the length in bits of the omi_mask_confidence[i][j][k] syntax element. A bitstream conformance requirement is that the value of omi_mask_confidence_length_minus1 shall be the same for all object_mask_info() syntax structures within a CLVS.

[0318] omi_object_depth_info_present_flag equal to 1 indicates the presence of the omi_object_depth[i][j][k] syntax element. omi_object_depth_info_present_flag equal to 0 indicates the absence of the omi_object_depth[i][j][k] syntax element. A bitstream conformance requirement is that the value of omi_object_depth_info_present_flag shall be the same for all object_mask_info() syntax structures within a CLVS.

[0319] omi_object_depth_length_minus1 plus 1 specifies the length in bits of the omi_object_depth[i][j][k] syntax element. A bitstream conformance requirement is that the value of omi_object_depth_length_minus1 shall be the same for all object_mask_info() syntax structures within a CLVS.

[0320] omi_mask_label_info_present_flag equal to 1 indicates that the omi_mask_label_language_present_flag and omi_mask_label[i][j][k] syntax elements are present. omi_mask_label_info_present_flag equal to 0 indicates that the omi_mask_label_language_present_flag and omi_mask_label[i][j][k] syntax elements are not present.

[0321] omi_mask_label_language_present_flag equal to 1 indicates that the omi_mask_label_language syntax element is present. omi_mask_label_language_present_flag equal to 0 indicates that the omi_mask_label_language syntax element is not present.

[0322] omi_bit_equal_to_zero should be equal to 0.

[0323] omi_mask_label_language contains a language tag specified by IETF RFC5646 followed by a null-terminating byte with a value equal to 0x00. The length of the omi_mask_label_language syntax element shall be less than or equal to 255 bytes, excluding the null-terminating byte. When not present, the language of the label is unspecified.

[0324] omi_mask_pic_update_flag[i][j] being equal to 1 indicates that the mask information of the j-th object mask auxiliary picture associated with the i-th primary picture is to be sent. omi_mask_pic_update_flag[i][j] being equal to 0 indicates that the mask information of the j-th object mask auxiliary picture associated with the i-th primary picture is not to be sent. When the mask information of the j-th object mask auxiliary picture associated with the i-th primary picture is not present, the persistence mechanism is used, i.e., the information is inherited from the last OMI SEI message that sent the mask information of the j-th object mask auxiliary picture associated with the i-th primary picture.

[0325] omi_num_mask_in_pic_update[i][j] indicates the number of object masks whose information is to be sent in the j-th auxiliary picture associated with the i-th primary picture. omi_num_mask_in_pic_update[i][j] shall be in the range of 0 to (1<<BitDepthY)-1 (including the end values), where BitDepthY is the bit depth of the samples of the luminance component. The variable omiNumMaskInPic[i][j] indicates that when the current SEI message is the first OMI SEI message in the current CLVS, the number of object masks in the j-th auxiliary picture associated with the i-th primary picture is set to omi_num_mask_in_pic_update[i][j].

[0326] The variable numAuxLayer[primaryLayerId] indicates the number of auxiliary image layers associated with the primary image layer whose nuh_layer_id is equal to primaryLayerId. The variable associatedAuxLayerId[primaryLayerId][i] indicates the value of the nuh_layer_id of the i-th auxiliary image layer, which is associated with the primary image layer whose nuh_layer_id is equal to primaryLayerId. numAuxLayer[primaryLayerId] and associatedAuxLayerId[primaryLayerId][i] are derived as follows:

[0327] omi_mask_id[i][j][k] indicates the identifier of the kth object mask in the jth object mask auxiliary image associated with the i-th primary image.

[0328] omi_aux_sample_value[i][j][k] indicates the sample value within the object mask whose identifier is equal to omi_mask_id[i][j][k].

[0329] The variable maskId[i][j][k] specifies the object mask identifier of the kth object mask in the jth object mask auxiliary picture associated with the i-th primary picture in the SEI message, derived as follows:

[0330] omi_mask_bounding_box_present_flag[i][j][k] equal to 1 indicates that the syntax elements omi_mask_top[i][j][k], omi_mask_left[i][j][k], omi_mask_width[i][j][k], and omi_mask_height[i][j][k] are present. omi_num_mask_in_pic_update[i][j][k] equal to 0 indicates that the syntax elements omi_mask_top[i][j][k], omi_mask_left[i][j][k], omi_mask_width[i][j][k], and omi_mask_height[i][j][k] are not present.

[0331] omi_mask_top[i][j][k], omi_mask_left[i][j][k], omi_mask_width[i][j][k] and omi_mask_height[i][j][k] indicate the coordinates of the top left corner and the width and height of the bounding box of the object mask with identifier equal to omi_mask_id[i][j][k] in the cropped decoded image, respectively, which are relative to the consistent cropping window specified by the activated SPS.

[0332] The value of omi_mask_left[i][j][k] should be in the range of 0 to (CroppedWidth / SubWidthC-1) (inclusive), where CroppedWidth and SubWidthC are related to the j-th object mask auxiliary image associated with the i-th primary image. When it does not exist, the value of omi_mask_left[i][j][k] is inferred to be 0.

[0333] The value of omi_mask_top[i][j][k] should be in the range of 0 to (CroppedWidth / SubWidthC-1) (inclusive), where CroppedHeight and SubHeightC are relative to the j-th object mask auxiliary image associated with the i-th primary image. When it does not exist, the value of omi_mask_top[i][j][k] is inferred to be 0.

[0334] The value of omi_mask_width[i][j][k] shall be in the range of 0 to (CroppedWidth / SubWidthC - omi_mask_left[i][j][k]), inclusive. When it is not present, the value of omi_mask_width[i][j][k] is inferred to be (CroppedWidth / SubWidthC - omi_mask_left[i][j][k]).

[0335] The value of omi_mask_height[i][j][k] should be in the range of 0 to (CroppedHeight / SubHeightC - omi_mask_top[i][j][k]), inclusive. When it is not present, the value of omi_mask_height[i][j][k] is inferred to be (CroppedHeight / SubWidthC - omi_mask_top[i][j][k]).

[0336] The identified object mask is within the bounding box containing the luma samples, with horizontal coordinates from SubWidthC*(ConfWinLeftOffset+omi_mask_left[i][j][k]) to SubWidthC*(ConfWinLeftOffset+omi_mask_left[i][j][k]+omi_mask_width[i][j][k])-1 (inclusive), and vertical coordinates from SubHeightC*(ConfWinTopOffset+omi_mask_top[i][j][k]) to SubHeightC*(ConfWinTopOffset+omi_mask_top[i][j][k]+omi_mask_height[i][j][k])-1 (inclusive).

[0337] The variable I[i][j][x][y] is the decoded value of the sample at the relative sample position (x, y) in the jth object mask auxiliary image associated with the i-th primary image. The following process is used to determine each mask region in each auxiliary image.

[0338] omi_mask_cancel[i][j][k] equal to 1 indicates canceling the persistent validity range of the object mask with the identifier equal to om_mask_id[i][j][k]. omi_mask_cancel[i][j][k] equal to 0 indicates sending the information of the object mask with the identifier equal to omi_mask_id[i][j].

[0339] A bitstream conformance requirement is that when an omi_mask_id[i][j][k] with a particular value is parsed for the first time in the current CLVS, the value of the corresponding omi_mask_cancel[i][j][k] shall be equal to 0.

[0340] omi_mask_confidence[i][j][k] indicates the confidence of the kth object mask in the jth object mask auxiliary image associated with the i-th primary image in 2 -(omi_mask_confidence_length_minus1+1) The confidence level is in units of omi_mask_confidence[i][j][k], so that higher omi_mask_confidence[i][j][k] values ​​indicate higher confidence levels. The length of the omi_mask_confidence[i][j][k] syntax element is omi_mask_confidence_length_minus1+1 bits.

[0341] omi_mask_depth[i][j][k] indicates the object depth associated with the kth object mask in the jth object mask auxiliary image associated with the i-th primary image. Smaller values ​​of omi_mask_depth indicate a shorter distance to the object. The length of the omi_mask_depth[i][j][k] syntax element is omi_object_depth_length_minus1+1 bits.

[0342] omi_mask_label[i][j][k] specifies the content of the label associated with the kth object mask in the jth object mask auxiliary picture associated with the i-th primary picture. The length of the omi_mask_label[i][j][k] syntax element shall be less than or equal to 255 bytes, excluding the null termination byte.

[0343] In some embodiments, methods for detecting an object are also provided. Figure 10 is a schematic diagram illustrating an exemplary method 1000 for detecting an object consistent with an embodiment of the present disclosure. Figure 10 As shown, the method 1000 may include steps 1002 to 1006, which may be performed by a decoder (e.g., Figure 1 Image / video decoder 144 or Figure 4 400 in the device).

[0344] In step 1002, the decoder may receive a bitstream. The bitstream may be encoded according to any of the above encoding methods.

[0345] In step 1004, the decoder may decode the encoded information of the bitstream to obtain a primary image and an auxiliary image. The auxiliary image may be used to indicate a mask of an object in the primary image. The mask of the object may be represented by sample values ​​of the auxiliary image.

[0346] In step 1006, the decoder may decode the encoded information of the bitstream to obtain a supplemental enhancement information (SEI) message associated with the primary image and applied to the auxiliary image. As described above, the SEI message may be used to indicate one or more attributes of the mask of the object.

[0347] In some embodiments, a non-transitory computer-readable storage medium storing a bitstream is further provided. The bitstream can be encoded and decoded according to the above method. Figure 11 is a schematic diagram showing the contents of an exemplary bitstream 1100. Figure 11As shown, a bitstream 1100 may be used to transmit a primary image 1101, an auxiliary image 1102, and a supplemental enhancement information (SEI) message 1103 (eg, Figure 5 ). Auxiliary image 1102 indicates a mask of the object in primary image 1101, wherein the mask of the object can be represented by sample values ​​of auxiliary image 1102. SEI message 1103 is associated with primary image 1101 and is applicable to auxiliary image 1102. SEI message 1103 can be used to indicate one or more attributes of the mask of the object, as described above.

[0348] In some embodiments, a non-transitory computer-readable storage medium comprising instructions is also provided, and the instructions can be executed by an apparatus (e.g., the disclosed encoder and decoder) for performing the above method. Common forms of non-transitory media include, for example, floppy disks, flexible magnetic disks, hard disks, solid-state drives, tapes or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a pattern of holes, RAM, PROM and EPROM, FLASH-EPROM or any other flash memory, NVRAM, caches, registers, any other memory chips or cassettes, and networked versions thereof. The apparatus may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memories.

[0349] The embodiments may be further described using the following terms: 1. A method for encoding a video sequence into a bitstream, the method comprising: receiving a video sequence; and Encoding one or more images of the video sequence to generate a bitstream, comprising: encoding an auxiliary image indicating a mask of an object in a primary image, the mask of the object being represented by sample values ​​of the auxiliary image; and A supplemental enhancement information (SEI) message is generated that indicates attributes of the mask for the object. 2. The method of clause 1, wherein generating the SEI message comprises: A cancellation flag is determined, the cancellation flag identifying whether the SEI message cancels the continuing validity of a previous SEI message. 3. A method according to clause 2, wherein the properties of the mask include: individual characteristics of the mask of the object indicated by the SEI message and common characteristics of multiple masks, and The generating of the SEI message further includes: In response to determining that the cancellation flag indicates that the SEI message does not cancel the continuing validity of information of the previous SEI message, the common characteristic and the individual characteristic are determined. 4. The method of clause 1, wherein the attributes of the mask include: individual characteristics of the mask of the object indicated by the SEI message and common characteristics of multiple masks, and the common characteristics include at least one of the following: an identifier of the auxiliary picture to which the SEI message applies; a number of bits used to encode an identifier of any mask in the plurality of masks; the bit depth of the sample values ​​of the auxiliary image; a confidence presence flag, which identifies whether the confidence information of the multiple masks is included in the SEI message; lengths of the confidence information of the plurality of masks, which are present when the confidence presence flag indicates that the confidence information of the plurality of masks is included in the SEI message; a depth presence flag, which identifies whether the depth information of the plurality of masks is included in the SEI message; lengths of the depth information of the plurality of masks, which are present when the depth presence flag indicates that the depth information of the plurality of masks is included in the SEI message; a label presence flag, which identifies whether the label information of the plurality of masks is included in the SEI message; a language presence flag, indicating whether the label language information of the plurality of masks is included in the SEI message, which is present when the label presence flag indicates that the label information of the plurality of masks is included in the SEI message; or The plurality of masked tag language information exists when the language presence flag identifies that the plurality of masked tag language information is included in the SEI message. 5. The method of clause 4, wherein the SEI message applies to a plurality of auxiliary pictures, and the common characteristic further comprises a number of the plurality of auxiliary pictures. 6. The method of clause 5, wherein the SEI message includes the individual features generated for a mask represented by the plurality of auxiliary images. 7. A method according to any of clauses 4 to 6, wherein the SEI message is associated with a plurality of primary pictures corresponding to the auxiliary pictures to which the SEI message applies. 8. The method of clause 7, wherein the common characteristics further comprise: a number of the plurality of master images and layer identifiers of the plurality of master images. 9. The method of clause 8, wherein generating the SEI message comprises: determining a second number of the auxiliary images corresponding to each of the plurality of primary images; and For each of the plurality of primary images, the individual features of the mask characterized by a second number of the auxiliary images are determined. 10. The method of any of clauses 1 to 9, wherein generating the SEI message further comprises: determining whether the mask of the object is different from a previous mask of the object represented by a previous auxiliary image; and In response to determining that the mask for the object is different than the previous mask for the object, the attributes of the mask for the object are encoded in the SEI message. 11. The method of clause 10, further comprising: In response to determining that the mask for the object is the same as the previous mask for the object, encoding the attributes of the mask for the object in the SEI message is skipped. 12. The method of any of clauses 1 to 11, wherein generating the SEI message further comprises: A mask cancel flag is determined, the mask cancel flag identifying whether the mask of the object cancels the ongoing effect of a previous mask of the object. 13. A method according to any of clauses 1 to 12, wherein generating the SEI message comprises: determining a bounding box of the mask surrounding the object; and The bounding box is encoded in the SEI message. 14. A method according to any of clauses 1 to 13, wherein the sample values ​​of the auxiliary image are encoded in a lossy manner. 15. A method according to any of clauses 1 to 13, wherein the mask of the object is indicated by one bit of the sample value of the auxiliary image. 16. A method according to any of clauses 1 to 13, wherein the mask of the object is indicated by sample values ​​of the auxiliary image. 17. The method of clause 16, wherein the sample value is included in the SEI message. 18. A method according to any one of clauses 1 to 13, wherein the auxiliary image comprises: a plurality of predetermined sample values, and the sample values ​​of the mask used to characterize the object are selected from the plurality of predetermined sample values ​​based on the value difference between the plurality of predetermined sample values. 19. A method for detecting an object, the method comprising: Receive bit stream; decoding the encoded information of the bitstream to obtain a main image and an auxiliary image, wherein the auxiliary image indicates a mask of an object in the main image, and the mask of the object is represented by sample values ​​of the auxiliary image; and The encoded information of the bitstream is decoded to obtain a Supplemental Enhancement Information (SEI) message, the SEI message indicating attributes of the mask of the object. 20. The method of clause 19, wherein decoding the encoded information of the bitstream to obtain the SEI message comprises: A cancellation flag is determined, the cancellation flag identifying whether the SEI message cancels the continuing validity of a previous SEI message. 21. A method according to clause 20, wherein the properties of the mask include individual characteristics of the mask of the object indicated by the SEI message and common characteristics of multiple masks, and The decoding of the encoded information of the bitstream to obtain the SEI message includes: In response to determining that the cancellation flag indicates that the SEI message does not cancel the continuing validity of information of the previous SEI message, the common characteristic and the individual characteristic are determined. 22. The method of clause 19, wherein the properties of the mask include individual characteristics of the mask of the object indicated by the SEI message and common characteristics of multiple masks, and the common characteristics include at least one of the following: The identifier of the auxiliary picture to which the SEI message applies; a number of bits used to encode an identifier of any mask in the plurality of masks; the bit depth of the sample values ​​of the auxiliary image; a confidence presence flag, which identifies whether the confidence information of the multiple masks is included in the SEI message; lengths of the confidence information of the plurality of masks, which are present when the confidence presence flag indicates that the confidence information of the plurality of masks is included in the SEI message; a depth presence flag, which identifies whether the depth information of the plurality of masks is included in the SEI message; lengths of the depth information of the plurality of masks, which are present when the depth presence flag indicates that the depth information of the plurality of masks is included in the SEI message; a label presence flag, which identifies whether the label information of the plurality of masks is included in the SEI message; a language presence flag, indicating whether the label language information of the plurality of masks is included in the SEI message, which is present when the label presence flag indicates that the label information of the plurality of masks is included in the SEI message; or The tag language information of the plurality of masks exists when the language presence flag identifies that the tag language information of the plurality of masks is included in the SEI message. 23. The method of clause 22, wherein the SEI message applies to a plurality of auxiliary pictures, and the common characteristic further comprises: a number of the plurality of auxiliary pictures. 24. The method of clause 23, wherein the SEI message comprises the individual features generated for a mask represented by the plurality of auxiliary images. 25. A method according to any of clauses 22 to 24, wherein the SEI message is associated with a plurality of primary pictures corresponding to the auxiliary pictures to which the SEI message applies. 26. The method of clause 25, wherein the common characteristics further comprise: a number of the plurality of master images and layer identifiers of the plurality of master images. 27. The method of clause 26, wherein decoding the encoded information of the bitstream to obtain the SEI message comprises: determining a second number of the auxiliary images corresponding to each of the plurality of primary images; and For each of the plurality of primary images, the individual features of the mask characterized by a second number of the auxiliary images are determined. 28. The method of any of clauses 19 to 27, wherein decoding the encoded information of the bitstream to obtain the SEI message further comprises: A mask cancel flag is determined, the mask cancel flag identifying whether the mask of the object cancels the ongoing effect of a previous mask of the object. 29. The method of any of clauses 19 to 28, wherein decoding the encoded information of the bitstream to obtain the SEI message further comprises: A bounding box surrounding the mask of the object is determined based on the SEI message. 30. A method according to any of clauses 19 to 29, wherein the mask of the object is indicated by one bit of the sample value of the auxiliary image. 31. A method according to any of clauses 19 to 29, wherein the mask of the object is indicated by sample values ​​of the auxiliary image. 32. The method of clause 31 , wherein the sample value is included in the SEI message. 33. The method of clause 32, wherein decoding the encoded information of the bitstream to obtain the SEI message further comprises: The sample values ​​of the auxiliary image are determined to represent the mask having the same identifier or the closest identifier in value. 34. A method according to any one of clauses 19 to 29, wherein the auxiliary image comprises a plurality of predetermined sample values, and the sample values ​​of the mask for characterizing the object are selected from the plurality of predetermined sample values ​​based on value differences between the plurality of predetermined sample values. 35. A non-transitory computer-readable storage medium storing a video bitstream, the bitstream comprising: A main image containing the object; an auxiliary image indicating a mask of the object, the mask of the object being represented by sample values ​​of the auxiliary image; and A supplemental enhancement information (SEI) message indicating attributes of the mask for the object. 36. The non-transitory computer-readable storage medium of clause 35, wherein the SEI message includes a cancel flag indicating whether the SEI message cancels the continuing validity of a previous SEI message. 37. A non-transitory computer-readable storage medium according to clause 36, wherein, in response to the cancellation flag indicating that the SEI message does not cancel the continuing effectiveness of the information of the previous SEI message, the attributes of the mask include: individual characteristics of the mask of the object indicated by the SEI message and common characteristics of multiple masks. 38. The non-transitory computer-readable storage medium of clause 35, wherein the attributes of the mask include individual characteristics of the mask of the object indicated by the SEI message and common characteristics of multiple masks, and the common characteristics include at least one of the following: an identifier of the auxiliary picture to which the SEI message applies; a number of bits used to encode an identifier of any mask in the plurality of masks; the bit depth of the sample values ​​of the auxiliary image; a confidence presence flag, which identifies whether the confidence information of the multiple masks is included in the SEI message; the length of the plurality of masked confidence information, which is present when the confidence presence flag indicates that the plurality of masked confidence information is included in the SEI message; a depth presence flag, which identifies whether the depth information of the plurality of masks is included in the SEI message; lengths of the depth information of the plurality of masks, which are present when the depth presence flag indicates that the depth information of the plurality of masks is included in the SEI message; a label presence flag, which identifies whether the label information of the plurality of masks is included in the SEI message; a language presence flag, indicating whether the label language information of the plurality of masks is included in the SEI message, which is present when the label presence flag indicates that the label information of the plurality of masks is included in the SEI message; or The plurality of masked tag language information exists when the language presence flag identifies that the plurality of masked tag language information is included in the SEI message. 39. The non-transitory computer-readable storage medium of clause 38, wherein the SEI message applies to a plurality of auxiliary pictures, and the common characteristic further comprises a number of the plurality of auxiliary pictures. 40. The non-transitory computer-readable storage medium of clause 39, wherein the SEI message comprises the individual features generated for a mask represented by the plurality of auxiliary images. 41. The non-transitory computer-readable storage medium of any of clauses 38 to 40, wherein the SEI message is associated with a plurality of primary pictures corresponding to the auxiliary pictures to which the SEI message applies. 42. The non-transitory computer-readable storage medium of clause 41, wherein the common characteristics further comprise: a number of the plurality of master images and layer identifiers of the plurality of master images. 43. The non-transitory computer-readable storage medium of clause 42, wherein the SEI message is generated based further on: determining a second number of the auxiliary images corresponding to each of the plurality of primary images; and For each of the plurality of primary images, the individual features of the mask characterized by a second number of the auxiliary images are determined. 44. The non-transitory computer-readable storage medium of any of clauses 35 to 43, wherein the SEI message is further generated based on: determining whether the mask of the object is different from a previous mask of the object represented by a previous auxiliary image; and In response to determining that the mask for the object is different than the previous mask for the object, the attributes of the mask for the object are encoded in the SEI message. 45. The non-transitory computer-readable storage medium of clause 44, wherein the SEI message is generated based further on: In response to determining that the mask for the object is the same as the previous mask for the object, encoding the attributes of the mask for the object in the SEI message is skipped. 46. ​​The non-transitory computer-readable storage medium of any of clauses 35 to 45, wherein the SEI message is generated based further on: A mask cancel flag is determined, the mask cancel flag identifying whether the mask of the object cancels the ongoing effect of a previous mask of the object. 47. The non-transitory computer-readable storage medium of any of clauses 35 to 46, wherein the SEI message is further generated based on: determining a bounding box of the mask surrounding the object; and The bounding box is encoded in the SEI message. 48. The non-transitory computer-readable storage medium of any of clauses 35 to 47, wherein the sample values ​​of the auxiliary image are encoded in a lossy manner. 49. The non-transitory computer-readable storage medium of any of clauses 35 to 47, wherein the mask of the object is indicated by one bit of the sample value of the auxiliary image. 50. The non-transitory computer-readable storage medium of any of clauses 35 to 47, wherein the mask of the object is indicated by sample values ​​of the auxiliary image. 51. The non-transitory computer-readable storage medium of clause 50, wherein the sample value is included in the SEI message. 52. A non-transitory computer-readable storage medium according to any one of clauses 35 to 47, wherein the auxiliary image includes: a plurality of predetermined sample values, and the sample values ​​of the mask used to characterize the object are selected from the plurality of predetermined sample values ​​based on the value difference between the plurality of predetermined sample values.

[0350] It should be noted that relational terms such as "first," "second," and the like in this document are used only to distinguish one entity or operation from another entity or operation, and do not require or imply any actual relationship or order between these entities or operations. In addition, the words "comprising," "having," "containing," and "including," and other similar forms are intended to be synonymous and open-ended, in that one or more items following any of these words is not intended to be an exhaustive list of such one or more items, nor is it intended to be limited to the listed one or more items.

[0351] As used herein, unless otherwise specifically stated, the term "or" encompasses all possible combinations unless not feasible. For example, if a database is specified to include either A or B, then unless otherwise specifically stated or not feasible, the database may include either A, or B, or A and B. As a second example, if a database is specified to include either A, B, or C, then unless otherwise specifically stated or not feasible, the database may include either A, or B, or C, or A and B, or A and C, or B and C, or A, B, and C.

[0352] It should be understood that the above embodiments can be implemented by hardware, or software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above-mentioned computer-readable medium. When executed by a processor, the software can execute the disclosed method. The computing unit and other functional units described in this disclosure can be implemented by hardware, or software, or a combination of hardware and software. Those of ordinary skill in the art will also understand that multiple modules / units in the modules / units described above can be combined into one module / unit, and each module / unit described above can be further divided into multiple sub-modules / sub-units.

[0353] In the above description, embodiments have been described with reference to many specific details, which may vary depending on the implementation. Certain adjustments and modifications may be made to the described embodiments. Other embodiments will be apparent to those skilled in the art by considering the description and practice of the invention disclosed herein. The description and examples are intended to be exemplary only, with the true scope and spirit of the invention being indicated by the appended claims. The order of steps shown in the figures is also intended to be for illustrative purposes only and is not intended to be limited to any particular order of steps. Therefore, it will be understood by those skilled in the art that these steps may be performed in different orders while implementing the same method.

[0354] In the drawings and the specification, exemplary embodiments have been disclosed. However, many variations and modifications may be made to these embodiments. Therefore, although specific terms are employed, these terms are used in a generic and descriptive sense only and not for purposes of limitation.

Claims

1. A method for encoding a video sequence into a bitstream, the method comprising: receiving a video sequence; as well as Encoding one or more images of the video sequence to generate a bitstream, comprising: encoding an auxiliary image indicating a mask of an object in a primary image, the mask of the object being represented by sample values ​​of the auxiliary image; as well as A supplemental enhancement information (SEI) message is generated that indicates attributes of the mask of the object.

2. The method according to claim 1, wherein Generating the SEI message includes: A cancellation flag is determined, the cancellation flag identifying whether the SEI message cancels the continuing validity of a previous SEI message.

3. The method according to claim 2, wherein: The attributes of the mask include: individual characteristics of the mask of the object indicated by the SEI message and common characteristics of multiple masks, and The generating of the SEI message further includes: In response to determining that the cancellation flag indicates that the SEI message does not cancel the continuing validity of information of the previous SEI message, the common characteristic and the individual characteristic are determined.

4. The method according to claim 1, wherein The attributes of the mask include: individual characteristics of the mask of the object indicated by the SEI message and common characteristics of multiple masks, and the common characteristics include at least one of the following: an identifier of the auxiliary picture to which the SEI message applies; a number of bits used to encode an identifier of any mask in the plurality of masks; the bit depth of the sample values ​​of the auxiliary image; a confidence presence flag, which identifies whether the confidence information of the multiple masks is included in the SEI message; lengths of the confidence information of the plurality of masks, which are present when the confidence presence flag indicates that the confidence information of the plurality of masks is included in the SEI message; a depth presence flag, which identifies whether the depth information of the plurality of masks is included in the SEI message; lengths of the depth information of the plurality of masks, which are present when the depth presence flag indicates that the depth information of the plurality of masks is included in the SEI message; a label presence flag, which identifies whether the label information of the plurality of masks is included in the SEI message; a language presence flag, indicating whether the tag language information of the plurality of masks is included in the SEI message, which is present when the tag presence flag indicates that the tag information of the plurality of masks is included in the SEI message; or The plurality of masked tag language information exists when the language presence flag identifies that the plurality of masked tag language information is included in the SEI message.

5. The method according to claim 1, wherein Generating the SEI message further includes: determining whether the mask of the object is different from a previous mask of the object represented by a previous auxiliary image; and In response to determining that the mask for the object is different than the previous mask for the object, the attributes of the mask for the object are encoded in the SEI message.

6. The method according to claim 5, further comprising: In response to determining that the mask for the object is the same as the previous mask for the object, encoding the attributes of the mask for the object in the SEI message is skipped.

7. The method according to claim 1, wherein Generating the SEI message further includes: A mask cancel flag is determined, the mask cancel flag identifying whether the mask of the object cancels the ongoing effect of a previous mask of the object.

8. The method according to claim 1, wherein Generating the SEI message includes: determining a bounding box of the mask surrounding the object; and The bounding box is encoded in the SEI message.

9. The method according to claim 1, wherein The sample values ​​of the auxiliary image are encoded in a lossy manner.

10. The method according to claim 1, wherein The auxiliary image includes a plurality of predetermined sample values, and the sample values ​​of the mask for representing the object are selected from the plurality of predetermined sample values ​​according to value differences between the plurality of predetermined sample values.

11. A method for detecting an object, the method comprising: Receive bit stream; decoding the encoded information of the bitstream to obtain a main image and an auxiliary image, wherein the auxiliary image indicates a mask of an object in the main image, and the mask of the object is represented by sample values ​​of the auxiliary image; as well as The encoded information of the bitstream is decoded to obtain a Supplemental Enhancement Information (SEI) message, the SEI message indicating attributes of the mask of the object.

12. The method according to claim 11, wherein Decoding the encoded information of the bitstream to obtain the SEI message includes: A cancellation flag is determined, the cancellation flag identifying whether the SEI message cancels the continuing validity of a previous SEI message.

13. The method according to claim 12, wherein: The attributes of the mask include: individual characteristics of the mask of the object indicated by the SEI message and common characteristics of multiple masks, and The decoding of the encoded information of the bitstream to obtain the SEI message includes: In response to determining that the cancellation flag indicates that the SEI message does not cancel the continuing validity of information of the previous SEI message, the common characteristic and the individual characteristic are determined.

14. The method according to claim 11, wherein The attributes of the mask include: individual characteristics of the mask of the object indicated by the SEI message and common characteristics of multiple masks, and the common characteristics include at least one of the following: an identifier of the auxiliary picture to which the SEI message applies; a number of bits used to encode an identifier of any mask in the plurality of masks; the bit depth of the sample values ​​of the auxiliary image; a confidence presence flag, which identifies whether the confidence information of the multiple masks is included in the SEI message; lengths of the confidence information of the plurality of masks, which are present when the confidence presence flag indicates that the confidence information of the plurality of masks is included in the SEI message; a depth presence flag, which identifies whether the depth information of the plurality of masks is included in the SEI message; lengths of the depth information of the plurality of masks, which are present when the depth presence flag indicates that the depth information of the plurality of masks is included in the SEI message; a label presence flag, which identifies whether the label information of the plurality of masks is included in the SEI message; a language presence flag, indicating whether the label language information of the plurality of masks is included in the SEI message, which is present when the label presence flag indicates that the label information of the plurality of masks is included in the SEI message; or The plurality of masked tag language information exists when the language presence flag identifies that the plurality of masked tag language information is included in the SEI message.

15. The method according to claim 11, wherein Decoding the encoded information of the bitstream to obtain the SEI message further comprises: A mask cancel flag is determined, the mask cancel flag identifying whether the mask of the object cancels the ongoing effect of a previous mask of the object.

16. The method according to claim 11, wherein Decoding the encoded information of the bitstream to obtain the SEI message further comprises: A bounding box surrounding the mask of the object is determined based on the SEI message.

17. The method according to claim 11, wherein The mask of the object is indicated by one bit of the sample value of the auxiliary image.

18. The method according to claim 11, wherein The mask of the object is indicated by sample values ​​of the auxiliary image.

19. The method according to claim 11, wherein The auxiliary image includes a plurality of predetermined sample values, and the sample values ​​of the mask for representing the object are selected from the plurality of predetermined sample values ​​according to value differences between the plurality of predetermined sample values.

20. A non-transitory computer-readable storage medium storing a bitstream of a video, the bitstream comprising: A main picture with the object; an auxiliary picture indicating a mask of the object, the mask of the object being represented by sample values ​​of the auxiliary picture; as well as A supplemental enhancement information (SEI) message indicating attributes of the mask for the object.