SEI messages in video encoding

By decoding SEI messages in the bitstream, targets in video images are identified, video encoding is optimized, and the problem of insufficient encoding efficiency in existing technologies is solved, achieving higher encoding efficiency and subjective quality.

CN115804088BActive Publication Date: 2025-12-02ALIBABA GROUP HOLDING LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202180044679.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-08-09
Filing Date
2021-09-14
Publication Date
2025-12-02
Estimated Expiration
2041-09-14

AI Technical Summary

Technical Problem

Existing video coding technologies, even in high-efficiency video coding standards such as HEVC/H.265 and VVC/H.266, have not fully utilized the target information in the video for coding optimization, resulting in a need to improve coding efficiency.

Method used

By decoding the SEI messages in the bitstream, especially the target representation SEI messages, the target in the image is determined, and the target information is used for video encoding optimization.

Benefits of technology

It improves the efficiency of video encoding, achieves higher subjective quality under the same bandwidth, and enhances the compression performance of video encoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115804088B_ABST
    Figure CN115804088B_ABST
Patent Text Reader

Abstract

This disclosure provides methods, apparatus, and non-transitory computer-readable media for processing video data. According to some disclosed embodiments, a method for determining a target in an image includes: decoding a message from a bitstream, including: decoding a first tag list; decoding a first index of a first tag associated with the target into the first tag list; and determining the target based on the message.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This disclosure claims priority to U.S. Provisional Application No. 63 / 084,116, filed September 28, 2020, which is incorporated herein by reference in its entirety. Technical Field

[0002] This disclosure relates generally to video processing, and more specifically to SEI (Supplemental Enhancement Information) messages in video coding. Background Technology

[0003] Video is a set of still images (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, video is compressed before storage or transmission and decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. Currently, various video coding formats exist that use standard video coding technologies, the most common being those based on prediction, transform, quantization, entropy coding, and loop filtering. Video coding standards such as High Efficiency Video Coding (HEVC / H.265), Universal Video Coding (VVC / H.266), and AVS specify particular video coding formats and are developed by standardization organizations. With the increasing application of advanced video coding technologies in video standards, the coding efficiency of new video coding standards is becoming increasingly higher. Summary of the Invention

[0004] Embodiments of this disclosure provide a method for determining a target in an image. The method includes: decoding a message from a bitstream, including: decoding a first tag list; decoding a first index of a first tag associated with the target into the first tag list; and determining the target based on the message.

[0005] Embodiments of this disclosure provide an apparatus for performing video data processing, the apparatus comprising: a memory configured to store instructions; and one or more processors configured to execute the instructions to cause the apparatus to perform: decoding a message from a bitstream, including: decoding a first tag list; and decoding a first index of a first tag associated with a target into the first tag list; and determining the target based on the message.

[0006] Embodiments of this disclosure provide a non-transitory computer-readable storage medium storing a set of instructions executable by one or more processors of a device to cause the device to begin performing a method for determining a target in an image, the method comprising: decoding a message from a bitstream, including: decoding a first tag list; and decoding a first index of a first tag associated with the target into the first tag list; and determining the target based on the message. Attached Figure Description

[0007] Embodiments and various aspects of this disclosure will be described in the following detailed description and accompanying drawings. The various features shown in the figures are not drawn to scale.

[0008] Figure 1 This is a schematic diagram of the structure of an exemplary video sequence according to some embodiments of the present disclosure.

[0009] Figure 2A This is a schematic diagram of the encoding process of an exemplary hybrid video encoding system consistent with embodiments of this disclosure.

[0010] Figure 2B This is a schematic diagram of the encoding process of another exemplary hybrid video encoding system consistent with embodiments of this disclosure.

[0011] Figure 3A This is a schematic diagram of the decoding process of an exemplary hybrid video coding system consistent with embodiments of this disclosure.

[0012] Figure 3B This is a schematic diagram of the decoding process of another exemplary hybrid video coding system consistent with embodiments of this disclosure.

[0013] Figure 4 This is a block diagram of an exemplary apparatus for encoding or decoding video according to some embodiments of the present disclosure.

[0014] Figure 5 This demonstrates an exemplary syntax for annotated regions (AR) SEI messages in a current HEVC environment.

[0015] Figure 6 This is a flowchart illustrating an exemplary method for video processing using target representation SEI messages according to some embodiments of this disclosure.

[0016] Figure 7A The syntax of an exemplary target representation SEI message according to some embodiments of this disclosure is shown.

[0017] Figure 7B Exemplary pseudocode according to some embodiments of the present disclosure is shown, which includes derivation of arrays ArBoundingPolygonVertexX[or_object_idx[i]][j] and ArBoundingPolygonVertexY[or_object_idx[i]][j].

[0018] Figure 8AThis is a flowchart illustrating an exemplary method for video processing using target representation SEI messages according to some embodiments of this disclosure.

[0019] Figure 8B Example portions of a syntax structure for adding signal transmission conditions to target information according to some embodiments of this disclosure are shown.

[0020] Figure 9A An example portion of the syntax structure for transmitting target location parameters and target tag information by signaling, according to some embodiments of the present disclosure, is shown.

[0021] Figure 9B Another example portion of the syntax structure for transmitting target location parameters and target tag information by signaling, according to some embodiments of the present disclosure, is shown.

[0022] Figure 10A A flowchart illustrating an exemplary method for a subordinate auxiliary label list according to some embodiments of the present disclosure is shown.

[0023] Figure 10B An example portion of the syntactic structure of a list of dependent auxiliary tags according to some embodiments of this disclosure is shown.

[0024] Figure 11A A flowchart illustrating an exemplary method for video processing using a combined tag list according to some embodiments of the present disclosure is shown.

[0025] Figure 11B An example portion of the syntax structure of a combined tag list according to some embodiments of this disclosure is shown.

[0026] Figure 11C Another example portion of the syntactic structure of a combined tag list according to some embodiments of this disclosure is shown.

[0027] Figure 12 A flowchart illustrating an exemplary method for video processing using target representation SEI messages according to some embodiments of the present disclosure is shown.

[0028] Figure 13 An example portion of the syntax structure for applying the same boundary method to all targets according to some embodiments of this disclosure is shown.

[0029] Figure 14A An example portion of the syntax structure for sending different coordinate values ​​of two connected vertices using signals, according to some embodiments of this disclosure, is shown.

[0030] Figure 14BExemplary pseudocode according to some embodiments of the present disclosure is shown, which includes derivation of arrays ArBoundingPolygonVertexX[or_object_idx[i]][j] and ArBoundingPolygonVertexY[or_object_idx[i]][j].

[0031] Figure 15 An example portion of the syntax structure using only boundary polygons is shown according to some embodiments of this disclosure.

[0032] Figure 16A Example portions of a syntax structure using fixed-length codes according to some embodiments of this disclosure are shown.

[0033] Figure 16B Example portions of a syntax structure using variable-length codes according to some embodiments of this disclosure are shown. Detailed Implementation

[0034] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. The following description will refer to the accompanying drawings, and unless otherwise stated, the same numbers in the different drawings represent the same or similar elements. The implementations illustrated in the exemplary embodiments described below do not represent all implementations consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with the aspects related to this disclosure listed in the appended claims. Some specific aspects of this disclosure will be described in more detail below. If there is any conflict between the terms and definitions provided herein and those incorporated by reference, the terms and definitions provided herein shall prevail.

[0035] The Joint Video Experts Group (JVET) of the ITU-T Video Coding Experts Group (ITU-T VCEG) and the ISO / IEC Moving Picture Experts Group (ISO / IEC MPEG) is currently developing the Universal Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, VVC aims to achieve the same subjective quality as HEVC / H.265 using half the bandwidth.

[0036] To achieve the same subjective quality as HEVC / H.265 using half the bandwidth, JVET has been developing techniques other than HEVC using a reference software called the Joint Exploration Model (JEM). With the integration of coding techniques into JEM, it achieves significantly higher coding performance than HEVC.

[0037] The recently developed VVC standard continues to include more coding techniques that provide better compression performance. VVC is based on the same hybrid video coding system used by modern video compression standards such as HEVC, H.264 / AVC, MPEG2, H.263, etc.

[0038] Video is a series of still images (or "frames") arranged in chronological order to store visual information. Video capture devices (e.g., cameras) are used to capture and store these images in a time-series manner, and video playback devices (e.g., televisions, computers, smartphones, tablets, video players, or any end-user terminal with a display capability) are used to display these images in a time-series manner. Furthermore, in some applications, video capture devices can transmit captured video in real time to video playback devices (e.g., computers with monitors), for example, for surveillance, conferencing, or live broadcasting.

[0039] To reduce the storage space and transmission bandwidth required for such applications, video can be compressed before storage and transmission, and decompressed before display. Compression and decompression can be implemented by software executed by a processor (e.g., a processor in a general-purpose computer) or by dedicated hardware. The module used for compression is typically called an "encoder," and the module used for decompression is typically called a "decoder." Encoders and decoders can be collectively referred to as a "codec." Encoders and decoders can be implemented as any of a variety of suitable hardware, software, or combinations thereof. For example, hardware implementations of encoders and decoders can include circuitry, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic devices, or any combination thereof. Software implementations of encoders and decoders can include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process embedded in a computer-readable medium. Video compression and decompression can be implemented using various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, H.26x series, etc. In some applications, the codec decompresses the video from a first encoding standard and recompresses the decompressed video using a second encoding standard; in this case, the codec is called a "code converter."

[0040] Video encoding processes identify and retain useful information that can be used to reconstruct images, while ignoring information that is not important for reconstruction. If the ignored, unimportant information cannot be fully reconstructed, this encoding process is called "lossy"; otherwise, it is called "lossless." Most encoding processes are lossy, a trade-off to reduce required storage space and transmission bandwidth.

[0041] Useful information about the image being encoded (referred to as the "current image") includes changes relative to a reference image (e.g., a previously encoded and reconstructed image). These changes can include variations in pixel position, brightness, or color, with positional changes being the most significant. The positional changes of a set of pixels representing a target can reflect the target's movement between the reference and current images.

[0042] An image encoded without referencing another image (i.e., it is its own reference image) is called an "I-image". An image is called a "P-image" if it uses intra-frame or inter-frame prediction of a reference image to predict (e.g., one-way prediction) some or all of the blocks in the image (e.g., blocks typically referring to portions of a video image). An image is called a "B-image" if it uses two reference images to predict at least one block in the image (e.g., two-way prediction).

[0043] Figure 1 The structure of an exemplary video sequence 100 according to some embodiments of the present disclosure is shown. The video sequence 100 may be live video or video that has already been captured and archived. The video sequence 100 may be real-life video, computer-generated video (e.g., computer game video), or a combination thereof (e.g., real-life video with augmented reality effects). The video sequence 100 may be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., a video file stored in a storage device), or a video feed interface (e.g., a video broadcast transceiver) to receive video from a video content provider.

[0044] like Figure 1 As shown, video sequence 100 includes a series of images arranged temporally along a timeline, including images 102, 104, 106, and 108. Images 102-106 are consecutive, with more images between images 106 and 108. Figure 1 In this diagram, image 102 is an I-image, and its reference image is image 102 itself. Image 104 is a P-image, and its reference image is image 102, as indicated by the arrow. Image 106 is a B-image, and its reference images are images 104 and 108, as indicated by the arrow. In some embodiments, the reference image of an image (e.g., image 104) may not immediately precede or follow that image. For example, the reference image of image 104 may be an image preceding image 102. It should be noted that the reference images of images 102-106 are merely examples, and this disclosure does not limit the embodiments of the reference images to specific cases. Figure 1 The example shown is shown in the image.

[0045] Typically, due to the complexity of the computational task, video codecs do not encode or decode the entire image at once. Instead, they segment the image into basic segments and encode or decode the image segment by segment. These basic segments are referred to herein as basic processing units (“BPUs”). For example, Figure 1 Structure 110 illustrates an example structure of a frame (e.g., any of frames 102-108) from video sequence 100. In structure 110, the frame is divided into 4×4 basic processing units, the boundaries of which are shown as dashed lines. In some embodiments, the basic processing unit is referred to as a “macroblock” in some video coding standards (e.g., MPEG series, H.261, H.263, or H.264 / AVC), or as a “coding tree unit” (“CTU”) in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing units in the frame have variable sizes, such as 128×128, 64×64, 32×32, 16×16, 4×8, 16×32, or any shape and number of pixels. The size and shape of the basic processing units for the frame can be chosen based on a trade-off between coding efficiency and the level of detail to be preserved by the basic processing units.

[0046] A basic processing unit can be a logical unit comprising a set of different types of video data stored in computer memory (e.g., a video frame buffer). For example, a basic processing unit for a color image includes a luminance component (Y) representing achromatic luminance information, one or more chrominance components (e.g., Cb and Cr) representing color information, and associated syntax elements, where the luminance and chrominance components may have the same size as the basic processing unit. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luminance and chrominance components are referred to as “code tree blocks” (“CTBs”). Any operation performed on a basic processing unit can be repeated on each of its luminance and chrominance components.

[0047] Video coding has multiple operation levels, examples of which are as follows: Figure 2A , Figure 2B , Figure 3A and Figure 3BAs shown. For each operational level, the size of the basic processing unit may still be too large for processing, and therefore can be further divided into segments referred to herein as "basic processing subunits". In some embodiments, the basic processing subunit is referred to as a "block" in some video coding standards (e.g., MPEG series, H.261, H.263, or H.264 / AVC), or as a "coding unit" ("CU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing subunit has the same or smaller size as the basic processing unit. Similar to the basic processing unit, the basic processing subunit is also a logical unit that includes a set of different types of video data (e.g., Y, Cb, Cr, and associated syntax elements) stored in computer memory (e.g., a video frame buffer). Any operation performed on a basic processing subunit can be repeated on each of its luminance and chrominance components. It should be noted that this division can be performed at higher levels as needed for processing. It should also be noted that different operational levels can use different schemes to divide the basic processing units.

[0048] For example, at the pattern decision level ( Figure 2B An example is shown in the figure. The encoder decides which prediction mode (e.g., intra-image prediction or inter-image prediction) to use for the basic processing unit, which may be too large to make such a decision. The encoder can divide the basic processing unit into multiple basic processing sub-units (e.g., CUs in H.265 / HEVC or H.266 / VVC) and determine a prediction mode for each individual basic processing sub-unit.

[0049] For another example, at the prediction level ( Figure 2A and 2B An example is shown in the figure. The encoder performs prediction operations at the level of basic processing subunits (e.g., CUs). However, in some cases, the basic processing subunits may still be too large to handle. The encoder can further divide the basic processing subunits into smaller segments (e.g., referred to as "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC) and perform prediction operations at that level.

[0050] For another example, in the transformation level ( Figure 2A-2BAn example is shown in the figure. The encoder performs transform operations on the remaining basic processing subunits (e.g., CUs). However, in some cases, the basic processing subunits may still be too large to process. The encoder can further divide the basic processing subunits into smaller segments (e.g., referred to as "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC) and perform transform operations at this level. It should be noted that the segmentation scheme of the same basic processing subunit can differ at the prediction level and the transform level. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU can have different sizes and numbers.

[0051] exist Figure 1 In structure 110, the basic processing unit 112 is further divided into 3×3 basic processing sub-units, the boundaries of which are shown by dashed lines. Different basic processing units of the same image can be divided into basic processing sub-units in different schemes.

[0052] In some implementations, to provide parallel processing and error recovery capabilities for video encoding and decoding, images are divided into regions for processing. Thus, for one region of an image, the encoding or decoding process can proceed independently without relying on information from any other region. In other words, each region of the image can be processed independently. Through this operation, the codec can process different regions of the image in parallel, thereby improving encoding efficiency. Furthermore, when data in one region is corrupted during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same image without relying on the corrupted or lost data, thus providing error recovery capabilities. In some video coding standards, an image is divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles." It should also be noted that different images in the video sequence 100 can be segmented into regions using different segmentation schemes.

[0053] For example, in Figure 1 In the diagram, structure 110 is divided into three regions 114, 116, and 118, whose boundaries are shown as solid lines within structure 110. Region 114 comprises four basic processing units. Each of regions 116 and 118 comprises six basic processing units. It should be noted that... Figure 1 The basic processing unit, basic processing subunit, and region of structure 110 are merely examples, and this disclosure does not limit its embodiments.

[0054] Figure 2A This is a schematic diagram of an exemplary encoding process 200A consistent with embodiments of this disclosure. For example, the encoding process 200A is performed by an encoder. Figure 2AAs shown, the encoder encodes the video sequence 202 into a video bitstream 228 according to process 200A. Similar to... Figure 1 Video sequences 100 and 202 in the dataset consist of a set of images arranged in chronological order (referred to as the "original images"). Similar to... Figure 1 In structure 110, each raw image of video sequence 202 is divided by the encoder into basic processing units, basic processing subunits, or regions for processing. In some embodiments, the encoder performs process 200A on each raw image of video sequence 202 at the level of basic processing units. For example, the encoder performs process 200A iteratively, wherein the encoder encodes the basic processing units in one iteration of process 200A. In some embodiments, the encoder performs process 200A in parallel on regions (e.g., regions 114-118) of each raw image of video sequence 202.

[0055] exist Figure 2A In this process, the encoder feeds the basic processing unit (referred to as the "raw BPU") of the raw images of video sequence 202 to prediction stage 204 to generate prediction data 206 and prediction BPU 208. The encoder subtracts prediction BPU 208 from the raw BPU to generate residual BPU 210. The encoder feeds residual BPU 210 to transform stage 212 and quantization stage 214 to generate quantization transform coefficients 216. The encoder feeds prediction data 206 and quantization transform coefficients 216 to binary coding stage 226 to generate video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 are referred to as the "forward path". In process 200A, after quantization stage 214, the encoder feeds quantization transform coefficients 216 to inverse quantization stage 218 and inverse transform stage 220 to generate reconstructed residual BPU 222. The encoder adds the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224, which is used in the next iteration of process 200A in prediction stage 204. Components 218, 220, 222, and 224 of process 200A are referred to as the "reconstruction path". The reconstruction path is used to ensure that both the encoder and decoder use the same reference data for prediction.

[0056] The encoder can iteratively execute process 200A (in the forward path) to encode each original BPU of the original image and (in the reconstruction path) generate a prediction reference 224 for encoding the next original BPU of the original image. After encoding all the original BPUs of the original image, the encoder can continue to encode the next image in the video sequence 202.

[0057] Referring to process 200A, the encoder receives a video sequence 202 generated by a video capture device (e.g., a camera). The term "receive" as used herein can refer to any action of receiving, inputting, acquiring, retrieving, obtaining, reading, accessing, or otherwise inputting data.

[0058] In prediction level 204, during the current iteration, the encoder receives the raw BPU and prediction reference 224, and performs prediction operations to generate prediction data 206 and prediction BPU 208. Prediction reference 224 can be generated from the reconstruction path of the previous iteration of process 200A. The purpose of prediction level 204 is to reduce information redundancy by extracting prediction data 206 from prediction data 206 and prediction reference 224, which is used to reconstruct the raw BPU into prediction BPU 208.

[0059] Ideally, the predicted BPU 208 is identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 typically differs slightly from the original BPU. To record this difference, after generating the predicted BPU 208, the encoder subtracts it from the original BPU to generate the residual BPU 210. For example, the encoder subtracts the pixel values ​​(e.g., grayscale or RGB values) of the predicted BPU 208 from the values ​​of the corresponding pixels in the original BPU. Each pixel in the residual BPU 210 has a residual value, which is the result of this subtraction between the corresponding pixels in the original BPU and the predicted BPU 208. Compared to the original BPU, the predicted data 206 and the residual BPU 210 have fewer bits, but they can be used to reconstruct the original BPU without significant quality degradation. Therefore, the original BPU is compressed.

[0060] To further compress the residual BPU 210, in transform stage 212, the encoder reduces the spatial redundancy of the residual BPU 210 by decomposing it into a set of two-dimensional “basic patterns,” each of which is associated with “transform coefficients.” The basic patterns can have the same size (e.g., the size of the residual BPU 210). Each basic pattern can represent a frequency component of the residual BPU 210 (e.g., the frequency of brightness variation). No single basic pattern can be regenerated from any combination of any other basic patterns (e.g., a linear combination). In other words, the decomposition breaks down the variation of the residual BPU 210 into the frequency domain. This decomposition is analogous to the discrete Fourier transform of a function, where the basic patterns are analogous to the fundamental functions of the discrete Fourier transform (e.g., trigonometric functions), and the transform coefficients are analogous to the coefficients associated with the fundamental functions.

[0061] Different transform algorithms can use different base patterns. Various transform algorithms, such as discrete cosine transform, discrete sine transform, etc., can be used in transform stage 212. The transform in transform stage 212 is reversible; that is, the encoder can recover the residual BPU 210 through the inverse operation of the transform (called the "inverse transform"). For example, to recover the pixels of the residual BPU 210, the inverse transform multiplies the values ​​of the corresponding pixels in the base pattern by their respective correlation coefficients and sums the products to produce a weighted sum. For video coding standards, the encoder and decoder can use the same transform algorithm (and therefore the same base pattern). Therefore, the encoder can record only the transform coefficients, and the decoder can reconstruct the residual BPU 210 based on the transform coefficients without receiving the base pattern from the encoder. Compared to the residual BPU 210, the transform coefficients have fewer bits, but they can be used to reconstruct the residual BPU 210 without significant quality degradation. Therefore, the residual BPU 210 is further compressed.

[0062] The encoder can further compress the transform coefficients at quantization level 214. During the transform process, different fundamental patterns represent different frequencies of change (e.g., brightness change frequencies). Because the human eye is generally better at recognizing low-frequency changes, the encoder can ignore information about high-frequency changes without causing significant quality degradation during decoding. For example, at quantization level 214, the encoder generates quantized transform coefficients 216 by dividing each transform coefficient by an integer value (called a "quantization scaling factor") and rounding the quotient to its nearest integer. After this operation, some transform coefficients of the high-frequency fundamental patterns are converted to zero, and the transform coefficients of the low-frequency fundamental patterns are converted to smaller integers. The encoder can ignore the zero-valued quantized transform coefficients 216, thus further compressing the transform coefficients. The quantization process is also reversible, where the quantized transform coefficients 216 can be reconstructed into transform coefficients in the inverse operation of quantization (called "inverse quantization").

[0063] Because the encoder does not consider the remainder of such division during rounding operations, quantization level 214 is lossy. Typically, quantization level 214 contributes the most to the information loss in process 200A. The greater the information loss, the fewer bits are needed for the quantization transform coefficients 216. To obtain different levels of information loss, the encoder can use different values ​​of the quantization parameters of the quantization process or any other parameter.

[0064] At binary coding level 226, the encoder encodes the prediction data 206 and quantization transform coefficients 216 using binary coding techniques (e.g., entropy coding, variable-length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm). In some embodiments, in addition to the prediction data 206 and quantization transform coefficients 216, the encoder encodes other information at binary coding level 226 (e.g., the prediction mode used in prediction level 204, parameters of the prediction operation, transform type of transform level 212, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), etc.). The encoder can use the output data of binary coding level 226 to generate a video bitstream 228. In some embodiments, the video bitstream 228 is further packaged for network transmission.

[0065] Following the reconstruction path of process 200A, in inverse quantization stage 218, the encoder performs inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. In inverse transform stage 220, the encoder generates reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder can add the reconstructed residual BPU 222 to the prediction BPU 208 to generate prediction reference 224, which will be used in the next iteration of process 200A.

[0066] It should be noted that other variations of process 200A can be used to encode video sequence 202. In some embodiments, the various levels of process 200A are performed by the encoder in different orders. In some embodiments, one or more levels of process 200A are combined into a single level. In some embodiments, a single level of process 200A is divided into multiple levels. For example, transform level 212 and quantization level 214 are combined into a single level. In some embodiments, process 200A includes additional levels. In some embodiments, process 200A is omitted. Figure 2A One or more levels in the system.

[0067] Figure 2B This is a schematic diagram of another exemplary encoding process 200B consistent with embodiments of this disclosure. Process 200B may be a modification of process 200A. For example, process 200B is used by an encoder conforming to a hybrid video coding standard (e.g., H.26x series). Compared to process 200A, the forward path of process 200B further includes a mode decision stage 230 and divides prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B also includes a loop filter stage 232 and a buffer 234.

[0068] Generally, prediction techniques can be divided into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-image prediction or intra-frame prediction) uses pixels from one or more already encoded neighboring BPUs within the same image to predict the current BPU. That is, the prediction reference 224 in spatial prediction includes neighboring BPUs. Spatial prediction can reduce the inherent spatial redundancy of images. Temporal prediction (e.g., inter-image prediction or inter-frame prediction) uses regions from one or more encoded images to predict the current BPU. That is, the prediction reference 224 in temporal prediction includes the encoded image. Temporal prediction can reduce the inherent temporal redundancy of images.

[0069] Referring to process 200B, in the forward path, the encoder performs prediction operations at spatial prediction level 2042 and temporal prediction level 2044. For example, in spatial prediction level 2042, the encoder performs intra-frame prediction. For the original BPU of the picture being encoded, prediction reference 224 includes one or more adjacent BPUs that have been encoded (in the forward path) and reconstructed (in the reconstruction path) in the same picture. The encoder can generate a predicted BPU 208 by inferring adjacent BPUs. Inference techniques can include, for example, linear inference or interpolation, polynomial inference or interpolation, etc. In some embodiments, the encoder performs inference at the pixel level, for example by inferring the value of the corresponding pixel for each pixel of the predicted BPU 208. The adjacent BPUs used for inference can be located relative to the original BPU from various directions, such as in the vertical direction (e.g., at the top of the original BPU), the horizontal direction (e.g., to the left of the original BPU), the diagonal direction (e.g., at the lower left, lower right, upper left, or upper right of the original BPU), or any direction defined in the video coding standard used. For intra-frame prediction, prediction data 206 may include, for example, the location (e.g., coordinates) of the neighboring BPUs used, the size of the neighboring BPUs used, the inferred parameters, the orientation of the neighboring BPUs used relative to the original BPU, etc.

[0070] In another example, at temporal prediction level 2044, the encoder performs inter-frame prediction. For the original BPU of the current image, prediction reference 224 may include one or more images (referred to as "reference images") that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference images are encoded and reconstructed using BPUs. For example, the encoder adds the reconstructed residual BPU 222 to the prediction BPU 208 to generate a reconstructed BPU. When all reconstructed BPUs for the same image have been generated, the encoder generates a reconstructed image as the reference image. The encoder may perform a "motion estimation" operation to search for matching regions within the range of the reference image (referred to as a "search window"). The position of the search window in the reference image may be determined based on the position of the original BPU in the current image. For example, the search window in the reference image may have a center at the same coordinates as the original BPU in the current image and extend outwards by a predetermined distance. When the encoder (e.g., by using a pixel recursive algorithm, block matching algorithm, etc.) identifies a region similar to the original BPU in the search window, the encoder identifies such a region as a matching region. The matching region can have a different size than the original BPU (e.g., less than, equal to, greater than, or with a different shape). This is because the reference image and the current image are temporally separated on the timeline (e.g., as...). Figure 1 As shown in the image, it can be assumed that the matching region "moves" to the original BPU's location over time. The encoder records the direction and distance of this movement as a "motion vector." When using multiple reference images (e.g., such as...), Figure 1 When working with image 106, the encoder can search for matching regions and determine the associated motion vector for each reference image. In some embodiments, the encoder assigns weights to the pixel values ​​of the matching regions of each matching reference image.

[0071] Motion estimation can be used to identify various types of motion, such as translation, rotation, scaling, etc. For inter-frame prediction, prediction data 206 may include, for example, the location (e.g., coordinates) of the matching region, the motion vector associated with the matching region, the number of reference images, the weights associated with the reference images, etc.

[0072] To generate the predicted BPU 208, the encoder can perform a "motion compensation" operation. Motion compensation is used to reconstruct the predicted BPU 208 based on the predicted data 206 (e.g., motion vectors) and the predicted reference 224. For example, the encoder moves the matching region of the reference image according to the motion vectors, where the encoder can predict the original BPU of the current image. When using multiple reference images (e.g., such as...), Figure 1When matching a reference image (image 106), the encoder can move the matching region of the reference image based on its respective motion vector and the average pixel value of the matching region. In some embodiments, if the encoder has already assigned weights to the pixel values ​​of the matching regions of each matching reference image, the encoder will sum the weighted sums of the pixel values ​​of the moved matching regions.

[0073] In some embodiments, inter-frame prediction can be unidirectional or bidirectional. Unidirectional inter-frame prediction can use one or more reference images in the same temporal direction relative to the current image. For example, Figure 1 Image 104 in the example is a unidirectional inter-frame prediction image, where the reference image (e.g., image 102) precedes image 104. Bidirectional inter-frame prediction can use one or more reference images in two temporal directions relative to the current image. For example, Figure 1 Image 106 in the image is a bidirectional inter-frame predicted image, where reference images (e.g., images 104 and 108) are in two temporal directions relative to image 104.

[0074] Still referring to the forward path of process 200B, after spatial prediction 2042 and temporal prediction stages 2044, at the mode decision stage 230, the encoder selects a prediction mode (e.g., one of intra-frame prediction or inter-frame prediction) for the current iteration of process 200B. For example, the encoder performs a rate-distortion optimization technique, where the encoder selects a prediction mode to minimize the value of the cost function based on the bit rate of the candidate prediction modes and the distortion of the reference image reconstructed under the candidate prediction modes. Based on the selected prediction mode, the encoder generates the corresponding prediction BPU 208 and prediction data 206.

[0075] In the reconstruction path of process 200B, if intra-prediction mode has been selected in the forward path, the encoder directly feeds prediction reference 224 to spatial prediction stage 2042 for later use (e.g., for inference of the next BPU in the current picture) after generating prediction reference 224 (e.g., the current BPU that has been encoded and reconstructed in the current picture). The encoder feeds prediction reference 224 to loop filter stage 232, where the encoder applies loop filters to prediction reference 224 to reduce or eliminate distortions (e.g., blocking artifacts) introduced during the encoding of prediction reference 224. The encoder can apply various loop filter techniques at loop filter stage 232, such as deblocking, sample adaptive offset, adaptive loop filter, etc. The loop-filtered reference picture can be stored in buffer 234 (or "decoded picture buffer") for later use (e.g., as an inter-frame prediction reference picture for future pictures of video sequence 202). The encoder may store one or more reference images in buffer 234 for use in the time prediction stage 2044. In some embodiments, the encoder encodes the parameters of the loop filter (e.g., loop filter strength), as well as the quantization transform coefficients 216, prediction data 206, and other information in the binary encoding stage 226.

[0076] Figure 3A This is a schematic diagram of an exemplary decoding process 300A consistent with embodiments of this disclosure. Process 300A corresponds to... Figure 2A The compression process 200A in the decompression process is described. In some embodiments, process 300A is similar to the reconstruction path of process 200A. The decoder decodes the video bitstream 228 into video stream 304 according to process 300A. Video stream 304 is very similar to video sequence 202. However, due to the compression and decompression processes (e.g., Figure 2A and 2B Information is lost in quantization level 214, and typically, video stream 304 is not the same as video sequence 202. Similar to... Figure 2A and Figure 2B In processes 200A and 200B, for each image already encoded in the video bitstream 228, the decoder can perform process 300A at the level of a basic processing unit (BPU). For example, the decoder performs process 300A iteratively, wherein the decoder decodes a basic processing unit in one iteration of process 300A. In some embodiments, the decoder performs process 300A in parallel for a region (e.g., region 114-118) of each image already encoded in the video bitstream 228.

[0077] exist Figure 3AIn this process, the decoder feeds a portion of the video bitstream 228 associated with a basic processing unit (referred to as the “encoded BPU”) of the encoded image to the binary decoding stage 302. At the binary decoding stage 302, the decoder can decode this portion into prediction data 206 and quantization transform coefficients 216. The decoder can feed the quantization transform coefficients 216 to the inverse quantization stage 218 and the inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder can feed the prediction data 206 to the prediction stage 204 to generate a prediction BPU 208. The decoder can add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224. In some embodiments, the prediction reference 224 is stored in a buffer (e.g., a decoded image buffer in computer memory). The decoder can feed the prediction reference 224 to the prediction stage 204 for performing a prediction operation in the next iteration of process 300A.

[0078] The decoder can iteratively execute process 300A to decode each encoded BPU of the encoded image and generate a prediction reference 224 for the next encoded BPU of the encoded image. After decoding all encoded BPUs of the encoded image, the decoder outputs the image to video stream 304 for display and continues decoding the next encoded image in video bitstream 228.

[0079] In binary decoding stage 302, the decoder performs the inverse operation of the binary encoding technique used by the encoder (e.g., entropy coding, variable-length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and quantization transform coefficients 216, the decoder decodes other information in binary decoding stage 302, such as prediction mode, prediction operation parameters, transform type, quantization process parameters (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), etc. In some embodiments, if the video bitstream 228 is transmitted over the network in packet form, the decoder unpacks it before feeding the video bitstream 228 to binary decoding stage 302.

[0080] Figure 3B This is a schematic diagram of another exemplary decoding process 300B consistent with embodiments of this disclosure. Process 300B may be a modification of process 300A. For example, process 300B is used by a decoder conforming to a hybrid video coding standard (e.g., H.26x series). Compared to process 300A, process 300B further divides prediction stage 204 into spatial prediction stage 2042 and temporal prediction stage 2044, and also includes a loop filter stage 232 and a buffer 234.

[0081] In process 300B, for the encoded basic processing unit (referred to as the "current BPU") of the encoded image being decoded (referred to as the "current image"), the prediction data 206 decoded by the decoder from the binary decoding stage 302 can include various types of data depending on the prediction mode used by the encoder to encode the current BPU. For example, if the encoder uses intra-frame prediction to encode the current BPU, the prediction data 206 includes a prediction mode indicator (e.g., a flag value) indicating intra-frame prediction, parameters for the intra-frame prediction operation, etc. Parameters for the intra-frame prediction operation include, for example, the positions (e.g., coordinates) of one or more neighboring BPUs used as references, the sizes of neighboring BPUs, inferred parameters, the orientation of neighboring BPUs relative to the original BPU, etc. As another example, if the encoder uses inter-frame prediction to encode the current BPU, the prediction data 206 includes a prediction mode indicator (e.g., a flag value) indicating inter-frame prediction, parameters for the inter-frame prediction operation, etc. The parameters for inter-frame prediction operations include, for example, the number of reference images associated with the current BPU, the weights associated with each reference image, the positions (e.g., coordinates) of one or more matching regions in each reference image, and one or more motion vectors associated with each matching region.

[0082] Based on the prediction mode indicator, the decoder determines whether to perform spatial prediction (e.g., intra-frame prediction) at spatial prediction level 2042 or temporal prediction (e.g., inter-frame prediction) at temporal prediction level 2044. The details of performing this spatial or temporal prediction are detailed in... Figure 2B As already described, and will not be repeated below. After performing such spatial or temporal prediction, the decoder generates prediction BPU 208. The decoder adds prediction BPU 208 and the reconstructed residual BPU 222 to generate prediction reference 224, as shown below. Figure 3A As shown.

[0083] In process 300B, the decoder feeds prediction reference 224 to spatial prediction stage 2042 or temporal prediction stage 2044 for performing prediction operations in the next iteration of process 300B. For example, if the current BPU is decoded using intra-frame prediction in spatial prediction stage 2042, the decoder feeds prediction reference 224 directly to spatial prediction stage 2042 for later use (e.g., for inferring the next BPU of the current image) after generating prediction reference 224 (e.g., the decoded current BPU). If the current BPU is decoded using inter-frame prediction in temporal prediction stage 2044, the decoder feeds prediction reference 224 to loop filter stage 232 after generating prediction reference 224 (e.g., a reference image where all BPUs have been decoded) to reduce or eliminate distortion (e.g., blocking artifacts). The decoder can be configured as follows: Figure 2BThe loop filter is applied to prediction reference 224 in the manner shown. The loop-filtered reference image can be stored in buffer 234 (e.g., a decoded image buffer in computer memory) for later use (e.g., as a reference image for inter-frame prediction of future encoded images of video bitstream 228). The decoder can store one or more reference images in buffer 234 for use in temporal prediction stage 2044. In some embodiments, the prediction data further includes parameters of the loop filter (e.g., loop filter strength). In some embodiments, the prediction data includes parameters of the loop filter when the prediction mode indicator of prediction data 206 indicates that inter-frame prediction is used to encode the current BPU.

[0084] Figure 4 This is a block diagram of an exemplary apparatus 400 for encoding or decoding video, consistent with embodiments of this disclosure. Figure 4 As shown, device 400 includes processor 402. When processor 402 executes the instructions described herein, device 400 becomes a dedicated machine for video encoding or decoding. Processor 402 is any type of circuit capable of manipulating or processing information. For example, processor 402 includes any number of central processing units (or “CPU”), graphics processing units (or “GPU”), neural processing units (“NPU”), microcontroller units (“MCU”), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), general-purpose array logic (GALs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), system-on-a-chip (SoCs), application-specific integrated circuits (ASICs), and any combination thereof. In some embodiments, processor 402 is a group of processors grouped into individual logic components. For example, such as Figure 4 As shown, processor 402 includes multiple processors, including processor 402a, processor 402b and processor 402n.

[0085] The device 400 also includes a memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, etc.). For example, as Figure 4As shown, the stored data includes program instructions (e.g., program instructions for implementing various levels in processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). Processor 402 (e.g., via bus 410) accesses the program instructions and data for processing and executes the program instructions to perform operations or manipulations on the data for processing. Memory 404 may include high-speed random access memory or non-volatile memory. In some embodiments, memory 404 includes any combination of any number of random access memories (RAM), read-only memories (ROM), optical discs, magnetic disks, hard disks, solid-state drives, flash drives, secure digital cards (SD cards), memory sticks, compact flash memory (CF cards), etc. Memory 404 may also be a group of memories grouped into single logical components. Figure 4 (Not shown in the image).

[0086] Bus 410 is a communication device for transmitting data between components within device 400, such as an internal bus (e.g., a CPU-memory bus), an external bus (e.g., a universal serial bus port, a high-speed peripheral component interconnect port), etc.

[0087] For ease of explanation and without ambiguity, processor 402 and other data processing circuitry are collectively referred to as "data processing circuitry" in this disclosure. The data processing circuitry may be implemented entirely in hardware, or as a combination of software, hardware, or firmware. Furthermore, the data processing circuitry may be a single, independent module, or may be wholly or partially integrated into any other component of device 400.

[0088] The device 400 also includes a network interface 406 for providing wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, etc.). In some embodiments, the network interface 406 includes any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, repeaters, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication (“NFC”) adapters, cellular network chips, etc.

[0089] In some embodiments, optionally, device 400 further includes a peripheral interface 408 to provide connectivity to one or more peripheral devices. Figure 4 As shown, peripheral devices may include, but are not limited to, cursor control devices (e.g., mouse, touchpad, or touchscreen), keyboards, displays (e.g., cathode ray tube displays, liquid crystal displays, or light-emitting diode displays), video input devices (e.g., cameras or input interfaces coupled to archived video), etc.

[0090] It should be noted that the video codec (e.g., the codec for executing processes 200A, 200B, 300A, or 300B) can be implemented as any combination of any software or hardware modules in device 400. For example, some or all of the operational levels of processes 200A, 200B, 300A, or 300B may be implemented as one or more software modules of device 400, such as program instructions loaded into memory 404. As another example, some or all of the operational levels of processes 200A, 200B, 300A, or 300B may be implemented as one or more hardware modules of device 400, such as dedicated data processing circuitry (e.g., FPGA, ASIC, NPU, etc.).

[0091] This disclosure provides for the encoder described above (e.g., via...). Figure 2A Process 200A or Figure 2B The process 200B) and the decoder (e.g., via Figure 3A Process 300A or Figure 3B The method for using SEI messages is described in Process 300B. SEI messages are intended to be transmitted within the encoded video bitstream in the manner specified in the video coding specification, or in other manner determined by the specification of the system utilizing such an encoded video bitstream. SEI messages contain various types of data that indicate the timing of video frames or describe various attributes of the encoded video or how it is used or enhanced. SEI messages may also contain arbitrary user-defined data. SEI messages do not affect the core of the decoding process but can indicate suggestions on how the video should be post-processed or displayed.

[0092] To specify SEI messages, the H.274 / VSEI standard was developed, which defines the syntax and semantics of Video Availability Information (VUI) parameters and SEI messages. VUI parameters and SEI messages are particularly preferred for encoded video bitstreams specified by the VVC standard. However, since VUI parameters and SEI messages do not affect the decoding process, SEI messages in H.274 / VSEI can also be used for other types of encoded video bitstreams, such as H.265 / HEVC and H.264 / AVC.

[0093] For target detection and tracking purposes, the current H.265 / HEVC standard employs Labeled Region (AR) SEI messages. These messages carry parameters describing the bounding boxes of detected or tracked targets in the compressed video bitstream. This allows the decoder-side device to avoid performing video analysis to identify the target if it has already been identified by the encoder, transcoder, or network node. This is beneficial for applications in decoder devices with limited computing resources and / or power supplies. Furthermore, the encoder can perform detection and tracking tasks using the raw video, which can have significantly higher quality than the reconstructed video recovered on the decoder side. Therefore, performing target detection and tracking on the encoder side and transmitting the information to the decoder helps improve the accuracy of detection and tracking.

[0094] In H.265 / HEVC labeled area SEI messages, in addition to the bounding boxes of detected or tracked targets, target labels and confidence levels associated with the targets can be provided. The target label provides information about the target, and the confidence level indicates the fidelity of the detected or tracked target within the bounding box. Additionally, a flag is provided to indicate whether the bounding box in the current SEI message represents the location of a target that may be occluded or partially occluded by other targets, or only represents the location of the visible portion of a target. Optionally, a flag can also be signaled for each bounding box, indicating whether the object represented by the current bounding box is only partially visible.

[0095] The syntax of labeled SEI messages uses parameter persistence to avoid the need to retransmit information already available in previous SEI messages within the same persistence range. For example, if a detected first target remains stationary in the current image relative to a previously encoded image, and a detected second target moves from one image to another, then only the bounding box information of the second target needs to be sent with a signal, and the position / bounding box information of the first target is copied from the previous SEI message.

[0096] Figure 5 The syntax 500 for an exemplary SEI message for an annotated region (AR) in the current HEVC is shown. The SEI message for an annotated region (AR) carries parameters that identify the annotated region using bounding boxes representing the size and location of the identified targets. The semantics of the syntax elements are shown below.

[0097] If the syntax element `ar_cancel_flag` is equal to 1, it indicates that the labeled region SEI message cancels the persistence of any previous labeled region SEI messages associated with one or more layers to which the labeled region SEI message was applied. If the syntax element `ar_cancel_flag` is equal to 0, it indicates that the following is annotated region information.

[0098] When the syntax element ar_cancel_flag is equal to 1 or when a new coding layer video sequence (CLVS) for the current layer is started, the variables LabelAssigned[i], ObjectTracked[i], and ObjectBoundingBoxAvail are set to 0 in the range of 0 to 255 (inclusive).

[0099] Assume picA is the current image. Each region identified in the labeled region SEI message persists in the current layer in output order until any of the following conditions are true: (i) a new CLVS for the current layer begins; (ii) the bitstream ends; or (ii) the current layer in the access unit outputs a picture picB containing a labeled region SEI message applicable to the current layer, where PicOrderCnt(picB) is greater than PicOrderCnt(picA), PicOrderCnt(picB) and PicOrderCnt(picA) are the values ​​of PicOrderCntVal for picB and picA, and the semantics of the labeled region SEI message for picB cancel the persistence of the regions identified in the labeled region SEI message for picA.

[0100] If the syntax element `ar_not_optimized_for_viewing_flag` is equal to 1, it indicates that the decoded image used in the SEI message for the labeled region is not optimized for user viewing, but rather for other purposes such as the performance of the algorithm's object classification. If the syntax element `ar_not_optimized_for_viewing_flag` is equal to 0, it indicates whether the decoded image used in the SEI message for the labeled region is optimized for user viewing or not.

[0101] If the syntax element `ar_true_motion_flag` is equal to 1, it indicates that motion information from the encoded image applied to the SEI message for the annotated region is selected, with the aim of accurately representing the motion of the target in the annotated region. If the syntax element `ar_true_motion_flag` is equal to 0, it indicates whether motion information from the encoded image applied to the SEI message for the annotated region is selected or not, with the aim of accurately representing the motion of the target in the annotated region.

[0102] If the syntax element `ar_occluded_object_flag` is equal to 1, it indicates that each of the syntax elements `ar_bounding_box_top[ar_object_idx[i]]`, `ar_bounding_box_left[ar_object_idx[i]]`, `ar_bounding_box_width[ar_object_idx[i]]`, and `ar_bounding_box_height[ar_object_idx[i]]` represents the size and position of a target or a portion of a target that is not visible or is only partially visible in the cropped decoded image. If the syntax element `ar_occluded_object_flag` is equal to 0, it indicates that the syntax elements `ar_bounding_box_top[ar_object_idx[i]]`, `ar_bounding_box_left[ar_object_idx[i]]`, `ar_bounding_box_width[ar_object_idx[i]]`, and `ar_bounding_box_height[ar_object_idx[i]]` represent the size and position of a target that is fully visible within the cropped decoded image. For all annotated_regions() syntax structures in CLVS, the value of ar_occluded_object_flag is the same, which is a requirement for bitstream consistency.

[0103] A syntax element `ar_partial_object_flag_present_flag` equal to 1 indicates the existence of the syntax element `ar_partial_object_flag[ar_object_idx[i]]`. A syntax element `ar_partial_object_flag_present_flag` equal to 0 indicates the non-existence of the syntax element `ar_partial_object_flag[ar_object_idx[i]]`. For all `annotated_regions()` syntax structures in CLVS, the value of `ar_partial_object_flag_present_flag` is the same; this is a requirement for bitstream consistency.

[0104] If the syntax element `ar_object_label_present_flag` is equal to 1, it indicates that label information corresponding to the target in the annotation area exists. If the syntax element `ar_object_label_present_flag` is equal to 0, it indicates that label information corresponding to the target in the annotation area does not exist.

[0105] A syntax element `ar_object_confidence_info_present_flag` equal to 1 indicates that the syntax element `ar_object_confidence[ar_object_idx[i]]` exists. A syntax element `ar_object_confidence_info_present_flag` equal to 0 indicates that the syntax element `ar_object_confidence[ar_object_idx[i]]` does not exist. The value of `ar_object_confidence_present_flag` is the same for all `annotated_regions()` syntax structures in CLVS; this is a requirement for bitstream consistency.

[0106] The syntax element ar_object_confidence_length_minus1+1 specifies the length (in bits) of the syntax element ar_object_confidence[ar_object_idx[i]]. The value of ar_object_confidence_length_minus1 is the same for all annotated_regions() syntax structures in CLVS, which is a requirement for bitstream consistency.

[0107] If the syntax element ar_object_label_language_present_flag is equal to 1, it means that the syntax element ar_object_label_language exists. If the syntax element ar_object_label_language_present_flag is equal to 0, it means that the syntax element ar_object_label_language does not exist.

[0108] The syntax element ar_bit_equal_to_zero is equal to zero.

[0109] The syntax element `ar_object_label_language` contains the language tag specified in IETF (Internet Engineering Task Force) RFC (Requests for Comments) 5646, followed by a terminating byte (null) equal to 0x00. The length of the syntax element `ar_object_label_language` is less than or equal to 255 bytes, excluding the terminating byte (null). If it does not exist, the language of the label is not specified.

[0110] The syntax element ar_num_label_updates indicates the total number of labels associated with the annotation region sent by the signal. The value of ar_num_label_updates is between 0 and 255, inclusive.

[0111] The syntax element ar_label_idx[i] indicates the index of the label that was signaled. The value of ar_label_idx[i] is in the range of 0 to 255, inclusive.

[0112] If the syntax element ar_label_cancel_flag is equal to 1, the persistent scope of the ar_label_idx[i]th label is canceled. If the syntax element ar_label_cancel_flag is equal to 0, it means that the ar_label_idx[i]th label has been assigned a signal value.

[0113] The syntax element ar_label[ar_label_idx[i]] specifies the content of the ar_label_idx[i]-th label. The length of the syntax element ar_label[ar_label_idx[i]] is less than or equal to 255 bytes, excluding the terminating byte (null).

[0114] The syntax element ar_num_object_updates indicates the number of target updates to be signaled. The range of the syntax element ar_num_object_updates is 0 to 255, inclusive.

[0115] The syntax element ar_object_idx[i] is the index of the target parameter to be sent with the signal. The range of the syntax element ar_object_idx[i] is 0 to 255, inclusive.

[0116] If the syntax element ar_object_cancel_flag is equal to 1, the persistent scope of the ar_object_idx[i]th target is canceled. If the syntax element ar_object_cancel_flag is equal to 0, it indicates that the parameters associated with the target tracked by the ar_object_idx[i]th target should be signaled.

[0117] If the syntax element ar_object_label_update_flag is equal to 1, it indicates that the target label is sent using a signal. If the syntax element ar_object_label_update_flag is equal to 0, it indicates that the target label is not sent using a signal.

[0118] The syntax element ar_object_label_idx[ar_object_idx[i]] indicates the index of the label corresponding to the ar_object_idx[i]th target. When the syntax element ar_object_label_idx[ar_object_idx[i]] is not present, its value is inferred from the previously output sequentially labeled region SEI messages (if any) in the same CLVS.

[0119] If the syntax element `ar_bounding_box_update_flag` is equal to 1, it indicates that the parameters of the target bounding box are sent using a signal. If the syntax element `ar_bounding_box_update_flag` is equal to 0, it indicates that the parameters of the target bounding box are not sent using a signal.

[0120] If the syntax element ar_bounding_box_cancel_flag is equal to 1, then the persistent scope of ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], ar_bounding_box_height[ar_object_idx[i]], ar_partial_object_flag[ar_object_idx[i]], and ar_object_confidence[ar_object_idx[i]] is canceled. If the syntax element ar_bounding_box_cancel_flag is equal to 0, it means that the syntax elements ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], ar_bounding_box_height[ar_object_idx[i]], ar_partial_object_flag[ar_object_idx[i]], and ar_object_confidence[ar_object_idx[i]] are sent by signal.

[0121] The syntax elements ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], and ar_bounding_box_height[ar_object_idx[i]] specify the top-left corner coordinates of the ar_object_idx[i]-th target and the width and height of the bounding box relative to the consistent cropping window specified by the active SPS in the cropped decoded image, respectively.

[0122] The value of ar_bounding_box_left[ar_object_idx[i]] is in the range of 0 to croppedWidth / SubWidthC-1, inclusive.

[0123] The value of ar_bounding_box_top[ar_object_idx[i]] is in the range of 0 to croppedHeight / SubHeightC-1, inclusive.

[0124] The value of ar_bounding_box_width[ar_object_idx[i]] is in the range of 0 to croppedWidth / SubWidthtC-ar_bounding_box_left[ar_object_idx[i]], including 0 and croppedWidth / SubWidthtC-ar_bounding_box_left[ar_object_idx[i]].

[0125] The value of ar_bounding_box_height[ar_object_idx[i]] is in the range of 0 to croppedHeight / SubHeightC-ar_bounding_box_top[ar_object_idx[i]], including 0 and croppedHeight / SubHeightC-ar_bounding_box_top[ar_object_idx[i]].

[0126] The identified target rectangle contains brightness samples with horizontal image coordinates from SubWidthC*(conf_win_left_offset+ar_bounding_box_left[ar_object_idx[i]]) to SubWidthC*(conf_win_left_offset+ar_bounding_box_left[ar_object_idx[i]]+ar_bounding_box_width[ar_object_idx[i]])–1, including SubWidthC*(conf_win_left_offset+ar_bounding_box_left[ar_object_idx[i]]) and SubWidthC*(conf_win_left_offset+ar_bounding_box_left[ar_object_idx[i]]+ar_bounding_box_width[ar_object_idx[i]]). [i]])–1, and has from SubHeightC*(conf_win_top_offset+ar_bounding_box_top[ar_object_idx[i]]) to SubHeightC*(conf_win_top_offset+ar_bounding_box_top[ar_object_idx[i]]+ar_bounding_box_height[ar_object_idx[i]])– Vertical picture coordinates of 1, including SubHeightC*(conf_win_top_offset+ar_bounding_box_top[ar_object_idx[i]]) and SubHeightC*(conf_win_top_offset+ar_bounding_box_top[ar_object_idx[i]]+ar_bounding_box_height[ar_object_idx[i]])–1.

[0127] For each value of ar_object_idx[i], the values ​​of ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], and ar_bounding_box_height[ar_object_idx[i]] remain unchanged in the output order within CLVS. When not present, the values ​​of ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], or ar_bounding_box_height[ar_object_idx[i]] are inferred from the previously output sequentially labeled region SEI messages in CLVS, if any.

[0128] If the syntax element ar_partial_object_flag[ar_object_idx[i]] is equal to 1, then the syntax elements ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], and ar_bounding_box_height[ar_object_idx[i]] indicate the size and position of the target that is only partially visible in the cropped decoded image. If the syntax element `ar_partial_object_flag[ar_object_idx[i]]` is equal to 0, then the syntax elements `ar_bounding_box_top[ar_object_idx[i]]`, `ar_bounding_box_left[ar_object_idx[i]]`, `ar_bounding_box_width[ar_object_idx[i]]`, and `ar_bounding_box_height[ar_object_idx[i]]` indicate the size and position of the target, which may or may not be partially visible in the cropped decoded image. When it does not exist, the value of `ar_partial_object_flag[ar_object_idx[i]]` is inferred from the previously output sequentially labeled SEI messages in CLVS (if any).

[0129] The syntax element ar_object_confidence[ar_object_idx[i]] indicates the confidence level associated with the ar_object_idx[i]-th target, in units of 2. -(ar_object_confidence_length_minus1+1) This makes a higher value of ar_object_confidence[ar_object_idx[i]] indicate a higher confidence level. The length of the syntax element ar_object_confidence[ar_object_idx[i]] is ar_object_confidence_length_minus1+1 bits. When it does not exist, the value of of_object_confidence[ar_object_idx[i]] is inferred from the previously output sequentially labeled SEI messages in CLVS (if any).

[0130] However, using labeled region SEI messages has some problems and limitations. To improve video processing, this disclosure provides a new SEI message called the object representation (OR) SEI message. Similar to labeled region SEI messages, object representation SEI messages use persistence mechanisms.

[0131] Figure 6 A flowchart illustrating an exemplary method 600 for video processing using a target representation SEI message according to some embodiments of the present disclosure is shown. Method 600 may be performed by an encoder (e.g., via...) Figure 2A Process 200A or Figure 2B The process 200B) is performed, or is performed by a device (e.g., Figure 4 The device 400 is executed by one or more software or hardware components. For example, one or more processors (e.g., Figure 4 The processor 402) executes method 600. In some embodiments, method 600 is implemented by a computer program product embodied in a computer-readable medium, the computer program product comprising components executed by a computer (e.g., processor 402). Figure 4 The device 400 executes computer-executable instructions, such as program code. (See reference...) Figure 6 Method 600 includes the following steps 602-608.

[0132] In step 602, it is determined whether to cancel the persistence of parameters of a previous target representation SEI message. For example, a cancellation flag (e.g., `or_cancel_flag`) is sent to indicate whether to cancel the persistence of a previous target representation SEI message. When the cancellation flag is equal to 1, the target representation SEI message is instructed to cancel the persistence of any parameters of any previous target representation SEI message associated with one or more layers to which the target representation SEI message is applied. When the cancellation flag is equal to 0, the target representation message is instructed to follow.

[0133] In step 604, in response to the fact that the persistence of the parameters in the previous target representation SEI message has not been cancelled (e.g., target representation information is retained), it is determined that a target parameter exists. For example, an existence flag, such as target depth, target confidence, target primary label, etc., is signaled to indicate the existence of the parameter. When the parameter exists, the length information of the parameter is further signaled to indicate the length of the parameter.

[0134] In step 606, tag information is sent using a signal to specify a tag associated with a target in the current image. The tag information may include tag control flags, tag language, a tag list, etc. Tag control flags include, but are not limited to, flags indicating whether to update tags, the number of tags, etc. The tag list may include all tags.

[0135] In step 608, target information is transmitted using signals based on the tag information. For example, the target information includes the target index, target tag index, target location parameters, and target confidence level.

[0136] Figure 7A Exemplary target representation syntax 700 for SEI messages according to some embodiments of this disclosure is shown. For example... Figure 7A As shown, the syntax comprises four parts: SEI cancellation flag section 710, rendering flags and syntax element length section 720, tag information section, and target information section. The tag information section further includes tag control flags section 731, tag language section 732, and tag list section 733. The target information section further includes target index section 741, target tag index section 742, target position parameter section 743, and target depth and confidence section 744.

[0137] The semantics of the syntax elements are shown below.

[0138] If the syntax element `or_cancel_flag` is equal to 1, it instructs the target representation SEI message to cancel the persistence of any previous target representation SEI messages associated with one or more layers to which the target representation SEI message is applied. If the syntax element `or_cancel_flag` is equal to 0, it indicates that the target representation message follows.

[0139] When the syntax element or_cancel_flag equals 1 or a new CLVS for the current layer begins, the variables ObjectTracked[i] and ObjectRegionAvail[i] are set to 0, i is in the range of 0 to 255 (inclusive), and the variables ObjectLabel[i] and ObjectLabel2[i] are cleared in the range of 0 to 255 (inclusive).

[0140] Assume picA is the current image. Each region identified in the target representation SEI message persists in the current layer in output order until any of the following conditions are true: (i) a new CLVS for the current layer begins; (ii) the bitstream ends; or (iii) the output contains an access unit containing a target representation SEI message applicable to the current layer, where PicOrderCnt(picB) is greater than PicOrderCnt(picA), PicOrderCnt(picB) and PicOrderCnt(picA) are the PicOrderCntVal values ​​of picB and picA, and the semantics of the target representation SEI message for picB cancel the persistence of the regions identified in the target representation SEI message for picA.

[0141] A syntax element `or_object_depth_present_flag` equal to 1 indicates the existence of the syntax element `or_object_depth[or_object_idx[i]]`. A syntax element `or_object_depth_present_flag` equal to 0 indicates the absence of the syntax element `or_object_depth[or_object_idx[i]]`. For all `object_representation()` syntax structures in CLVS, the value of `or_object_depth_present_flag` is the same; this is a requirement for bitstream consistency.

[0142] A syntax element `or_object_confidence_info_present_flag` equal to 1 indicates the existence of the syntax element `or_object_confidence[or_object_idx[i]`. A syntax element `or_object_confidence_info_present_flag` equal to 0 indicates the non-existence of the syntax element `or_object_confidence[or_object_idx[i]`. The value of `or_object_confidence_present_flag` is the same for all `object_representation()` syntax structures in CLVS; this is a requirement for bitstream consistency.

[0143] If the syntax element `or_object_primary_label_present_flag` is equal to 1, it indicates that primary label information exists corresponding to the represented target. If the syntax element `or_object_primary_label_present_flag` is equal to 0, it indicates that primary label information does not exist corresponding to the represented target. The value of `or_object_primary_label_present_flag` is the same for all `object_representation()` syntax structures in CLVS, which is a requirement for bitstream consistency.

[0144] The syntax element or_object_depth_length_minus1+1 specifies the length (in bits) of the syntax element or_object_depth[or_object_idx[i]]. The value of or_object_depth_length_minus1 is the same for all object_representation() syntax constructs in CLVS, which is a requirement for bitstream consistency.

[0145] The syntax element `or_object_confidence_length_minus1+1` specifies the length (in bits) of the syntax element `or_object_confidence[or_object_idx[i]]`. The value of `or_object_confidence_length_minus1` is the same for all `object_representation()` syntax constructs in CLVS, which is a requirement for bitstream consistency.

[0146] If the syntax element `or_object_secondary_label_present_flag` is equal to 1, it indicates that auxiliary label information exists corresponding to the represented target. If the syntax element `or_object_secondary_label_present_flag` is equal to 0, it indicates that no auxiliary label information exists corresponding to the represented target. The value of `or_object_secondary_label_present_flag` is the same for all `object_representation()` syntax structures in CLVS, which is a requirement for bitstream consistency.

[0147] If the syntax element `or_object_primary_label_update_allow_flag` is equal to 1, it indicates that the primary label information corresponding to the represented target can be updated. If the syntax element `or_object_primary_label_update_allow_flag` is equal to 0, it indicates that the primary label information corresponding to the represented target cannot be updated. The value of `or_object_primary_label_update_allow_flag` is the same for all `object_representation()` syntax structures in CLVS, which is a requirement for bitstream consistency.

[0148] If the syntax element or_object_label_language_present_flag is equal to 1, it indicates that the syntax element or_object_label_language exists. If the syntax element or_object_label_language_present_flag is equal to 0, it indicates that the syntax element or_object_label_language does not exist.

[0149] The syntax element `or_num_primary_label` indicates the total number of primary labels associated with the target being signaled. The value of `or_num_primary_label` is in the range of 0 to 255 (inclusive).

[0150] The syntax element or_num_secondary_label indicates the total number of secondary labels associated with the target being signaled. The value of or_num_secondary_label is in the range of 0 to 255 (inclusive).

[0151] If the syntax element `or_object_secondary_label_update_allow_flag` is equal to 1, it indicates that the secondary label information corresponding to the represented target can be updated. If the syntax element `or_object_secondary_label_update_allow_flag` is equal to 0, it indicates that the secondary label information corresponding to the represented target cannot be updated. The value of `or_object_secondary_label_update_allow_flag` is the same for all `object_representation()` syntax structures in CLVS, which is a requirement for bitstream consistency.

[0152] The syntax element or_bit_equal_to_zero is equal to zero.

[0153] The syntax element `or_object_label_language` contains the language tag specified by IETF RFC 5646, followed by a terminating byte (null) equal to 0x00. The length of the syntax element `or_object_label_language` is less than or equal to 255 bytes, excluding the terminating byte (null). If it does not exist, the language of the label is not specified.

[0154] The syntax element or_primary_label[i] specifies the content of the i-th primary label. The length of the syntax element or_primary_label[i] is less than or equal to 255 bytes, excluding the terminating byte (null).

[0155] The syntax element or_secondary_label[i] specifies the content of the i-th secondary label. The length of the syntax element or_secondary_label[i] is less than or equal to 255 bytes, excluding the terminating byte (null).

[0156] The syntax element or_num_object_updates indicates the number of target updates to be signaled. or_num_object_updates ranges from 0 to 255, inclusive.

[0157] The syntax element or_object_idx[i] is the index of the target, and the parameters associated with that target are sent or canceled by a signal. or_object_idx[i] is in the range of 0 to 255, inclusive.

[0158] When the syntax element or_object_cancel_flag[or_object_idx[i]] equals 1, it indicates that the persistent scope of the or_object_idx[i]-th target is canceled. When the syntax element or_object_cancel_flag[or_object_idx[i]] equals 0, it indicates that the parameters associated with the or_object_idx[i]-th target are signaled.

[0159] If the syntax element `or_object_primary_label_update_flag[or_object_idx[i]]` is equal to 1, it means that the primary label associated with the `or_object_idx[i]`-th target has been updated. If the syntax element `or_object_primary_label_update_flag[or_object_idx[i]]` is equal to 0, it means that the primary label associated with the `or_object_idx[i]`-th target has not been updated.

[0160] The syntax element or_object_primary_label_idx[or_object_idx[i]] indicates the index of the primary label associated with the or_object_idx[i]th target.

[0161] If the syntax element `or_object_secondary_label_update_flag[or_object_idx[i]]` is equal to 1, it indicates that the auxiliary label associated with the `or_object_idx[i]`-th target has been updated. If the syntax element `or_object_secondary_label_update_flag[or_object_idx[i]]` is equal to 0, it indicates that the auxiliary label associated with the `or_object_idx[i]`-th target has not been updated.

[0162] The syntax element or_object_secondary_label_idx[or_object_idx[i]] indicates the index of the secondary label associated with the or_object_idx[i]th target.

[0163] If the syntax element or_object_pos_parameter_update_flag[or_object_idx[i]] equals 1, it indicates that the position parameter associated with the or_object_idx[i]-th target has been updated. If the syntax element or_object_pos_parameter_update_flag[or_object_idx[i]] equals 0, it indicates that the position parameter associated with the or_object_idx[i]-th target has not been updated.

[0164] If the syntax element or_object_pos_parameter_cancel_flag[or_object_idx[i]] is equal to 1, it indicates the cancellation of the persistence range of the target parameters, including or_bounding_box_top[or_object_idx[i]], or_bonding_box_left[or_object_idx[i]], or_bounding_box_width[or_bject_idx[i]], or_bounding_box_height[or_object_idx[i]], or_bounding_polygon_vertex_num_minus3[or_object_idx[i]], or_bounding_polygon_vertex_x[or_object_idx[i]][j], or_bounding_polygon_vertex_y[or_object_idx[i]][j] (for j, where j ranges from 0 to or_bounding_polygon_vertex_num_minus3[or_object_idx[i]] + 2, including 0 and or_bounding_polygon_vertex_num_minus3[or_object_idx[i]] + 2), or_object_depth[or_object_idx[i]] and or_object_confidence[or_object_idx[i]].The syntax element or_bounding_box_cancel_flag[or_object_idx[i]] is equal to 0, which means or_bounding_box_top[or_object_idx[i]], or_bonding_box_left[or_object_idx[i]], or_bounding_box_ width[or_object_idx[i]], or_boundng_box_height[or_object_idx[i]], or_bounding_polygon_vertex_num_minus3[or_bject_idx[i], or or_bounding_polygon_vertex_x[ `or_object_idx[i]][j]` and `or_bounding_polygon_vertex_y[or_object_idx[i]][j]` (where, for j, j is in the range from 0 to `or_bounding_polygon_vertex_num_minus3[or_object_idx[i]]+2, inclusive)`) are signaled, and the syntax elements `or_object_depth[or_object_idx[i]]` and `or_object_confidence[or_object_id[i]]` are signaled.

[0165] If the syntax element or_object_region_flag[or_object_idx[i]] is equal to 1, it is specified that or_bounding_box_top[or_object_idx[i]], or_bounding_box_left[or_object_idx[i]], or_bounding_box_width[or_object_idx[i]], and or_bounding_box_height[or_object_idx[i]] exist, and or_bounding_polygon_vertex_num_minus3[or_object_idx[i]], or_bounding_polygon_vertex_x[or_object_idx[i]][j], or_bounding_polygon_vertex_y[or_object_idx[i]][j] (where, for j, j ranges from 0 to or_bounding_polygon_vertex_num_minus3[or_object_idx[i]] + 2, including 0 and or_bounding_polygon_vertex_num_minus3[or_object_idx[i]] + 2) do not exist.The syntax element or_object_region_flag[or_object_idx[i]] is equal to 0, indicating or_bounding_box_top[or_object_idx[i]], or_bounding_box_left[or_object_idx[i]] , or_bounding_box_width[or_object_idx[i]], or_bounding_box_height[or_object_idx[i]] does not exist, or_bounding_polygon_vertex_num_minus3[or_o bject_idx[i]], or_bounding_polygon_vertex_x[or_object_idx[i]][j], or_bounding_polygon_vertex_y[or_object_idx[i]][j] (where, for j, j is in the range from 0 to or_bounding_polygon_vertex_num_minus3[or_object_idx[i]]+2, inclusive of 0 and or_bounding_polygon_vertex_num_minus3[or_object_idx[i]]+2).

[0166] The syntax elements or_bounding_box_top[or_object_idx[i]], or_bounding_box_left[or_object_idx[i]], or_bounding_box_width[or_object_idx[i]], and or_bounding_box_height[or_object_idx[i]] specify the top-left corner coordinates of the or_object_idx[i]-th target in the cropped decoded image, as well as the width and height of the bounding box, relative to the consistent cropping window specified by the active SPS.

[0167] Suppose that croppedWidth and croppedHeight are the width and height of the cropped decoded image, respectively, in units of luminance samples.

[0168] The value of or_bounding_box_left[or_object_idx[i]] is in the range of 0 to croppedWidth / SubWidthC-1 (inclusive).

[0169] The value of or_bounding_box_top[or_object_idx[i]] is in the range of 0 to croppedHeight / SubHeightC-1, inclusive.

[0170] The value of or_bounding_box_width[or_object_idx[i]] is in the range of 0 to croppedWidth / SubWidthtC-or_bounding_box_left[or_object_idx[i]], inclusive of 0 and croppedWidth / SubWidthtC-or_bounding_box_left[or_object_idx[i]].

[0171] The value of or_bounding_box_height[or_object_idx[i]] is in the range of 0 to croppedHeight / SubHeightC-or_bounding_box_top[or_object_idx[i]], including 0 and croppedHeight / SubHeightC-or_bounding_box_top[or_object_idx[i]].

[0172] For each or_object_idx[i] value associated with the bounding box, the values ​​of or_bounding_box_top[or_object_idx[i]], or_bounding_box_left[or_object_idx[i]], or_bounding_box_width[or_object_idx[i]], and or_bounding_box_height[or_object_idx[i]] remain unchanged in the output order in CLVS.

[0173] The syntax element or_bounding_polygon_vertex_num_minus3[or_object_idx[i]] plus 3 specifies the number of vertices of the bounding polygon associated with the or_object_idx[i]-th target in the cropped decoded image relative to the consistent cropping window specified by the active SPS.

[0174] The syntax elements or_bounding_polygon_vertex_x[or_object_idx[i]][j] and or_bounding_polygon_vertex_y[or_object_idx[i]][j] specify the coordinates of the j-th vertex of the bounding polygon associated with the or_object_idx[i] target in the clipped decoded image relative to the consistent clipping window specified by the active SPS.

[0175] The value of or_bounding_polygon_vertex_x[or_object_idx[i]][j] is in the range of 0 to croppedWidth / subwidthC-1, inclusive.

[0176] The value of or_bounding_polygon_vertex_y[or_object_idx[i]][j] is in the range of 0 to croppedHeight / SubHeightC-1, inclusive.

[0177] For each value of or_object_idx[i] associated with the boundary polygon, the values ​​of or_bounding_polygon_vertex_x[or_object_idx[i]][j] and or_bounding_polygon_vertex_y[or_object_idx[i]][j] remain unchanged in the output order in CLVS.

[0178] Figure 7B Exemplary pseudocode according to some embodiments of the present disclosure is shown, which includes derivation of arrays ArBoundingPolygonVertexX[or_object_idx[i]][j] and ArBoundingPolygonVertexY[or_object_idx[i]][j].

[0179] The derived arrays ArBoundingPolygonVertexX[or_object_idx[i]][j] and ArBoundingPolygonVertexY[or_object_idx[i]][j] are as follows: Figure 7B As shown.

[0180] The value of ArBoundingPolygonVertexX[or_object_idx[i]][j] is in the range of 0 to croppedWidth / SubWidthC-1, inclusive.

[0181] The value of ArBoundingPolygonVertexY[or_object_idx[i]][j] is in the range of 0 to croppedHeight / SubHeightc-1, inclusive.

[0182] The syntax element `or_object_depth[or_object_idx[i]]` specifies the depth associated with the `or_object_idx[i]`-th target. If it does not exist, the value of `of_object_depth[or_object_idx[i]]` is inferred from the previously output sequentially generated target representation SEI messages in CLVS (if any).

[0183] The syntax element `or_object_confidence[or_object_idx[i]]` indicates the confidence level associated with the `or_object_idx[i]`-th target, in units of 2. -(or_object_confidence_length_minus1+1) This ensures that a higher value of `or_object_confidence[or_object_idx[i]]` indicates a higher confidence level. The length of the syntax element `or_object_confidence[or_object_idx[i]]` is `or_object_confidence_length_minus1+1` bits. When it does not exist, the value of `of_object_confidence[or_object_idx[i]]` is inferred from the sequentially output target representation SEI messages in CLVS (if any).

[0184] In the current Labeled Area SEI message, a persistence mechanism is used when signaling tag information. If the tag list is changed, only the changed tags are notified in the new Labeled Area SEI message. The current syntax supports canceling tags that are no longer in use and adding new tags that are used for the first time. However, in common cases, the number of tags in the CLVS is relatively small, meaning that signaling all tags in the new Labeled Area SEI message (even if only some tags are changed) does not incur much signaling overhead. Using OR SEI messages, according to some embodiments of this disclosure, a more direct way for tag signaling is provided, which can be expressed with fewer syntax elements.

[0185] In some embodiments, in step 606, if it is uncertain whether to update the tags, all tags are signaled. In this embodiment, the entire tag list is signaled, including tags to be updated and tags not to be updated.

[0186] like Figure 7A As shown, with Figure 5 Compared to syntax 500, the syntax for tag information has been simplified, and (see reference) Figure 5 Box 510 in the document no longer requires a label cancellation flag (e.g., or_label_cancel_flag), a label index (e.g., or_label_indx[]), and an array LabelAssigned[].

[0187] By sending all tags with signals without checking whether the tags need to be updated, video processing is simplified by sending fewer syntax elements with signals.

[0188] For some common use cases, the label is a category of the target, such as "person" or "vehicle". Therefore, in these cases, it is not necessary to change the target's label information. However, in the current labeled region SEI message, if the target has not been cancelled, the syntax element ar_object_label_update_flag 520 (e.g., ...) is always sent to indicate whether the target's label information should be updated. Figure 5 (As shown).

[0189] In some embodiments, step 606 of method 600 further includes determining whether updating the label is allowed before updating the label. Return to Reference Figure 7AThe system sends two flags, indicating whether updating the primary and secondary label information of a target is permitted, respectively. For example, the syntax elements `or_object_primary_label_update_allow_flag` 7311 and `or_object_secondary_label_update_allow_flag` 7312 are sent using signals in the label control flags section 731. If updating the target's primary or secondary label is permitted, the corresponding label information of the target can be updated in the OR SEI message below. Otherwise, the target's label information should not be changed within the CLVS. In some embodiments, in applications where the labels remain unchanged, the encoder (e.g., ...) Figure 2A Process 200A or Figure 2B The process (200B) allows setting restrictions on updating the primary and secondary label information of a target. For example, the encoder can set `or_object_primary_label_update_allow_flag 7311` and `or_object_secodnary_label_update_allow_flag 7312` to 0. In this case, label information can only be updated when the label is allowed to be updated; if the label is not allowed to be updated, no update information is sent via signaling. Therefore, since labels are not frequently updated, signaling is reduced.

[0190] In the current SEI message for the labeled region, when sending the target's parameters using a signal, ar_object_cancel_flag 540 is sent using a signal (e.g., Figure 5 (As shown) is a parameter indicating whether to cancel a target. This flag is still signaled even for newly added targets in the current SEI message and can be equal to 1. Canceling a newly appearing target in the current image is meaningless. Furthermore, in the current syntax of labeled area SEI messages, assigning labels or defining bounding boxes for new targets is not allowed. In this case, the decoder only knows that there is a new target in the image, but has no information about that target.

[0191] This disclosure provides embodiments of signal transmission conditions for target information.

[0192] Figure 8A A flowchart illustrating an exemplary method 800A for video processing using target representation SEI messages according to some embodiments of the present disclosure is shown. Method 800A may be performed by an encoder (e.g., by...) Figure 2A Process 200A or Figure 2B The process 200B) is performed, or is performed by a device (e.g., Figure 4The device 400 is executed by one or more software or hardware components. For example, one or more processors (e.g., Figure 4 The processor 402) executes method 800A. In some embodiments, method 800A is implemented by a computer program product embodied in a computer-readable medium, the computer program product comprising components implemented by a computer (e.g., processor 402). Figure 4 The device 400 executes computer-executable instructions, such as program code. (See reference...) Figure 8A Method 800A includes the following steps 802A and 804A.

[0193] In step 802A, in response to a new target in the current SEI message, the determination of whether to cancel the persistence of parameters representing the previous target in the SEI message is skipped. That is, for a new target in the current SEI message, the cancellation flag is skipped from being signaled. The cancellation flag is only signaled if the target previously existed, meaning that the target was being tracked.

[0194] In step 804A, tag information and location parameters are directly signaled for new targets in the current SEI message. Therefore, for new targets in the current SEI message, the signaling of flags indicating parameter and tag updates is skipped. Flags indicating parameter and tag updates are only signaled if the target previously existed.

[0195] Figure 8B Example portions of a syntax structure 800B for adding signal transmission conditions to target information according to some embodiments of this disclosure are shown. Syntax structure 800B can be used in method 800A. Syntax structure 800B only shows changes made to syntax structure 700. Changes to syntax structure 700 are shown in boxes 810B-830B.

[0196] Referring to 810B, the syntax element `or_object_cancel_flag[or_object_idx[i]]` in 811B is signaled only if the target already exists in the current SEI message (e.g., `ObjectTracked[or_object_idx[i]]` equals 1). Therefore, for new targets, the signaling syntax element `or_object_cancel_flag` in 811B is not used. Referring to 820B and 830B, signaling conditions for signaling the target index and target position parameters are added. When the target is new (e.g., `ObjectTracked[or_object_idx[i]]` equals 0), the target information is directly signaled. When the target already exists in the current SEI message (e.g., `ObjectTracked[or_object_idx[i]]` equals 1), the update flag is signaled. For example, when the target is new (e.g., ObjectTracked[or_object_idx[i]] equals 0), the syntax elements or_object_primary_label_idx[or_object_idx[i]]822B and or_object_region_flag[or_object_idx[i]]832B are directly signaled. The syntax elements or_object_primary_label_update_flag[or_object_idx[i]]821B and or_object_pos_parameter_update_flag[or_object_idx[i]] are signaled only when the target already exists in the current SEI message (e.g., ObjectTracked[or_object_idx[i]] equals 1).

[0197] exist Figure 7AIn the embodiment shown in syntax structure 700, when target information is signaled, target tag information is signaled, followed by target position parameters. When target tag information is signaled, a flag indicating whether the tag information associated with the target has been updated is first signaled. If the tag information associated with the target has been updated, the new tag index is signaled. Similarly, when target position parameters are signaled, a flag indicating whether the position parameters have been updated is first signaled. If the position parameters have been updated, the updated target position parameters are signaled. Syntax structure 700 allows for not updating either the target tag information or the position parameters. However, the labeled region SEI message uses a persistence mechanism, so only targets that need updating are signaled. That is, it is allowed to signal targets that need updating, but whose tag information and position have not actually been updated. This is a very strange case.

[0198] In some embodiments of this disclosure, target tag information is transmitted via signaling based on target location parameters. Therefore, the target location parameters are transmitted via signaling before the target tag information. When the target location parameters are not updated, the signaling of a flag indicating whether to update the tag information is skipped, and the tag information is updated directly. This ensures that for a target to be updated, at least one of the target tag information and the target location parameters is updated.

[0199] Figure 9A An example portion of a syntax structure 900A for transmitting target location parameters and target tag information by signaling, according to some embodiments of the present disclosure, is shown. Syntax structure 900A shows only the changes made to syntax structure 700. Changes to syntax structure 700 are shown in blocks 910A and 920A.

[0200] Reference Figure 9A Following the target position parameter section 910A, the target label index section 920A is signaled. The syntax elements `or_object_primary_label_update_allow_flag` and `or_object_secondary_label_update_allow_flag` are not signaled, nor were they determined for signaling the target label index. Therefore, the syntax is simplified.

[0201] Typically, a target's label is more stable than its location. In particular, when the target's location remains unchanged, the likelihood of its label changing is quite small.

[0202] In some embodiments, this disclosure proposes removing the flag indicating whether the target location parameters have been updated, and instead directly updating the target's parameters. By doing so, it is also unnecessary to check whether the target location parameters have been updated when transmitting target tag information via signaling, since it is assumed that the target location parameters are always updated.

[0203] Figure 9B Another example portion of a syntax structure 900B for transmitting target location parameters and target tag information by signaling, according to some embodiments of the present disclosure, is shown. Syntax structure 900B shows only the changes made to syntax structure 900A. The changes to syntax structure 900A are shown in boxes 910B and 920B.

[0204] Reference Figure 9B As shown in 910B, the syntax element `or_object_pos_paramter_update_flag[or_object_idx[i]]` is not signaled. Furthermore, as shown in 920B, the values ​​of `or_object_pos_paramter_update_flag[or_object_idx[i]]` and `or_object_primary_label_update_flag[or_object_idx[i]]` are not determined in order to signal `or_object_secondary_label_update_flag[or_object_idx[i]]` and `or_object_secondary_label_idx[or_object_idx[i]]`. Therefore, the syntax is further simplified.

[0205] Currently, only single labels are supported in SEI (Search Engine Identification) messages. However, in practical applications, multiple labels need to be assigned to a single target. For example, some applications need to detect "people" and "vehicles" in street views. Simultaneously, it needs to distinguish between people lying on the street and people walking on the street, as the former may indicate an accident requiring medical attention. In the vehicle example, it's desirable to differentiate by color. Typically, it might be desirable to be able to attach more than one label to a target. For example, the first label dimension might be "people" and "vehicles"; the second label dimension might be "lying down," "standing," and "walking"; the third label dimension might be "red," "yellow," "blue," and so on.

[0206] Return to reference Figure 7AIn some embodiments, multiple labels are provided for a target. For example, a primary label (e.g., or_primary_label[i]7331) and secondary labels (e.g., or_secondary_label[i]7332) are applied to a target. Secondary labels exist only if the primary label exists. For example, the primary labels are “person” and “vehicle”, and secondary labels are “lying down,” “standing,” and “walking” for “person,” or “red,” “yellow,” and “blue” for “vehicle.” A target can have one or more labels. If there is only one primary label, the target has only one label. With both primary and secondary labels present, each target has two labels. In some embodiments, a third label is applied. For example, the third label is “male” and “female” for a “walking” person. The third label can be independent of the second label. In this case, the target’s two labels can be the primary label and the second label or the third label. For example, a target has the labels “person” and “male.” In some embodiments, the third label depends on the second label. In this case, the third label exists only if the second label exists. For example, a target might be labeled "person," "walking," and "male." The number of labels depends on the target's required level of precision.

[0207] Targets with multiple labels are represented more accurately, thereby improving the accuracy of video processing.

[0208] In the above embodiments, for example, to support two labels for a target, a total of two label lists are sent using signals. Therefore, all targets share the same primary label list and the same secondary label list. That is, regardless of the primary label, each target has the same secondary label space. However, in practice, targets with different primary labels may have different secondary labels. For example, for "person," action or posture is important information for image processing; for "vehicle," shape or color is important information for image processing. That is, for a target with the primary label "person," the secondary list could be "walking," "standing," "lying down," "sitting," while for a target with the primary label "vehicle," the secondary labels could be "red," "blue," "yellow," etc.

[0209] Therefore, in some embodiments according to this disclosure, auxiliary tags associated with the main tags are used. For each main tag in the main tag list, there is a separate corresponding auxiliary tag list.

[0210] Figure 10A A flowchart illustrating an exemplary method 1000A for a dependent auxiliary tag list according to some embodiments of the present disclosure is shown. Method 1000A may be generated by an encoder (e.g., via...). Figure 2A Process 200A or Figure 2B The process 200B) is performed, or is performed by a device (e.g., Figure 4 The device 400 is executed by one or more software or hardware components. For example, one or more processors (e.g., Figure 4 The processor 402) executes method 1000A. In some embodiments, method 1000A is implemented by a computer program product embodied in a computer-readable medium, the computer program product comprising components implemented by a computer (e.g., processor 402). Figure 4 The device 400 executes computer-executable instructions, such as program code. (See reference...) Figure 10A Method 1000A includes the following steps 1002A and 1004A.

[0211] In step 1002A, a first-level tag list, including the main tag, is sent by a signal. For example, the first-level tag list includes multiple tags, such as "person", "vehicle", etc.

[0212] In step 1004A, a second-level tag list associated with a primary tag in the first-level tag list is transmitted via signal. Each primary tag can have a separate corresponding second-level tag list. Each second-level tag list can include multiple tags. For example, for the primary tag "person," the associated second-level tag list includes tags such as "walking," "standing," "lying down," and "sitting." For the primary tag "vehicle," the associated second-level tag list includes tags such as "red," "blue," and "yellow." Then, when transmitting auxiliary tags of the target via signal, the target's auxiliary tags are selected from the second-level tag list associated with the target's primary tag and transmitted via signal. Therefore, the efficiency of transmitting the target's auxiliary tags via signal is improved.

[0213] Figure 10B A portion of the grammatical structure 1000B of an exemplary dependent auxiliary tag list according to some embodiments of this disclosure is shown. The grammatical structure 1000B can be used in method 1000A. Only changes made to the grammatical structure 800B are shown in grammatical structure 1000B. Changes to grammatical structure 800B are shown in boxes 1010B-1030B. The semantics of the updated grammatical structure 1000B are as follows.

[0214] If the syntax element `or_object_secondary_label_present_flag[i]` is equal to 1, it indicates that auxiliary label information corresponding to the target represented by the i-th main label exists. If the syntax element `or_object_secondary_label_present_flag` is equal to 0, it indicates that auxiliary label information corresponding to the target represented by the i-th main label does not exist. The value of `or_object_secondary_label_present_flag` is the same for all `object_representation()` syntax structures in CLVS, which is a requirement for bitstream consistency.

[0215] The syntax element or_num_secondary_label[i] indicates the number of secondary labels associated with the target represented by the i-th primary label. The value of or_num_secondary_label[i] is in the range of 0 to 255 (inclusive).

[0216] If the syntax element `or_object_secondary_label_update_allow_flag[i]` is equal to 1, it indicates that the secondary label information corresponding to the target with the i-th primary label can be updated. If the syntax element `or_object_secondary_label_update_allow_flag[i]` is equal to 0, it indicates that the secondary label information corresponding to the target with the i-th primary label should not be updated. For all `object_representation()` syntax structures in CLVS, the value of `or_object_secondary_label_update_allow_flag[i]` is the same, which is a requirement for bitstream consistency.

[0217] The syntax element `or_secondary_label[j][i]` specifies the content of the `i`-th secondary label associated with the target that has the `j`-th primary label. The length of the syntax element `or_secondary_label[j][i]` is less than or equal to 255 bytes, excluding the terminating byte (null).

[0218] refer to Figure 10B As shown in 1010B, the tag control flags of the auxiliary tags associated with the main tag are sent by signal. As shown in 1020B, a separate list of auxiliary tags is sent by signal for the corresponding main tag. Then, as shown in 1030B, auxiliary tags are sent or updated by signal from the list of auxiliary tags associated with the main tag.

[0219] In some embodiments, to support two tags for a target, two tag lists are signaled. This disclosure also provides embodiments in which only one tag list is signaled, and both the primary tag and the secondary tag for the target are extracted from that tag list.

[0220] Figure 11A A flowchart illustrating an exemplary method 1100A for video processing using a combined tag list according to some embodiments of the present disclosure is shown. Method 1100A can be performed by an encoder (e.g., via...) Figure 2A Process 200A or Figure 2B The process 200B) is performed, or is performed by a device (e.g., Figure 4 The device 400 is executed by one or more software or hardware components. For example, one or more processors (e.g., Figure 4 The processor 402) executes method 1100A. In some embodiments, method 1100A is implemented by a computer program product embodied in a computer-readable medium, the computer program product comprising components executed by a computer (e.g., processor 402). Figure 4 The device 400 executes computer-executable instructions, such as program code. (See reference...) Figure 11A Method 1100A includes the following steps 1102A and 1104A.

[0221] In step 1102A, a list of labels, including both primary and secondary labels, is sent using a signal. For example, in a street view, the primary labels are {"person", "vehicle"}. For a person, it is necessary to describe actions such as "standing", "lying down", or "walking", and for a vehicle, it is necessary to describe color. Therefore, for a person, the secondary labels could be {"standing", "lying down", "walking"}, and for a vehicle, the secondary labels could be {"red", "yellow", "blue"}. Figure 7A In the syntax of the embodiment shown in -7C, a main label list such as {"person", "vehicle"} and an auxiliary label list such as {"standing", "lying down", "walking", "red", "yellow", "blue"} are sent using signals. Figure 10B In the syntax of the subordinate auxiliary label list shown in 10C, the main label list is {"person", "vehicle"}, and the two auxiliary label lists {"standing", "lying down", "walking"} and {"red", "yellow", "blue"} corresponding to each of the two main labels are sent by signals. In the embodiment of the combined label list, only one combined label list is sent by signal, such as {"person", "vehicle", "standing", "lying down", "walking", "red", "yellow", "blue"}.

[0222] In step 1104A, two tag indices of the tag list are transmitted for each target using a signal. These two tag indices correspond to the primary and secondary tags, respectively. Typically, these two tag indices are different.

[0223] Figure 11B A portion of the syntax structure 1100B for an exemplary list of combined tags according to some embodiments of this disclosure is shown. Syntax structure 1100B can be used in method 1100A. Syntax structure 1100B only shows changes made to syntax structure 800B. Changes to syntax structure 800B are shown in boxes 1110B-1140B.

[0224] Reference Figure 11B If the syntax element `or_object_primary_label_present_flag` is equal to 1, it indicates that `or_object_primary_label_idx` exists. If the syntax element `or_object_label_present_flag` is equal to 0, it indicates that the syntax element `or_object_primary_label_idx` does not exist. For all `object_representation()` syntax structures in CLVS, the value of `or_object_primary_label_present_flag` is the same; this is a requirement for bitstream consistency.

[0225] The syntax element or_object_primary_label_idx[or_object_idx[i]] indicates the index of the primary label associated with the or_object_idx[i]th target.

[0226] If the syntax element `or_object_secondary_label_present_flag` is equal to 1, it indicates that `or_object_secondary_label_idx` exists. If the syntax element `or_object_secondary_label_present_flag` is equal to 0, it indicates that `or_object_secondary_label_idx` does not exist. The value of `or_object_secondary_label_present_flag` is the same for all `object_representation()` syntax structures in CLVS; this is a requirement for bitstream consistency.

[0227] The syntax element or_object_secondary_label_idx[or_object_idx[i]] indicates the index of the secondary label associated with the or_object_idx[i]th target.

[0228] Reference Figure 11B As shown in 1110B and 1120B, a list of labels including all labels (e.g., or_label[i]) is sent by signal.

[0229] In some embodiments, as shown in 1130B and 1140B, all targets share a flag indicating the presence of secondary labels. For example, the syntax element `or_object_secondary_label_present_flag` 1130B is signaled to indicate the presence of secondary labels for all targets. If `or_object_secondary_label_present_flag` 1130B equals 1, then all labels have secondary labels. Therefore, the index of the secondary label is sent for each target. If `or_object_secondary_label_present_flag` 1130B equals 0, then the target's secondary label does not exist. Therefore, the index of the secondary label is not signaled.

[0230] In some embodiments, an auxiliary tag presence flag is signaled for each target, so the encoder can decide individually whether to signal an auxiliary tag for each target.

[0231] Figure 11C A portion of a syntax structure 1100C for another exemplary combined tag list according to some embodiments of this disclosure is shown. Syntax structure 1100C can be used in method 1100A. Syntax structure 1100C only shows changes made to syntax structure 1100B. Changes to syntax structure 1100B are shown in boxes 1110C-1130C.

[0232] Reference Figure 11CIf the syntax element `or_object_secondary_label_present_flag[or_object_idx[i]]` is equal to 1, it indicates that the `or_object_secondary_label_idx` of the `or_object_idx[i]`-th target exists. If `or_object_secondary_label_present_flag` is equal to 0, it indicates that the `or_object_secondary_label_idx` of the `or_object_idx[i]`-th target does not exist. The value of `or_object_secondary_label_present_flag` is the same for all `object_representation()` syntax structures in CLVS, which is a requirement for bitstream consistency.

[0233] As shown in 1110C, with Figure 11B In contrast, the syntax element or_object_secondary_label_present_flag 1130B in the label control flag section is not signaled. Instead, for each target in the target label index section, the syntax element or_object_secondary_label_prsent_flag[or_object_idx[i]]1120C is signaled, and based on the determination of or_object_secondary_label_prsent_flag[or_object_idx[i]]1120C, the syntax element or_object_secondary_label_idx[or_object_idx[i]]1130C is signaled for each target.

[0234] In addition, in such Figure 11B and Figure 11CIn the embodiment of the combined label list shown, if both a primary label and a secondary label exist, the primary label and the secondary label share the syntax elements `or_object_label_update_allow_flag` and `or_object_label_update_flag`. However, in other embodiments of this disclosure, separate flags exist for the primary label and the secondary label. For example, `or_object_primary_label_update_allow_flag` and `or_object_primary_label_update_flag` are used for the primary label, and `or_object_secondary_label_update_allow_flag` and `or_object_secondary_label_update_flag` are used for the secondary label.

[0235] In the current SEI message of the labeled region, detected or tracked targets are represented by bounding boxes. While the location information of a target can be described using bounding boxes, its shape information cannot. For applications using segmentation to enhance features such as virtual backgrounds, a more accurate description of the target's shape information is required. Furthermore, the energy consumption of performing target segmentation is a significant burden on mobile devices. Once target segmentation is performed, it is desirable to carry this information as supplementary information in the video bitstream. Figure 5 The syntax of the SEI message for the currently labeled region does not carry such information.

[0236] In order to more accurately describe the shape information of the target, in addition to the bounding box, according to some embodiments of this disclosure, a boundary polygon in the form of a vertex set is proposed. Figure 12 A flowchart illustrating an exemplary method 1200 for video processing using target representation SEI messages according to some embodiments of the present disclosure is shown. Method 1200 may be performed by an encoder (e.g., via...) Figure 2A Process 200A or Figure 2B The process 200B) is performed, or is performed by a device (e.g., Figure 4 The device 400 is executed by one or more software or hardware components. For example, one or more processors (e.g., Figure 4 The processor 402) executes method 1200. In some embodiments, method 1200 is implemented by a computer program product embodied in a computer-readable medium, the computer program product comprising components executed by a computer (e.g., processor 402). Figure 4 The device 400 executes computer-executable instructions, such as program code. (See reference...) Figure 12 Method 1200 includes the following steps 1202-1206.

[0237] In step 1202, a representation method is determined to describe the shape and location of the target. The representation method can be a bounding box or a bounding polygon. A signal transmission flag can be used to indicate whether a bounding box or a bounding polygon is used to describe the shape and location of the target. In some embodiments, the representation method is a bounding circle, and a signal transmission index indicates which representation method is used.

[0238] In step 1204, in response to using the boundary polygon, the number of vertices is determined. The number of vertices is not fixed; the encoder can determine the number of vertices based on the shape of the target and the required precision depending on the application's description. For targets with simple shapes (e.g., triangles or rectangles) or applications that do not require precise shape information, a small number of vertices are determined to store bits, while for targets with complex shapes or applications that require a precise representation of the target's shape (e.g., video conferencing applications that use boundary information to provide virtual background functionality), a large number of vertices are determined to represent the target's boundaries.

[0239] In step 1206, the number of vertices and the position parameters of each vertex are transmitted via signals. The boundary polygon can be determined based on the number of vertices and the position parameters. In some embodiments, the position parameters include the coordinates of the vertices.

[0240] The proposed bounding box and boundary polygon also employ a persistence mechanism, so only the boundary information of the moving target is retransmitted. The minimum number of vertices in the boundary polygon is set to 3.

[0241] Return to reference Figure 7A As shown in the target location parameter section 743, the signaling syntax element or_object_region_flag 7431 indicates the use of a bounding box or a bounding polygon. If a bounding box is used, the bounding box parameters describing the target's location are signaled. If a bounding polygon is used, the number of vertices of the bounding polygon is signaled, and further, the coordinates of each vertex are signaled.

[0242] In some embodiments, a flag `or_object_region_flag[or_object_idx[i]]` is signaled for each target, allowing different targets to be represented in different ways (either using bounding boxes or bounding polygons). In some applications, all tracked targets in an image or the entire sequence use the same target representation. Therefore, signaling a flag for each target is inefficient. Therefore, according to some embodiments of this disclosure, a switch between bounding boxes and bounding polygons is provided, wherein the flag `or_object_region_flag` is signaled for all targets updated in the current target representation SEI message, and this flag is restricted to have the same value throughout the CLVS. Thus, all targets in the CLVS have the same representation.

[0243] Figure 13 Example portions of a syntax structure 1300, which applies the same representation method to all targets according to some embodiments of this disclosure, are shown. Syntax structure 1300 only shows changes made to syntax structure 800B. The main changes to syntax structure 800B are shown in block 1310.

[0244] Reference Figure 13 If the syntax element `or_object_region_flag 1320` equals 1, it indicates that for i in the range 0 to `or_num_object_updates-1`, `or_bounding_box_top[or_object_idx[i]]`, `or_bounding_box_left[or_object_idx[i]]`, `or_bounding_box_width[or_object_idx[i]]`, and `or_bounding_box_height[or_object_idx[i]]` exist.

[0245] For i in the range of 0 to or_num_object_updates-1, or_bounding_polygon_vertex_num_minus3[or_object_idx[i]], or_bounding_polygon_vertex_x[or_object_idx[i]][j], and or_bounding_polygon_vertex_y[or_object_idx[i]][j] do not exist. If the syntax element or_object_region_flag 1320 equals 0, it indicates that for i in the range of 0 to or_num_object_updates-1, or_bounding_box_top[or_object_idx[i]], or_bounding_box_left[or_object_idx[i]], or_bounding_box_width[or_object_idx[i]], or_bounding_box_height[or_object_idx[i]] do not exist, while for i in the range of 0 to or_num_object_updates-1, or_bounding_polygon_vertex_num_minus3[or_object_idx[i]], or_bounding_polygon_vertex_x[or_object_idx[i]][j], and or_bounding_polygon_vertex_y[or_object_idx[i]][j] exist.

[0246] The syntax element `or_object_region_flag` 1320 is signaled to indicate the representation method of the target. As shown in box 1310, when the syntax element `or_object_region_flag` 1320 equals 1, the parameters of the bounding box method are signaled; otherwise, the parameters of the bounding polygon are signaled. In this way, the same representation method is applied to all targets. It is not necessary to determine the representation method for each target, thus improving efficiency.

[0247] In some embodiments, the absolute values ​​of vertex coordinates are signaled. For a polygon with many vertices, this represents a significant signaling overhead. As an alternative signaling method proposed in this disclosure, different coordinate values ​​of two connected vertices are signaled to save signal bits.

[0248] Figure 14AAn example portion of a syntax structure 1400 for signaling different coordinate values ​​of two connected vertices according to some embodiments of this disclosure is shown. Syntax structure 1400 is shown only for changes made to syntax structure 1300. Variations to syntax structure 1300 are shown in box 1410.

[0249] Reference Figure 14A The syntax elements `or_bounding_polygon_vertex_diff_x[or_object_idx[i]][j]1411` and `or_bounding_polygon_vertex_diff_y[or_object_idx[i]][j]1412` specify the coordinate difference between the j-th vertex and the (j-1)-th vertex of the boundary polygon associated with the or_object_idx[i] target in the cropped decoded image, relative to the consistent cropping window specified by the active SPS, when j is greater than 0; the syntax elements `or_bounding_polygon_vertex_diff_x[or_object_idx[i]][0]` and `or_bounding_polygon_vertex_diff_y[or_object_idx[i]][0]` specify the coordinates of the 0th vertex of the boundary polygon associated with the or_object_idx[i] target in the cropped decoded image, relative to the consistent cropping window specified by the active SPS.

[0250] Figure 14B Exemplary pseudocode according to some embodiments of the present disclosure is shown, which includes derivation of arrays ArBoundingPolygonVertexX[or_object_idx[i]][j] and ArBoundingPolygonVertexY[or_object_idx[i]][j].

[0251] The derived arrays ArBoundingPolygonVertexX[or_object_idx[i]][j] and ArBoundingPolygonVertexY[or_object_idx[i]][j] are as follows: Figure 14B As shown.

[0252] Assume that croppedWidth and croppedHeight are the width and height of the cropped decoded image, respectively, in units of luminance samples.

[0253] The value of ArBoundingPolygonVertexX[or_object_idx[i]][j] is in the range of 0 to croppedWidth / SubWidthC-1 (inclusive).

[0254] The value of ArBoundingPolygonVertexY[or_object_idx[i]][j] is in the range of 0 to croppedHeight / SubHeightC-1 (inclusive).

[0255] For each or_object_idx[i] value, the values ​​of ArBoundingPolygonVertexX[or_object_idx[i]][j] and ArBoundingPolygonVertexY[or_object_idx[i]][j] remain unchanged in the output order in CLVS.

[0256] As shown in box 1410, the syntax elements `or_bounding_polygon_vertex_diff_x[or_object_idx[i]][j]1411` and `or_bounding_polygon_vertex_diff_y[or_object_idx[i]][j]1412` are signaled instead of `or_bounding_polygon_vertex_x[or_object_idx[i]][j]` and `or_bounding_polygon_vertex_y[or_object_idx[i]][j]`. Therefore, signaling the different coordinate values ​​of the two connected vertices saves signal bits.

[0257] Considering the special case where the bounding box is a bounding polygon, in some embodiments, only the bounding polygon is used to represent the target. Therefore, the syntax is simplified in the following embodiments concerning the removal of the bounding box.

[0258] Figure 15 An example portion of a syntax structure 1500 using only boundary polygons according to some embodiments of this disclosure is shown. Syntax structure 1500 only shows changes made to syntax structure 800B. The changes to syntax structure 800B are shown in box 1510.

[0259] Reference Figure 15Since only the boundary polygons are used for all targets, the signaling syntax elements or_objet_region_flag, or_bounding_box_top[], or_bounding_box_left[], or_bounding_box_width[], and or_bounding_box_height[] are not used in this embodiment. That is, the reference is returned. Figure 12 Step 1202 can be skipped, thus further simplifying the syntax.

[0260] In the SEI message for the current labeled region, the syntax element ar_partial_object_flag 530 (e.g.) Figure 5 The bounding box (as shown) indicates whether a target represented by a bounding box is partially or fully visible. However, when a target is partially visible, there are no parameters to tell the decoder which parts are visible and which are occluded. Therefore, the syntax element `ar_partial_object_flag 530` itself does not provide the decoder with much information to determine the visible and invisible areas of a target. Target depth information, on the other hand, provides a better mechanism to describe the relative positions of different targets in an image with respect to the camera. This information can be directly used to infer which parts of which targets are occluded or not.

[0261] In some embodiments, the depth of the target is signaled to indicate its relative position (e.g., whether part of the target is visible, partially visible, or completely occluded). Therefore, when two bounding boxes or boundary polygons overlap, the decoder can easily determine which parts of the target are visible based on the target's depth. For example, as... Figure 7A As shown, the syntax element or_object_depth[or_object_idx[i]]7441 is sent using a signal.

[0262] In some embodiments, a variable-length code u(v) is used to encode the depth of the target. The code length is determined by the encoder and signaled in the bitstream. This truly gives the encoder flexibility. Thus, when there are many targets with different depths, the encoder can use more bits to fully represent all levels of depth, and when there are not many targets with different depths, the encoder can use fewer bits to save signal overhead.

[0263] However, in common use cases, there are usually not many different depths associated with the target. Even if fixed-length codes are used to encode depth, it wouldn't consume many bits. Therefore, as an alternative encoding method, fixed-length codes are used for depth in some embodiments.

[0264] Figure 16A Example portions of a syntax structure 1600A using fixed-length codes according to some embodiments of this disclosure are shown. Syntax structure 1600A only shows changes made to syntax structure 700.

[0265] As Figure 16A In the example shown, the depth of each target is encoded using an 8-bit code u(8)1601A, thus supporting up to 256 different depths. However, this embodiment does not limit the depth code length to 8 bits; other lengths can also be used. The precision of the depth depends on the depth code length.

[0266] In some cases, the codes u(v) and u(8) used to encode depth are of equal length. Therefore, codes with different depth values ​​have the same length, even for non-overlapping targets.

[0267] Figure 16B Example portions of a syntax structure 1600B using variable-length codes according to some embodiments of this disclosure are shown. Syntax structure 1600B shows only the changes made to syntax structure 700.

[0268] like Figure 16B As shown, a variable-length code, such as ue(v)1601B, is used to encode depth. Since depth is encoded using unsigned integer exponent Golomb codes, the code length differs for different depth values. Using ue(v) encoding, shorter codes are assigned to smaller values, and longer codes are assigned to larger values. Therefore, the encoding of the target depth length is more flexible.

[0269] It should be understood that in some embodiments, methods 600, 800A, 1000A (or 1100A), and 1200 can be performed in any combination. In some embodiments, syntax structures 800B, 900A (or 900B), 1000B, 1100B (or 1100C), 1300, 1400, 1500, and 1600A (or 1600B) can be applied in any combination by modifying syntax structure 700.

[0270] It should be understood that although this disclosure provides various syntactic elements based on values ​​equal to 0 or 1 to provide inference, these values ​​(e.g., 1 or 0) can be configured in any way to provide appropriate inference.

[0271] The following sentences can be used to further describe the embodiments:

[0272] 1. A method for indicating a target in an image using multiple parameters, comprising:

[0273] Send the first tag list using a signal; and

[0274] The first index of the first tag associated with the target is sent to the first tag list using a signal.

[0275] 2. The method according to claim 1, further comprising:

[0276] A second index of a second tag associated with the target is sent to the first tag list via a signal, wherein the second index is different from the first index.

[0277] 3. The method according to claim 1, further comprising:

[0278] A second tag list is sent by signaling, wherein the first tag list and the second tag list do not contain the same tags; and

[0279] The second index of the second tag associated with the target is sent to the second tag list via a signal.

[0280] 4. The method according to claim 1, further comprising:

[0281] Each of the following uses a signal to send a second tag list corresponding to a tag in the first tag list; and

[0282] The second index of the second tag associated with the target is sent to the second tag list via a signal.

[0283] 5. The method according to any one of claims 1 to 4, further comprising:

[0284] If it is uncertain whether the tag needs to be updated, the tags in the first tag list are sent by signal.

[0285] 6. The method according to claim 5, further comprising:

[0286] In response to a new target in the image, and without being certain whether to cancel the persistence of the parameters, a first index of the first tag associated with the target is signaled.

[0287] 7. The method according to claim 5 or 6, further comprising:

[0288] In response to a new target in the image, and without knowing whether to update the first tag associated with the target, the first index of the first tag associated with the target is signaled.

[0289] 8. The method according to any one of claims 1 to 7, further comprising:

[0290] The depth of the target is transmitted by signal to indicate the target's relative position.

[0291] 9. The method according to any one of claims 1 to 8, further comprising:

[0292] Send target position parameters using signals; and

[0293] Based on the target location parameters, the first index of the first tag associated with the target is transmitted by signal.

[0294] 10. The method according to any one of claims 1 to 9, further comprising:

[0295] Use signals to send polygons to indicate the shape and location of objects in the image.

[0296] 11. A method for indicating a target in an image using multiple parameters, comprising:

[0297] Use signals to send polygons to indicate the shape and location of objects in the image.

[0298] 12. The method of claim 11, wherein sending a polygon with a signal to indicate the shape and position of a target in the image comprises:

[0299] The number of vertices of the polygon is sent by signal; and

[0300] The coordinates of each vertex of the polygon are sent using a signal.

[0301] 13. The method of claim 11, wherein, before transmitting the polygon with a signal to indicate the shape and position of the target in the image, the method further comprises:

[0302] Signal a marker to indicate whether the target should be indicated by a polygon or a rectangle; and

[0303] In response to the sign indicating that the target is indicated by a rectangle, the coordinates of the four vertices of the rectangle are sent by signal.

[0304] 14. The method according to any one of claims 11 to 13, further comprising:

[0305] When it is uncertain whether a tag needs to be updated, the tag is sent using a signal.

[0306] 15. The method of claim 14, further comprising:

[0307] In response to a new target in the image, and without being certain whether to cancel the persistence of the parameters, tag information associated with the target is sent by signal.

[0308] 16. The method according to any one of claims 11 to 15, further comprising:

[0309] The depth of the target is transmitted by signal to indicate the relative position of the target.

[0310] 17. The method according to any one of claims 11 to 16, further comprising:

[0311] Send target position parameters using signals; and

[0312] Based on the target location parameters, target tag information is transmitted using signals.

[0313] 18. A method for indicating a target in an image using multiple parameters, comprising:

[0314] The depth of the target is transmitted by signal to indicate the relative position of the target.

[0315] 19. The method of claim 18, wherein the code length for the depth of the target is fixed.

[0316] 20. The method of claim 18, wherein the depth of the target is encoded using unsigned integer exponent Golomb code.

[0317] 21. A method for determining a target in an image, comprising:

[0318] Decoding messages from a bitstream includes:

[0319] Decode the first tag list; and

[0320] Decode the first index of the first tag associated with the target into the first tag list; and

[0321] The target is determined based on the message.

[0322] 22. The method according to claim 21, wherein the decoding of the message from the bitstream further comprises:

[0323] The second index of the second tag associated with the target is decoded into the first tag list, wherein the second index is different from the first index.

[0324] 23. The method according to claim 21, wherein the decoding of the message from the bitstream further comprises:

[0325] Decode the second tag list, wherein the first tag list and the second tag list do not contain the same tags; and

[0326] The second index of the second tag associated with the target is decoded into the second tag list.

[0327] 24. The method according to claim 21, wherein the decoding of the message from the bitstream further comprises:

[0328] Decode the second tag list corresponding to the tags in the first tag list; and

[0329] The second index of the second tag associated with the target is decoded into the second tag list.

[0330] 25. The method according to any one of claims 21 to 24, wherein the decoding of the message from the bitstream further comprises:

[0331] If it is uncertain whether to update the first tag, decode the tags in the first tag list.

[0332] 26. The method according to claim 25, wherein the decoding of the message from the bitstream further comprises:

[0333] In response to a new target in the image, without being certain whether to cancel the persistence of the parameters, the first index of the first tag associated with the target is decoded.

[0334] 27. The method according to claim 25 or 26, wherein the decoding of the message from the bitstream further comprises:

[0335] In response to a new target in the image, without knowing whether to update the first tag associated with the target, the first index of the first tag associated with the target is decoded.

[0336] 28. The method according to any one of claims 21 to 27, wherein the decoding of the message from the bitstream further comprises:

[0337] Decode the depth of the target to indicate its relative position.

[0338] 29. The method according to any one of claims 21 to 28, wherein the decoding of the message from the bitstream further comprises:

[0339] Decode the target position parameters; and

[0340] Based on the target location parameters, decode the first index of the first tag associated with the target.

[0341] 30. The method according to any one of claims 21 to 29, wherein the decoding of the message from the bitstream further comprises:

[0342] Decode the polygons to indicate the shape and location of the target described in the image.

[0343] 31. A method for determining a target in an image, comprising:

[0344] Decoding messages from a bitstream includes:

[0345] Decoding the polygon used to indicate the shape and position of the target in the image; and

[0346] The target is determined based on the message.

[0347] 32. The method of claim 31, wherein decoding the polygon used to indicate the shape and position of the target in the image further comprises:

[0348] Decode the number of vertices of the polygon; and

[0349] Decode the coordinates of each vertex of the polygon.

[0350] 33. The method of claim 31, wherein decoding the message from the bitstream before decoding the polygon used to indicate the shape and position of the target in the image further comprises:

[0351] Decoding indicates whether the target is indicated by a polygon or a rectangle; and

[0352] In response to a marker indicating the target by a rectangle, decode the coordinates of the rectangle's four vertices.

[0353] 34. The method according to any one of claims 31 to 33, wherein the decoding of the message from the bitstream further comprises:

[0354] Decode the tag when it is uncertain whether the tag needs to be updated.

[0355] 35. The method according to Article 34, wherein the decoding of the message from the bitstream further includes:

[0356] In response to a new target in the image, without being certain whether to cancel the persistence of parameters, the tag information associated with the target is decoded.

[0357] 36. The method according to any one of claims 31 to 35, wherein the decoding of the message from the bitstream further comprises:

[0358] Decode the depth of the target to indicate its relative position.

[0359] 37. The method according to any one of claims 31 to 36, wherein the decoding of the message from the bitstream further comprises:

[0360] Decode the target position parameters; and

[0361] The target tag information is decoded based on the target location parameters.

[0362] 38. A method for determining a target in an image, comprising:

[0363] Decoding messages from a bitstream includes:

[0364] Decode the depth of the target to indicate its relative position; and

[0365] The target in the image is determined based on the message.

[0366] 39. The method of claim 38, wherein the code length of the depth of the target is fixed.

[0367] 40. The method of claim 38, wherein the depth of the target is encoded using unsigned integer exponent Golomb code.

[0368] 41. An apparatus for indicating a target in an image, the apparatus comprising:

[0369] The memory is configured to store instructions; and

[0370] One or more processors are configured to execute the instructions to cause the device to perform:

[0371] Send the first tag list using a signal; and

[0372] Send a signal to the first tag list the first index of the first tag associated with the target.

[0373] 42. The apparatus of claim 41, wherein the one or more processors are further configured to execute the instructions to cause the apparatus to perform:

[0374] A second index of a second tag associated with the target is sent to the first tag list via a signal, wherein the second index is different from the first index.

[0375] 43. The apparatus of claim 41, wherein the one or more processors are further configured to execute the instructions to cause the apparatus to perform:

[0376] A second tag list is sent by signaling, wherein the first tag list and the second tag list do not contain the same tags; and

[0377] The second index of the second tag associated with the target is sent to the second tag list via a signal.

[0378] 44. The apparatus of claim 41, wherein the one or more processors are further configured to execute the instructions to cause the apparatus to perform:

[0379] Each of the following uses a signal to send a second tag list corresponding to a tag in the first tag list; and

[0380] The second index of the second tag associated with the target is sent to the second tag list via a signal.

[0381] 45. The apparatus of claim 41, wherein the one or more processors are further configured to execute the instructions to cause the apparatus to perform:

[0382] If it is uncertain whether the tag needs to be updated, the tags in the first tag list are sent by signal.

[0383] 46. ​​The apparatus of claim 45, wherein the one or more processors are further configured to execute the instructions to cause the apparatus to perform:

[0384] In response to a new target in the image, and without being certain whether to cancel the persistence of the parameters, a first index of the first tag associated with the target is signaled.

[0385] 47. The apparatus of claim 45, wherein the one or more processors are further configured to execute the instructions to cause the apparatus to perform:

[0386] In response to a new target in the image, and without knowing whether to update the first tag associated with the target, the first index of the first tag associated with the target is signaled.

[0387] 48. An apparatus for indicating a target in an image, the apparatus comprising:

[0388] The memory is configured to store instructions; and

[0389] One or more processors are configured to execute the instructions to cause the device to perform:

[0390] Use signals to send polygons to indicate the shape and location of objects in the image.

[0391] 49. The apparatus of claim 48, wherein the step of signaling a polygon to indicate the shape and position of a target in an image comprises:

[0392] The number of vertices of the polygon is sent by signal; and

[0393] Send the coordinates of each vertex of the polygon using a signal.

[0394] 50. The apparatus of claim 48, wherein, before transmitting the polygon with a signal to indicate the shape and position of the target in the image, the one or more processors are further configured to execute the instructions to cause the apparatus to perform:

[0395] Signals are sent to indicate whether the target is indicated by a polygon or a rectangle; and

[0396] In response to the sign indicating a target using a rectangle, the coordinates of the four vertices of the rectangle are signaled.

[0397] 51. An apparatus for indicating a target in an image, the apparatus comprising:

[0398] The memory is configured to store instructions; and

[0399] One or more processors are configured to execute the instructions to cause the device to perform:

[0400] The depth of the target is transmitted by signal to indicate the target's relative position.

[0401] 52. The apparatus of claim 51, wherein the code length for the depth of the target is fixed.

[0402] 53. The apparatus of claim 51, wherein the depth of the target is encoded using unsigned integer exponent Golomb code.

[0403] 54. An apparatus for determining a target in an image, the apparatus comprising:

[0404] The memory is configured to store instructions; and

[0405] One or more processors are configured to execute the instructions to cause the device to perform:

[0406] Decoding messages from a bitstream includes:

[0407] Decode the first tag list; and

[0408] Decode the first index of the first tag associated with the target into the first tag list; and

[0409] The target is determined based on the message.

[0410] 55. The apparatus of claim 54, wherein the one or more processors are further configured to execute the instructions to cause the apparatus to perform:

[0411] The second index of the second tag associated with the target is decoded into the first tag list, wherein the second index is different from the first index.

[0412] 56. The apparatus of claim 54, wherein the one or more processors are further configured to execute instructions to cause the apparatus to perform:

[0413] Decode the second tag list, wherein the first tag list and the second tag list do not contain the same tags; and

[0414] The second index of the second tag associated with the target is decoded into the second tag list.

[0415] 57. The apparatus of claim 54, wherein the one or more processors are further configured to execute instructions to cause the apparatus to perform:

[0416] Decode the second tag list corresponding to the tags in the first tag list; and

[0417] The second index of the second tag associated with the target is decoded into the second tag list.

[0418] 58. The apparatus of claim 54, wherein the one or more processors are further configured to execute instructions to cause the apparatus to perform:

[0419] If it is uncertain whether to update the first tag, decode the tags in the first tag list.

[0420] 59. The apparatus of claim 58, wherein the one or more processors are further configured to execute instructions to cause the apparatus to perform:

[0421] In response to a new target in the image, without being certain whether to cancel the persistence of the parameters, the first index of the first tag associated with the target is decoded.

[0422] 60. The apparatus of claim 58, wherein the one or more processors are further configured to execute instructions to cause the apparatus to perform:

[0423] In response to a new target in the image, without knowing whether to update the first tag associated with the target, the first index of the first tag associated with the target is decoded.

[0424] 61. An apparatus for determining a target in an image, the apparatus comprising:

[0425] The memory is configured to store instructions; and

[0426] One or more processors are configured to execute the instructions to cause the device to perform:

[0427] Decoding messages from a bitstream includes:

[0428] Decoding polygons used to indicate the shape and location of objects in an image; and

[0429] The target is determined based on the message.

[0430] 62. The apparatus of claim 61, wherein the one or more processors are further configured to execute instructions to cause the apparatus to perform:

[0431] Decode the number of vertices of the polygon; and

[0432] Decode the coordinates of each vertex of the polygon.

[0433] 63. The apparatus of claim 61, wherein, prior to decoding the polygon used to indicate the shape and position of a target in an image, the one or more processors are further configured to execute the instructions to cause the apparatus to perform:

[0434] Decoding indicates whether the target is indicated by a polygon or a rectangle; and

[0435] In response to the flag indicating that the target is indicated by a rectangle, the coordinates of the four vertices of the rectangle are decoded.

[0436] 64. The apparatus of claim 61, wherein the one or more processors are further configured to execute the instructions to cause the apparatus to perform:

[0437] Decode the tag when it is uncertain whether the tag needs to be updated.

[0438] 65. The apparatus of claim 64, wherein the one or more processors are further configured to execute the instructions to cause the apparatus to perform:

[0439] In response to a new target in the image, without being certain whether to cancel the persistence of parameters, the tag information associated with the target is decoded.

[0440] 66. The apparatus of claim 61, wherein the one or more processors are further configured to execute instructions to cause the apparatus to perform:

[0441] Decode the target position parameters; and

[0442] The target tag information is decoded based on the target location parameters.

[0443] 67. An apparatus for determining a target in an image, the apparatus comprising:

[0444] The memory is configured to store instructions; and

[0445] One or more processors are configured to execute the instructions to cause the device to perform:

[0446] Decoding messages from a bitstream includes:

[0447] Decode the target's depth to indicate its relative position; and

[0448] The target in the image is determined based on the message.

[0449] 68. A non-transitory computer-readable medium storing a set of instructions, the instructions being executed by one or more processors of a device to cause the device to initiate a method for indicating a target in an image, the method comprising:

[0450] Send the first tag list using a signal; and

[0451] Send a signal to the first tag list the first index of the first tag associated with the target.

[0452] 69. The non-transitory computer-readable medium of claim 68, wherein the instructions are executed by one or more processors of the device to cause the device to further perform:

[0453] A second index of a second tag associated with the target is sent to the first tag list via a signal, wherein the second index is different from the first index.

[0454] 70. The non-transitory computer-readable medium of claim 68, wherein the instructions are executed by one or more processors of the device to cause the device to further perform:

[0455] A second tag list is sent by signaling, wherein the first tag list and the second tag list do not contain the same tags; and

[0456] The second index of the second tag associated with the target is sent to the second tag list using a signal.

[0457] 71. The non-transitory computer-readable medium of claim 68, wherein the instructions are executed by one or more processors of the device to cause the device to further perform:

[0458] Each of the following uses a signal to send a second tag list corresponding to a tag in the first tag list; and

[0459] The second index of the second tag associated with the target is sent to the second tag list using a signal.

[0460] 72. The non-transitory computer-readable medium of claim 68, wherein the instructions are executed by one or more processors of the device to cause the device to further perform:

[0461] If it is uncertain whether the tag needs to be updated, the tags in the first tag list are sent by signal.

[0462] 73. The non-transitory computer-readable medium of claim 72, wherein the instructions are executed by one or more processors of the device to cause the device to further perform:

[0463] In response to a new target in the image, and without being certain whether to cancel the persistence of the parameters, the first index of the first tag associated with the target is signaled.

[0464] 74. The non-transitory computer-readable medium of claim 73, wherein the instructions are executed by one or more processors of the device to cause the device to further perform:

[0465] In response to a new target in the image, without knowing whether to update the first tag associated with the target, the first index of the first tag associated with the target is sent.

[0466] 75. A non-transitory computer-readable medium storing a set of instructions, said instructions being executed by one or more processors of a device to cause the device to initiate a method for indicating a target in an image, said method comprising:

[0467] Use signals to send polygons to indicate the shape and location of objects in the image.

[0468] 76. The non-transitory computer-readable medium of claim 75, wherein the instructions are executed by one or more processors of the device to cause the device to further perform:

[0469] The number of vertices of the polygon is sent by signal; and

[0470] Send the coordinates of each vertex of the polygon using a signal.

[0471] 77. The non-transitory computer-readable medium of claim 75, wherein, prior to signaling the polygon to indicate the shape and position of the target in the image, the instructions are executed by one or more processors of the device to cause the device to further perform:

[0472] The signal is used to indicate whether the target is indicated by a polygon or a rectangle; and

[0473] In response to the sign indicating the target with a rectangle, the coordinates of the four vertices of the rectangle are sent by signal.

[0474] 78. A non-transitory computer-readable medium storing a set of instructions, said instructions being executed by one or more processors of a device to cause the device to initiate a method for indicating a target in an image, said method comprising:

[0475] The depth of the target is transmitted by signal to indicate the target's relative position.

[0476] 79. The non-transitory computer-readable medium of claim 78, wherein the code length of the depth of the target is fixed.

[0477] 80. The non-transitory computer-readable medium of claim 78, wherein the depth of the target is encoded using unsigned integer exponent Golomb code.

[0478] 81. A non-transitory computer-readable medium storing a set of instructions, the instructions being executed by one or more processors of a device to cause the device to initiate a method for determining a target in an image, the method comprising:

[0479] Decoding messages from a bitstream includes:

[0480] Decode the first tag list; and

[0481] Decode the first index of the first tag associated with the target into the first tag list; and

[0482] The target is determined based on the message.

[0483] 82. The non-transitory computer-readable medium of claim 81, wherein the instructions are executed by one or more processors of the device to cause the device to further perform:

[0484] The second index of the second tag associated with the target is decoded into the first tag list, wherein the second index is different from the first index.

[0485] 83. The non-transitory computer-readable medium of claim 81, wherein the instructions are executed by one or more processors of the device to cause the device to further perform:

[0486] Decode the second tag list, wherein the first tag list and the second tag list do not contain the same tags; and

[0487] The second index of the second tag associated with the target is decoded into the second tag list.

[0488] 84. The non-transitory computer-readable medium of claim 81, wherein the instructions are executed by one or more processors of the device to cause the device to further perform:

[0489] Decode the second tag list corresponding to the tags in the first tag list; and

[0490] The second index of the second tag associated with the target is decoded into the second tag list.

[0491] 85. The non-transitory computer-readable medium of claim 81, wherein the instructions are executed by one or more processors of the device to cause the device to further perform:

[0492] If it is uncertain whether to update the first tag, decode the tags in the first tag list.

[0493] 86. The non-transitory computer-readable medium of claim 85, wherein the instructions are executed by one or more processors of the device to cause the device to further perform:

[0494] In response to a new target in the image, without being certain whether to cancel the persistence of the parameters, the first index of the first tag associated with the target is decoded.

[0495] 87. The non-transitory computer-readable medium of claim 86, wherein the instructions are executed by one or more processors of the device to cause the device to further perform:

[0496] In response to a new target in the image, without knowing whether to update the first tag associated with the target, the first index of the first tag associated with the target is decoded.

[0497] 88. A non-transitory computer-readable medium storing a set of instructions, said instructions being executed by one or more processors of a device to cause the device to initiate a method for determining a target in an image, said method comprising:

[0498] Decoding messages from a bitstream includes:

[0499] Decoding the polygon indicating the shape and position of the target in the image; and

[0500] The target is determined based on the message.

[0501] 89. The non-transitory computer-readable medium of claim 88, wherein the instructions are executed by one or more processors of the device to cause the device to further perform:

[0502] Decode the number of vertices of the polygon; and

[0503] Decode the coordinates of each vertex of the polygon.

[0504] 90. The non-transitory computer-readable medium of claim 88, wherein, prior to decoding the polygon indicating the shape and position of the target in the image, the instructions are executed by one or more processors of the device to cause the device to further perform:

[0505] Decoding indicates whether the target is indicated by a polygon or a rectangle; and

[0506] In response to the target being indicated by a rectangle, the coordinates of the rectangle's four vertices are decoded.

[0507] 91. The non-transitory computer-readable medium of claim 88, wherein the instructions are executed by one or more processors of the device to cause the device to further perform:

[0508] Decode the tag when it is uncertain whether the tag needs to be updated.

[0509] 92. The non-transitory computer-readable medium of claim 91, wherein the instructions are executed by one or more processors of the device to cause the device to further perform:

[0510] In response to a new target in the image, without being certain whether to cancel the persistence of parameters, the tag information associated with the target is decoded.

[0511] 93. The non-transitory computer-readable medium of claim 88, wherein the instructions are executed by one or more processors of the device to cause the device to further perform:

[0512] Decode the target position parameters; and

[0513] The target tag information is decoded based on the target location parameters.

[0514] 94. A non-transitory computer-readable medium storing a set of instructions, the instructions being executed by one or more processors of a device to cause the device to initiate a method for determining a target in an image, the method comprising:

[0515] Decoding messages from a bitstream includes:

[0516] Decode the depth of the target used to indicate its relative position; and

[0517] The target in the image is determined based on the message.

[0518] In some embodiments, a non-transitory computer-readable storage medium including instructions is also provided, and these instructions can be executed by means of means for performing the methods described above (e.g., disclosed encoders and decoders). Common forms of non-transitory media include, for example, floppy disks, floppy disk drives, hard disks, solid-state drives, magnetic tape or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a perforated pattern, RAM, PROMs and EPROMs, flash memory—EPROM or any other flash memory, NVRAM, caches, registers, any other memory chips or cassette tapes and their network versions. The means may include one or more processors, input / output interfaces, network interfaces, and / or memory.

[0519] It should be noted that the relational terms used here, such as “first” and “second”, are used only to distinguish one entity or operation from another, and do not require or imply any actual relationship or order between these entities or operations. Furthermore, the words “including,” “having,” “containing,” and “comprising,” as well as other similar forms, are semantically equivalent and open-ended, because one or more items following any of these words do not imply an exhaustive list of such items or items, or that the list is limited to one or more of the listed items.

[0520] As used herein, unless otherwise specified, the term "or" includes all possible combinations, unless impractical. For example, if a database is declared to include A or B, then unless otherwise stated or impractical, the database may include A, or B, or A and B. As a second example, if a database is declared to include A, B, or C, then unless otherwise stated or impractical, the database may include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.

[0521] It should be understood that the above embodiments can be implemented by hardware or software (program code) or a combination of hardware and software. If implemented by software, it can be stored on the above-described computer-readable medium. When executed by a processor, the software can perform the disclosed methods. The computing units and other functional units described in this disclosure can be implemented by hardware or software or a combination of hardware and software. Those skilled in the art will understand that multiple modules / units in the above modules / units can be combined into one module / unit, and each of the above modules / units can be further divided into multiple sub-modules / sub-units.

[0522] In the foregoing specification, numerous specific details have been described with reference to embodiments, which may vary depending on the implementation. Certain adaptations and modifications can be made to the described embodiments. Other embodiments will be apparent to those skilled in the art in consideration of the specification and practice disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the claims. The sequence of steps shown in the figures is also intended for illustrative purposes only and is not intended to be limited to any particular sequence of steps. Therefore, those skilled in the art will understand that these steps may be performed in different orders while achieving the same method.

[0523] Exemplary embodiments have been disclosed in the accompanying drawings and description. However, many variations and modifications can be made to these embodiments. Therefore, although specific terms are used, they are used only in a general and descriptive sense and not for limiting purposes.

Claims

1. A method for determining a target in an image, comprising: Decoding messages from a bitstream includes: Decode the first tag list; and Decode the first index of the first tag associated with the target into the first tag list, and decode the second index of the second tag associated with the target into the first tag list, wherein the second index is different from the first index; and The target is determined based on the message.

2. The method according to claim 1, wherein, The message decoded from the bitstream also includes: Decode the depth of the target used to indicate its relative position.

3. The method according to claim 1, wherein, The message decoded from the bitstream also includes: Decode the target position parameters; and Based on the target location parameters, decode the first index of the first tag associated with the target.

4. The method according to claim 1, wherein, The message decoded from the bitstream also includes: Decode the polygons that indicate the shape and location of the target in the image.

5. The method of claim 4, wherein, The polygon indicating the shape and position of the target in the image also includes: Decode the number of vertices of the polygon; and Decode the coordinates of each vertex of the polygon.

6. An apparatus for determining a target in an image, the apparatus comprising: The memory is configured to store instructions; as well as One or more processors are configured to execute the instructions to cause the device to perform: Decoding messages from a bitstream includes: Decode the first tag list; and Decode the first index of the first tag associated with the target into the first tag list, and decode the second index of the second tag associated with the target into the first tag list, wherein the second index is different from the first index; and The target is determined based on the message.

7. The apparatus of claim 6, wherein the one or more processors are further configured to execute the instructions to cause the apparatus to perform: Decode the polygons that indicate the shape and location of the target in the image.

8. The apparatus of claim 7, wherein the one or more processors are further configured to execute the instructions to cause the apparatus to perform: Decode the number of vertices of the polygon; and Decode the coordinates of each vertex of the polygon.

9. The apparatus of claim 6, wherein the one or more processors are further configured to execute the instructions to cause the apparatus to perform: Decode the depth of the target used to indicate its relative position.

10. A non-transitory computer-readable medium storing a set of instructions, said instructions being executed by one or more processors of a device to cause the device to initiate a method for determining a target in an image, said method comprising: Decoding messages from a bitstream includes: Decode the first tag list; and Decode the first index of the first tag associated with the target into the first tag list, and decode the second index of the second tag associated with the target into the first tag list, wherein the second index is different from the first index; and The target is determined based on the message.

11. The non-transitory computer-readable medium of claim 10, wherein, The instructions are executed by one or more processors of the device to cause the device to further perform: Decode the polygons that indicate the shape and location of the target in the image.

12. The non-transitory computer-readable medium of claim 11, wherein, The instructions are executed by one or more processors of the device to cause the device to further perform: Decode the number of vertices of the polygon; and Decode the coordinates of each vertex of the polygon.

13. The non-transitory computer-readable medium of claim 10, wherein, The instructions are executed by one or more processors of the device to cause the device to further perform: Decode the depth of the target used to indicate its relative position.

Citation Information

Patent Citations

  • Encoding method, decoding method, encoding device, and decoding device

    CN110741635A

  • Video analytics encoding for improved efficiency of video processing and compression

    US20190141340A1

  • Method for processing image data by using information generated from external electronic device, and electronic device

    US20200252555A1