Moving image encoding device and decoding device

The hierarchical image decoding device and object mask information decoding device enhance video encoding and decoding systems for machines by accurately representing object mask images, improving computational efficiency and detection accuracy.

WO2026083770A1PCT designated stage Publication Date: 2026-04-23SHARP KK
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SHARP KK
Filing Date
2025-09-25
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Existing methods for encoding and decoding video for machine image recognition fail to correctly match object mask images, leading to computational inefficiencies and difficulties in applying these methods to video encoding and decoding systems for machines.

Method used

A hierarchical image decoding device that decodes encoded data of image information including object mask images, represented by rectangles with a confidence level greater than 0, and an object mask information decoding device that decodes object mask information, ensuring accurate object detection and tracking.

Benefits of technology

This configuration enables a video encoding and decoding system for machines with improved computational efficiency and accuracy in object detection and tracking, even after encoding and decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025033693_23042026_PF_FP_ABST
    Figure JP2025033693_23042026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention addresses the problem that an object mask image cannot be correctly collated with an object mask image in order to efficiently achieve image recognition in a moving image encoding and decoding system. A moving image decoding device according to one aspect of the present invention is characterized by comprising: a hierarchical image decoding device that decodes encoded data of image information composed of a plurality of image signals including an object mask image for image recognition; and an object mask information decoding device that decodes encoded data of object mask information, wherein when an object mask region is expressed as a rectangle and the object mask information is collated with the object mask image, a state in which the object mask information is not cancelled is considered.
Need to check novelty before this filing date? Find Prior Art

Description

Moving image encoding device, decoding device

[0007]

[0001] Embodiments of the present invention relate to a moving image encoding device and a decoding device.

[0002] In order to efficiently transmit or record a moving image, a moving image encoding device that generates encoded data by encoding an image, and a moving image decoding device that generates a decoded image by decoding the encoded data are used.

[0003] Specific moving image encoding methods include, for example, the H.266 / VVC (Versatile Video Coding) method.

[0004] In such traditional image encoding methods, the image is divided and encoded / decoded. First, a predicted image is generated based on a local decoded image obtained by encoding / decoding the input image / decoding the encoded data. Next, the prediction error (sometimes referred to as "difference image" or "residual image") obtained by subtracting the predicted image from the input image (original image) is encoded / decoded.

[0005] In Non-Patent Document 1, as a technology for moving image encoding and decoding, an auxiliary enhancement information SEI (Supplemental Enhancement Information) message for transmitting the properties of an image, display method, timing, etc. simultaneously with the encoded data is defined. There is a scalability dimension information SEI message (STD: Scalability Dimension Information SEI message) that indicates information of each layer in scalable encoding. By using this SEI and the hierarchical encoding method, an alpha map and a depth map can be encoded and decoded.

[0006] In recent years, in moving image encoding and decoding systems, not only the encoding, transmission, and decoding of video content for human viewing but also the study of encoding and decoding systems for moving images for machines for the purpose of image recognition has been carried out.

[0007] The Object Mask Information (OMI) SEI message defined in Non-Patent Document 2 can be used to notify information about the shape of an object detected in an image. When using the Object Mask Information SEI message to notify object shape information, the Object Mask Information is notified as an auxiliary image using the functionality provided by the H.266 / VVC hierarchical coding scheme and the aforementioned Scalable Dimension Information SEI message.

[0008] The base image contains encoded video content, and the auxiliary image associated with each base image contains encoded object mask images that show the shape of the respective object. The Object Mask Information SEI message provides information about the identifier and pixel values ​​of each object mask in the auxiliary image. Similar to the Object Mask Information SEI message, attributes of the object mask image, such as the label, depth value, and confidence level associated with tracking and detection, can optionally be notified in the Object Mask Information SEI message. Since object detection and tracking are computationally intensive tasks, the Object Mask Information SEI message and the associated auxiliary images (object mask images) notified in the encoded base image reduce the computational load on the receiving system, allowing for the detection and tracking of object shape information based on the decoded video.

[0009] ITU-T Recommendation H.274 (03 / 24) Versatile supplemental enhancement information messages for coded video bitstreamsJ. Boyce, J. Chen, S. Deshpande, MM Hannuksela, S. McCarthy, GJ Sullivan, H. Tan, Y.-K. Wang, Additional SEI messages for VSEI version 4 (Draft 3), JVET-AI2006, Sapporo, 2024.

[0010] The method disclosed in Non-Patent Document 2 had the problem of not being able to correctly match object mask images. Therefore, it was difficult to apply this method to the encoding and decoding of video for machines.

[0011] A motion image decoding device according to one aspect of the present invention comprises a hierarchical image decoding device that decodes encoded data of image information consisting of a plurality of image signals including an object mask image for image recognition, and an object mask information decoding device that decodes encoded data of object mask information, wherein the object mask information is characterized in that the object mask region is represented by a rectangle, and when comparing it with the object mask image, it is taken into consideration that the object mask information has not been canceled.

[0012] Furthermore, a motion image decoding device according to one aspect of the present invention comprises a hierarchical image decoding device that decodes encoded data of image information consisting of a plurality of image signals including an object mask image for image recognition, and an object mask information decoding device that decodes encoded data of object mask information, wherein the object mask information has a confidence level greater than 0.

[0013] A motion image encoding device according to one aspect of the present invention comprises a hierarchical image encoding device for encoding image information consisting of a plurality of image signals including an object mask image for image recognition, and an object mask information encoding device for encoding object mask information, wherein the object mask information is characterized in that the object mask region is represented by a rectangle, and when comparing it with the object mask image, it is taken into consideration that the object mask information has not been canceled.

[0014] This configuration solves the challenge of realizing a video encoding and decoding system for machines.

[0015] This is a schematic diagram showing the configuration of the image transmission system according to this embodiment. This is a diagram showing an example of a block diagram of the pre-image processing device according to this embodiment. This is a diagram showing an example of a block diagram of the hierarchical image coding device according to this embodiment. This is a diagram showing an example of a block diagram of the object mask information creation device according to this embodiment. This is a diagram showing an example of a block diagram of the hierarchical image decoding device according to this embodiment. This is a diagram showing an example of a block diagram of the object mask information decoding device according to this embodiment. This is a diagram showing the syntax of the Scalability Dimension Information SEI message from Non-Patent Literature 1. This is a diagram showing the auxiliary information from Non-Patent Literature 1 and Non-Patent Literature 2. This is a diagram showing the relationship between the variable ChromaFormatIdc and the variables SubWidthC and SubHightC from Non-Patent Literature 2. This is a diagram (1) showing the syntax of the Object Mask Information SEI message from Non-Patent Literature 2. This is a diagram (2) showing the syntax of the Object Mask Information SEI message from Non-Patent Literature 2. This is a diagram (1) showing the syntax of the Object Mask Information SEI message of this embodiment. This is a diagram (2) showing the syntax of the Object Mask Information SEI message of this embodiment.

[0016] (First Embodiment) Figure 1 is a conceptual diagram showing the configuration of the image transmission system for machines according to this embodiment.

[0017] The image transmission system 1 consists of a video encoding device 10, a transmission network 20, a video decoding device 30, an image display device 40, and an image recognition device 50.

[0018] The video encoding device 10 takes an input image signal T as input and outputs encoded data Te.

[0019] The transmission network 20 transmits encoded data Te from the video encoding device 10 to the video decoding device 30. The transmission network 20 is the Internet, a wide area network (WAN), a local area network (LAN), or a combination thereof. The network 20 is not necessarily limited to a bidirectional communication network; it may also be a unidirectional communication network that transmits broadcast waves such as terrestrial digital broadcasting or satellite broadcasting. Furthermore, the transmission network 20 may be replaced by a storage medium that records encoded data Te, such as a DVD (Digital Versatile Disc: trademark) or a BD (Blu-ray Disc: registered trademark).

[0020] The video decoding device 30 takes encoded data Te as input, outputs basic image information Td, and sends it to the image display device 40. The video decoding device 30 also sends the basic image information, object mask image information, and object mask information to the image recognition device 50.

[0021] The image display device 40 displays all or part of the generated image Td output from the video decoding device 30. The image display device 40 includes a display device such as a liquid crystal display or an organic EL (electro-luminescence) display. Examples of display forms include stationary, mobile, and HMD (head-mounted display). Furthermore, if the video decoding device 30 has high processing power, it displays a high-quality image, and if it has lower processing power, it displays an image that does not require high processing power or display power.

[0022] The image recognition device 50 performs image recognition processing on the basic image information Td using the transmitted and decoded basic image information Td and the image recognition output of the object mask information decoder. Specifically, it identifies what the objects in the image are, and performs processes such as tracking the identified objects when they are converted into moving images, and segmenting each region within the screen.

[0023] The video encoding device 10 consists of a pre-image processing device 101, a hierarchical video encoding device 102, an object mask information creation device 103, and an object mask information encoding device 104.

[0024] The pre-image processing device 101 receives an input image signal T and outputs image information consisting of a basic image signal and an object mask image signal.

[0025] The hierarchical image encoding device 102 encodes the image information created by the pre-image processing device and creates encoded data Te.

[0026] The object mask information creation device 103 takes the input image signal T and the object mask image signal from the pre-image processing device as inputs, creates object mask information, and sends it to the object mask information encoding device 104.

[0027] The object mask information encoding device 104 encodes the object mask information to create encoded data Te.

[0028] The video decoding device 30 consists of a hierarchical image decoding device 301, an object mask information decoding device 302, and an image recognition device 303.

[0029] The hierarchical image decoding device 301 receives live encoded data Te transmitted via the transmission network 20 as input, performs decoding processing, sends object mask image information to the object mask information decoding device 302, and sends basic image information Td to the image display device 40 and the image recognition device 50.

[0030] The object mask decoding device 302 decodes the encoded data Te based on its syntax, compares it with the object mask image information to perform post-image processing, and sends the image recognition information to the image recognition device 50.

[0031] Figure 2 is a conceptual diagram showing the configuration of the pre-image processing device of this embodiment. The pre-image processing device 101 consists of an object mask image processing unit 1011, which receives an input image signal and outputs image information consisting of a basic image signal identical to the input image signal and an object mask image signal, which is auxiliary image information indicating the image recognition region. In Figure 2, an example is shown where there is one basic image signal and one object mask image signal, but there is not necessarily one basic image signal and one object mask image signal; there may be two or more of each.

[0032] The object mask image processing unit 1011 analyzes the input image signal, creates an image that masks pixels that may be the target of image recognition, and outputs an object mask image signal.

[0033] Figure 3 is a block diagram showing the configuration of the hierarchical image coding device 102 in this embodiment. In the hierarchical image coding device 102 in this embodiment, image information created by the pre-image processing device 101 is compressed and coded, and coded data is created to be transmitted to the video decoding device 30 via the transmission network 20. The hierarchical image coding device 102 consists of a basic image coding unit 1021 and an object mask image coding unit 1022. The basic image coding unit 1021 codes the basic image signal from the image information from the pre-image processing device 101 and outputs basic image coded data. The object mask image coding unit 1022 codes the object mask image signal from the image information from the pre-image processing device 101 and outputs object mask image coded data. In Figure 3, an example is shown where there is one basic image coding unit and one object mask image coding unit, but the basic image coding unit 1021 and the object mask image coding unit 1022 are not limited to one each, and there may be two or more of each.

[0034] The encoded data output by the hierarchical image encoding device 102 is configured such that object mask image encoding data is associated with the basic image encoding data. Specifically, it uses a method that extends the Scalability Dimension Information SEI message described in Non-Patent Document 1 (described later) as described in Non-Patent Document 2.

[0035] Figure 4 is a block diagram showing the configuration of the object mask information creation device 103 in this embodiment. In the object mask information creation device 103 in this embodiment, the input image signal and the object mask image signal created by the pre-image processing device are input, object mask information is output, and it is sent to the object mask information encoding device 104.

[0036] The object mask information creation device 103 consists of a pre-image recognition unit 1031 and an object mask information creation unit 1032. The pre-image recognition unit 1031 uses the input image signal and the object mask signal to narrow down the pixels that may be the target of image recognition, then performs image recognition processing and sends the recognition result to the object mask information creation unit 1032.

[0037] The object mask information creation unit 1032 creates object mask information based on the recognition results of the pre-image recognition unit 1031. Specifically, it creates: an identifier (ID) for the object mask region; pixel values ​​of the object mask image signal for the object mask region; coordinates, width, and height of the top left corner of the image when the object mask region is represented as a rectangle; confidence level of the object mask region; depth information for the object mask region; and string label information for the object mask region. The created object mask information is sent to the object mask encoding device 104.

[0038] The object mask coding device 104 encodes object mask information based on the Object Mask Information (OMI) SEI described in Non-Patent Document 2 below, and outputs encoded object mask information data.

[0039] The basic image coding data, object mask image coding data, and object mask information coding data are output to the transmission path 20 as part of the coded data Te of the video coding device 10. The video decoding device 30 consists of a hierarchical image decoding device 301 and an object mask information decoding device 302.

[0040] Figure 5 is a conceptual diagram showing the configuration of the hierarchical image decoding device of this embodiment. In the hierarchical image decoding device 301 of this embodiment, the basic image encoding data and object mask image encoding data are decoded from the encoded data encoded by the video encoding device 10 and sent over the transmission network 20 to create image information. The image information output by the hierarchical image decoding device 301 consists of basic image information and object mask image information.

[0041] The hierarchical image decoding device 301 consists of a basic image decoding unit 3011 and an object mask image decoding unit 3012. The basic image decoding unit 3011 decodes the basic image encoded data from the encoded data to create basic image information and sends it to the image display device 40 and the image recognition processing device 50. The object mask image decoding unit 3012 decodes the object mask image encoded data from the encoded data to create object mask image information and sends it to the image recognition device 303.

[0042] In this embodiment, the hierarchical image encoding device 102 and the hierarchical image decoding device 301 are realized by applying a multi-layer encoding and decoding method that can be implemented with HEVC, VVC, etc., and encoding and decoding the basic image encoding and decoding method and the object mask image encoding and decoding method as separate layers.

[0043] Note that independent image encoding and decoding methods can be applied to the basic image encoding unit 1021 and the basic image decoding unit 3011, and the object mask image encoding unit 1022 and the object mask image decoding unit 3012. For example, video encoding and decoding methods such as AVC, HEVC, and VVC may be applied respectively.

[0044] The object mask information decoding device 302 decodes the object mask information encoded data in the encoded data according to the specifications of the object mask information SEI message described later, creates the object mask information, and sends it to the image recognition processing device 50.

[0045] FIG. 6 is a block diagram showing the configuration of the object mask information decoding device 302. The object mask information decoding device 302 includes an object mask information decoding unit 3021 and a post-image processing unit 3022, and decodes the object mask information encoded data in the encoded data Te based on the syntax of the object mask information SEI described later. Using the decoded object mask information and the object mask image information, the post-image processing unit 3022 performs the semantics and decoding process of the object mask information SEI described later, and sends it to the image recognition device 50 as image recognition information.

[0046] In the image recognition device 50, since the target image recognition process can be performed based on the image recognition information pre-processed by the post-image processing unit 3022, even for the basic image information Td deteriorated by the encoding and decoding processes, a high recognition rate and accuracy can be ensured with a small amount of calculation.

[0047] In the present embodiment, based on the syntax described later, it is assumed that encoding and decoding are performed as a plurality of SEI (Supplemental Enhancement Information) messages. Note that the encoding and decoding methods are not limited to the SEI message, and encoding and decoding may be performed as the syntax in the video encoding and decoding methods.

[0048] Next, the scalability dimension information (SDI: Scalability Dimension Information) SEI message described in Non-Patent Document 1 will be explained.

[0049] The scalability dimension information SEI message indicates the information of each layer in hierarchical coding. Specifically, when there are multiple viewpoints, it includes the viewpoint identifier of each layer, and when there is auxiliary information (such as depth or alpha) carried by one or more layers, it includes the auxiliary identifier of each layer. This SEI message is information in the CVS (Coded Video Sequence) unit, which is a single unit of the coded data, and exists in the first AU (Access Unit).

[0050] Note that the meaning of the notation of Descriptor in the following syntax table is interpreted as follows.

[0051] ・b(8): Represents the value of a byte having an arbitrary pattern of a bit string (8 bits).

[0052] ・f(n): Represents a bit string of a fixed pattern that uses n bits written in order from the left bit (from left to right).

[0053] ・se(v): Represents a syntax element obtained by encoding a signed integer using the 0th-order Exp-Golomb code.

[0054] ・st(v): Represents a string encoded in UTF-8 and terminated with null.

[0055] ・u(n): Represents an unsigned integer that uses n bits. In the syntax table, when n is "v", the number of bits changes according to the value of other syntax elements.

[0056] ・ue(v): Represents a syntax element obtained by encoding an unsigned integer using the 0th-order Exp-Golomb code (the left bit is the first).

[0057] Hereinafter, the syntax elements in FIG. 7 will be described.

[0058] The syntax element sdi_max_layers_minus1 + 1 indicates the maximum number of layers in the current CVS.

[0059] A value of 1 for the syntax element sdi_multiview_info_flag indicates that the current CVS may have multiple views and that the syntax element sdi_view_id_val[] is present in this SEI message. A value of 0 for sdi_multiview_info_flag indicates that the current CVS does not have multiple views and that the syntax element sdi_view_id_val[] is not present in this SEI message.

[0060] A value of 1 for the syntax element sdi_auxiliary_info_flag indicates that one or more layers in the current CVS may be auxiliary layers that transmit auxiliary information, and that the syntax element sdi_aux_id[] is present in this SEI message. A value of 0 for sdi_auxiliary_info_flag indicates that there are no auxiliary layers in the current CVS, and that the syntax element sdi_aux_id[] is not present in this SEI message.

[0061] The syntax element sdi_view_id_len_minus1 + 1 specifies the length of the syntax element sdi_view_id_val[i] in bits.

[0062] The syntax element sdi_layer_id[i] specifies the layer identifier of the i-th possible layer in the current CVS. Note that in HEVC and VVC, the layer identifier is the value of the syntax element nuh_layer_id in the NAL unit header.

[0063] sdi_view_id_val[i] specifies the view identifier for the i-th tier of the current CVS in the case of multiple views. The length of the sdi_view_id_val[i] syntax element is sdi_view_id_len_minus1 + 1 bits. The variable NumViews, which specifies the number of views in the current CVS, and the list ViewId, which specifies the view identifiers of the views in the current CVS, are derived as follows:

[0064] ) { ViewId[0] = sdi_view_id_val[0] if( i = 1; i <= sdi_max_layers_minus1; i++ ) { newViewFlag = 1 if( j = 0; j < i; j++ ) if( sdi_view_id_val[i] = = sdi_view_id_val[j] ) newViewFlag = 0 if( newViewFlag ) { ViewId[NumViews] = sdi_view_id_val[i] NumViews++}}} The syntax element sdi_aux_id[i] indicates what information is contained in the i-th layer of the current CVS. If the value of the syntax element sdi_aux_id[i] is 0, it indicates that the i-th layer of the current CVS does not contain any auxiliary pictures. If the value of sdi_aux_id[i] is greater than 0, it indicates the type of auxiliary picture at the i-th level of the current CVS. Note that if the value of the syntax element sdi_auxiliary_info_flag is equal to 0, the value of sdi_aux_id[i] is inferred to be equal to 0.

[0065] Figure 8 shows the content of the auxiliary information based on the value of the syntax element sdi_aux_id[i]. If the value of sdi_aux_id[i] is 0, the i-th layer is the primary layer; otherwise, the i-th layer is an auxiliary layer. If the value of sdi_aux_id[i] is 1, the i-th layer is an auxiliary layer for the alpha map. If the value of sdi_aux_id[i] is 2, the i-th layer is an auxiliary layer for depth information. Non-patent document 2 adds that if the value of sdi_aux_id[i] is 3, the i-th layer is an auxiliary layer for the object mask image.

[0066] Next, we will describe the Object Mask Information SEI message from non-patent literature. The Object Mask Information (OMI) SEI message provides object mask information for the object mask image of the auxiliary hierarchy related to the current main hierarchy in which the SEI message exists.

[0067] If an object mask information SEI message exists, it must be present in the primary hierarchy. A primary hierarchy can be associated with one or more secondary hierarchies. The number of secondary hierarchies associated with the current primary hierarchy is equal to the value of the variable omiNumAuxLayer, and the hierarchy identifier of the j-th associated secondary hierarchy is equal to the value of the variable omiAuxLayerId[j]. If the value of omiAuxLayerId[j] is equal to the value of sdi_layer_id[i], then for values ​​of i ranging from 0 to sid_max_layers_minus1, and for each value of j ranging from 0 to omiNumAuxLayer - 1, the value of the variable sdi_aux_id[i] must be equal to AUX_OBJECT_MASK. If a scalability dimension information SEI message does not exist in the current CVS, the object mask information SEI message is ignored.

[0068] To use this SEI message, the following variables must be defined: • BitDepthY, the pixel bit length of the luminance pixel array of the current image.

[0069] - CroppedWidth and CroppedHeight represent the width and height of the image cropped in the conformance window, expressed in units of brightness pixels.

[0070] - Conformance cropping window left offset (ConfWinLeftOffset).

[0071] - Conformance cropping window top offset (ConfWinTopOffset).

[0072] The variable ChromaFormatId indicates the color difference format; a value of 0 indicates monochrome, a value of 1 indicates 4:2:0, a value of 2 indicates 4:2:2, and a value of 3 indicates 4:4:4.

[0073] The variables SubWidthC and SubHeightC are determined by the value of ChromaFormatIdc. As shown in Figure 9, if ChromaFormatIdc is valued at 0 (monochrome), both SubWidthC and SubHeightC are 1. If ChromaFormatIdc is valued at 1 (4:2:0), both SubWidthC and SubHeightC are 2. If ChromaFormatIdc is valued at 2 (4:2:2), SubWidthC is 2 and SubHeightC is 1. If ChromaFormatIdc is valued at 3 (4:4:4), both SubWidthC and SubHeightC are 1.

[0074] The variables omiNumAuxLayer and omiAuxLayerId[k], which contain the number of images in the object mask image and the ID information of the object mask image, are derived as follows:

[0075] for( i = 0; i <= sdi_max_layers_minus1; i++ ) if( sdi_aux_id[ i ] = = AUX_OBJECT_MASK ) for( j = 0; j <= sdi_num_associated_primary_layers_minus1[ i ]; j++ ) if( sdi_layer_id[ sdi_associated_primary_layer_idx[ i ][ j ] ] = = omiPrimaryLayerId ) { omiAuxLayerId[ omiNumAuxLayer ] = sdi_layer_id[ i ] omiNumAuxLayer++;} Here, the variable omiPrimaryLayerId is the layer identifier of the current primary layer.

[0076] Figure 10 shows the syntax of the object mask information SEI message from Non-Patent Document 2.

[0077] If the syntax element omi_cancel_flag is 1, this SEI message indicates that, if an object mask information SEI message exists in the same hierarchy layer, the persistence of the previous object mask information SEI message will be canceled in the order of output.

[0078] The syntax element omi_persistence_flag specifies the persistence of object mask information provided in this SEI message. If omi_persistence_flag is 1, it specifies that the object mask information applies only to the current picture. If omi_persistence_flag is 1, it specifies that the object mask information applies to the current picture and all subsequent pictures on the same layer in output order until one or more of the following conditions are true.

[0079] - A new CLVS (Coded Layer Video Sequence) will begin in the current layer.

[0080] - The bitstream ends.

[0081] - The picture at the current hierarchy within the PU (Picture Unit) containing the object mask information SEI message will be output in output order, following the current picture.

[0082] If the CVS does not contain a scalability dimension information SEI message for at least one value of i where sdi_aux_id[i] is equal to AUX_OBJECT_MASK, the object mask information SEI message is ignored.

[0083] If the AU contains at least one value i for both the scalability dimension information SEI message and the object mask information SEI message, where sdi_aux_id[i] is equal to AUX_OBJECT_MASK, then the scalability dimension information SEI message must be decoded before the object mask information SEI message.

[0084] The value obtained by adding 1 to the syntax element omi_num_aux_pic_layer_minus1 indicates the number of auxiliary layers associated with the current primary layer. The value of omi_num_aux_pic_layer_minus1 + 1 must be equal to omiNumAuxLayer.

[0085] The value obtained by adding 1 to the syntax element omi_mask_id_length_minus1 specifies the bit length of the syntax element omi_mask_id[i][j]. The value of omi_mask_id_length_minus1 must be in the range of 0 to 31.

[0086] The value obtained by adding 8 to the syntax element omi_mask_sample_value_length_minus8 specifies the bit length of the syntax element omi_aux_sample_value[i][j]. The value of omi_mask_sample_value_length_minus8 must be in the range of 0 to 8. The value obtained by adding 8 to omi_mask_sample_value_length_minus8 must be less than or equal to BitDepthY.

[0087] If the syntax element omi_mask_confidence_info_present_flag is 1, it indicates that the syntax element omi_mask_confidence[i][j] exists. If omi_mask_confidence_info_present_flag is 0, it indicates that omi_mask_confidence[i][j] does not exist.

[0088] The value obtained by adding 1 to the syntax element omi_mask_confidence_length_minus1 specifies the length of the syntax element omi_mask_confidence[i][j] in bits.

[0089] If the syntax element omi_mask_depth_info_present_flag is 1, it indicates that the syntax element omi_mask_depth[i][j] exists. If omi_mask_depth_info_present_flag is 0, it indicates that omi_mask_depth[i][j] does not exist.

[0090] The value obtained by adding 1 to the syntax element omi_mask_depth_length_minus1 specifies the bit length of the syntax element omi_mask_depth[i][j] in bits.

[0091] The values ​​of the syntax elements omi_num_aux_pic_layer, omi_mask_id_length_minus1, omi_mask_sample_value_length_minus8, omi_mask_confidence_info_present_flag, omi_mask_confidence_length_minus1 (if present), omi_mask_depth_info_present_flag, and omi_mask_depth_length_minus1 (if present) must be the same for all object_mask_info() syntax structures within CLVS.

[0092] If the syntax element omi_mask_label_info_present_flag is 1, it indicates that the syntax elements omi_mask_label_language_present_flag and omi_mask_label[i][j] exist. If omi_mask_label_info_present_flag is 0, it indicates that the syntax elements omi_mask_label_language_present_flag and omi_mask_label[i][j] do not exist.

[0093] If the syntax element omi_mask_label_language_present_flag is 1, it indicates that the syntax element omi_mask_label_language exists. If omi_mask_label_language_present_flag is 0, it indicates that the syntax element omi_mask_label_language does not exist.

[0094] The syntax element omi_bit_equal_to_zero is set to 0, and the syntax element omi_mask_label_languag is inserted so that it is positioned byte by byte on the bitstream.

[0095] The syntax element omi_mask_label_language contains a language tag as defined in IETF RFC 5646, followed by a null-terminating byte with the value 0x00. The length of the omi_mask_label_language syntax element, excluding the null-terminating byte, must be 255 bytes or less. If it does not exist, the label language is not specified.

[0096] If the syntax element omi_mask_pic_update_flag[i] is 1, it indicates that the object mask information of the object mask image of the i-th auxiliary hierarchy associated with the current main hierarchy may be updated. If omi_mask_pic_update_flag[i] is 0, it indicates that there is no change in the mask information of the object mask image of the i-th auxiliary hierarchy associated with the current main hierarchy. If omi_mask_pic_update_flag[i] is 0, the persistence mechanism is used. In other words, the object mask information of the object mask image of the i-th auxiliary hierarchy associated with the current main hierarchy is inherited from the last object mask information SEI message present in the same hierarchy in decoding order.

[0097] The syntax element omi_num_mask_in_pic_update[i] specifies the number of object masks in the i-th auxiliary hierarchy object mask image associated with the current primary hierarchy. The value of omi_num_mask_in_pic_update[i] must be in the range of 0 to (1 << (omi_mask_id_length_minus1 + 1)) - 1.

[0098] The syntax element omi_mask_id[i][j] represents the ID, or identifier, of the j-th object mask information for the i-th auxiliary hierarchy object mask image associated with the current primary hierarchy. The length of the omi_mask_id[i][j] syntax element is omi_mask_id_length_minus1 + 1 bits.

[0099] The variable maskId[i][j], which specifies the ID of the j-th object mask information for the i-th auxiliary hierarchy object mask image associated with the current main hierarchy, is derived as follows:

[0100] for( i = 0; i < omi_num_aux_pic_layer; i++ ) for( j = 0; j < omi_num_mask_in_pic_update[ i ]; j++ ) maskId[ i ][ j ] = omi_mask_id[ i ][ j ] + ( 1 << ( omi_mask_id_length_minus1 + 1 ) ) * i omiA is defined as an object mask information SEI message containing the mask object image objectMaskA with maskId[ i0 ][ j0 ], and omiB is defined as the first object mask information SEI message following omiA in the output order within the same CLVS, containing the mask object image objectMaskB with maskId[ i1 ][ j1 ], and whose maskId[ i0 ][ j0 ] is equal to maskId[ i1 ][ j1 ]. objectMaskA and objectMaskB are object masks of the same object if both of the following conditions are true:

[0101] - The value of omi_mask_cancel[i0][j0] in omiA is equal to 0.

[0102] - Within the same CLVS where omi_cancel_flag is equal to 1, there is no object mask information SEI message that follows omiA and precedes omiB in output order.

[0103] If the syntax element omi_mask_cancel[i][j] is 1, it cancels the persistence of the j-th object mask of the i-th auxiliary hierarchy object mask image associated with the current base image. If omi_mask_cancel[i][j] is 0, it indicates that the information of the j-th object mask of the i-th auxiliary hierarchy object mask image associated with the current main hierarchy is encoded.

[0104] If an omi_mask_id[i][j] with a specific value is decrypted for the first time within the current CLVS, the corresponding value of omi_mask_cancel[i][j] must be 0.

[0105] The syntax element omi_aux_sample_value[i][j] indicates the pixel value within the region of the j-th object mask information of the object mask image of the i-th auxiliary hierarchy associated with the current primary hierarchy.

[0106] If the syntax element omi_mask_bounding_box_present_flag[i][j] is 1, it indicates that the syntax elements omi_mask_top[i][j], omi_mask_left[i][j], omi_mask_width[i][j], and omi_mask_height[i][j] exist. If omi_mask_bounding_box_present_flag[i][j] is 0, it indicates that the syntax elements omi_mask_top[i][j], omi_mask_left[i][j], omi_mask_width[i][j], and omi_mask_height[i][j] do not exist.

[0107] The syntax elements omi_mask_top[i][j], omi_mask_left[i][j], omi_mask_width[i][j], and omi_mask_height[i][j] indicate the coordinates of the top-left corner, width, and height of the bounding box of the j-th object mask information of the i-th auxiliary hierarchy of the extracted decoded object mask image associated with the current primary hierarchy for the fitted cutout window specified by the active SPS.

[0108] The value of the syntax element omi_mask_left[i][j] must be in the range of 0 to (CroppedWidth / SubWidthC - 1). If it does not exist, the value of omi_mask_left[i][j] is assumed to be 0.

[0109] The value of the syntax element omi_mask_top[i][j] ranges from 0 to (CroppedHeight / SubHeightC - 1), where CroppedHeight and SubHeightC are associated with the object mask image of the i-th auxiliary hierarchy associated with the current primary hierarchy. If this does not exist, the value of omi_mask_top[i][j] is assumed to be 0.

[0110] The value of the syntax element omi_mask_width[i][j] is within the range of 0 to (CroppedWidth / SubWidthC - omi_mask_left[i][j]). If it does not exist, the value of omi_mask_width[i][j] is assumed to be (CroppedWidth / SubWidthC - omi_mask_left[i][j]).

[0111] The value of the syntax element omi_mask_height[i][j] is in the range of 0 to (CroppedHeight / SubHeightC - omi_mask_top[i][j]). If it does not exist, the value of omi_mask_height[i][j] is assumed to be (CroppedHeight / SubWidthC - omi_mask_top[i][j]).

[0112] The identified object mask information is defined as follows: horizontal coordinates are in the range from SubWidthC * ( ConfWinLeftOffset + omi_mask_left[ i ][ j ] ) to SubWidthC * ( ConfWinLeftOffset + omi_mask_left[ i ][ j ] + omi_mask_width[ i ][ j ] ) - 1; and vertical coordinates are in the range from SubHeightC * ( ConfWinTopOffset + omi_mask_top[ i ][ j ] ) to SubHeightC * ( ConfWinTopOffset + omi_mask_top[ i ][ j ] + omi_mask_height[ i ][ j ] ) - 1.

[0113] The variable pI[i][x][y] is the decoded value of the pixel at the relative pixel position (x, y) in the cropped object mask image of the i-th auxiliary hierarchy associated with the current main hierarchy. The following process is used to determine the mask region within the object mask image.

[0114] for( i = 0; i < omi_num_aux_pic_layer; i++ ) for( j = 0; j < omi_num_mask_in_pic_update[ i ]; j++ ) if( pI[ i ][ x ][ y ] = = omi_aux_sample_value [ i ][ j ] && x >= omi_mask_left[ i ][ j ] && x < omi_mask_left[ i ][ j ] + omi_mask_width[ i ][ j ] && y >= omi_mask_top[ i ][ j ] && y < omi_mask_top[ i ][ j ] + omi_mask_height[ i ][ j ] ) The pixel at position (x, y) is associated with object mask information with the ID maskId[ i ][ j ].

[0115] If omi_mask_label_info_present_flag is 1, the syntax element omi_mask_confidence[i][j] exists. omi_mask_confidence[i][j] specifies the confidence associated with the j-th object mask of the i-th auxiliary hierarchy object mask image associated with the current primary hierarchy, in units of 2 to the power of -(omi_mask_confidence_length_minus1 + 1), where a higher value for omi_mask_confidence[i][j] indicates higher confidence. The length of the syntax element omi_mask_confidence[i][j] is omi_mask_confidence_length_minus1 + 1 bits.

[0116] If omi_mask_depth_info_present_flag is 1, the syntax element omi_mask_depth[i][j] exists. omi_mask_depth[i][j] indicates the depth value of the object associated with the j-th object mask information of the i-th auxiliary hierarchy object mask image associated with the current primary hierarchy. A smaller value of omi_mask_depth indicates a shorter distance to the object. The length of the syntax element omi_mask_depth[i][j] is omi_mask_depth_length_minus1 + 1 bits.

[0117] If omi_mask_label_info_present_flag is 1, the syntax element omi_mask_label[i][j] exists. omi_mask_label[i][j] indicates the content of the label associated with the j-th object mask information of the i-th auxiliary hierarchy object mask image associated with the current primary hierarchy. The length of the syntax element omi_mask_label[i][j] must be 255 bytes or less, excluding the null-terminating byte.

[0118] One problem with Non-Patent Document 2 is that the syntax element omi_mask_pic_update_flag[i] is defined as f(1). f(n) represents a fixed-pattern bit string that uses n bits written sequentially from left to right. A fixed-bit pattern cannot function as a syntax element for a flag.

[0119] In this embodiment, the syntax element omi_mask_pic_update_flag[i] is described using u(1) instead of f(1).

[0120] Another problem with Non-Patent Document 2 is that the range of values ​​for the syntax element omi_mask_depth_length_minus1 is not defined.

[0121] Therefore, in this embodiment, the value of the syntax element omi_mask_depth_length_minus1 is set to be between 0 and 15.

[0122] In another form of this implementation, the value of the syntax element omi_mask_depth_length_minus1 is set to be greater than or equal to 0 (BitDepthY-1).

[0123] In another form of this implementation, the syntax element may be represented as omi_mask_depth_length_minus8, with a value between 0 and (BitDepthY-8), and expressed as u(3) in 3 bits. Then, the variable MaskDepthLength = omi_mask_depth_length_minus8 + 8 may be used.

[0124] Alternatively, instead of explicitly defining syntax elements, a variable MaskDepthLength = BitDepthY can be used to represent the bit length of the unsigned integer depth value of the mask object. In this case, the syntax omi_mask_depth_length_minus1, which indicates the bit length of omi_mask_depth[i][j], becomes unnecessary.

[0125] Another problem with Non-Patent Document 2 is that the range of values ​​for the syntax element omi_mask_id_length_minus1 is too wide. omi_mask_id_length_minus1 is defined as a value between 0 and 31. If the value of omi_mask_id_length_minus1 is 31, then omi_num_mask_in_pic_update[i] would range from 0 to (1<<(31+1))-1, that is, from 0 to 2 to the power of 32 minus 1. omi_num_mask_in_pic_update[i] represents the number of object masks, but this maximum value is greater than the number of pixels in an 8K resolution (7680x4320) image, making it impractical.

[0126] Similarly, the maximum value of the syntax element omi_mask_id[i][j] is 2 to the power of 32 minus 1. Furthermore, the array size also has a maximum value of 2 to the power of 32 for j, meaning a very large amount of memory must be allocated.

[0127] Furthermore, the maximum value of the variable maskId[i][j] is ((1<<32)-1) + ((omi_num_aux_pic_layer-1)<<32), and since this value exceeds 32 bits, there is an implementation problem.

[0128] Therefore, in this embodiment, the value of omi_mask_id_length_minus1 is set to be between 0 and 15. In this way, the value of omi_mask_id[i][j] becomes a 16-bit unsigned integer, and the maximum size of the array becomes 2 to the power of 16. Also, the maximum value of the variable maskId[i][j] becomes ((1<<16)-1) + ((omi_num_aux_pic_layer-1)<<16), so the value fits within 32 bits, and the problem is solved.

[0129] Note that the maximum value of omi_mask_id_length_minus1 can be any number other than 15; it can also be a positive integer such as 10, 11, or 12.

[0130] Another problem with Non-Patent Document 2 is that the byte alignment position of the syntax element omi_mask_label[i][j] is inappropriate. Since omi_mask_label[i][j] is a syntax element of a string, it needs to be written in an 8-bit unit position. The syntax of Non-Patent Document 2 is as shown in Figure 11, as follows: while( !byte_aligned( ) ) omi_bit_equal_to_zero f(1) if( omi_mask_label_info_present_flag ) omi_mask_label[ i ][ j ] st(v) In this case, if the syntax element omi_mask_label_info_present_flag is 0, byte alignment will occur even if omi_mask_label does not exist.

[0131] Therefore, in this embodiment, as shown in Figure 13, the syntax is as follows.

[0132] if( omi_mask_label_info_present_flag ) { while( !byte_aligned( ) ) omi_bit_equal_to_zero f(1) omi_mask_label[ i ][ j ] st(v)} This syntax prevents unnecessary byte alignment.

[0133] Another problem with Non-Patent Document 2 is the ambiguity of the confidence value associated with the object mask. In the Non-Patent Document, omi_mask_confidence[i][j] specifies the confidence associated with the j-th object mask of the i-th auxiliary hierarchy object mask image associated with the current main hierarchy in units of 2 to the power of -(omi_mask_confidence_length_minus1 + 1), with a larger value of omi_mask_confidence[i][j] indicating higher confidence, and the length of the syntax element of omi_mask_confidence[i][j] is omi_mask_confidence_length_minus1 + 1 bits representing an unsigned integer. Despite omi_mask_confidence[i][j] being represented as an unsigned integer, it is impossible to specify it in units of 2 to the power of -(omi_mask_confidence_length_minus1 + 1). Furthermore, omi_mask_confidence[i][j] can take a value of 0, but a confidence level of 0 indicates that it is not an object mask area, which is redundant.

[0134] Therefore, in this embodiment, a real number variable MaskConfidence[i][j] that takes a value greater than 0 and less than or equal to 1 is prepared, and MaskConfidence[i][j] = (omi_mask_confidence[i][j] + 1) ÷ (1 << (omi_mask_confidence_length_minus1 + 1)) is defined, and this is taken as the confidence level associated with the object mask.

[0135] This configuration allows for an appropriate level of confidence to be assigned to the object mask area.

[0136] Another problem with Non-Patent Literature 2 is that when determining the mask region in the auxiliary image and the variable pI[i][x][y], which is the decoded value of a pixel at a relative pixel position (x, y) in the cropped object mask image of the i-th auxiliary hierarchy associated with the current main hierarchy, the state of the canceled object mask region is not taken into consideration. Therefore, the following correspondence is necessary. Specifically, the condition that it is an object mask region is increased by adding that omi_mask_cancel[i][j] is not true (or omi_mask_cancel[i][j] is not 1, or omi_mask_cancel[i][j] is 0).

[0137] for( i = 0; i < omi_num_aux_pic_layer; i++ ) for( j = 0; j < omi_num_mask_in_pic_update[ i ]; j++ ) if( !omi_mask_cancel[ i ][ j ] ) if( pI[ i ][ x ][ y ] = = omi_aux_sample_value [ i ][ j ] && x >= omi_mask_left[ i | This is associated with object mask information that has the identifier ].

[0138] Another way to add conditions is to use a conditional expression that combines several if statements, as shown below, which will produce equivalent results.

[0139] for( i = 0; i < omi_num_aux_pic_layer; i++ ) for( j = 0; j < omi_num_mask_in_pic_update[ i ]; j++ ) if( ( !omi_mask_cancel[ i ][ j ] ) & & pI[ i ][ x ][ y ] = = omi_aux_sample_value [ i ][ j ] ) && x >= omi_mask_left[ i ][ j ] && x < omi_mask_left[ i ][ j ] + omi_mask_width[ i ][ j ] && y >= omi_mask_top[ i ][ j ] && y < omi_mask_top[ i ][ j ] + omi_mask_height[ i ][ j ] ) The pixel at position (x, y) is maskId[ i ][ j This is associated with object mask information that has the identifier ].

[0140] When object mask information is matched with an object mask image, the presence or absence of cancellation in the object mask information ensures that the matching is performed correctly.

[0141] In this embodiment, an object mask information encoding and decoding method using SEI is shown, but it can also be implemented by encoding and decoding the encoded data as an internal syntax of the video encoding and decoding method.

[0142] In this embodiment, we demonstrated that a machine-oriented video encoding and decoding image transmission system can be realized by encoding and decoding image information using hierarchical image encoding, and encoding and decoding object max information using SEI messages.

[0143] Furthermore, some or all of the video encoding device 10 and video decoding device 30 in the above-described embodiment may be implemented using a computer. In that case, the program for implementing this control function may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be loaded into a computer system and executed. Here, "computer system" refers to a computer system built into either the video encoding device 10 or the video decoding device 30, and includes hardware such as an OS and peripheral devices. Furthermore, "computer-readable recording medium" refers to portable media such as flexible disks, magneto-optical disks, ROMs, CD-ROMs, and storage devices such as hard disks built into a computer system. Moreover, "computer-readable recording medium" may also include those that dynamically hold programs for a short period of time, such as communication lines used when transmitting programs via networks such as the Internet or communication lines such as telephone lines, and those that hold programs for a certain period of time, such as volatile memory inside a computer system that acts as a server or client in such a case. Furthermore, the above-mentioned program may be for implementing some of the functions described above, and may also be able to implement the above-mentioned functions in combination with programs already recorded in the computer system.

[0144] Furthermore, some or all of the video encoding device 10 and video decoding device 30 in the above-described embodiment may be implemented as an integrated circuit such as an LSI (Large Scale Integration). Each functional block of the video encoding device 10 and video decoding device 30 may be individually implemented as a processor, or some or all of them may be integrated into a single processor. In addition, the method of implementing the integrated circuit is not limited to LSIs; it may also be implemented using dedicated circuits or general-purpose processors. Furthermore, if an integrated circuit technology that can replace LSIs emerges due to advances in semiconductor technology, an integrated circuit using that technology may be used.

[0145] Although one embodiment of this invention has been described in detail above with reference to the drawings, the specific configuration is not limited to that described above, and various design changes can be made without departing from the spirit of this invention.

[0146] The embodiments of the present invention are not limited to those described above, and various modifications are possible within the scope of the claims. That is, embodiments obtained by combining technical means that have been appropriately modified within the scope of the claims are also included in the technical scope of the present invention.

[0147] Embodiments of the present invention can be suitably applied to a video decoding device that decodes encoded data in which an image signal has been encoded, and a video encoding device that generates encoded data in which image data has been encoded. Furthermore, they can be suitably applied to the data structure of encoded data generated by the video encoding device and referenced by the video decoding device.

[0148] 1 Image transmission system 10 Video encoding device 101 Pre-image processing device 1011 Object mask image processing device 102 Hierarchical image encoding device 1021 Basic image encoding unit 1022 Object mask image encoding unit 103 Object mask information creation device 1031 Pre-image recognition unit 1032 Object mask information creation unit 104 Object mask information encoding device 20 Transmission network 30 Video decoding device 301 Hierarchical image decoding device 3011 Basic image decoding unit 3012 Object mask image decoding unit 302 Object mask information decoding device 3021 Object mask information decoding unit 3022 Post-image processing unit 40 Image display device 50 Image recognition device

Claims

1. A video decoding device comprising a hierarchical image decoding device that decodes encoded data of image information consisting of multiple image signals including an object mask image for image recognition, and an object mask information decoding device that decodes encoded data of object mask information, wherein the object mask information represents the object mask region as a rectangle, and when comparing it with the object mask image, the state in which the object mask information has not been canceled is taken into consideration.

2. A video decoding device comprising a hierarchical image decoding device that decodes encoded data of image information consisting of multiple image signals including an object mask image for image recognition, and an object mask information decoding device that decodes encoded data of object mask information, wherein the object mask information has a confidence level greater than 0.

3. A video encoding device comprising a hierarchical image encoding device for encoding image information consisting of multiple image signals including an object mask image for image recognition, and an object mask information encoding device for encoding object mask information, wherein the object mask information represents the object mask region as a rectangle, and when comparing it with the object mask image, it takes into account the state in which the object mask information has not been canceled.