Video encoding device and decoding device

The video decoding device uses URIs to efficiently encode and decode generative AI information, reducing code size and ensuring byte-aligned character interpretation, addressing inefficiencies in traditional video encoding methods.

JP2025161009APending Publication Date: 2025-10-24SHARP KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024063828
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-11
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing video encoding methods using generative AI require large amounts of information for full-scale post-image processing, leading to excessive coding and inefficiencies due to the need for multiple models and control parameters, and character information not being byte-aligned for immediate decoding.

Method used

A video decoding device that decodes encoded data using a URI-based approach to identify generation information, allowing efficient encoding and decoding of image generation techniques by specifying generation information through URIs, reducing the need for direct transmission of large amounts of text information.

Benefits of technology

This method enables efficient video encoding and decoding with reduced code size by utilizing URIs to specify generation information, addressing the inefficiencies of traditional methods and ensuring byte-aligned character interpretation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025161009000001_ABST
    Figure 2025161009000001_ABST
Patent Text Reader

Abstract

To solve a problem in which, using generative AI makes it possible to process images using text information provided by prompts; however, when attempting to perform full-scale post-image processing, not only the prompts but also the number of models and control parameters required to define the processing increases, and thereby the amount of information is increased and the amount of coding for the information is increased.SOLUTION: A video decoding device according to the present invention includes an image decoding device that decodes coded data of an image signal, a generated information decoding device that decodes coded data of a URI that identifies generated information for image generation, and an image generating device that generates an image from the image information decoded by the video decoding device and the generated information decoded by the generated image decoding device, and the generated information decoding device decodes the coded data to determine the encoding method of text information indicating the generated information.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a video encoding device and a video decoding device. [Background technology]

[0002] In order to efficiently transmit or record moving images, a moving image encoding device is used that generates encoded data by encoding an image, and a moving image decoding device is used that generates a decoded image by decoding the encoded data.

[0003] A specific example of a video encoding method is the H.266 / VVC (Versatile Video Coding) method.

[0004] In such traditional image coding methods, an image is divided into parts for encoding / decoding. First, a predicted image is generated based on a locally decoded image obtained by encoding an input image / decoding the encoded data. Next, the predicted image is subtracted from the input image (original image) to obtain a prediction error (sometimes called a "difference image" or "residual image"), which is then coded / decoded.

[0005] Recently, a generative AI technique called Stable Diffusion, which uses a diffusion model as an image generation method using a neural network, has been disclosed. This technique can generate images based on text entered by the user, called a prompt.

[0006] Non-Patent Document 1 proposes a Supplemental Enhancement Information (SEI) message that can specify text information of a prompt for image processing in a generation AI for video encoded data.

[0007] Furthermore, Non-Patent Document 2 defines a Supplemental Enhancement Information (SEI) message as a video encoding and decoding technique for transmitting image properties, display methods, timing, etc. simultaneously with encoded data. It also presents a Neural-Network Post-filter Activation SEI message that indicates the application of post-filter processing based on a neural network. [Prior art documents] [Non-patent literature]

[0008] [Non-Patent Document 1] A. Hinds, G. Teniou, and S. Wenger, “AHG9: Text prompt for generative AI SEI,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29, JVET-AG0167, Jan. 2024. [Non-patent document 2] ITU-T Rec. H.274 V3 "Versatile supplemental enhancement information messages for coded video bitstreams" Summary of the Invention [Problem to be solved by the invention]

[0009] In the method disclosed in Non-Patent Document 1, image processing by generative AI can be realized using text information provided by prompts. However, when trying to perform full-scale post-image processing, not only prompts but also many models and control parameters are required to define the processing. This leads to a problem that the amount of information becomes enormous, which results in a large amount of coding. Also, character information that is not aligned in byte units cannot be used immediately after decoding. [Means for solving the problem]

[0010] A video decoding device according to one aspect of the present invention includes: an image decoding device that decodes encoded data of an image signal; a generation information decoding device that decodes coded data of a URI that identifies generation information for image generation; an image generating device that generates an image from image information decoded by the video decoding device and generated information decoded by the generated image decoding device; The generated information decoding device is characterized in that it decodes the encoding method of text information indicating the generated information from the encoded data. [Effects of the Invention]

[0011] By adopting such a configuration, it is possible to solve the problem of implementing video encoding and decoding with a small amount of code using an image generation technique. [Brief explanation of the drawings]

[0012] [Figure 1] 1 is a schematic diagram showing the configuration of an image transmission system according to an embodiment of the present invention. [Figure 2] FIG. 1 is a block diagram illustrating an example of an image generation processing apparatus according to an embodiment of the present invention. [Figure 3] FIG. 1 is a diagram illustrating an example of a block diagram of a generation information creating device according to an embodiment of the present invention. [Figure 4] FIG. 1 is a diagram illustrating an example of a block diagram of a generated information encoding device according to an embodiment of the present invention. [Figure 5] FIG. 2 is a diagram illustrating an example of a block diagram of a generated information decoding device according to the present embodiment. [Figure 6] FIG. 10 is a diagram showing the syntax of the AI ​​Text Data SEI message of Reference 1. [Figure 7] FIG. 10 is a diagram showing an example of an extension of an AI text data SEI message according to this embodiment. [Figure 8] FIG. 10 is a diagram showing another example of an extension of the AI ​​text data SEI message according to this embodiment. [Figure 9] FIG. 10 is a diagram showing another example of an extension of the AI ​​text data SEI message according to this embodiment. [Figure 10] FIG. 10 is a diagram showing an example of a text character string (example of JSON) indicating generation information according to the present embodiment. [Figure 11] FIG. 10 is a diagram showing a graph indicated by a text character string indicating generation information according to the present embodiment. [Figure 12] FIG. 10 is a diagram showing another example of an extension of the AI ​​text data SEI message according to this embodiment. [Figure 13] FIG. 10 is a diagram showing another example of an extension of the AI ​​text data SEI message according to this embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0013] (First embodiment) FIG. 1 is a conceptual diagram showing the configuration of an image transmission system according to this embodiment.

[0014] The image transmission system 1 comprises a video encoding device 10, a transmission network 20, a video decoding device 30, and an image display device 40.

[0015] The video encoding device 10 receives an input image signal T and outputs encoded data Te.

[0016] The transmission network 20 transmits the encoded data Te from the video encoding device 10 to the video decoding device 30. The transmission network 20 is the Internet, a wide area network (WAN), a local area network (LAN), or a combination of these. The network 20 is not necessarily limited to a bidirectional communication network, but may also be a unidirectional communication network that transmits broadcast waves such as terrestrial digital broadcasting and satellite broadcasting. Furthermore, the transmission network 20 may be replaced by a storage medium on which the encoded data Te is recorded, such as a DVD (Digital Versatile Disc: trademark) or a BD (Blue-ray Disc: registered trademark).

[0017] The video decoding device 30 receives the coded data Te as input, outputs a generated image Td, and sends it to the image display device 40.

[0018] The image display device 40 displays all or part of the generated image Td output from the video decoding device 30. The image display device 40 includes a display device such as a liquid crystal display or an organic EL (Electro-luminescence) display. The display may be in the form of a stationary display, a mobile display, an HMD, or the like. Furthermore, if the video decoding device 30 has high processing power, it displays high-quality images, and if it has only low processing power, it displays images that do not require high processing power or display power.

[0019] The video encoding device 10 comprises an image encoding device 101, a generated information creating device 102, and a generated information encoding device 103.

[0020] The image encoding device 101 encodes an input image signal T to generate encoded data Te, and sends the decoded image information to the generated information creating device 102 .

[0021] The generated information creating device 102 receives the input image signal T, external model data, and decoded image information from the video encoding device, creates generated information, and sends it to the generated information encoding device 103 .

[0022] The generated information encoding device 103 encodes the generated information and saves the generated information as URI data in a specified URI (Uniform Resource Identifier) ​​on a server or a specific storage location on a network, and generates supplemental extension information encoded data including the URI. A URI is a character string for identifying an abstract or physical resource, and a URI specified in RFC2396 or RFC3986 may be used. The URI may be the name of the information, such as a Uniform Resource Name (URN), or the location of the information, such as a Uniform Resource Locator (URL).

[0023] The video decoding device 30 comprises an image decoding device 301 , an image generation processing device 302 , and a generated information decoding device 303 .

[0024] The image decoding device 301 receives the live encoded data Te transmitted via the transmission network 20 as input, decodes the image information, and transmits it to the image generating device 302 .

[0025] The generation information decoding device 302 decodes the auxiliary extension information of the encoded data Te based on the syntax, loads the URI data from the storage location (a server on the network or a specific storage location) based on the decoded URI, creates generation information, and sends the generation information to the image generation device 301.

[0026] The image generation processing device 303 performs image generation processing using the image information decoded by the image decoding device 301, the generation information decoded by the generation information decoding device 303, and model data from outside, generates a generated image Td, and outputs it to the image display device 40.

[0027] In this embodiment, the image encoding device 101 and the image decoding device 301 are realized by applying a general-purpose video encoding and decoding method such as AVC, HEVC, or VVC.

[0028] Fig. 2 is a conceptual diagram showing the configuration of an image generation processing device according to this embodiment. The image generation processing device according to this embodiment uses a generation image processing method based on so-called image generation AI, which is configured using a neural network such as a diffusion model. Image information, generation information, and model data are input, and a generated image is output.

[0029] The image generation processing device 303 consists of an image generation unit 3031, a control unit 3032, and a control image generation unit 3033. The image generation unit 3031 uses a generated image processing method configured with a stable diffusion neural network. The control unit 3032 uses a control method configured with a neural network called a control net. The control image generation unit 3033 generates a control image signal from image information.

[0030] The control unit 3032 receives as input the image information, the control parameter information in the generation information, and the model data specified by the control parameters, and outputs control image information (control image signal) to be input to the image generation unit 3031. Here, the image information is the locally decoded image signal output by the image encoding device 101, or the decoded image signal output by the image decoding device 301.

[0031] Both are image information obtained by encoding and decoding an input image signal.

[0032] The control image signal is generated from image information by the control image generation unit 3033. Specifically, the control image signal uses the following information: The identification of these is included in the control parameters. Contour (Canny) image Soft edge image Sketch image Line art images Normal map image Depth map image Segmentation image Open Pose Image Wireframe (MLSD) images Inpainted image Reference image Both are monochrome or color image information created using an image. Note that the control image signal is not limited to one image, and multiple different control image signals may exist for the same image information.

[0033] The generation information consists of control parameter information, model information, model parameter information, prompt information, etc. The generation information may be coded and decoded by individually coding and decoding each piece of information, or the entire generation information may be coded and decoded as text information.

[0034] The control parameter information is a parameter for controlling the control unit 3302 described above, and includes identification of the basic image of the image information, identification of the control image, model information of the control unit, and the like.

[0035] The model information is the name of a neural network model to be used for image generation by the image generation unit 3301. The model data indicated by the model information is input as model data from outside to the image generation processing device 303. Alternatively, the video encoding device 10 and the video decoding device 30 may have the same model data.

[0036] The model parameters are parameters for controlling the neural network, and are various kinds of numerical values ​​and character string information (text information) such as intensity values, number of steps, type of sampler, and seed information.

[0037] The prompt information is character string information that indicates the content of the image to be generated. The prompt information includes a prompt that indicates the content that is desired to be generated and a negative prompt that indicates the content that is not desired to be generated.

[0038] 3 is a block diagram showing the configuration of generation information creation device 102 according to this embodiment. Generation information creation device 102 according to this embodiment receives an input image signal, image information created by a video encoding unit, and external model data, outputs generation information, and sends it to generation information encoding device 103.

[0039] The generated information creation device 102 is composed of a generated information creation unit 1021, an encoding control unit 1022, and an image generation processing device 1023. The image generation processing device 1023 is the same as the image generation processing device 303 described above, and outputs a generated image from the generated information, image information, and model data. The encoding control unit 1032 selects generated information based on two indices: an evaluation criterion D for image similarity, such as mean square error, absolute sum of error, SSIM (Structural Similarity), MS-SSIM (Multi-Scale Structural Similarity), or LPIPS (Learned Perceptual Image Patch Similarity), which is based on a comparison between the generated image result of the image generation device 303 and the input image signal; and the code amount R of the generated information created by the generated information creation unit 1021, and outputs the optimal one.

[0040] The generation information creating unit 1021 generates generation information by exchanging information with the encoding control unit 1022 and sends it to the image generation processing device 1023 .

[0041] The generation information encoding device 103 encodes the generation information created by the generation information creation device 102 and sends the auxiliary extension information encoded data together with the encoded data output by the image encoding device 101 as encoded data Te to the created transmission network 20.

[0042] The generated information decoding device 302 decodes the auxiliary extension encoded data out of the encoded data Te sent from the transmission network 20, and sends the decoding result to the image generation processing device 303 as generated information.

[0043] In this embodiment, encoding and decoding are performed as an SEI (Supplemental Enhancement Information) message based on a syntax described later. Note that the encoding and decoding method is not limited to the SEI message, and encoding and decoding may also be performed as a syntax in a video encoding and decoding method, such as an APS (Adaptation Parameter Set).

[0044] 4 is a block diagram showing the configuration of generation information encoding device 103 of this embodiment. Generation information encoding device 103 of this embodiment includes a supplemental extension information encoding unit 1031, a URI data encoding unit 1032, and a URI data saving unit 1033.

[0045] The supplementary extension information encoding unit 1031 defines an identifier for the generation information generated by the generation information creation device 102 as a URI (Uniform Resource Identifier), creates supplementary extension information encoded data as an SEI message described later, and sends it to the transmission network 20 as part of the encoded data Te.

[0046] The URI data encoding unit 1032 encodes the content of the generated information, for which the identifier is defined by the supplementary extension information encoding unit 1031, as text information or compressed text data, and generates the URI data. and sends it to the URI data save unit 1033.

[0047] The URI data saving unit 1033 saves the URI data encoded by the URI data encoding unit 1032 to the URI defined by the auxiliary extension information encoding unit 1031 in the location indicated by the URI (a server on the network or a specific storage location).

[0048] 5 is a block diagram showing the configuration of the generation information decoding device 302 according to this embodiment. The generation information decoding device 302 according to this embodiment includes a supplemental extension information decoding unit 3021, a URI data decoding unit 3022, and a URI data loading unit 3023.

[0049] The auxiliary extension information decoding unit 3021 decodes the auxiliary extension information encoded data out of the encoded data Te received from the transmission network 20. The auxiliary extension encoded data is decoded as an SEI message (described later), and the decoded URI is sent to the URI data loading unit 3023.

[0050] The URI data loading unit 3023 loads the encoded URI data from the location where it is stored (a server on the network or a specific storage location) based on the URI decoded by the supplemental extension information decoding unit 3021.

[0051] The URI data decoding unit 3022 decodes the generation information from the loaded URI data and sends it to the auxiliary extension decoding unit 3021. The auxiliary extension information decoding unit 3021 combines the generation information with the decoding result of the auxiliary extension encoded data and outputs it to the image generation processing device 303 as the final result.

[0052] 6, 7, 8, 9, 12, and 13 show the syntax of the generated information coded data that is coded and decoded by generated information coding device 103 and generated information decoding device 302 in this embodiment.

[0053] The meaning of the Descriptor notation in the following syntax tables is interpreted as follows: · b(8): Represents a byte value with any pattern of bit string (8 bits). f(n): represents a fixed-pattern bit string using n bits written left to right. se(v): represents a syntax element obtained by encoding a signed integer into 0th-order Exp-Golomb code. st(v): Represents a null-terminated string encoded in UTF-8. u(n): represents an unsigned integer using n bits. If n is "v" in the syntax table, the number of bits varies depending on the values ​​of other syntax elements. ·ue(v): represents the syntax element (left bit first) encoded as an unsigned integer with order 0 Exp-Golomb coding.

[0054] Figure 6 shows the syntax of the AI ​​Text Data SEI message in Non-Patent Document 1. In this embodiment, the specifications of this SEI message are extended. This SEI message sends prompt information as text, enabling post-image processing by the generation AI.

[0055] The syntax elements in FIG. 6 will be explained below.

[0056] A value of 1 in the syntax element ait_cancel_flag indicates that the SEI message cancels the persistence of the previous AI text data SEI message in the output order. A value of 0 in ait_cancel_flag indicates that AI text data information follows.

[0057] The syntax element ait_data_persistence_flag specifies the persistence of the AI ​​text data SEI message of the current layer. If the value of ait_data_persistence_flag is 0, the AI ​​text data ait_data_persistence_flag with a value of 1 specifies that the AI ​​text data SEI message applies to the current decoded picture only. ait_data_persistence_flag with a value of 1 specifies that the AI ​​text data SEI message applies to the current decoded picture and persists in output order for all subsequent pictures in the current hierarchical layer until one or more of the following conditions become true: A new CLVS for the current hierarchy is started. The bitstream ends. ·The picture of the current layer in the AU associated with the AI ​​text data SEI message is output following the current picture in output order.

[0058] The syntax element ait_data_string is a text string containing a command prompt to be interpreted by the generative AI engine. The text prompt is encoded as specified in ISO / IEC 10646: Information technology - Universal Coded Character Set (UCS). UTF-8 of the UCS may be used here, as specified by st(v). ait_data_string may be text information representing generation information. It may also be prompt information.

[0059] In the AI ​​Test Data SEI message in Reference 1, image processing by the generating AI was realized by sending text information. However, when trying to perform full-scale post-image processing, not only prompts but also the number of models and control parameters required to define the processing increases, resulting in an enormous amount of text information and a large code size for the SEI.

[0060] Therefore, in this embodiment, the syntax of the AI ​​text data SEI message is extended to not only transmit text information directly as SEI, but also define a URI to specify the generation information, as in the syntax of Figure 7, and define the URI as SEI. The generation information is then specified by the URI. The information indicated by the URI may be stored in advance in the generation information encoding device 103 and the generation information decoding device 302, or may be stored in a location indicated by the URI. It may also be stored in a location known to the generation information encoding device 103 and the generation information decoding device 302. For example, information stored on the Internet does not need to be transmitted again once it has been loaded. In this way, there is no need to directly transmit large amounts of data, and the problem can be solved.

[0061] Hereinafter, in FIG. 7, syntax elements added from FIG. 6 will be explained.

[0062] If the value of the syntax element ait_data_mode_idc is 0, it indicates that the generated information is encoded as text information by the syntax element ait_data_string. If the value of ait_data_mode_idc is 1, it indicates that the generated information is identified by the URI indicated by ait_data_uri in the format identified by the tag URI, ait_data_tag_uri.

[0063] The byte_aliged() function returns whether the current encoded data is in byte units. If it is not in byte units, it inserts the syntax element ait_data_bit_equal_to_zero to adjust the bit position so that the next element is positioned on a byte boundary. ait_data_alignment_zero_bit_a is assumed to be equal to 0.

[0064] In the AI ​​Test Data SEI message in Reference 1, there was an issue that the character information could not be used immediately after being decoded because it was not aligned in byte units before the ait_data_string.

[0065] Therefore, in this embodiment, byte_aligned() is inserted before the character string ait_data_string for generating AI, and if it is not byte aligned, a predetermined bit is inserted to have the effect that ait_data_string can be easily interpreted as a character. By putting byte_aliged() before i, the URI strings ait_data_tag_uri and ait_data_uri can be easily interpreted as characters.

[0066] The syntax element ait_data_tag_uri contains a tag URI with syntax and semantics specified in IETF RFC 4151 - The 'tag' URI Scheme, which identifies the data type of the generated information and related information. Using ait_data_tag_uri makes it possible to uniquely identify the data type of the generated information specified by ait_data_uri without the need for a registration authority. For example, if ait_data_tag_uri is equal to "tag:stable_diffusion:deforum", it indicates that this is the Deforum configuration file for the Stable Diffusion identified by ait_data_uri. ait_data_uri contains a URI with syntax and semantics specified in IETF RFC 3986 on Uniform Resource Identifiers (URI), which identifies the generated information.

[0067] Note that ait_data_tag_uri and ait_data_uri may specify information related to the generated information, such as copyright, terms of use, etc. As described above, by identifying the generated information with a URI, it is possible to handle even large amounts of generated information for image generation.

[0068] As another alternative, a zip-compressed UCS string may be specified as a possible value of ait_data_mode_idc. For example, the following syntax configuration may be used when the value of ait_data_mode_idc is 2. Here, when ait_data_mode_idc is a specific value, the zip-compressed UCS string indicated by ait_zip_data_string, for example, a zip-compressed UCS-8 string, is decoded. The zip encoding method may be the method specified in ISO / IEC 21320-1.

[0069] if (ait_data_mode_idc == 0) ait_data_string else if (ait_data_mode_idc == 1) ait_data_tag_uri ait_data_uri else if (ait_data_mode_idc == 2) ait_zip_data_string As described above, by encoding the generation information using ZIP, it is possible to deal with even large amounts of generation information for image generation.

[0070] FIG. 8 shows an embodiment in which, in addition to the syntax of FIG. 7, information required for displaying the generated image to be output on the image display device 40 is added and described.

[0071] The syntax element ait_output_parameters_present_flag is a flag indicating whether information about the generated image to be output exists. If the value of ait_output_parameters_present_flag is 1, information about the generated image to be output exists. If the value of ait_output_parameters_present_flag is 0, information about the generated image to be output does not exist.

[0072] The value of the syntax element ait_output_pic_width_minus1 plus 1 indicates the number of pixels in the horizontal direction of the luminance signal of the generated image that is output.

[0073] The value of the syntax element ait_output_pic_height_minus1 plus 1 indicates the number of pixels in the vertical direction of the luminance signal of the output generated image.

[0074] The syntax element ait_output_chroma_format_idc indicates the chrominance format of the generated image to be output. If the value of ait_output_chroma_format_idc is 0, it indicates a monochrome image (4:0:0 format), if the value of ait_output_chroma_format_idc is 1, it indicates the 4:2:0 format, and if the value of ait_output_chroma_format_idc is 2, it indicates the 4:2:2 format. If the value of ait_output_chroma_format_idc is 3, it indicates the 4:4:4 format.

[0075] The value of the syntax element ait_output_luma_bitdepth_minus8 plus 8 indicates the pixel bit length of the luminance signal of the resulting image that is output.

[0076] The value of the syntax element ait_output_chroma_bitdepth_minus8 plus 8 indicates the pixel bit length of the color difference signal of the output resulting image.

[0077] The value of the syntax element ait_output_interpolated_pic indicates the number of pictures interpolated between input pictures.

[0078] The syntax element ait_vui_parameters_present_flag is a flag indicating whether or not VUI (Video Usability Information) parameter information exists for the generated image to be output. If the value of ait_vui_parameters_present_flag is 1, this indicates that VUI parameters exist for the generated image to be output. If the value of ait_vui_parameters_present_flag is 0, VUI parameters do not exist for the generated image to be output.

[0079] The syntax element ait_vui_payload_size_minus1 plus 1 indicates the number of bytes of raw data of the information inside the syntax structure of vui_payload().

[0080] The byte_aliged() function returns whether the current encoded data is in byte units. If it is not in byte units, it inserts the syntax element ait_vui_bit_equal_to_zero to adjust the bit position so that the next element is positioned on a byte boundary. ait_vui_alignment_zero_bit_a is assumed to be equal to 0.

[0081] The vui_payload() that specifically describes the YUI of the generated image to be output is the same as the vui_payload() of the VUI disclosed in Non-Patent Document 2.

[0082] In Fig. 9, we will explain the syntax elements added from Fig. 6. Text information indicating generation information (information of ait_data_string, information pointed to by ait_data_uri) may include control parameter information, model information, model parameter information, and prompt information.

[0083] The syntax element ait_data_compress_idc indicates the encoding method for the text information indicating the generation information. 0 unencoded 1 ZIP 2 or more reserved Here, "reserved" means that the value is not specified in this version, but may be used in a subsequent update.

[0084] The syntax element ait_data_type_idc indicates the format of the text information indicating the generation information. 0 Plain Text 2 CSV Comma-separated values. Defined in RFC 4180. text / csv 3 JSON JavaScript Object Notation. A text format defined in RFC8259, ECMA-404, and ISO / IEC 21778:2017. application / json 4 XML Extensible Markup Language. Markup language defined in RFC7303. text / xml 5 YAML Yet Another Markup Language. application / yaml 6 or more reserved The syntax element ait_data_platform_flag indicates whether or not the platform used by the text information indicating the generation information is specified. 0 Platform-agnostic 1 Platform specific The syntax element ait_data_platform_string is a string that indicates the platform used by the generation information, and is decoded or encoded if ait_data_platform_flag is true. It is encoded in UTF-8 as indicated by st(v). For example, the following string may be used: StableDiffusion WebUI is a platform that uses StableDiffusion as a base model and processes generated information and objects specified by text information. ComfyUI is a platform that uses StableDiffusion as a base model and processes generated information and objects specified in JSON format. According to the above configuration, the compression method of the text information indicating the generation information is specified, so that the volume of the text information can be reduced.

[0085] According to the above configuration, the format of the text information indicating the generated information is specified, so that even if the text information is structured, the parser can be reliably operated to check the content.

[0086] According to the above configuration, since a platform on which the generated information is to be operated is specified, the text information indicating the specified generated information can be reliably operated.

[0087] Figure 10 shows the generation information in JSON format. This generation information has a graph structure and indicates the node ID, node input, node output, node content (type), and attributes. Here, ID, input, output, type, and widgets_values ​​are used as JSON keys, but they are not limited to these and other keys may be included. In XML, keys are expressed using tags. For example, widgets_values, which indicates attributes, may be expressed as attr, value, attribute_value, etc. Here, LoadImage, VAELoader, CheckpointLoaderSimple, VAEEncode, CLIPTextEncod, Ksampler, VAEDecode, and SaveImage are used for the node content. Each of these shows image input, loading of a VAE neural network model (VAE model), loading of a basic neural network model (basic model), conversion of an image to features (encoding) using the VAE model, encoding of text information using CLIP (Contrastive Language-Image Pre-Training) and a neural network model, sampler, inverse conversion (decoding) from features to an image using the VAE model, and image output. Also, LoadImage, VAELoader, CheckpointLoaderSimple, VAEEncode, CLIPTextEncod, Ksampler, VAEDecode, and SaveImage each have attributes, such as the image name (input_image.png) shown in OBJ1, the VAE model name (vae.safetensors) shown in M1 and M2, and the base model's main_natural.safetensors.As shown in T1 and T2, the attributes of CLIPTextEncode contain text prompts, and the two nodes contain the strings “Two girls with long brown hairs are blowing soap bubbles, drinking straws, teddy bears, present boxes with ribbon, realistic, ponytails, brick wall” and “worst quality, low quality, normal quality, bad face, bad anatomy, missing limbs , missing fingers, extra limbs, extra fingers , text, ugly”, respectively. The former is a (positive) prompt, and the latter is a negative prompt.

[0088] FIG. 11 shows a graph of text information representing the generation information of FIG. 10. Boxes represent JSON nodes of the text information. In the generation information, the order of processing is indicated for each node, and is shown as S1301 to S1307. An image is loaded by LoadImage (S1301), a VAE model is loaded by VAELoader (S1302), and input to VAEEncode. The image is converted into VAE features by VAE using the input model (S1303). The basic model loaded by CheckpointLoaderSimple is input to CLIPTextEncode (S1303) and CLIPTextEncode (S1304), where the text prompt is converted into CLIP features. The VAE features and CLIP features are input to KSampler, where the DiffusionModel converts the VAE features into image processing data in accordance with the input text prompt (S1305). The converted VAE features are input to VAEDecode, where the VAE features are decoded into an image (S1306). Finally, the data is input to SaveImage, and the image is output (S1307).

[0089] The syntax in the syntax configuration shown in FIG. 9 will be further explained.

[0090] The ait_signal_model_flag indicates that model information is specified and transmitted via a URI as supplementary information to the text information indicating the generation information.

[0091] If ait_signal_model_flag is true, it contains the base model ait_num_base_model and the supplemental model ait_num_supplemental_model. The supplemental model may be a VAE model used as pre-processing or post-processing for the base model, or a low-dimensional parameter LoRA (Low-Rank Adaptation) model that is used in parallel with the base model and added to individualize the processing. The number of ait_num_base_models includes ait_base_model_tag_uri[i] and ait_base_model_uri[i], and the number of ait_num_supplemental_models includes ait_base_supplemental_tag_uri[i] and ait_supplemental_model_uri[i]. Each ait_base_model_uri[i] contains a URI indicating the ith base model, and each ait_supplemental_model_uri[i] contains a URI indicating the ith supplemental model. Here, a distinction is made between base and supplemental, but it is not necessary to distinguish between them. ait_num_model, ait_base_tag_uri[i], and ait_model_uri[i] may be used. When ait_signal_model_flag is true, the supplemental extension information decoding unit 3021 decodes ait_base_model_tag_uri[i] and ait_base_model_uri[i] for the number of ait_num_base_models, and decodes ait_base_supplemental_tag_uri[i] and ait_supplemental_model_uri[i] for the number of ait_num_supplemental_models. The order of the models specified in the generation information of ait_data_string or ait_data_uri corresponds to the i-th order of the models indicated by the above ait_base_supplemental_tag_uri[i] and ait_supplemental_model_uri[i]. In other words, the first, second, and third models listed in the generation model are designated by the models listed in the URIs of i=0, 1, and 2.

[0092] In the above configuration, when the name of the neural network model indicated in the attribute of the creation information is insufficient, the neural network model can be specified by the URI. In the above configuration, when the object indicated by the attribute of the generation information is insufficient, an image or video can be specified by a URI. The supplemental extension information decoding unit 3021 decodes ait_base_model_tag_uri[i] and ait_base_model_uri[i] for the number of ait_num_base_models, and decodes ait_base_supplemental_tag_uri[i] and ait_supplemental_model_uri[i] for the number of ait_num_supplemental_models. The description order of the models specified in the generation information of ait_data_string or ait_data_uri is indicated by the above ait_base_supplemental_tag_uri[i] and ait_supplemental_model_uri[i]. In other words, the description order of the models specified in the generation information of ait_data_string or ait_data_uri corresponds to the i-th order of the models indicated by the above ait_base_supplemental_tag_uri[i] and ait_supplemental_model_uri[i]. That is, the first, second, and third models listed in the generative model are designated by the models listed in the URIs i=0, 1, and 2.

[0093] In addition, this structure has the effect of minimizing the addition of bits for byte alignment by placing the URI, which is byte-level text information indicated by st(v), after the bit-level syntax elements indicated by u(v) and ue(v).

[0094] Fig. 12 shows another example of syntax configuration. When ait_signal_model_flag is true, the number of models, ait_num_model, is included. The model may be a basic model, a VAE model, a LoRA model, etc. When ait_signal_model_flag is true, the auxiliary extension information decoding unit 3021 decodes ait_model_text[i] and ait_model_uri[i] for the number of ait_num_model. ait_model_text is the model name in the text information indicating the generation information of ait_data_string or ait_data_uri. This is text (label) that indicates the name. In this text information, for example, XXX_model (e.g., vae.safemodel) and YYY_model (e.g., main_natural_image.safemodel) are shown as attributes of the node. What the XXX_model and YYY_model actually are is specified by a URI. For example, ait_model_text[0] = “vae.safemodel” ait_model_uri[0] = “https: / / 192.168.0.2 / general_model / vae.safemodel” ait_model_text[1] = “main_natural_image.safemodel” ait_model_uri[1] = “https: / / 192.168.0.2 / general_model / main_natural_image.safemodel” By doing so, it is possible to determine that the vae.safemodel and main_natural_image.safemodel listed in the generation information are the models vae.safemodel and main_natural_image.safemodel specified by the above URI. Alternatively, the files may be actually stored in the above locations and downloaded. Furthermore, tag information ait_model_tag_uri[i] indicating the type of each URI may also be included.

[0095] ait_num_object indicates the number of object information items to be input to the model as supplementary information to the text information indicating the generation information. The object information items may be control image information items.

[0096] ait_object_source indicates the source of the object information.

[0097] 0: Image input by external means 1: Images identified by layer number 2: Images identified by URI is. If ait_object_source==1, object information (such as control image information) is provided by an external means. If ait_object_source==1, the layer number (ait_nuh_layer_id) is also encoded and decoded. The decoded image of the layer specified by ait_nuh_layer_id is used as object information. If ait_object_source==2, the layer number (ait_object_uri) is further encoded and decoded. The decoded image of the layer specified by ait_object_uri is set as object information.

[0098] ait_object_text is text (label) that indicates the object name in the creation information of ait_data_string or ait_data_uri. In the text information that indicates the creation information of ait_data_string or ait_data_uri, for example, as an attribute of the node, it specifies what image.png actually is. For example, ait_object_text[0] = “image.png” ait_object_nuh_layer_id[0] = 1 In this case, the entity of image.png shown as an attribute of LoadImage in the JSON in Figure 10 is It can be specified that the image is decoded with layer ID=1.

[0099] FIG. 13 shows another example of a syntax configuration. This syntax includes ait_data_optional_flag, which indicates whether optional generation information that is not required to be used is included. When ait_data_optional_flag is true, the syntax further includes, for example, ait_data_string, which is required text information of high importance, followed by ait_data_optional_string, which is less important and may not be required. The supplemental extension information decoding unit 3021 decodes ait_data_optional_flag, and when ait_data_mode_idc == false, may perform the above operation of decoding optional generation information according to the value of the ait_data_optional_flag= flag (if true). Furthermore, when ait_data_optional_flag is true, the syntax further includes, for example, ait_data_uri, which is a required URI of high importance, followed by ait_data_optional_uri, which is a less important URI that may not be required. When ait_data_mode_idc==true, the auxiliary extension information decoder 3021 performs optional decoding according to the value of the ait_data_optional_flag= flag (if true). Alternatively, the URI of the generated information may be decoded. As in the above configuration, by decoding two or more pieces of generated information with different levels of importance according to the flag, it is possible to switch whether or not to use all generated information, including options, depending on the capabilities of the device and the user's preferences. Furthermore, while less important generated information cannot be changed along the transmission path, making it possible to change less important generated information along the transmission path allows for the application of generated information with a high degree of freedom.

[0100] As described above, in this embodiment, it has been shown that an image transmission system using a video encoding and decoding method that uses an image generation method can be realized by encoding and decoding using an SEI message that identifies generation information by a URI.

[0101] Note that part or all of the video encoding device 10 and the video decoding device 30 in the above-described embodiments may be implemented by a computer. In this case, a program for implementing the control functions may be recorded on a computer-readable recording medium, and the program may be loaded into a computer system and executed. Note that the term "computer system" as used herein refers to a computer system built into either the video encoding device 10 or the video decoding device 30, including hardware such as an OS and peripheral devices. Furthermore, the term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into a computer system. Furthermore, the term "computer-readable recording medium" may also include media that dynamically store programs for a short period of time, such as communication lines used when transmitting programs via networks such as the Internet or telephone lines, or media that store programs for a fixed period of time, such as volatile memory within a computer system that serves as a server or client in such cases. Furthermore, the program may be a program for implementing part of the above-described functions, or may be a program that can be implemented in combination with a program already stored in the computer system.

[0102] Furthermore, part or all of the video encoding device 10 and the video decoding device 30 in the above-described embodiments may be realized as an integrated circuit such as an LSI (Large Scale Integration). Each functional block of the video encoding device 10 and the video decoding device 30 may be individually implemented as a processor, or part or all of them may be integrated into a processor. Furthermore, the integrated circuit implementation method is not limited to LSI, and may be implemented using a dedicated circuit or a general-purpose processor. Furthermore, if an integrated circuit implementation technology that can replace LSI emerges due to advances in semiconductor technology, an integrated circuit based on that technology may be used.

[0103] One embodiment of the present invention has been described in detail above with reference to the drawings, but the specific configuration is not limited to that described above, and various design changes and the like are possible within the scope that does not deviate from the gist of the present invention.

[0104] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. In other words, embodiments obtained by combining technical means modified appropriately within the scope of the claims are also included in the technical scope of the present invention. [Industrial Applicability]

[0105] The embodiments of the present invention can be suitably applied to a video decoding device that decodes coded data obtained by coding an image signal, and a video coding device that generates coded data obtained by coding image data, and can also be suitably applied to the data structure of coded data generated by a video coding device and referenced by the video decoding device. [Explanation of symbols]

[0106] 1. Image transmission system 10 Video Encoding Device 101 Image encoding device 102 Generative information creation device 1021 Generation Information Creation Department 1023, 303 Image generation processing device 1022 Encoding control unit 103 Generated information encoding device 1031 Supplementary Extension Code 1032 URI data encoding part 1033 URI data save section 20 Transmission Network 30 Video decoding device 301 Image decoding device 302 Generated information decoding device 3021 Supplementary extension information decoding unit 3022 URI data decoding unit 3023 URI Data Loading Unit 303 Image generation processing device 3031 Image Generation Unit 3032 Control Unit 3033 Control image generation unit 40 Image display device

Claims

1. an image decoding device that decodes encoded data of an image signal; a generation information decoding device that decodes coded data of a URI that identifies generation information for image generation; an image generating device that generates an image from image information decoded by the image decoding device and generated information decoded by the generated image decoding device; The generated information decoding device is a video decoding device characterized in that the generated information decoding device decodes the encoding method of text information indicating the generated information from the encoded data.

2. an image decoding device that decodes encoded data of an image signal; a generation information decoding device that decodes coded data of a URI that identifies generation information for image generation; an image generating device that generates an image from image information decoded by the image decoding device and generated information decoded by the generated image decoding device; The generated information decoding device is a video decoding device characterized in that the generated information decoding device decodes a format of text information indicating the generated information from the coded data.

3. an image decoding device that decodes encoded data of an image signal; a generation information decoding device that decodes coded data of a URI that identifies generation information for image generation; an image generating device that generates an image from image information decoded by the image decoding device and generated information decoded by the generated image decoding device; The generation information decoding device is a video decoding device characterized in that the generation information decoding device decodes a platform used by the generation information from encoded data.

4. an image encoding device that encodes an image signal; a generation information creation device that creates generation information for generating an image from image information decoded by the image encoding device; a generated information encoding device that encodes a URI that identifies the generated information; The generation information encoding device is a moving image encoding device characterized in that the generation information encoding device encodes text information indicating the generation information using a method for encoding the text information.