Decoding device, encoding device, decoding method, and encoding method

The decoder and encoder system addresses geometric attribute errors in face motion images by using a concealment parameter for error recovery, enhancing encoding efficiency and image quality.

WO2025154665A1PCT designated stage expired Publication Date: 2025-07-24PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/000612
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-16
Filing Date
2025-01-10
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

Current video coding technologies face challenges in efficiently encoding and decoding face motion images, particularly in handling geometric attributes and error recovery during face reproduction, leading to distorted or incorrect frames due to inappropriate geometric information acquisition.

Method used

A decoder and encoder system that includes a circuit and memory to decode geometric information and a concealment parameter for error recovery, using a generation model to generate face motion images, and adaptively utilizing stored geometric information for error recovery control.

Benefits of technology

Improves encoding efficiency, reduces processing amount, and enhances image quality by effectively handling geometric attribute errors, ensuring natural output images even with insufficient information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025000612_24072025_PF_FP_ABST
    Figure JP2025000612_24072025_PF_FP_ABST
Patent Text Reader

Abstract

This decoding device (200) comprises a circuit (251), and a memory (252) connected to the circuit (251). In operation, the circuit (251): decodes, from a bitstream, base data of a face image pertaining to a face moving image, and geometric information which corresponds to each of a plurality of frames of the face moving image and indicates the geometric attributes in a region including the face of a person (S701); further decodes, from the bitstream, concealment parameters pertaining to an error recovery control when the geometric information is not appropriately acquired in an encoding device (S702); and uses a generation model to generate the face moving image from the base data, the geometric information, and the concealment parameters (S703).
Need to check novelty before this filing date? Find Prior Art

Description

Decoding device, encoding device, decoding method, and encoding method

[0001] The present disclosure relates to a decoding device and the like.

[0002] Video coding technology has progressed from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Codec). With this progress, there is a constant need to provide improvements and optimizations in video coding technology to handle the ever-increasing amount of digital video data in various applications. This disclosure relates to further advances, improvements, and optimizations in video coding.

[0003] Non-Patent Document 1 relates to an example of a conventional standard related to the above-mentioned video coding technology.

[0004] H. 265 (ISO / IEC 23008-2 HEVC) / HEVC (High Efficiency Video Coding)

[0005] With regard to the above-mentioned encoding methods, it is desirable to propose new methods to improve encoding efficiency, improve image quality, reduce the amount of processing, reduce the circuit scale, or appropriately select elements or operations such as filters, block sizes, motion vectors, reference pictures or reference blocks.

[0006] The present disclosure provides a configuration or method that can contribute to one or more of, for example, improved coding efficiency, improved image quality, reduced processing amount, reduced circuit size, improved processing speed, and appropriate selection of elements or operations, etc. Note that the present disclosure may include a configuration or method that can contribute to benefits other than those described above.

[0007] For example, a decoding device according to one aspect of the present disclosure includes a circuit and a memory connected to the circuit, and in operation, the circuit decodes, from a bit stream, base data of a facial image related to a facial video and geometric information, which corresponds to each of multiple frames of the facial video and indicates geometric attributes within an area including a person's face, further decodes, from the bit stream, concealment parameters related to error recovery control in the event that the geometric information is not properly acquired in the encoding device, and generates the facial video from the base data, the geometric information, and the concealment parameters using a generative model.

[0008] Each embodiment of the present disclosure, or a partial configuration or method thereof, enables at least one of, for example, improved coding efficiency, improved image quality, reduced encoding / decoding processing volume, reduced circuit size, or improved encoding / decoding processing speed. Alternatively, each embodiment of the present disclosure, or a partial configuration or method thereof, enables appropriate selection of components / operations such as filters, block sizes, motion vectors, reference pictures, and reference blocks in encoding and decoding. Note that the present disclosure also includes disclosure of configurations or methods that may provide benefits other than those described above. For example, a configuration or method that improves coding efficiency while suppressing an increase in processing volume.

[0009] Further advantages and benefits of certain aspects of the present disclosure will become apparent from the specification and drawings. While such advantages and / or benefits may be obtained by several embodiments and features described in the specification and drawings, not all of them necessarily need to be provided to obtain one or more advantages and / or benefits.

[0010] These general or specific aspects may be realized by a system, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be realized by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.

[0011] A configuration or method according to an aspect of the present disclosure may contribute to, for example, one or more of improved coding efficiency, improved image quality, reduced processing amount, reduced circuit size, improved processing speed, and appropriate selection of elements or operations, etc. Note that a configuration or method according to an aspect of the present disclosure may also contribute to benefits other than those described above.

[0012] FIG. 1 is a block diagram showing the configuration of an encoding / decoding system in a reference example. FIG. 2 is a block diagram showing the configuration of an encoding device in a reference example. FIG. 3 is a block diagram showing the configuration of a decoding device in a reference example. FIG. 4 is a conceptual diagram showing an example of a reference image. FIG. 5 is a conceptual diagram showing an example of a geometric attribute set. FIG. 6 is a conceptual diagram showing an example of a facial moving image. FIG. 7 is a block diagram showing an example of the configuration of an encoding / decoding system in an embodiment. FIG. 8 is a diagram showing an example of a hierarchical structure of data in a stream. FIG. 9 is a block diagram showing an example of the configuration of a decoding device in an embodiment. FIG. 10 is a flowchart showing an example of the operation of a decoding device in an embodiment. FIG. 11 is a conceptual diagram showing an example of the operation performed in accordance with concealment parameters. FIG. 12 is a conceptual diagram showing another example of the operation performed in accordance with concealment parameters. FIG. 13 is a conceptual diagram showing an example of the operation when a geometric attribute set is not acquired. FIG. 14 is a conceptual diagram showing an example of the operation of a decoding device when the number of attributes in the geometric attribute set is different from the original number of attributes. FIG. 15 is a conceptual diagram showing an example of the operation in which storage of the geometric attribute set is controlled. FIG. 16 is a conceptual diagram showing an example of an operation performed in a decoding device according to the reliability of a geometric attribute set. FIG. 17 is a conceptual diagram showing an example of an operation in which a geometric attribute is replaced in a decoding device. FIG. 18 is a conceptual diagram showing another example of an operation in which a geometric attribute is replaced in a decoding device. FIG. 19 is a conceptual diagram showing an example of a geometric attribute set stored in a decoding device. FIG. 20 is a conceptual diagram showing an example of a geometric attribute set decoded in a decoding device. FIG. 21 is a conceptual diagram showing an example of a geometric attribute set derived in a decoding device. FIG. 22 is a block diagram showing another example of the configuration of a decoding device in an embodiment. FIG. 23 is a block diagram showing an example of the configuration of an encoding device in an embodiment. FIG. 24 is a flowchart showing an example of the operation of an encoding device in an embodiment. FIG. 25 is a conceptual diagram showing a specific example of a geometric attribute set. FIG. 26 is a conceptual diagram showing another specific example of a geometric attribute set. FIG. 27 is a conceptual diagram showing yet another specific example of a geometric attribute set. FIG. 28 is a conceptual diagram showing an example of the operation of the encoding device when the number of attributes in the geometric attribute set is different from the original number of attributes.FIG. 29 is a conceptual diagram showing an example of a face partially hidden by occlusion. FIG. 30 is a conceptual diagram showing an example of a face having an extreme yaw posture. FIG. 31 is a conceptual diagram showing an example of a face having an extreme roll posture. FIG. 32 is a conceptual diagram showing an example of a face having blur. FIG. 33 is a conceptual diagram showing an example of a face with a low number of detected landmarks. FIG. 34 is a conceptual diagram showing an example of an operation performed in an encoding device according to the reliability of a geometric attribute set. FIG. 35 is a conceptual diagram showing an example of an operation in which geometric attributes are replaced in an encoding device. FIG. 36 is a conceptual diagram showing another example of an operation in which geometric attributes are replaced in an encoding device. FIG. 37 is a conceptual diagram showing an example of a geometric attribute set stored in an encoding device. FIG. 38 is a conceptual diagram showing an example of a geometric attribute set detected in an encoding device. FIG. 39 is a conceptual diagram showing an example of a geometric attribute set derived in an encoding device. FIG. 40 is a syntax diagram showing an example of a syntax structure related to concealment parameters. FIG. 41 is a syntax diagram showing another example of a syntax structure related to concealment parameters.

[0093] Fig. 42 is a syntax diagram showing yet another example syntax structure for concealment parameters. Fig. 43 is a syntax diagram showing yet another example syntax structure for concealment parameters. Fig. 44 is a syntax diagram showing yet another example syntax structure for concealment parameters. Fig. 45 is a syntax diagram showing yet another example syntax structure for concealment parameters. Fig. 46 is a syntax diagram showing yet another example syntax structure for concealment parameters. Fig. 47 is a syntax diagram showing yet another example syntax structure for concealment parameters. Fig. 48 is a syntax diagram showing yet another example syntax structure for concealment parameters. Fig. 49 is a syntax diagram showing yet another example syntax structure for concealment parameters. Fig. 50 is a diagram showing examples of various neural networks that can be used as generative models. Fig. 51 is a block diagram showing an example configuration for an encoding device according to an embodiment to encode moving images. Fig. 52 is a block diagram showing an example configuration for a decoding device according to an embodiment to decode moving images.FIG. 53 is a block diagram showing an example of implementation of an encoding device in an embodiment. FIG. 54 is a flowchart showing an example of basic operation of an encoding device in an embodiment. FIG. 55 is a flowchart showing another example of basic operation of an encoding device in an embodiment. FIG. 56 is a block diagram showing an example of implementation of a decoding device in an embodiment. FIG. 57 is a flowchart showing an example of basic operation of a decoding device in an embodiment. FIG. 58 is a flowchart showing another example of basic operation of a decoding device in an embodiment. FIG. 59 is a diagram showing the overall configuration of a content supply system that realizes a content distribution service. FIG. 60 is a diagram showing an example of a display screen of a web page. FIG. 61 is a diagram showing an example of a display screen of a web page. FIG. 62 is a diagram showing an example of a smartphone. FIG. 63 is a block diagram showing an example of the configuration of a smartphone.

[0013] Introduction Facial reconstruction refers to the process of mapping the pose and facial expression of one or more source persons to one or more target persons while ensuring that the identities of the target persons are preserved. Currently, facial reconstruction techniques are used in a variety of applications, including video conferencing, the entertainment industry, and social media. This description can be used in any multimedia data coding context for facial reconstruction techniques that enhance the photo-like aspects of the resulting output video.

[0014] Fig. 1 is a block diagram showing the configuration of an encoding / decoding system according to a reference example. The encoding / decoding system shown in Fig. 1 includes an encoding device 700 and a decoding device 800. The model architecture of the current research can be represented as a framework of the encoding device 700 and the decoding device 800, as shown in Fig. 1.

[0015] The model architecture may receive input in the form of a reference image (source image) of a target person and one or more frames of driving video, which are encoded and compressed into one or more bitstreams by the encoding device 700. The compressed bitstreams are then transmitted to the decoding device 800 using a transmission channel. Finally, the decoding device 800 reconstructs the output video from the received bitstreams.

[0016] 2 is a block diagram showing the configuration of an encoding device 700 according to a reference example. The encoding device 700 includes compressors 731 and 733 and a deriver 732.

[0017] The encoding device 700 first encodes one or more reference images into a bitstream via a compressor 731. Here, each reference image may be one or more first frames of a driving video or at least one image or avatar containing the face of a target person. The reference image may also be referred to as a face image.

[0018] Subsequent frames of the driving video are sent to a deriver 732 to derive a set of geometric attributes for the frame, which are then encoded into a bitstream via a compressor 733. Each driving frame may include a person's face. Compressor 733 may or may not be the same as compressor 731.

[0019] The bitstream is transmitted via a transmission channel to the decoding device 800. The bitstream may be a single bitstream or may consist of multiple sub-bitstreams.

[0020] 3 is a block diagram showing the configuration of a decoding device 800 according to a reference example. The decoding device 800 includes decompressors 831 and 833, a deriver 832, and a generator 834.

[0021] The decoding device 800 first decodes one or more reference images from the bitstream. Each reference image may include a human face. The reference images are passed to a deriver 832, which extracts related information of the reference images as a reference attribute set. The reference attribute set may include a geometric attribute set of the reference images. The data of the reference images and the reference attribute set is data commonly used for multiple pictures in the video to be generated, and is also referred to as base data.

[0022] For example, generator 834 may be provided with the reference image obtained by decompressor 831 as well as the reference attribute set obtained by deriver 832. Alternatively, the configuration and processing of deriver 832 may be omitted, and generator 834 may be provided with the reference image without being provided with the reference attribute set.

[0023] The geometric attribute set for each frame is then decoded from the bitstream using the decompressor 833. Each geometric attribute set contains the geometric attributes of a person in a driving frame.

[0024] The generator 834 then takes as input one or more reference images obtained by the decompressor 831 and the reference attribute set obtained by the deriving unit 832, along with the geometric attribute set obtained by the decompressor 833, to generate a facial animation. The generator 834 may use only one of the reference images obtained by the decompressor 831 and the reference attribute set obtained by the deriving unit 832.

[0025] The generator 834 includes a neural network, which may be a generative network such as a Generative Adversarial Network (GAN). Typically, the face reconstruction process is performed through the GAN using a set of geometric attributes and a reference image. Each frame of the output may contain a generated face or no face at all.

[0026] 4 is a conceptual diagram showing an example of a reference image. As in this example, the reference image is an image that includes a face.

[0027] 5 is a conceptual diagram showing an example of a geometric attribute set. In this example, the geometric attributes included in the geometric attribute set are facial landmarks. For example, a geometric attribute set is derived for each frame of a captured video. Note that the geometric attribute set may be simply expressed as geometric attributes, or may be expressed as geometric information, facial geometric attributes, facial attributes, facial attribute parameters, or the like.

[0028] 6 is a conceptual diagram showing an example of a facial motion image. As shown in this example, the facial motion image is a motion image including a face. In the facial motion image, for each of a plurality of frames, the geometric attributes of the frame are reflected in the reference image. This provides movement to the reference image.

[0029] The amount of information for the multiple geometric attributes corresponding to the reference image and multiple frames is smaller than the amount of information for the multiple frames included in the captured video. Therefore, by encoding the multiple geometric attributes corresponding to the reference image and multiple frames, the amount of code is reduced compared to encoding the multiple frames included in the captured video. Furthermore, each geometric attribute imparts movement to the reference image. This imparts movement to the face of the display target, enabling richer expression.

[0030] In current research, a geometric attribute set is required when generating each frame of a facial video. However, the geometric attribute set is not always appropriate. For example, it is not easy for the decoding device 800 to evaluate whether the geometric information set decoded from the bitstream is appropriate. If the geometric attribute set is inappropriate, it may result in the generation of distorted or incorrect frames.

[0031] Therefore, the decoding device of Example 1 includes a circuit and a memory connected to the circuit, and in operation, the circuit decodes, from a bit stream, base data of facial images related to a facial video and geometric information, which corresponds to each of multiple frames of the facial video and is information indicating geometric attributes within an area including a person's face, further decodes, from the bit stream, concealment parameters related to error recovery control in case the geometric information is not properly acquired in the encoding device, and generates the facial video from the base data, the geometric information, and the concealment parameters using a generative model.

[0032] This may enable the encoding device to notify the decoding device of information regarding error recovery control when the encoding device has not properly acquired geometric information, which may enable the decoding device to properly perform error recovery control when the encoding device has not properly acquired geometric information.

[0033] The decoding device of Example 2 may be the decoding device of Example 1, in which the circuitry decodes the concealment parameters from a header region in the bitstream.

[0034] This may enable the encoding device to notify the decoding device of information regarding error recovery control when the encoding device is unable to properly acquire geometric information via the header area of ​​the bitstream, thereby enabling the decoding device to properly perform error recovery control based on the information in the header area.

[0035] Furthermore, the decoding device of Example 3 may be the decoding device of Example 1 or 2, in which the concealment parameter indicates, depending on the value of the concealment parameter, that the geometric information corresponding to the current frame, obtained by the encoding device, has low reliability.

[0036] This may allow the encoding device to notify the decoding device that the geometric information acquired by the encoding device is low in reliability, which may enable the decoding device to appropriately perform error recovery control.

[0037] Furthermore, the decoding device of Example 4 may be any of the decoding devices of Examples 1 to 3, in which the concealment parameter indicates, depending on the value of the concealment parameter, that the geometric information corresponding to the current frame has not been obtained in the encoding device.

[0038] This may allow the encoding device to notify the decoding device that the geometry information was not acquired in the encoding device, which may allow the decoding device to appropriately perform error recovery control.

[0039] Furthermore, the decoding device of Example 5 may be any of the decoding devices of Examples 1 to 4, in which the concealment parameter indicates, depending on the value of the concealment parameter, that the number of attributes in the geometric information corresponding to the current frame, obtained by the encoding device, is different from the original number of attributes.

[0040] This may allow the encoding device to notify the decoding device that the number of attributes in the geometric information acquired by the encoding device is inappropriate, which may enable the decoding device to perform appropriate error recovery control.

[0041] Furthermore, the decoding device of Example 6 may be any of the decoding devices of Examples 1 to 5, in which the concealment parameter indicates, depending on the value of the concealment parameter, that stored geometric information is to be applied to the current frame of the facial video image instead of the geometric information decoded from the bitstream.

[0042] This may allow the encoding device to notify the decoding device that the stored geometric information should be applied to the current frame, thereby enabling the decoding device to perform appropriate error recovery control.

[0043] Furthermore, the decoding device of Example 7 may be the decoding device of any one of Examples 1 to 6, in which the concealment parameter indicates, depending on the value of the concealment parameter, that the geometric information corresponding to the current frame is to be stored for a subsequent frame.

[0044] This may allow the encoding device to inform the decoding device that geometric information corresponding to the current frame is to be stored, thereby enabling the decoding device to appropriately perform error recovery control for subsequent frames.

[0045] Furthermore, the decoding device of Example 8 may be the decoding device of any of Examples 1 to 7, wherein when the concealment parameters indicate that the geometric information corresponding to the current frame obtained by the encoding device has low reliability or that the number of attributes in the geometric information is different from the original number of attributes, the circuit corrects the geometric information corresponding to the current frame decoded from the bitstream using stored geometric information, and applies the corrected geometric information to the geometric information for generating the facial image corresponding to the current frame in the facial video.

[0046] As a result, if the encoding device does not properly acquire the geometric information, it may be possible to correct the improperly acquired geometric information using the stored geometric information, thereby enabling appropriate error recovery control.

[0047] Furthermore, the decoding device of Example 9 may be the decoding device of any of Examples 1 to 7, wherein when the concealment parameters indicate that the geometric information corresponding to the current frame obtained by the encoding device has low reliability, that the geometric information was not obtained, that the number of attributes in the geometric information is different from the original number of attributes, or that stored geometric information should be applied to the geometric information corresponding to the current frame, the circuit applies the stored geometric information to the geometric information for generating the facial image corresponding to the current frame in the facial video.

[0048] This may allow the encoding device to use stored geometric information in place of the geometric information that was not properly acquired, thereby enabling appropriate error recovery control.

[0049] Furthermore, the decoding device of Example 10 may be any of the decoding devices of Examples 1 to 7, wherein when the concealment parameters indicate that the geometric information corresponding to the current frame obtained by the encoding device has low reliability, that the geometric information was not obtained, or that the number of attributes in the geometric information is different from the original number of attributes, the circuit does not use the generative model to generate the facial image corresponding to the current frame in the facial video, but instead applies the facial image that has already been generated to the facial image corresponding to the current frame.

[0050] This may allow the encoding device to use an already generated facial image for the current frame instead of generating a new one if the encoding device does not properly obtain geometric information, thereby enabling appropriate error recovery control.

[0051] Furthermore, the decoding device of Example 11 may be any of the decoding devices of Examples 6, 8, and 9, in which the stored geometric information is geometric information corresponding to a past frame decoded from the bitstream.

[0052] This may make it possible to appropriately perform error recovery control using geometric information corresponding to past frames.

[0053] Furthermore, the decoding device of Example 12 may be the decoding device of any one of Examples 6, 8, and 9, in which the stored geometric information is geometric information of a predefined reference.

[0054] This may enable appropriate error recovery control using predefined reference geometric information.

[0055] Moreover, the decoding device of Example 13 includes a circuit and a memory connected to the circuit, and in operation, the circuit obtains base data of an image included in a video, decodes facial attribute parameters indicating a face included in the image from a bitstream, and inputs the base data and the facial attribute parameters into a generative model to generate an output image corresponding to the image, the bitstream includes a reliability parameter related to the reliability of the facial attribute parameter, and the reliability parameter indicates that the reliability is low depending on the value of the reliability parameter.

[0056] This may allow the encoding device to notify the decoding device that the reliability of the face attribute parameter is low, when the reliability of the face attribute parameter is low, and thus may allow the decoding device to appropriately determine that the reliability of the face attribute parameter is low.

[0057] Furthermore, the decoding device of Example 14 may be the decoding device of Example 13, in which the images correspond to each of a plurality of pictures included in the moving image, the base data is data common to the plurality of pictures, and the bitstream includes the reliability parameter and the face attribute parameter for each of the plurality of pictures.

[0058] This may enable the encoding device to notify the decoding device that the reliability of the face attribute parameter is low for each picture, so that the decoding device may be able to appropriately determine that the reliability of the face attribute parameter is low for each picture.

[0059] Furthermore, the decoding device of Example 15 may be the decoding device of Example 14, in which the bitstream includes the reliability parameter before the face attribute parameter for each of the plurality of pictures.

[0060] This may enable the encoding device to notify the decoding device of the reliability of the face attribute parameter before notifying the face attribute parameter, which may enable the decoding device to process the face attribute parameter after appropriately determining that the reliability of the face attribute parameter is low.

[0061] Furthermore, the decoding device of Example 16 may be any of the decoding devices of Examples 13 to 15, in which the facial attribute parameter is stored if the reliability parameter does not indicate that the reliability of the facial attribute parameter is low.

[0062] This may make it possible to suppress the storage of face attribute parameters with low reliability, and therefore to suppress the reuse of face attribute parameters with low reliability.

[0063] Moreover, the encoding device of Example 17 includes a circuit and a memory connected to the circuit, and in operation, the circuit encodes, into a bit stream, base data of facial images related to a facial video sequence and geometric information, which corresponds to each of a plurality of frames of the facial video sequence and indicates geometric attributes within an area including a person's face, determines whether the geometric information has been properly acquired, and further encodes, into the bit stream, concealment parameters related to error recovery control in a decoding device in the event that the geometric information has not been properly acquired.

[0064] This may enable the encoding device to notify the decoding device of information regarding error recovery control when the encoding device has not properly acquired geometric information, which may enable the decoding device to properly perform error recovery control when the encoding device has not properly acquired geometric information.

[0065] The encoding device of Example 18 may be the encoding device of Example 17, in which the circuit encodes the concealment parameters into a header region in the bitstream.

[0066] This may enable the encoding device to notify the decoding device of information regarding error recovery control when the encoding device is unable to properly acquire geometric information via the header area of ​​the bitstream, thereby enabling appropriate error recovery control based on the information in the header area.

[0067] Furthermore, the encoding device of Example 19 may be the encoding device of Example 17 or 18, in which the concealment parameter indicates, depending on the value of the concealment parameter, that the geometric information corresponding to the current frame obtained by the encoding device has low reliability.

[0068] This may allow the encoding device to notify the decoding device that the geometric information acquired by the encoding device is low in reliability, which may enable the decoding device to appropriately perform error recovery control.

[0069] Furthermore, the encoding device of Example 20 may be any of the encoding devices of Examples 17 to 19, in which the concealment parameter indicates, depending on the value of the concealment parameter, that the geometric information corresponding to the current frame has not been obtained in the encoding device.

[0070] This may allow the encoding device to notify the decoding device that the geometry information was not acquired in the encoding device, which may allow the decoding device to appropriately perform error recovery control.

[0071] Furthermore, the encoding device of Example 21 may be any of the encoding devices of Examples 17 to 20, in which the concealment parameter indicates, depending on the value of the concealment parameter, that the number of pieces of geometric information corresponding to the current frame acquired in the encoding device is different from the original number.

[0072] This may allow the encoding device to notify the decoding device that the number of attributes in the geometric information acquired by the encoding device is inappropriate, which may enable the decoding device to perform appropriate error recovery control.

[0073] Furthermore, the encoding device of Example 22 may be any of the encoding devices of Examples 17 to 21, in which the concealment parameter indicates, depending on the value of the concealment parameter, that stored geometric information is to be applied to the current frame of the facial video image instead of the geometric information decoded from the bitstream.

[0074] This may allow the encoding device to notify the decoding device that the stored geometric information should be applied to the current frame, thereby enabling the decoding device to perform appropriate error recovery control.

[0075] Furthermore, the encoding device of Example 23 may be the encoding device of any one of Examples 17 to 22, in which the concealment parameter indicates, depending on the value of the concealment parameter, that the geometric information corresponding to the current frame is to be stored for a subsequent frame.

[0076] This may allow the encoding device to inform the decoding device to store the geometric information corresponding to the current frame, and thus may allow for appropriate error recovery control for subsequent frames.

[0077] Furthermore, the encoding device of Example 24 may be any of the encoding devices of Examples 17 to 23, wherein the circuit sets the concealment parameter to a value indicating that the reliability of the geometric information corresponding to the current frame obtained by the encoding device is low when the reliability of the geometric information is lower than a threshold.

[0078] This may allow the encoding device to notify the decoding device that the reliability of the geometric information is low when the reliability of the geometric information is lower than a threshold, thereby enabling the decoding device to appropriately perform error recovery control.

[0079] Furthermore, the encoding device of Example 25 may be the encoding device of any one of Examples 17 to 24, wherein the circuit does not encode the geometric information corresponding to the current frame into the bitstream and sets the concealment parameter to a value indicating that the geometric information was not acquired, when the reliability of the geometric information corresponding to the current frame acquired by the encoding device is lower than a threshold value or when the current frame acquired by the encoding device does not include a face.

[0080] This may allow the encoding device to notify the decoding device that the geometric information was not acquired if the reliability of the geometric information is lower than a threshold or if a face is not included, which may allow the decoding device to appropriately perform error recovery control.

[0081] Furthermore, the encoding device of Example 26 may be the encoding device of any one of Examples 17 to 25, wherein the circuit sets the concealment parameter to a value indicating that the number of attributes in the geometric information corresponding to the current frame obtained by the encoding device is different from a predefined number of attributes.

[0082] This may allow the encoding device to notify the decoding device that the number of attributes in the geometric information is different from the original number when the number of attributes in the geometric information is different from the predefined number of attributes, thereby enabling the decoding device to appropriately perform error recovery control.

[0083] Furthermore, the encoding device of Example 27 may be the encoding device of any one of Examples 17 to 26, wherein the circuit sets the concealment parameter to a value indicating that stored geometric information is to be applied to the geometric information corresponding to the current frame when the reliability of the geometric information acquired by the encoding device and corresponding to the current frame is lower than a threshold or when the current frame acquired by the encoding device does not include a face.

[0084] This may allow the encoding device to notify the decoding device that the stored geometric information should be applied to the current frame if the reliability of the geometric information is lower than a threshold or if no face is included, thereby enabling the decoding device to appropriately perform error recovery control.

[0085] Furthermore, the encoding device of Example 28 may be the encoding device of any one of Examples 17 to 27, wherein the circuit sets the concealment parameter to a value indicating that the geometric information corresponding to a current frame, acquired by the encoding device, is stored for a subsequent frame when the reliability of the geometric information is equal to or greater than a specific threshold.

[0086] This may allow the encoding device to notify the decoding device that the geometric information will be stored if the reliability of the geometric information is equal to or greater than a threshold, which may allow the decoding device to appropriately perform error recovery control for subsequent frames.

[0087] Furthermore, the encoding device of Example 29 may be the encoding device of any one of Examples 17 to 28, wherein, when the reliability of the geometric information corresponding to the current frame acquired by the encoding device is lower than a threshold, the circuit corrects the acquired geometric information using stored geometric information and encodes the corrected geometric information into the bitstream.

[0088] As a result, when the encoding device does not properly acquire geometric information, it may be possible to correct the improperly acquired geometric information using the stored geometric information, which may enable the encoding device to properly perform error recovery control.

[0089] Furthermore, the encoding device of Example 30 may be the encoding device of any one of Examples 17 to 28, wherein when the reliability of the geometric information corresponding to the current frame obtained by the encoding device is lower than a threshold, the circuit encodes stored geometric information into the bitstream as the geometric information corresponding to the current frame.

[0090] This may allow the encoding device to use stored geometric information in place of the geometric information that was not properly acquired, thereby enabling the encoding device to perform appropriate error recovery control.

[0091] Furthermore, the encoding device of Example 31 may be any of the encoding devices of Examples 22, 27, 29, and 30, in which the stored geometric information is geometric information corresponding to a past frame acquired in the encoding device.

[0092] This may make it possible to appropriately perform error recovery control using geometric information corresponding to past frames.

[0093] Furthermore, the encoding device of Example 32 may be any of the encoding devices of Examples 22, 27, 29, and 30, in which the stored geometric information is geometric information of a predefined reference.

[0094] This may enable appropriate error recovery control using predefined reference geometric information.

[0095] Moreover, the encoding device of Example 33 includes a circuit and a memory connected to the circuit, and the circuit, in operation, encodes base data of an image included in a video, encodes facial attribute parameters indicating a face included in the image into a bitstream, and further encodes a confidence parameter relating to the confidence of the facial attribute parameter into the bitstream, the confidence parameter indicating that the confidence is low depending on the value of the confidence parameter.

[0096] For example, if the reliability of the face attribute parameters is unknown in the decoding device, the decoding device may infer the reliability and perform processing, which may result in the generation of an inappropriate image. With the above configuration, if the reliability of the face attribute parameters is low, it may be possible for the encoding device to notify the decoding device that the reliability of the face attribute parameters is low.

[0097] Therefore, it may be possible for the decoding device to appropriately determine whether the reliability of the facial attribute parameter is low, and further, it may be possible for the encoding device to control the decoding device's determination of whether the reliability of the facial attribute parameter is low.

[0098] Furthermore, the encoding device of Example 34 may be the encoding device of Example 33, in which the images correspond to each of a plurality of pictures included in the moving image, the base data is data common to the plurality of pictures, and the bitstream includes the reliability parameter and the face attribute parameter for each of the plurality of pictures.

[0099] This may enable the encoding device to notify the decoding device that the reliability of the face attribute parameter is low for each picture, so that the decoding device may be able to appropriately determine that the reliability of the face attribute parameter is low for each picture.

[0100] Also, the encoding device of Example 35 may be the encoding device of Example 34, wherein the bitstream includes the reliability parameter before the facial attribute parameter for each of the plurality of pictures.

[0101] This may enable the encoding device to notify the decoding device of the reliability of the face attribute parameter before notifying the face attribute parameter, which may enable the decoding device to process the face attribute parameter after appropriately determining that the reliability of the face attribute parameter is low.

[0102] Furthermore, the encoding device of Example 36 may be any of the encoding devices of Examples 33 to 35, in which the facial attribute parameter is stored if the reliability parameter does not indicate that the reliability of the facial attribute parameter is low.

[0103] This may make it possible to suppress the storage of face attribute parameters with low reliability, and therefore to suppress the reuse of face attribute parameters with low reliability.

[0104] Furthermore, the decoding method of Example 37 is a decoding method that decodes, from a bit stream, base data of a facial image related to a facial moving image and geometric information, which corresponds to each of multiple frames of the facial moving image and is information indicating geometric attributes within an area including a person's face, further decodes, from the bit stream, concealment parameters related to error recovery control in the event that the geometric information is not properly acquired in the encoding device, and generates the facial moving image from the base data, the geometric information, and the concealment parameters using a generative model.

[0105] This may enable the encoding device to notify the decoding device of information regarding error recovery control when the encoding device has not properly acquired geometric information, which may enable the decoding device to properly perform error recovery control when the encoding device has not properly acquired geometric information.

[0106] Furthermore, the decoding method of Example 38 is a decoding method that acquires base data of an image included in a moving image, decodes facial attribute parameters indicating a face included in the image from a bitstream, and generates an output image corresponding to the image by inputting the base data and the facial attribute parameters into a generative model, the bitstream including a reliability parameter related to the reliability of the facial attribute parameter, and the reliability parameter indicates that the reliability is low depending on the value of the reliability parameter.

[0107] This may allow the encoding device to notify the decoding device that the reliability of the face attribute parameter is low, when the reliability of the face attribute parameter is low, and thus may allow the decoding device to appropriately determine that the reliability of the face attribute parameter is low.

[0108] In addition, the encoding method of Example 39 is an encoding method that encodes base data of a facial image related to a facial moving image and geometric information, which is information corresponding to each of multiple frames of the facial moving image and indicates geometric attributes within an area including a person's face, into a bit stream, determines whether the geometric information has been properly acquired, and further encodes into the bit stream concealment parameters related to error recovery control in a decoding device in the event that the geometric information has not been properly acquired.

[0109] This may allow the encoding device to notify the decoding device that the reliability of the face attribute parameter is low, when the reliability of the face attribute parameter is low, and thus may allow the decoding device to appropriately determine that the reliability of the face attribute parameter is low.

[0110] Moreover, the encoding method of Example 40 is an encoding method that encodes base data of an image included in a moving image, encodes face attribute parameters indicating a face included in the image into a bit stream, and further encodes reliability parameters relating to the reliability of the face attribute parameters into the bit stream, and the reliability parameters indicate that the reliability is low depending on the value of the reliability parameter.

[0111] For example, if the reliability of the face attribute parameters is unknown in the decoding device, the decoding device may infer the reliability and perform processing, which may result in the generation of an inappropriate image. With the above configuration, if the reliability of the face attribute parameters is low, it may be possible for the encoding device to notify the decoding device that the reliability of the face attribute parameters is low.

[0112] Therefore, it may be possible for the decoding device to appropriately determine whether the reliability of the facial attribute parameter is low, and further, it may be possible for the encoding device to control the decoding device's determination of whether the reliability of the facial attribute parameter is low.

[0113] Furthermore, these comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0114] In an example of the present disclosure, errors in realistic face generation are concealed by introducing one or more flags that control the storage of relevant information in the decoder. An example of relevant information may be previously decoded geometric attribute information that can be used to reconstruct a current or future frame. The storage of information may provide the decoder with the flexibility to use previously decoded information in generating the current frame, thereby reducing errors in generating frames when the information is insufficient.

[0115] The decoding device does not need to acquire all the information before starting the reconstruction process of the current frame, and may adaptively utilize past information.

[0116] For example, the overall flexibility of the decoder is improved by adaptively utilizing previously decoded information in reproducing the current frame. As a result, errors in frames generated when insufficient information is received are concealed, and the output video appears more natural. Furthermore, the operation of the GAN can be guaranteed at each step of the facial reconstruction process.

[0117] [Definition of Terms] As an example, each term may be defined as follows.

[0118] (1) Image: A unit of data made up of a set of pixels, consisting of pictures or blocks smaller than pictures, and includes both moving images and still images.

[0119] (2) Picture: A processing unit of an image composed of a set of pixels, and is sometimes called a frame or field.

[0120] (3) Block: A processing unit for a set containing a specific number of pixels, and can be named anything, as shown in the following examples. It can also be shaped anything, including, for example, a rectangle made up of M×N pixels, a square made up of M×M pixels, a triangle, a circle, or any other shape.

[0121] (Examples of blocks) Slice / tile / brick CTU / superblock / basic division unit VPDU / hardware processing division unit CU / processing block unit / prediction block unit (PU) / orthogonal transform block unit (TU) / unit Sub-block

[0122] (4) Pixel / Sample A pixel / sample is a minimum unit point that constitutes an image, and includes not only pixels at integer positions but also pixels at decimal positions generated based on pixels at integer positions.

[0123] (5) Pixel Value / Sample Value: A value inherent to a pixel, including not only brightness value, color difference value, and RGB gradation, but also depth value or binary values ​​of 0 and 1.

[0124] (6) Flags: In addition to one bit, flags may be multi-bit, for example, parameters or indexes of two or more bits. In addition, flags may be multi-valued using other bases as well as two values ​​using binary numbers.

[0125] (7) Signal: A signal that is symbolized or coded to transmit information, including discrete digital signals as well as analog signals that take continuous values.

[0126] (8) Stream / Bitstream: A digital data string or flow. A stream / bitstream may consist of a single stream or multiple streams divided into multiple layers. It also includes transmission by serial communication over a single transmission path as well as transmission by packet communication over multiple transmission paths.

[0127] (9) Difference / Difference In the case of scalar quantities, in addition to simple difference (x-y), it is sufficient to include difference calculations, including absolute value of difference (|x-y|), squared difference (x^2-y^2), square root of difference (√(x-y)), weighted difference (ax-by: a, b are constants), and offset difference (x-y+a: a is an offset).

[0128] (10) Sum In the case of a scalar quantity, in addition to simple sum (x + y), it is sufficient if a sum operation is included, including the absolute value of the sum (|x + y|), sum of squares (x^2 + y^2), square root of the sum (√(x + y)), weighted sum (ax + by: a, b are constants), and offset sum (x + y + a: a is an offset).

[0129] (11) Based on: This includes cases where factors other than the one being based on are taken into consideration. It also includes cases where a result is obtained directly or via an intermediate result.

[0130] (12) Using (used, using) This includes cases where elements other than the target of use are taken into account. It also includes cases where a result is obtained directly or via an intermediate result.

[0131] (13) Prohibit (forbid) This can be rephrased as not being allowed. Also, not prohibiting or being allowed does not necessarily mean obligation.

[0132] (14) Limit (restriction / restrict / restricted) This can be rephrased as not being permitted. Also, not prohibiting something or being permitted does not necessarily mean that it is an obligation. Furthermore, it is sufficient if something is partially prohibited in terms of quantity or quality, and it also includes cases where something is completely prohibited.

[0133] (15) Chroma: An adjective, denoted by the symbols Cb and Cr, that specifies that a sample array or a single sample represents one of two color difference signals associated with a primary color. Instead of the term chroma, the term chrominance can also be used.

[0134] (16) Luma: An adjective, denoted by the symbol or subscript Y or L, that specifies that a sample array or a single sample represents a monochrome signal associated with a primary color. Instead of the term luma, the term luminance may also be used.

[0135] [Description] In the drawings, the same reference numerals refer to the same or similar elements, and the sizes and relative positions of the elements in the drawings are not necessarily drawn to scale.

[0136] Hereinafter, embodiments will be described in detail with reference to the drawings. Note that the embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, the arrangement and connection of the components, steps, and the relationship and order of the steps shown in the following embodiments are merely examples and are not intended to limit the scope of the claims.

[0137] Below, embodiments of an encoding device and a decoding device will be described. The embodiments are examples of encoding devices and decoding devices to which the processes and / or configurations described in each aspect of the present disclosure can be applied. The processes and / or configurations can also be implemented in encoding devices and decoding devices different from the embodiments. For example, with regard to the processes and / or configurations applied to the embodiments, any of the following may be implemented.

[0138] (1) Any of the multiple components of the encoding device or decoding device of the embodiments described in each aspect of the present disclosure may be replaced or combined with other components described in any of the aspects of the present disclosure.

[0139] (2) In the encoding device or decoding device according to the embodiment, the functions or processes performed by some of the components of the encoding device or decoding device may be changed in any way, such as by adding, replacing, or deleting a function or process. For example, any function or process may be replaced with or combined with another function or process described in any of the aspects of the present disclosure.

[0140] (3) In the method implemented by the encoding device or decoding device according to the embodiment, some of the processes included in the method may be arbitrarily modified, such as by addition, replacement, deletion, etc. For example, any process in the method may be replaced with or combined with another process described in any of the aspects of the present disclosure.

[0141] (4) Some of the components constituting the encoding device or decoding device of the embodiment may be combined with components described in any of the aspects of the present disclosure, or may be combined with components having some of the functions described in any of the aspects of the present disclosure, or may be combined with components that perform some of the processing performed by the components described in each aspect of the present disclosure.

[0142] (5) A component having part of the functionality of the encoding device or decoding device of an embodiment, or a component that performs part of the processing of the encoding device or decoding device of an embodiment, may be combined or replaced with a component described in any of the aspects of the present disclosure, a component having part of the functionality described in any of the aspects of the present disclosure, or a component that performs part of the processing described in any of the aspects of the present disclosure.

[0143] (6) In the method implemented by the encoding device or decoding device of the embodiment, any of the multiple processes included in the method may be replaced or combined with the process described in any of the aspects of the present disclosure or any similar process.

[0144] (7) Some of the processes included in the method implemented by the encoding device or decoding device of the embodiment may be combined with the processes described in any of the aspects of the present disclosure.

[0145] (8) The implementation of the processes and / or configurations described in each aspect of the present disclosure is not limited to the encoding device or decoding device of the embodiments. For example, the processes and / or configurations may be implemented in a device used for a purpose other than video encoding or video decoding disclosed in the embodiments.

[0146] [Configuration of Encoding / Decoding System] Fig. 7 is a block diagram showing an example configuration of an encoding / decoding system according to this embodiment. For example, the encoding / decoding system includes an encoding device 100 and a decoding device 200. The example in Fig. 7 is similar to the example in Fig. 1, but the specific configuration and processing of the encoding device 100, the specific configuration and processing of the decoding device 200, and the bitstream are different from those in the example in Fig. 1.

[0147] For example, the reference image is an image including a face and may also be referred to as a face image, a source image, or an identity image. The reference image represents static visual features for reconstructing a face video. The driving video is a video including a face and is captured by a camera. The driving video serves to impart motion to the reference image. The bit stream may also be simply referred to as a stream. Furthermore, the use of one bit stream is not limited, and multiple bit streams may be used.

[0148] The person included in the reference image and the person included in the driving video may or may not be the same person.

[0149] The encoding / decoding system according to the present embodiment can be applied to video conferencing, video generation and editing in the entertainment industry, social media, the e-commerce industry, etc. However, the scope of application is not limited to these.

[0150] [Data Structure] Figure 8 is a diagram showing an example of a hierarchical structure of data in a stream. The stream includes, for example, a video sequence. This video sequence includes, for example, a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), supplemental enhancement information (SEI), and multiple pictures, as shown in (a) of Figure 8.

[0151] In a video composed of multiple layers, the VPS includes coding parameters common to multiple layers, and coding parameters related to multiple layers included in the video or to each individual layer.

[0152] The SPS includes parameters used for the sequence, i.e., encoding parameters that the decoding device 200 refers to in order to decode the sequence. For example, the encoding parameters may indicate the width or height of a picture. Note that there may be multiple SPSs.

[0153] The PPS includes parameters used for a picture, i.e., encoding parameters referenced by the decoding device 200 to decode each picture in a sequence. For example, the encoding parameters may include a reference value of the quantization width used in decoding the picture and a flag indicating the application of weighted prediction. Note that there may be multiple PPSs. Furthermore, the SPS and PPS may be simply referred to as parameter sets.

[0154] A picture may include a picture header and one or more slices, as shown in (b) of Fig. 8. The picture header includes coding parameters that the decoding device 200 references to decode the one or more slices.

[0155] As shown in (c) of Fig. 8, a slice includes a slice header and one or more bricks. The slice header includes coding parameters that are referenced by the decoding device 200 to decode the one or more bricks.

[0156] As shown in FIG. 8(d), a brick includes one or more coding tree units (CTUs).

[0157] Note that a picture may not contain slices, but may instead contain tile groups, where a tile group contains one or more tiles, and a brick may contain slices.

[0158] A CTU is also called a superblock or a basic division unit. Such a CTU includes a CTU header and one or more CUs (Coding Units), as shown in (e) of Fig. 8. The CTU header includes coding parameters that the decoding device 200 references to decode the one or more CUs.

[0159] A CU may be divided into multiple small CUs. Furthermore, as shown in (f) of FIG. 8, a CU includes a CU header, prediction information, and residual coefficient information. The prediction information is information for predicting the CU, and the residual coefficient information is information indicating a prediction residual, which will be described later. A CU is basically the same as a PU (Prediction Unit) and a TU (Transform Unit), but may include multiple TUs smaller than the CU, for example, in an SBT, which will be described later. A CU may also be processed for each VPDU (Virtual Pipeline Decoding Unit) that constitutes the CU. A VPDU is a fixed unit that can be processed in one stage, for example, when performing pipeline processing in hardware.

[0160] Note that the stream may not have some of the layers shown in FIG. 8 . The order of these layers may be changed, or some layers may be replaced with other layers. A picture currently being processed by a device such as the encoding device 100 or the decoding device 200 is referred to as a current picture. If the processing is encoding, the current picture is synonymous with a picture to be encoded, and if the processing is decoding, the current picture is synonymous with a picture to be decoded. A block, such as a CU or CU, currently being processed by a device such as the encoding device 100 or the decoding device 200 is referred to as a current block. If the processing is encoding, the current block is synonymous with a block to be encoded, and if the processing is decoding, the current block is synonymous with a block to be decoded.

[0161] Here, a region in which parameters used for encoding and decoding are described may be referred to as a header region. For example, the header region is a region including an SEI. The header region may further include a VPS, an SPS, a PPS, an SEI, a picture header, a slice header, a CTU header, and a CU header.

[0162] [Configuration and Processing of Decoding] Fig. 9 is a block diagram showing an example configuration of a decoding device 200 according to this embodiment. In this example, the decoding device 200 includes decompressors 231, 233, and 235, derivers 232 and 236, a generator 234, and a buffer 237. Each component is, for example, an electric circuit that performs information processing. Two or more of the decompressors 231, 233, and 235 may be integrated together.

[0163] The decompressor 231 decodes a reference image from the bitstream. The deriver 232 derives a reference attribute set from the reference image. For example, the generator 234 may be provided with the reference image obtained by the decompressor 231 as well as the reference attribute set obtained by the deriver 232. Alternatively, the configuration and processing of the deriver 232 may be omitted, and the generator 234 may be provided with the reference image without being provided with a reference attribute set.

[0164] The decompressor 233 decodes a geometry attribute set for each picture from the bitstream. The geometry attribute set decoded from the bitstream may be simply referred to as a decoded geometry attribute set. The decompressor 235 decodes concealment parameters from the bitstream.

[0165] The concealment parameters are parameters related to error recovery control when an appropriate geometric attribute set cannot be obtained in the encoding device 100, and are decoded for each picture, for example. The concealment parameters can also be expressed as error recovery control parameters.

[0166] The buffer 237 stores a geometric attribute set. The geometric attribute set stored in the buffer 237 may be simply referred to as a stored geometric attribute set. The stored geometric attribute set may be a geometric attribute set that has been previously decoded, or may be a predefined geometric attribute set. The buffer 237 may store multiple stored geometric attribute sets corresponding to multiple pictures.

[0167] The deriver 236 derives a geometric attribute set based on the decoded geometric attribute set, the stored geometric attribute set, and the concealment parameters. The geometric attribute set derived by the deriver 236 may be simply referred to as a derived geometric attribute set. The deriver 236 may derive a geometric attribute set based on multiple stored geometric attribute sets corresponding to multiple pictures. The stored geometric attribute set may be an average of the multiple stored geometric attribute sets.

[0168] The generator 234 generates a facial motion image based on the reference attribute set and the derived geometric attribute set. For example, the generator 234 has a generative model that outputs an image in response to the input of the reference attribute set and the derived geometric attribute set, and generates a facial motion image from the reference attribute set and the derived geometric attribute set using the generative model. The generative model may be a neural network such as a Generative Adversarial Network (GAN).

[0169] Also, a reference image may be used instead of the reference attribute set, or both the reference attribute set and the reference image may be used.

[0170] Note that decoding information from the bitstream corresponds to decompressing the compressed information in the bitstream, and a portion of the geometric attribute set corresponding to a frame may be stored in or retrieved from buffer 237.

[0171] 10 is a flowchart showing an example of the operation of the decoding device 200 according to this embodiment. For example, the components of the decoding device 200 shown in FIG. 9 operate in accordance with the flowchart of FIG.

[0172] First, the decompressor 235 decodes the concealment parameter from the bitstream (S201). Then, the deriver 236 determines whether the concealment parameter is equal to a predetermined value (S202). If the concealment parameter is equal to the predetermined value, the deriver 236 derives a geometric attribute set as a derived geometric attribute set using the stored geometric attribute set (S203). If the concealment parameter is not equal to the predetermined value, the deriver 236 may derive the decoded geometric attribute set as the derived geometric attribute set.

[0173] The generator 234 then generates a facial motion image based on the derived geometric attribute set (S204).

[0174] In the above operation, for example, the deriver 236 determines an error recovery method for the decoded geometric attribute set according to the concealment parameters decoded from the bitstream.

[0175] An example of a geometric attribute set is facial landmarks derived from frames of a driving video at a given time. These landmarks indicate the locations of points on key areas of the face, including the facial contours, eyes, eyebrows, nose, mouth, lips, and chin. This allows for the interpretation of facial attributes and allows for easy modification of the facial attributes to create desired emotions and facial expressions. These landmarks may be in a 2D or 3D spatial coordinate system. After decoding, the geometric attribute set may be stored in buffer 237.

[0176] Below are several examples (1) to (5) of concealment parameters decoded from the bitstream.

[0177] (1) Example of concealment parameter indicating that a geometric attribute set was not acquired For example, the concealment parameter indicates that a geometric attribute set was not acquired in the encoding device 100. Specifically, a geometric attribute set may not be acquired because no face is present in the driving frame captured by the encoding device 100. In such a case, the concealment parameter may indicate that a geometric attribute set was not acquired. Such a concealment parameter may indicate that no geometric attribute is present in the bitstream.

[0178] (2) Example of a concealment parameter indicating that the number of attributes is different from the original number of attributes For example, the concealment parameter indicates that the number of attributes in the geometric attribute set is different from the original number of attributes. The original number of attributes may be the expected number of attributes, or the number of attributes that is assumed.

[0179] Specifically, the number of attributes in the geometry attribute set decoded for the current frame may differ from the number of attributes in the geometry attribute set decoded for the previous frame, or the number of attributes in the geometry attribute set decoded for the current frame may differ from the predefined number of attributes or the number of attributes that should be included in each geometry attribute set.

[0180] Such cases may occur in situations where only part of the face is enclosed in the driving frame, or where the encoding device 100 is unable to partially detect (extract) the geometric attribute set from the current frame due to accuracy, environment, etc.

[0181] In the above case, the encoding device 100 may transmit the geometry attribute set as is without performing error concealment on the geometry attribute set, and may further signal a concealment parameter indicating that the number of attributes is different from the original number of attributes.

[0182] (3) Example of a concealment parameter indicating use of a stored geometry attribute set For example, the concealment parameter indicates use of a geometry attribute set stored in the buffer 237 of the decoding device 200. This makes it possible for the encoding device 100 to instruct the decoding device 200 to use a stored geometry attribute set. Specifically, the user of the encoding device 100 may select a stored geometry attribute set. Then, the encoding device 100 may request the decoding device 200 to acquire the selected stored geometry attribute set. In this case, detection of the geometry attribute set may be omitted in the encoding device 100.

[0183] (4) Example of concealment parameter indicating storage of decoded geometry attribute set For example, the concealment parameter may indicate storage of a decoded geometry attribute set. Specifically, the concealment parameter may indicate when the decoding device 200 starts and stops decoding the geometry attribute set from the bitstream and storing it in the buffer 237.

[0184] (5) Examples of Hiding Parameters Indicating Low Confidence in a Geometric Attribute Set Hiding parameters may indicate low confidence (confidence score) in a geometric attribute set. This can arise from any of the following scenarios:

[0185] Specifically, temporary occlusion of a portion of the face may occur due to a facial mask, an eye patch, body movement, facial jewelry, or other objects. Furthermore, the head may have extreme poses. Furthermore, head movement may cause motion blur. In such cases, the encoding device 100 may detect (extract) some or all of the geometric attribute set with a low confidence score and signal an occlusion parameter indicating the low confidence score of the geometric attribute set.

[0186] The concealment parameters may indicate that the geometric attribute set was not properly obtained in the encoding device 100 based on the above examples (1) to (5).

[0187] The deriver 236 of the decoding device 200 determines a process for deriving a geometric attribute set by error concealment in the decoding device 200 by determining whether the value of the concealment parameter decoded from the bitstream is equal to a predetermined value (S202).

[0188] If the concealment parameters decoded from the bitstream are equal to the predetermined values, the deriver 236 uses the stored geometric attribute set to derive a geometric attribute set for generating a facial video (S203). On the other hand, if the concealment parameters decoded from the bitstream are not equal to the predetermined values, the deriver 236 applies the geometric attribute set decoded from the bitstream to the geometric attribute set for generating a facial video. That is, in this case, the derived geometric attribute set is the decoded geometric attribute set.

[0189] In one example, the predetermined value is decoded from the bitstream. In another example, the predetermined value is a fixed value. In another example, the predetermined value is either 0 or 1.

[0190] If it is determined that the concealment parameter decoded from the bitstream is equal to the predetermined value, the deriver 236 derives a geometric attribute set for generating a facial moving image (S203) using at least the geometric attribute set stored in the buffer 237. The stored attribute set, which is the geometric attribute set stored in the buffer 237, is, for example, a geometric attribute set previously decoded from the bitstream.

[0191] Here, a geometric attribute set for generating a facial motion image may be derived based on attributes and values ​​of the occlusion parameters and at least one of the stored geometric attribute set and the decoded geometric attribute set, where the decoded geometric attribute set is a geometric attribute set corresponding to a current frame.

[0192] 11 is a conceptual diagram illustrating an example of operations performed according to the concealment parameters. For example, the decompressors 233 and 235 apply arithmetic decoding to the bitstream to obtain a decoded geometric attribute set and the concealment parameters. The deriver 236 then determines whether the concealment parameters are equal to a predetermined value. If the concealment parameters are equal to the predetermined value, the deriver 236 obtains a stored geometric attribute set.

[0193] The deriver 236 then applies an inverse affine transformation to the decoded geometric attribute set and the stored geometric attribute set to derive a geometric attribute set for generating a facial video image. In the inverse affine transformation process, the decoded geometric attribute set and the stored geometric attribute set may be combined. Alternatively, the decoded geometric attribute set may be combined with multiple stored geometric attribute sets.

[0194] Furthermore, if the occlusion parameter is not equal to the predetermined value, the deriver 236 may derive the geometric attribute set for generating the facial moving image only from the decoded geometric attribute set, without using the stored geometric attribute set. Furthermore, the deriver 236 may apply the decoded geometric attribute set directly to the derived geometric attribute set without using the inverse affine transformation.

[0195] Furthermore, when the occlusion parameter is equal to a predetermined value, the deriver 236 may derive the geometric attribute set for generating the facial moving image only from the stored geometric attribute set, without using the decoded geometric attribute set. Furthermore, the deriver 236 may apply the stored geometric attribute set directly to the derived geometric attribute set without using the inverse affine transformation.

[0196] Here, the inverse affine transformation corresponds to linear transformation such as enlargement, reduction, rotation, and translation. Also, here, the inverse affine transformation may be an affine transformation.

[0197] 12 is a conceptual diagram showing another example of an operation performed in accordance with the concealment parameters. In this example, the deriver 236 applies an inverse affine transformation to the decoded geometric attribute set. The deriver 236 then combines the result obtained by applying the inverse affine transformation to the decoded geometric attribute set with the stored geometric attribute set to derive a geometric attribute set for generating a facial video image. In other words, as in this example, the inverse affine transformation does not need to be applied to the stored geometric attribute set.

[0198] It should be noted that the inverse affine transformations shown in FIGS. 11 and 12 may be replaced with other types of transformations.

[0199] The stored geometry attribute set may be a geometry attribute set previously decoded from the bitstream and stored in the buffer 237. In one example, the stored geometry attribute set may be a geometry attribute set decoded and stored for a frame prior to the current frame. In another example, the stored geometry attribute set may be a geometry attribute set decoded and stored for the first (intra) frame of a group of pictures.

[0200] In yet another example, the stored geometric attribute set may be a geometric attribute set that exists locally in both the encoding device 100 and the decoding device 200. Specifically, the geometric attribute set may be derived from a common image that exists locally in both the encoding device 100 and the decoding device 200. In yet another example, the stored geometric attribute set may be a predetermined geometric attribute set.

[0201] In yet another example, the stored geometric attribute set may correspond to a portion of a complete geometric attribute set representing an entire face, such as the left eye group in the complete geometric attribute set.

[0202] To prevent memory buffer overflow, the decoding device 200 may retain only the most recently derived geometry attribute sets. Alternatively, the decoding device 200 may retain only geometry attribute sets that are less than a preset threshold in the sorted historical geometry attribute sets and discard less recent geometry attribute sets. This keeps one or more geometry attribute sets in the buffer 237 of the decoding device 200 up to date with the latest changes to the driving frame scene.

[0203] Note that in a design example including a concealment parameter indicating "low_confidence," only geometric attribute sets with high predicted confidence scores are stored in the buffer 237 and serve as reference attributes required for drawing other frames. In this case, geometric attribute sets with low confidence scores are not stored in the buffer 237. This allows only geometric attribute sets with high confidence scores to be used as baselines. This prevents error-prone geometric attribute sets from being stored and resulting in error accumulation.

[0204] Below are examples (1) to (5) of derived geometric attribute sets obtained by the derivator 236. The derived geometric attribute set examples (1) to (5) may correspond to the obscurance parameter examples (1) to (5) above.

[0205] (1) Case where the concealment parameter indicates that a geometry attribute set has not been acquired For example, the concealment parameter indicates that a geometry attribute set has not been acquired. In this case, the decoding device 200 can skip decoding the geometry attribute set. Then, the decoding device 200 can derive the geometry attribute set by directly using the stored geometry attribute set acquired from the buffer 237. For example, if a geometry attribute set has not been acquired in the encoding device 100, the concealment parameter is set to 1, indicating that a geometry attribute set has not been acquired.

[0206] 13 is a conceptual diagram showing an example of operation when a geometric attribute set is not acquired. In this example, the hiding parameter is equal to 1. In this case, the deriving unit 236 acquires a stored geometric attribute set from the buffer 237. Then, the deriving unit 236 sets the stored geometric attribute set to a derived geometric attribute set. Then, the deriving unit 236 outputs the derived geometric attribute set. That is, in this case, the deriving unit 236 outputs the stored geometric attribute set as the derived geometric attribute set.

[0207] (2) When the concealment parameter indicates that the number of attributes is different from the original number of attributes For example, the concealment parameter indicates that the number of attributes is different from the original number of attributes. Specifically, in this case, the number of attributes in the decoded geometry attribute set is different from the expected number of attributes. Therefore, in this case, the decoding device 200 obtains the stored geometry attribute set from the buffer 237 and uses the decoded geometry attribute set and the stored geometry attribute set in combination. For example, when the concealment parameter is set to 1, the concealment parameter indicates that the number of attributes is different from the original number of attributes.

[0208] If the number of attributes in the geometry attribute set decoded from the bitstream is fewer than expected, the missing attributes in the decoded geometry attribute set may be filled in or replaced by corresponding attributes in the stored geometry attribute set. If the number of attributes in the geometry attribute set decoded from the bitstream is more than expected, the excess attributes may be deleted from the decoded geometry attribute set.

[0209] 14 is a conceptual diagram showing an example of the operation of the decoding device 200 when the number of attributes in the geometry attribute set is different from the original number of attributes. In this example, missing attributes in the decoded geometry attribute set are filled in with corresponding attributes in the stored geometry attribute set.

[0210] In another example, the geometric attribute set decoded from the bitstream may be ignored, in which case the decoding device 200 directly uses the stored geometric attribute set obtained from the buffer 237 to derive the geometric attribute set for generating the facial video image.

[0211] Also, a corresponding attribute may be selected from a plurality of corresponding attributes in a plurality of stored geometry attribute sets corresponding to a plurality of pictures and applied to the missing attribute in the decoded geometry attribute set.

[0212] (3) When the concealment parameter indicates that a stored geometric attribute set is to be used: For example, the concealment parameter indicates that a stored geometric attribute set is to be used. Therefore, in this case, the decoding device 200 may skip the decoding process of the geometric attribute set and directly use the stored geometric attribute set acquired from the buffer 237 to generate a facial moving image. For example, when the concealment parameter is set to 1, this indicates that a stored geometric attribute set is to be used.

[0213] (4) Case where the concealment parameter indicates that the decoded geometry attribute set is to be stored For example, the concealment parameter indicates that the decoded geometry attribute set is to be stored in the buffer 237, depending on the value of the concealment parameter. This allows multiple geometry attribute sets decoded from the bitstream to be stored in the buffer 237. Then, the concealment parameter may indicate that the stored geometry attribute set in the buffer 237 is to be used, depending on the value of the concealment parameter. Specifically, the concealment parameter may indicate that multiple geometry attribute sets are to be retrieved from the buffer 237 in a loop manner.

[0214] Therefore, if a loop is required by the concealment parameters, the decoding device 200 directly uses the previously decoded and stored sets of geometric attributes to perform loop processing between the start_loop timestamp and the end_loop timestamp, which may be specified by the concealment parameters.

[0215] For example, if the concealment parameters indicate to start storing a geometric attribute set, the decoding device 200 starts storing the decoded geometric attribute set in the buffer 237 for each frame, and simultaneously applies the decoded geometric attribute set to the derived geometric attribute set.

[0216] If the concealment parameters indicate to stop storing geometric attribute sets, the decoding device 200 stops storing decoded geometric attribute sets in the buffer 237. In subsequent frames, the decoding device 200 continuously repeats in a time loop to retrieve geometric attribute sets from the buffer 237, from the first stored geometric attribute set to the last stored geometric attribute set.

[0217] The decoding device 200 repeats the loop until the concealment parameter indicates that the loop should be stopped and the decoded geometric attribute set should be applied to the derived geometric attribute set without being stored in the buffer 237 .

[0218] For example, a user of encoding device 100 may repeatedly perform natural movements such as blinking or nodding, in which case driving frame capture may be temporarily disabled and a natural movement such as blinking every three seconds may be simulated in a time loop.

[0219] 15 is a conceptual diagram showing an example of an operation in which storage of a geometric attribute set is controlled. For example, when the hiding parameter has a value of 1, it indicates that storage of a geometric attribute set in the buffer 237 is to begin.

[0220] Furthermore, when the hiding parameter has a value of 2, it indicates to stop storing geometric attribute sets in the buffer 237 and start retrieving geometric attribute sets from the buffer 237. When the hiding parameter has a value of 0, it indicates to stop retrieving geometric attribute sets from the buffer 237 and apply the decoded geometric attribute set to the derived geometric attribute set.

[0221] In one example, buffer 237 is a first-in-first-out (FIFO) queue, and when the obscurance parameter is equal to 2, multiple stored geometric attribute sets may be retrieved from buffer 237 in the order in which they were stored. In another example, buffer 237 is a last-in-first-out (LIFO) stack, and multiple stored geometric attribute sets may be retrieved from buffer 237 in the reverse order of their storage.

[0222] (5) When the concealment parameters indicate that the reliability of the geometric attribute set is low: For example, when the concealment parameters indicate that the reliability (reliability score) of the geometric attribute set is low, the decoding device 200 may perform one or more of the following processes for error concealment:

[0223] (a) Temporally smoothing the decoded geometry attribute set based on the decoded geometry attribute set and the stored geometry attribute set; (b) directly combining the decoded geometry attribute set with the stored geometry attribute set (substituting or replacing attributes); and (c) correcting the distances between various parts of the face (such as the eyes and mouth) in the decoded geometry attribute set using the stored geometry attribute set as a reference.

[0224] The obscuration parameter may indicate a low confidence level for the geometric attribute set when set to 1, 2, or 3. In this case, the obscuration parameter set to 1, 2, or 3 may correspond to the above-described operations (a), (b), or (c).

[0225] The occlusion parameters may indicate low confidence in units of geometric attribute sets, i.e., frames, or alternatively, the occlusion parameters may indicate low confidence in units of attributes in the geometric attribute set, or alternatively, the occlusion parameters may indicate low confidence in units of different parts of the face in the geometric attribute set.

[0226] In one example, for each attribute in the set of geometric attributes, a confidence score may be calculated and compared to a predetermined threshold, and for attributes below the threshold, a predetermined condition may be determined to be true, i.e., for attributes below the threshold, a low confidence level may be determined.

[0227] In another example, the set of geometric attributes may be divided into different groups, each corresponding to a key facial feature. Then, the confidence scores of all attributes in each group may be averaged and compared with a predetermined threshold. Then, for groups whose average confidence score is below the threshold, a predetermined condition may be determined to be true. That is, for groups whose average confidence score is below the threshold, the confidence may be determined to be low.

[0228] 16 is a conceptual diagram showing an example of an operation performed in the decoding device 200 according to the reliability of the geometry attribute set. When the concealment parameter is set to 0, the concealment parameter does not indicate that the reliability of the geometry attribute set is low. For example, in this case, the reliability of the geometry attribute set is high, so the decoded geometry attribute set is stored as the stored geometry attribute set and is derived as the derived geometry attribute set.

[0229] If the concealment parameter is not set to 0, the concealment parameter indicates that the reliability of the geometric attribute set is low. In this case, the stored geometric attribute set is obtained from the buffer 237. Then, a derived geometric attribute set is obtained from the decoded geometric attribute set and the stored geometric attribute set. The derived geometric attribute set is then used to generate the facial video sequence. If the concealment parameter is not set to 0, the concealment parameter may indicate how to derive the geometric attribute set for generating the facial video sequence depending on the value of the concealment parameter.

[0230] In another example, the decoded geometric attribute set obtained from the bitstream may be ignored, and the decoding device 200 may directly use the stored geometric attribute set obtained from the buffer 237 as the derived geometric attribute set for generating the facial video image.

[0231] Furthermore, multiple decoded geometry attribute sets corresponding to multiple pictures may be stored as multiple stored geometry attribute sets in the buffer 237. Then, multiple stored geometry attribute sets corresponding to multiple pictures may be used in deriving the geometry attribute set.

[0232] The derived geometric attribute set can be obtained based on the above examples (1) to (5) corresponding to multiple examples of concealment parameters. Specific examples of (a) temporal smoothing, (b) combining the decoded geometric attribute set with the stored geometric attribute set, and (c) distance correction for obtaining the derived geometric attribute set are given below.

[0233] (a) Temporal Smoothing For example, when the concealment parameters indicate that the reliability of the geometric attribute set is low, the deriver 236 obtains a derived geometric attribute set by temporally smoothing the decoded geometric attribute set based on the decoded geometric attribute set and the stored geometric attribute set. Specifically, the deriver 236 may obtain the derived geometric attribute set using first-order exponential smoothing. More specifically, the deriver 236 obtains the derived geometric attribute set by smoothing x derived = (1-α) × x decoded +α×x stored may be used to obtain the derived geometric attribute set.

[0234] Here, x derived refers to the derived geometric attribute set, and x decoded refers to the decoded geometric attribute set, and x stored where α denotes a stored geometric attribute set. Furthermore, α is a smoothing parameter that indicates the level of emphasis given to the stored geometric attribute set. The value of α may be predetermined or may be decoded from the bitstream. For example, the smoothing parameter α may take a value between 0% and 100% and indicate the percentage weight assigned to the stored geometric attribute set.

[0235] For example, x derived refers to the coordinate value of an attribute (landmark) in the derived geometric attribute set, and x decoded refers to the coordinate value of the attribute (landmark) in the decoded geometric attribute set, and x stored refers to the coordinate value of an attribute (landmark) in the stored geometric attribute set. The coordinate value for the derived geometric attribute set is calculated by a weighted average of the coordinate value for the decoded geometric attribute set and the coordinate value for the stored geometric attribute set.

[0236] In the above, a weighted average of one decoded geometry attribute set and one stored geometry attribute set is used, however, a weighted average of one decoded geometry attribute set and multiple stored geometry attribute sets may also be used.

[0237] (b) Combining the decoded geometric attribute set with the stored geometric attribute set For example, if the concealment parameters indicate that the geometric attribute set has low reliability, the deriver 236 obtains a derived geometric attribute set by combining the decoded geometric attribute set with the stored geometric attribute set.

[0238] 17 is a conceptual diagram illustrating an example of the operation of replacing geometric attributes in the decoding device 200. Specifically, attributes with low reliability in the decoded geometry attribute set are replaced with attributes with high reliability in the stored geometry attribute set. For example, the reliability threshold for replacement is set to 0.7. Therefore, as shown in FIG. 17 , attributes in the decoded geometry attribute set with reliability lower than the threshold are replaced with corresponding attributes in the stored geometry attribute set.

[0239] 18 is a conceptual diagram illustrating another example of the operation of replacing geometric attributes in the decoding device 200. For example, in the encoding device 100, unreliable attributes are identified and deleted. The decoding device 200 replaces the missing attributes with corresponding attributes in the stored geometric attribute set. That is, the decoding device 200 fills in the missing attributes with corresponding attributes in the stored geometric attribute set.

[0240] A corresponding attribute may be selected from among the corresponding attributes in the stored geometric attribute sets based on the confidence levels to be applied to the missing attribute.

[0241] The reliability may be determined in units of a geometric attribute set, an attribute, or a group including multiple attributes in a geometric attribute set. Then, a geometric attribute set or an attribute may be selected in units for which the reliability is determined.

[0242] (c) Distance Correction For example, if the concealment parameters indicate that the geometric attribute set is not reliable, the derivator 236 obtains a derived geometric attribute set by referring to the stored geometric attribute set and correcting the distances between parts in the decoded geometric attribute set.

[0243] Specifically, a reference distance set is derived and stored by dividing the memory geometric attribute set into different groups, each representing a key facial feature, deriving the center point of each group, and calculating the relative distance from the derived center point to a reference point (such as the tip of the nose).

[0244] Then, a distance is derived for each group by similarly dividing the decoded geometric attribute set into groups, deriving a center point for each group, and calculating the relative distance from the derived center point to the reference point. If the derived distance differs from the distance in the reference distance set by more than a threshold, the entire group is shifted so that the derived distance matches the distance in the reference distance set.

[0245] Here, the geometric attribute set refers to the geometric attribute set corresponding to the frame. That is, the geometric attribute set refers to the complete geometric attribute set for the entire face. Also, for example, the center point of the group can be derived by averaging the x- and y-coordinates of all attributes (points) in the group. Alternatively, the center point of the group can be determined by deriving a bounding box that surrounds all points in the group based on the minimum and maximum x- and y-values ​​of all attributes (points) in the group, and setting the center of the bounding box as the center point of the group.

[0246] FIG. 19 is a conceptual diagram showing an example of a stored geometric attribute set, which is a geometric attribute set stored in the decoding device 200. In this example, the center point of the "right eye" group in the stored geometric attribute set is (x stored_re , y stored_re ) and the nose point of the memory geometry attribute set is expressed as (x stored_n , y stored_n ) is expressed as

[0247] FIG. 20 is a conceptual diagram showing an example of a decoded geometry attribute set that is a geometry attribute set decoded by the decoding device 200. In this example, the center point of the "right eye" group in the decoded geometry attribute set is expressed as (xdecoded_re, ydecoded_re). In addition, the nose point of the decoded geometry attribute set is expressed as (x decoded_n , y decoded_n ) is expressed as

[0248] FIG. 21 is a conceptual diagram showing an example of a derived geometric attribute set, which is a geometric attribute set derived in the decoding device 200. In this example, the center point of the "right eye" group in the derived geometric attribute set is expressed as (xderived_re, yderived_re). In addition, the nose tip point of the derived geometric attribute set is expressed as (x derived_n , y derived_n ) is expressed as

[0249] The method for setting the derived geometry attribute set in FIG. 21 based on the stored geometry attribute set in FIG. 19 and the decoded geometry attribute set in FIG. 20 is as follows.

[0250] (1) First, the deriving unit 236 calculates the distance d from the nose tip point to the center point of the right eye group in the stored geometric attribute set. stored_re is calculated according to the following formula:

[0251]

[0252] (2) Next, the deriver 236 calculates the distance ddecoded_re from the nose tip point to the center point of the right eye group in the decoded geometric attribute set according to the following formula:

[0253]

[0254] (3) Next, the deriver 236 calculates |ddecoded_re-d stored_re |> ε, then dderived_re = d stored_re , and (x derived_n , y derived_n ) = (x decoded_n , y decoded_n ) to set the

[0255] That is, the deriver 236 determines whether the difference between the distance from the nose tip point to the center point of the right eye group in the stored geometric attribute set and the distance from the nose tip point to the center point of the right eye group in the decoded geometric attribute set is greater than a threshold.

[0256] If the difference is greater than the threshold, the deriving unit 236 sets the distance from the nose tip point in the derived geometric attribute set to the center point of the right eye group to the distance from the nose tip point in the stored geometric attribute set to the center point of the right eye group. In this case, the deriving unit 236 also sets the nose tip point in the derived geometric attribute set to the nose tip point in the decoded geometric attribute set.

[0257] The deriver 236 calculates |ddecoded_re-d stored_re If |≦ε, the right-eye group in the decoded geometric attribute set is set to the right-eye group in the derived geometric attribute set, and the following processes (4), (5) and (6) are skipped.

[0258] (4) Next, the deriving unit 236 converts (x derived_re, y derived_re) into d derived_re = d stored_re That is, the deriver 236 sets the center point of the right-eye group in the derived geometric attribute set to a point where the distance between the nose tip point and the center point of the right-eye group is equal to the distance in the stored geometric attribute set and where the center point is closest to the center point of the right-eye group in the decoded geometric attribute set.

[0259] (5) Next, the deriver 236 derives a transformation to map the center point of the right-eye group (xdecoded_re, ydecoded_re) in the decoded geometric attribute set to the center point of the right-eye group (xderived_re, yderived_re) in the derived geometric attribute set.

[0260] (6) Next, the deriver 236 shifts the entire right-eye group using the same transformation derived in process (5) above.

[0261] The deriving unit 236 sets a derived geometric attribute set by performing the above processes (1) to (6) for each of all other groups, thereby adjusting the distance between groups so that they are not too far apart.

[0262] A generator 234 generates an image from the derived geometric attribute set via a neural network.

[0263] For example, a neural network is a generative network. The generative network may also be a generative adversarial network (GAN), a variational autoencoder (VAE), an autoregressive model, a diffusion model, etc. The generative network may be a machine learning framework that generates new data based on a provided dataset.

[0264] Generative networks, also known as generative models, analyze and learn from the underlying distribution of a dataset, ensuring that the resulting new dataset is similar to the original dataset.

[0265] The decoding device 200 generates an image based on the concealment parameters. Target applications may include, but are not limited to, video generation, editing, and playback in the video conferencing, entertainment, social media, and e-commerce industries.

[0266] Note that the processing of the decoding device 200 can also be performed in the encoding device 100. Furthermore, not all of the components in the present disclosure are necessarily required, and only some of the components may be implemented.

[0267] 22 is a block diagram showing another example configuration of the decoding device 200 according to this embodiment. In this example, the decoding device 200 includes a video decoder 401, an entropy decoder 402, an inverse affine transformer 403, a determiner 404, a geometric attribute buffer 405, and a generator 406. The entropy decoder 402 may be included in the video decoder 401.

[0268] The video decoder 401 applies a video decoding process to the bitstream received from the encoding device 100. Next, the entropy decoder 402 applies an entropy decoding process, such as CABAC or VLC, to the SEI data obtained from the bitstream to obtain concealment parameters and, optionally, a current decoded geometric attribute set.

[0269] Next, the inverse affine transformer 403 performs a transform, such as an inverse affine transform, on the decoded geometric attribute set. Note that the inverse affine transformer 403 may be replaced by another transformer used to transform the decoded geometric attribute set. Specifically, if affine parameters or other transformation parameters are decoded, the inverse affine transform may be performed. Otherwise, the inverse affine transform may be replaced by another compatible process.

[0270] Then, the determiner 404 determines whether the concealment parameter is equal to a predetermined value. If it is determined that the concealment parameter is equal to the predetermined value, the inverse affine transformer 403 retrieves the stored geometric attribute set in the geometric attribute buffer 405 and combines it with the decoded geometric attribute set to set a derived geometric attribute set. Then, the generator 406 generates frames of the facial video image based on the derived geometric attribute set set by error concealment in the decoding device 200.

[0271] With the above configuration, the decoding device 200 can perform error concealment based on the concealment parameters and generate a facial moving image.

[0272] Also, in the above examples, a memory geometric attribute set is used in error concealment. However, the decoding device 200 may not use a memory geometric attribute set for error concealment. For example, the decoding device 200 may apply a previous frame in a facial video sequence to a current frame in the facial video sequence. In other words, the decoding device 200 may apply an already generated facial image to a facial image corresponding to the current frame, without using a generative model to generate a facial image corresponding to the current frame in the facial video sequence.

[0273] [Encoding Configuration and Processing] Fig. 23 is a block diagram showing an example configuration of an encoding device 100 according to this embodiment. In this example, the encoding device 100 includes compressors 131, 133, and 134, a deriver 132, and a buffer 135. Each component is, for example, an electric circuit that performs information processing. Two or more of the compressors 131, 133, and 134 may be integrated together.

[0274] The compressor 131 encodes a reference image into a bitstream. The deriver 132 derives a geometric attribute set for each picture from the driving video. The deriver 132 also generates (derives) concealment parameters in deriving the geometric attribute set. The compressor 133 encodes the geometric attribute set for each picture into a bitstream. The geometric attribute set derived and encoded from the driving video may also be referred to as a derived geometric attribute set, a target geometric attribute set, or an encoded geometric attribute set.

[0275] The compressor 134 encodes the concealment parameters into a bitstream. The concealment parameters are parameters related to error recovery control when the encoding device 100 is unable to obtain an appropriate geometric attribute set. The concealment parameters may also be expressed as error recovery control parameters. For example, the concealment parameters are derived and encoded for each picture.

[0276] The buffer 135 stores a geometric attribute set. The geometric attribute set stored in the buffer 135 may be simply referred to as a stored geometric attribute set. The stored geometric attribute set may be a geometric attribute set derived in the past, or may be a predefined geometric attribute set. The buffer 135 may store multiple stored geometric attribute sets corresponding to multiple pictures.

[0277] The deriver 132 may use the stored geometric attribute set in the buffer 135 to derive the geometric attribute set. For example, the deriver 132 detects a geometric attribute set for each picture from the driving video. The geometric attribute set detected from the driving video may also be referred to as a detected geometric attribute set or an extracted geometric attribute set. The deriver 132 then establishes a derived geometric attribute set based on the detected geometric attribute set and the stored geometric attribute set.

[0278] The deriver 132 may derive the geometric attribute set based on a plurality of stored geometric attribute sets corresponding to a plurality of pictures, and may use an average of the plurality of stored geometric attribute sets as the stored geometric attribute set.

[0279] The concealment parameters may be generated by the detected geometric attribute set, or may be generated by the derived geometric attribute set, or may be generated by the detected geometric attribute set and the derived geometric attribute set. Also, for example, the concealment parameters may be generated based on the detected geometric attribute set, the derived geometric attribute set may be set based on the concealment parameters, and the concealment parameters may be updated based on the derived geometric attribute set.

[0280] Note that encoding the information into the bitstream corresponds to compressing the information and including the compressed information in the bitstream. Also, a portion of the geometric attribute set corresponding to the frame may be stored in or retrieved from buffer 135.

[0281] 24 is a flowchart showing an example of the operation of the encoding device 100 according to this embodiment. For example, the components of the encoding device 100 shown in FIG. 23 operate in accordance with the flowchart in FIG.

[0282] First, the deriver 132 detects a geometric attribute set for each picture in the driving video (S101). Then, the deriver 132 generates concealment parameters based on the detected geometric attribute set (S102). Then, the compressor 134 encodes the concealment parameters into a bitstream (S103).

[0283] The detection of the geometric attribute set (S101) may correspond to the detection of a face in an image, where the image is, for example, a frame (picture) of a driving video. Specifically, the face in the image may be detected by a face detection algorithm.

[0284] The occlusion parameters may be generated based on the face detection results. Alternatively, the face detection results in the image may be reflected in the geometric attribute set detection results. For example, if no face is detected in the image, the occlusion parameters may indicate that no geometric attribute set is detected.

[0285] For example, the geometric attributes may be facial landmarks that indicate the locations of points on key areas of the face, including the facial contours, eyes, eyebrows, nose, mouth, lips, and chin, thereby enabling interpretation of facial expressions and allowing the facial expressions to be easily modified to produce appropriate facial expressions. These landmarks may be in a two-dimensional spatial coordinate system or a three-dimensional spatial coordinate system.

[0286] 25 is a conceptual diagram showing a specific example of a geometric attribute set. In this example, the geometric attribute set is a facial landmark set covering the eyes, eyebrows, nose, mouth, lips, chin, and facial contour extracted from an image.

[0287] Fig. 26 is a conceptual diagram showing another specific example of a geometric attribute set. In this example, the geometric attribute set is a face landmark set that covers the eyes, eyebrows, nose, mouth, lips, chin, and facial contour extracted from an image, and covers more points than the example of Fig. 25.

[0288] Fig. 27 is a conceptual diagram showing yet another specific example of a geometric attribute set. In this example, the geometric attribute set is a face landmark set that covers the eyes, eyebrows, nose, mouth, lips, chin, cheeks, and facial contours extracted from an image, and covers more points than the examples in Fig. 25 and Fig. 26.

[0289] If a geometric attribute set is detected, the geometric attribute set may be stored in a buffer 135 .

[0290] Additionally, the set of geometric attributes coded into the bitstream may be a subset of the set of geometric attributes for the entire face, or one or more groups within the set of geometric attributes for the entire face, each group representing a key facial feature.

[0291] For example, the deriver 132 may predict a confidence score for a geometric attribute set. Here, predicting the confidence score corresponds to deriving, calculating, evaluating, determining, or obtaining the confidence score. The confidence score may be predicted for a geometric attribute set, a geometric attribute, or a group in the geometric attribute set. Then, a concealment parameter may be generated based on the confidence score.

[0292] Examples of concealment parameters (1) to (5) are shown below.

[0293] (1) Example of a concealment parameter indicating that a geometric attribute set has not been acquired During detection of a geometric attribute set, there is a possibility that the detection of the geometric attribute set may fail, resulting in the geometric attribute set not being detected. Therefore, the encoding device 100 may signal a concealment parameter based on the detection result of the geometric attribute set. In other words, the concealment parameter may indicate that a geometric attribute set has not been acquired.

[0294] If the concealment parameter indicates that the geometric attribute set has not been obtained, the encoding device 100 skips encoding the geometric attribute set and signals the concealment parameter to the decoding device 200. In this case, the decoding device 200 may directly use the stored geometric attribute set in the buffer 237 for face reconstruction based on the concealment parameter. For example, the concealment parameter may indicate that the geometric attribute set has not been obtained when set to 1.

[0295] Such a case may occur in a situation where no face is detected in the driving frame captured by the encoding device 100. This may result in an empty set of geometric attributes in the bitstream, which may result in the face not being reproduced or in erroneous distortion in the output image. The concealment parameters can suppress such errors.

[0296] (2) Example of a concealment parameter indicating that the number of attributes is different from the original number of attributes For example, the concealment parameter may indicate that the number of attributes in the geometric attribute set is different from the original number of attributes. The original number of attributes may be the expected number of attributes.

[0297] Specifically, the number of attributes in the geometric attribute set corresponding to the current frame may be different from the number of attributes in the geometric attribute set corresponding to the previous frame, or the number of attributes in the geometric attribute set corresponding to the current frame may be different from the predefined number of attributes or the number of attributes to be included in each geometric attribute set.

[0298] In the above case, the encoding device 100 may retrieve a stored geometry attribute set from the buffer 135. The stored geometry attribute set may be a geometry attribute set derived from a previous driving frame and previously coded into the bitstream. The stored geometry attribute set may then be used in combination with the detected geometry attribute set before being coded into the bitstream.

[0299] For example, if the detected geometry attribute set is smaller than expected, missing attributes in the detected geometry attribute set may be filled in or replaced by corresponding attributes in the stored geometry attribute set. If the number of attributes in the detected geometry attribute set is greater than expected, excess attributes may be deleted from the detected geometry attribute set.

[0300] 28 is a conceptual diagram illustrating an example of the operation of the encoding device 100 when the number of attributes in the geometry attribute set is different from the original number of attributes. In this example, missing attributes in the detected geometry attribute set are filled with corresponding attributes in the stored geometry attribute set. Note that corresponding attributes may be selected from multiple corresponding attributes in multiple stored geometry attribute sets corresponding to multiple pictures and applied to the missing attributes in the detected geometry attribute set.

[0301] In another example, the encoding device 100 may encode and transmit the detected geometry attribute set as a derived geometry attribute set and further signal a concealment parameter indicating that the number of attributes in the derived geometry attribute set is different from expected. In yet another example, the encoding device 100 may ignore the detected geometry attribute set and encode the stored geometry attribute set obtained from the buffer 135 into the bitstream as the derived geometry attribute set.

[0302] Note that, if the number of attributes coded into the bitstream after filling in missing attributes or deleting excess attributes becomes equal to the predefined number of attributes or the number of attributes to be included in each geometric attribute set, the encoding device 100 does not need to set a value indicating that the number of attributes is different from the original number of attributes as the concealment parameter. Instead, the encoding device 100 may signal in the bitstream a concealment parameter indicating that the geometric attribute set has low reliability.

[0303] Such a case may occur when only a portion of a face is included in the driving frame. In this case, the encoding device 100 may additionally transmit the position of the center of the face in the driving frame to the decoding device 200 to inform the decoding device 200 of which part of the frame the face is located in. This may assist in signaling the direction and proportion of the occluded face to the decoding device 200.

[0304] Furthermore, even if the entire face is present in the driving frame, the deriver 132 may not be able to partially detect (extract) a set of geometric attributes from the driving frame due to accuracy, environment, etc. Alternatively, the deriver 132 may not be able to detect (extract) a consistent number of attributes across multiple frames.

[0305] The decoding device 200 may be able to recognize the difference in the number of attributes based on the concealment parameters notified by the encoding device 100 and perform face reproduction according to the difference in the number of attributes.

[0306] (3) Example of Concealment Parameter Indicating Use of a Stored Geometry Attribute Set For example, a geometry attribute set is detected from a current frame of a driving video. However, it may be advantageous to use a stored geometry attribute set regardless of the detected geometry attribute set. Therefore, the encoding device 100 may signal a concealment parameter indicating use of a stored geometry attribute set.

[0307] If the concealment parameter indicates to use the stored geometric attribute set, encoding device 100 skips encoding the geometric attribute set and signals the concealment parameter to decoding device 200. In this case, decoding device 200 may directly use the stored geometric attribute set in buffer 237 for face reconstruction based on the concealment parameter. For example, the concealment parameter may indicate to use the stored geometric attribute set when set to 1.

[0308] Such a case may occur when a user of the encoding device 100 selects a stored geometric attribute set and the encoding device 100 requests the decoding device 200 to perform face reconstruction using the selected stored geometric attribute set. In this case, detection of the geometric attribute set may be omitted in the encoding device 100.

[0309] (4) Example of concealment parameter indicating storage of decoded geometry attribute set For example, the concealment parameter may indicate storage of a decoded geometry attribute set. Specifically, the encoding device 100 may signal the concealment parameter to the decoding device 200 to notify the timing of the start and end of storing the geometry attribute set transmitted in the bitstream. This makes it possible to use the stored geometry attribute sets in a loop, thereby simulating natural movements such as blinking or nodding in the loop. During the loop, the encoding device 100 does not need to transmit the geometry attribute set.

[0310] Therefore, the encoding device 100 may signal to the decoding device 200 the selection information of the stored geometric attribute set to be used for face reconstruction, without signaling the detected geometric attribute set.

[0311] (5) When the occlusion parameters indicate that the confidence level of the geometric attribute set is low, for example, when a face is included in a driving frame and a geometric attribute set is detected from the driving frame, a confidence score of the geometric attribute set may be predicted. Specifically, a part of the face may be temporarily hidden by wearing a mask that partially or completely covers the mouth, an eye patch that covers one or both eyes, a hand, body movement, or other object that temporarily obstructs part of the face, or facial decorations and accessories such as sunglasses.

[0312] 29 is a conceptual diagram showing an example of a face partially hidden by occlusion. In such a case, the encoding device 100 can predict the attribute (location of landmark) that was not detected due to the occlusion by using neighboring attributes (location of landmark) that are not affected by the occlusion. For this purpose, a set of geometric attributes derived and stored from a previous frame in the driving video is useful.

[0313] For these cases, encoding device 100 may signal concealment parameters indicating that these attributes have low confidence scores due to the concealment.

[0314] In another example, a face in a driving frame may be in an extreme head pose, resulting in parts of the face being obscured.

[0315] 30 is a conceptual diagram illustrating an example of a face with an extreme yaw pose. In such cases, encoding device 100 may signal occlusion parameters indicating that these landmarks are occluded and have a low confidence score. Furthermore, for architectures that use head pose values ​​during the generation process, a preset limit may be set in encoding device 100 on the maximum allowable head pose angle, such as a 45-degree limit for yaw from the front.

[0316] 31 is a conceptual diagram showing an example of a face having an extreme roll pose. In such a case, face reconstruction may not be performed properly. Therefore, even in such a case, the encoding device 100 may signal concealment parameters indicating a low confidence score.

[0317] Fig. 32 is a conceptual diagram showing an example of a blurred face. A face in a driving frame may suffer from motion blur due to the overall facial movement being greater than the capture rate of the camera.

[0318] 33 is a conceptual diagram illustrating an example of a face with a low number of landmark detections. In this example, parts of the face are occluded. Some landmarks may be detected with low confidence scores. In such a case, the encoding device 100 may signal concealment parameters indicating that these landmarks have low confidence scores.

[0319] In the above examples, the encoding device 100 predicts some or all of the attributes in the geometric attribute set with low confidence scores. Then, the encoding device 100 encodes a concealment parameter indicating the low confidence score of the geometric attribute set. Specifically, the concealment parameter may indicate the low confidence score for the geometric attribute set, for a group in the geometric attribute set, or for an attribute in the geometric attribute set.

[0320] If the concealment parameters indicate that the geometric attribute set is not reliable, the encoding device 100 may perform one or more of the following processes for error concealment.

[0321] (a) Temporally smoothing the derived geometric attribute set based on the detected geometric attribute set and the stored geometric attribute set; (b) directly combining the detected geometric attribute set with the stored geometric attribute set (substituting or replacing attributes); (c) correcting the distances between various parts of the face (such as the eyes and mouth) in the detected geometric attribute set using the stored geometric attribute set as a reference.

[0322] The obscuration parameter may indicate a low confidence level for the geometric attribute set when set to 1, 2, or 3. In this case, the obscuration parameter set to 1, 2, or 3 may correspond to the above-described operations (a), (b), or (c).

[0323] In one example, for each attribute in the set of geometric attributes, a confidence score may be calculated and compared to a predetermined threshold, and for attributes below the threshold, a predetermined condition may be determined to be true, i.e., for attributes below the threshold, a low confidence level may be determined.

[0324] In another example, the set of geometric attributes may be divided into different groups, each corresponding to a key facial feature. Then, the confidence scores of all attributes in each group may be averaged and compared with a predetermined threshold. Then, for groups whose average confidence score is below the threshold, a predetermined condition may be determined to be true. That is, for groups whose average confidence score is below the threshold, the confidence may be determined to be low.

[0325] 34 is a conceptual diagram illustrating an example of an operation performed in the encoding device 100 according to the reliability of the geometric attribute set. When the concealment parameter is set to 0, the concealment parameter does not indicate that the reliability of the geometric attribute set is low. For example, in this case, the reliability of the geometric attribute set is high, so the detected geometric attribute set is stored as the stored geometric attribute set and is derived as the derived geometric attribute set.

[0326] If the concealment parameter is not set to 0, the concealment parameter indicates low confidence in the geometric attribute set. In this case, the stored geometric attribute set is obtained from the buffer 135. Then, a derived geometric attribute set is obtained from the detected geometric attribute set and the stored geometric attribute set. The derived geometric attribute set is then encoded into the bitstream. If the concealment parameter is not set to 0, the concealment parameter may indicate how to derive the geometric attribute set for generating the facial animation, depending on the value of the concealment parameter.

[0327] In another example, the detected geometric attribute set obtained from the driving frame may be ignored, and the encoding device 100 may directly use the stored geometric attribute set obtained from the buffer 135 as the derived geometric attribute set to be coded into the bitstream.

[0328] Furthermore, a plurality of detected geometry attribute sets corresponding to a plurality of pictures may be stored as a plurality of stored geometry attribute sets in the buffer 135. Then, the plurality of stored geometry attribute sets corresponding to a plurality of pictures may be used in deriving the geometry attribute set.

[0329] In yet another example, the detected geometric attribute set obtained from the driving frame may be directly encoded into the bitstream as a derived geometric attribute set, along with concealment parameters indicative of a low confidence score, in which case the above-described process for error concealment is instead performed in the decoding device 200.

[0330] It should be noted that rather than a signal indicating a low reliability score as a concealment parameter, the reliability score itself may be signaled in the bitstream.

[0331] Below, specific examples of (a) temporal smoothing, (b) combining a decoded geometry attribute set with a stored geometry attribute set, and (c) distance correction are given for obtaining a derived geometry attribute set.

[0332] (a) Temporal Smoothing For example, the deriver 132 obtains the derived geometric attribute set by temporally smoothing the detected geometric attribute set based on the detected geometric attribute set and the stored geometric attribute set. Specifically, the deriver 132 may obtain the derived geometric attribute set using first-order exponential smoothing. More specifically, the deriver 132 obtains the derived geometric attribute set by smoothing the detected geometric attribute set based on the stored geometric attribute set. derived = (1-α) × x detected +α×x stored may be used to obtain the derived geometric attribute set.

[0333] Here, x derived refers to the derived geometric attribute set that is encoded into the bitstream, and x detected refers to the detection geometric attribute set obtained from the driving frame, and x stored refers to the stored geometric attribute set obtained from buffer 135.

[0334] Furthermore, α is a smoothing parameter that indicates the level of emphasis given to the stored geometric attribute set. The value of α may be predetermined or may be dynamically determined and coded into the bitstream. For example, the smoothing parameter α may take a value between 0% and 100% and indicates the percentage weight assigned to the stored geometric attribute set.

[0335] For example, x derived refers to the coordinate value of an attribute (landmark) in the derived geometric attribute set, and x detected refers to the coordinate value of the attribute (landmark) in the detection geometric attribute set, and x stored refers to the coordinate values ​​of the attributes (landmarks) in the stored geometric attribute set. The coordinate values ​​for the derived geometric attribute set are calculated by taking a weighted average of the coordinate values ​​for the detected geometric attribute set and the coordinate values ​​for the stored geometric attribute set.

[0336] In the above, a weighted average of one detected geometry attribute set and one stored geometry attribute set is used, but a weighted average of one detected geometry attribute set and multiple stored geometry attribute sets may also be used.

[0337] (b) Combining the Decoded Geometry Attribute Set with the Stored Geometry Attribute Set For example, the deriver 132 obtains the derived geometry attribute set by combining the decoded geometry attribute set with the stored geometry attribute set.

[0338] 35 is a conceptual diagram illustrating an example of an operation in which geometry attributes are replaced in the encoding device 100. Specifically, attributes with low reliability in the detected geometry attribute set are replaced with attributes with high reliability in the stored geometry attribute set. For example, the reliability threshold for replacement is set to 0.7. Therefore, as shown in FIG. 35 , attributes in the detected geometry attribute set with reliability lower than the threshold are replaced with corresponding attributes in the stored geometry attribute set.

[0339] 36 is a conceptual diagram illustrating another example of the operation of replacing geometric attributes in the encoding device 100. For example, the encoding device 100 identifies and deletes unreliable attributes. In the decoding device 200, the missing attributes are replaced with corresponding attributes in the stored geometric attribute set. That is, in the decoding device 200, the missing attributes are filled in with corresponding attributes in the stored geometric attribute set.

[0340] A corresponding attribute may be selected from among the corresponding attributes in the stored geometric attribute sets based on the confidence levels to be applied to the missing attribute.

[0341] The reliability may be determined in units of a geometric attribute set, an attribute, or a group including multiple attributes in a geometric attribute set. Then, a geometric attribute set or an attribute may be selected in units for which the reliability is determined.

[0342] (c) Distance Correction For example, the deriver 132 obtains a derived geometric attribute set by referring to the stored geometric attribute set and correcting the distances between parts in the detected geometric attribute set.

[0343] Specifically, a reference distance set is derived and stored by dividing the memory geometric attribute set into different groups, each representing a key facial feature, deriving the center point of each group, and calculating the relative distance from the derived center point to a reference point (such as the tip of the nose).

[0344] Then, a distance is derived for each group by similarly dividing the set of detected geometric attributes into multiple groups, deriving a center point for each group, and calculating the relative distance from the derived center point to the reference point. If the derived distance differs from the distance in the reference distance set by more than a threshold, the entire group is shifted so that the derived distance matches the distance in the reference distance set.

[0345] Here, the geometric attribute set refers to the geometric attribute set corresponding to the frame. That is, the geometric attribute set refers to the complete geometric attribute set for the entire face. Also, for example, the center point of the group can be derived by averaging the x- and y-coordinates of all attributes (points) in the group. Alternatively, the center point of the group can be determined by deriving a bounding box that surrounds all points in the group based on the minimum and maximum x- and y-values ​​of all attributes (points) in the group, and setting the center of the bounding box as the center point of the group.

[0346] FIG. 37 is a conceptual diagram showing an example of a stored geometric attribute set, which is a geometric attribute set stored in the encoding device 100. In this example, the center point of the "right eye" group in the stored geometric attribute set is (x stored_re , y stored_re ) and the nose point of the memory geometry attribute set is expressed as (x stored_n , y stored_n ) is expressed as

[0347] 38 is a conceptual diagram showing an example of a detected geometry attribute set, which is a geometry attribute set detected by the encoding device 100. In this example, the center point of the "right eye" group in the detected geometry attribute set is represented by (xdetected_re, ydetected_re). Also, the nose tip point in the detected geometry attribute set is represented by (xdetected_n, ydetected_n).

[0348] FIG. 39 is a conceptual diagram showing an example of a derived geometric attribute set, which is a geometric attribute set derived in the encoding device 100. In this example, the center point of the "right eye" group in the derived geometric attribute set is expressed as (xderived_re, yderived_re). In addition, the nose tip point in the derived geometric attribute set is expressed as (x derived_n , y derived_n ) is expressed as

[0349] The method for setting the derived geometric attribute set in FIG. 39 based on the stored geometric attribute set in FIG. 37 and the detected geometric attribute set in FIG. 38 is as follows.

[0350] (1) First, the deriving unit 132 calculates the distance d from the nose tip point to the center point of the right eye group in the stored geometric attribute set. stored_re is calculated according to the following formula:

[0351]

[0352] (2) Next, the deriver 132 calculates the distance ddetected_re from the nose tip point to the center point of the right eye group in the detected geometric attribute set according to the following formula:

[0353]

[0354] (3) Next, the deriving unit 132 calculates |ddetected_red stored_re |> ε, then dderived_re = d stored、re , and (x derived_n , y derived_n )=(xdetected_n, ydetected_n).

[0355] That is, the deriver 132 determines whether the difference between the distance from the nose tip point to the center point of the right eye group in the stored geometric attribute set and the distance from the nose tip point to the center point of the right eye group in the detected geometric attribute set is greater than a threshold value.

[0356] If the difference is greater than the threshold, the deriving unit 132 sets the distance from the nose tip point in the derived geometric attribute set to the center point of the right eye group as the distance from the nose tip point in the stored geometric attribute set to the center point of the right eye group. In this case, the deriving unit 132 also sets the nose tip point in the derived geometric attribute set to the nose tip point in the detected geometric attribute set.

[0357] The deriver 132 calculates |ddetected_red stored_re If |≦ε, the right-eye group in the detected geometric attribute set is set to the right-eye group in the derived geometric attribute set, and the following processes (4), (5) and (6) are skipped.

[0358] (4) Next, the deriving unit 132 converts (x derived_re, y derived_re) into d derived_re = d stored_re and is closest to (xdetected_re, ydetected_re). That is, the deriver 132 sets the center point of the right eye group in the derived geometric attribute set to a point where the distance between the nose tip point and the center point of the right eye group is equal to the distance in the stored geometric attribute set and is closest to the center point of the right eye group in the detected geometric attribute set.

[0359] (5) Next, the deriver 132 derives a transformation to map the center point of the right-eye group (xdetected_re, ydetected_re) in the detected geometric attribute set to the center point of the right-eye group (xderived_re, yderived_re) in the derived geometric attribute set.

[0360] (6) Next, the deriver 132 shifts the entire right-eye group using the same transformation as that derived in process (5) above.

[0361] The deriving unit 132 sets a derived geometric attribute set by performing the above processes (1) to (6) for each of all other groups, thereby adjusting the distance between the groups so that they are not too far apart.

[0362] The stored geometric attribute set may be a geometric attribute set previously coded into the bitstream and stored in the buffer 135. In one example, the stored geometric attribute set may be a geometric attribute set coded and stored for a frame prior to the current frame. In another example, the stored geometric attribute set may be a geometric attribute set coded and stored for the first (intra) frame of a group of pictures.

[0363] In yet another example, the stored geometric attribute set may be a geometric attribute set that exists locally in both the encoding device 100 and the decoding device 200. Specifically, the geometric attribute set may be derived from a common image that exists locally in both the encoding device 100 and the decoding device 200. In yet another example, the stored geometric attribute set may be a predetermined geometric attribute set.

[0364] In yet another example, the stored geometric attribute set may correspond to a portion of a complete geometric attribute set representing an entire face, such as the left eye group in the complete geometric attribute set.

[0365] To prevent memory buffer overflow, the encoding device 100 may retain only the most recently derived geometric attribute sets. Alternatively, the encoding device 100 may retain only geometric attribute sets in the reverse-ordered historical geometric attribute sets that are below a preset threshold and discard less recent geometric attribute sets. This keeps one or more geometric attribute sets in the buffer 135 of the encoding device 100 up to date with the latest changes to the driving frame scene.

[0366] Note that in an example design including a concealment parameter indicating "low_confidence," only geometric attribute sets with high predicted confidence scores are stored in the buffer 135 to serve as reference attributes required for drawing other frames. In this case, geometric attribute sets with low confidence scores are not stored in the buffer 135. This allows only geometric attribute sets with high confidence scores to be used as baselines. This prevents error-prone geometric attribute sets from being stored and resulting in error accumulation.

[0367] The compressor 134 encodes the concealment parameters into the bitstream. The concealment parameters may indicate an error recovery method for the geometric attribute set. For example, not only the concealment parameters but also the geometric attribute set are encoded into the bitstream.

[0368] The geometric attribute set to be encoded may be a detected geometric attribute set, which is a geometric attribute set detected from the driving frame, or may be a derived geometric attribute set, which is a geometric attribute set derived using a stored geometric attribute set.

[0369] The encoding device 100 may encode the detected geometry attribute set as a derived geometry attribute set and may further encode concealment parameters related to error resilience control of the geometry attribute set, which may allow the decoding device 200 to apply appropriate error resilience control to the geometry attribute set based on the concealment parameters.

[0370] Alternatively, the encoding device 100 may encode a derived geometry attribute set obtained using the detected geometry attribute set and the stored geometry attribute set, and further encode the concealment parameters, which may allow the decoding device 200 to identify the validity of the geometry attribute set based on the concealment parameters and to apply more appropriate error recovery control to the geometry attribute set based on the concealment parameters.

[0371] In one example, the concealment parameters may be encoded in a supplementary enhancement information (SEI) message. The SEI message in which the concealment parameters are encoded may be an existing SEI message. Specifically, the SEI message may be a generative face video compression SEI message.

[0372] In another example, the concealment parameters may be encoded into the bitstream by being coded into a separate SEI message.

[0373] Furthermore, the encoding device 100 may perform the same processing as the decoding device 200, or may be provided with a configuration for performing the same processing as the decoding device 200. More specifically, the encoding device 100 may further perform processing (S202, S203, and S204) for generating a facial moving image based on the occlusion parameters.

[0374] For example, the encoding device 100 determines whether the occlusion parameter is equal to a predetermined value. Next, if it is determined that the occlusion parameter is equal to the predetermined value, the encoding device 100 obtains a derived geometric attribute set using the stored geometric attribute set obtained from the buffer 135. Thereafter, the encoding device 100 generates a facial moving image using a neural network based on the derived geometric attribute set.

[0375] The generated facial moving image is used for checking the appearance by the user of the encoding device 100. If the generated facial moving image is not required by the encoding device 100, the encoding device 100 does not need to perform processing to generate a facial moving image based on the occlusion parameters.

[0376] [Syntax and Semantics] Several examples of syntax and corresponding semantics are shown below. Note that in the following description, a face attribute parameter is a parameter corresponding to a geometric attribute set, and the face attribute parameter and the face attribute set may be read as a geometric attribute set.

[0377] 40 is a syntax diagram showing an example syntax structure for concealment parameters. In this example, retrieve_from_buffer equals 0 to indicate that the facial attribute parameters are present in the bitstream. Retrieve_from_buffer equals 1 to indicate that the facial attribute parameters are not present in the bitstream and are retrieved from buffer 237.

[0378] 41 is a syntax diagram illustrating another example syntax structure for concealment parameters. In this example, attribute_present indicates that the facial attribute parameters are present in the bitstream when equal to 1. Also, attribute_present indicates that the facial attribute parameters are not present in the bitstream and are obtained from buffer 237 when equal to 0.

[0379] 42 is a syntax diagram illustrating yet another example syntax structure for concealment parameters. In this example, attribute_present indicates that the facial attribute parameters are present in the bitstream when equal to 1. Also, attribute_present indicates that the facial attribute parameters are not present in the bitstream and are obtained from buffer 237 when equal to 0.

[0380] Furthermore, max_buffer_idx indicates the maximum number of face attribute sets to be stored. The face attribute sets to be stored are face attribute sets that have been previously encoded or decoded. Furthermore, buffer_idx specifies the index of a face attribute set to be acquired from the face attribute sets, within the range from 0 to max_buffer_idx minus 1.

[0381] 43 is a syntax diagram illustrating yet another example syntax structure for concealment parameters. In this example, attribute_present indicates that the facial attribute parameters are present in the bitstream when equal to 1. Also, attribute_present indicates that the facial attribute parameters are not present in the bitstream and are obtained from buffer 237 when equal to 0.

[0382] Furthermore, store_to_buffer, when equal to 1, indicates that the facial attribute set of the bitstream is stored in the buffer 237. On the other hand, store_to_buffer, when equal to 0, indicates that the facial attribute set of the bitstream is not stored in the buffer 237.

[0383] Furthermore, max_buffer_idx indicates the maximum number of face attribute sets to be stored. The face attribute sets to be stored are face attribute sets that have been previously encoded or decoded. Furthermore, buffer_idx indicates the index of a face attribute set to be acquired from the face attribute sets, ranging from 0 to max_buffer_idx minus 1.

[0384] Also, when max_buffer_idx is greater than 1, store_to_buffer, when greater than 0, may indicate a buffer index position where the face attribute set is stored in the buffer 237. For example, when store_to_buffer is greater than 0, the face attribute set is stored at the buffer index position indicated by store_to_buffer minus 1. Also, when store_to_buffer is equal to 0, the face attribute set is not stored.

[0385] 44 is a syntax diagram illustrating yet another example syntax structure for concealment parameters. In this example, attribute_present indicates that the facial attribute parameters are present in the bitstream when equal to 1. Also, attribute_present indicates that the facial attribute parameters are not present in the bitstream and are obtained from buffer 237 when equal to 0.

[0386] Furthermore, store_to_buffer, when equal to 1, indicates that the bitstream facial attribute set is stored in the buffer 237. On the other hand, store_to_buffer, when equal to 0, indicates that the bitstream facial attribute set is not stored in the buffer 237.

[0387] Also, store_buffer_idx indicates the buffer index position where the face attribute set of the bitstream is stored in the buffer 237 .

[0388] Furthermore, max_buffer_idx indicates the maximum number of face attribute sets to be stored. The face attribute sets to be stored are face attribute sets that have been previously encoded or decoded. Furthermore, buffer_idx indicates the index of a face attribute set to be acquired from the face attribute sets, ranging from 0 to max_buffer_idx minus 1.

[0389] FIG. 45 is a syntax diagram illustrating yet another example syntax structure for concealment parameters.

[0390] In this example, when attributes_low_confidence is equal to 0, it indicates that the facial attribute parameters of the bitstream have high confidence and should be stored in the buffer 237. When attribute_low_confidence is equal to 1, it indicates that the facial attribute parameters of the bitstream have low confidence and should be corrected by a method such as combining with a stored facial attribute set obtained from the buffer 237.

[0391] Here, a method for combining the stored face attribute set and the less reliable decoded face attribute set to conceal errors may be predefined in the decoding device 200 .

[0392] FIG. 46 is a syntax diagram illustrating yet another example syntax structure for concealment parameters.

[0393] In this example, when attributes_low_confidence is equal to 0, it indicates that the facial attribute parameters of the bitstream have high confidence and should be stored in buffer 237. Also, when attribute_low_confidence is not equal to 0, it indicates that the facial attribute parameters of the bitstream have low confidence and should be combined with the stored facial attribute set retrieved from buffer 237.

[0394] Here, the method for combining the stored face attribute set and the low-reliability decoded face attribute set to conceal errors may be defined by the value of attribute_low_confidence.

[0395] Furthermore, smoothing_parameter indicates the degree of emphasis given to the stored face attribute set for performing temporal smoothing by first-order exponential smoothing when the value of attribute_low_confidence is 1.

[0396] FIG. 47 is a syntax diagram illustrating yet another example syntax structure for concealment parameters.

[0397] In this example, num_attributes_not_equal_expected equals 0 to indicate that the number of attributes in the facial attribute parameters of the bitstream is equal to the expected number of facial attributes, and num_attributes_not_equal_expected equals 1 to indicate that the number of attributes in the facial attribute parameters of the bitstream is not equal to the expected number of facial attributes, implying that missing or additional attributes may be present.

[0398] Figure 48 is a syntax diagram showing yet another example syntax structure for concealment parameters. In this example, max_num_attributes refers to the maximum number of attributes for each frame in the bitstream. Also, num_attributes indicates the number of attributes for the current frame in the bitstream. If num_attributes is not equal to max_num_attributes, it indicates that the number of attributes does not match the expected number.

[0399] FIG. 49 is a syntax diagram illustrating yet another example syntax structure for concealment parameters.

[0400] In this example, looping_around_stored_attributes, when equal to 0, indicates that there is no need to loop, store to buffer 237, and retrieve from buffer 237. Therefore, the facial attribute parameters in the bitstream are used directly as the derived facial attribute set.

[0401] Furthermore, when looping_around_stored_attributes is equal to 1, it indicates the start of writing a facial attribute set to the buffer 237. Therefore, from this frame onwards, facial attribute parameters of the bitstream are stored in the buffer 237 until the value of looping_around_stored_attributes changes.

[0402] Furthermore, when looping_around_stored_attributes is equal to 2, it indicates the end of writing face attribute sets to the buffer 237. From this frame onwards, until the value of looping_around_stored_attributes changes, the decoding device 200 repeatedly retrieves face attribute sets from the buffer 237, from the first face attribute set stored to the last face attribute set stored, and all face attribute parameters in the bitstream are ignored.

[0403] Furthermore, max_buffer_idx indicates the maximum number of face attribute sets to be stored. Furthermore, store_buffer_idx indicates the buffer index position at which the face attribute sets are stored, within the range from 0 to max_buffer_idx minus 1. Furthermore, buffer_idx indicates the buffer index position of the face attribute set to be acquired, within the range from 0 to max_buffer_idx minus 1.

[0404] The buffer index position for storing or retrieving the face attribute set may be controlled using only an internal counter, without relying on store_buffer_idx and buffer_idx, and therefore the signaling of store_buffer_idx and buffer_idx may be omitted.

[0405] In the above examples, the operation of the decoding device 200 corresponding to the syntax is shown, but the encoding device 100 may perform the same operation as the decoding device 200.

[0406] Also, in the above examples, a memory geometric attribute set is used in error concealment. However, the decoding device 200 may not use a memory geometric attribute set for error concealment. For example, the decoding device 200 may apply a previous frame in a facial video sequence to a current frame in the facial video sequence. In other words, the decoding device 200 may apply an already generated facial image to a facial image corresponding to the current frame, without using a generative model to generate a facial image corresponding to the current frame in the facial video sequence.

[0407] [Examples of Generative Models] Fig. 50 is a diagram showing examples of various models that can be used as generative models. For example, neural networks can be used as generative models. Specifically, Fig. 50 shows a generative adversarial network, a variational autoencoder, a flow-based generative model, and a diffusion model.

[0408] In a generative adversarial network, new data instances similar to input data are generated through learning the features of the input data. Specifically, the unsupervised task in the generative model is converted into a supervised task by two types of submodels.

[0409] For example, the generator sub-model generates fake samples, the classifier sub-model discriminates between real inputs and fake samples generated by the generator sub-model, and an output image is generated through a minimax game that maximizes the classification probability of the classifier sub-model in assigning correct labels to the real inputs and fake samples while minimizing the distribution difference between the real inputs and fake samples.

[0410] In a variational autoencoder, the input data is first compressed into a multivariate latent distribution that reconstructs the data from the latent space as accurately as possible, resulting in efficient data compression and dimensionality reduction. In a flow-based generative model, the source distribution is transformed into the training data distribution through a sequence of one or more reversible transformations, which allows learning of the data distribution and accurate computation of the end-goal likelihood.

[0411] Diffusion models also generate new data instances similar to the training data. First, they degrade the structure of the training data through repeated injection of perturbations and noise before initiating a denoising process in an attempt to recover the original data. As a result, the data is iteratively mapped to a latent distribution via a Markov chain, where the latent state at each step depends only on the latent state at the previous step. The data is then recovered through hierarchical denoising.

[0412] For example, the neural network may be a face picture generation neural network that can be used to generate an output picture using geometric information represented in a fixed format for face parameters and a picture, i.e., the neural network corresponds to a process of generating a plurality of sample values ​​that constitute an output picture, which is a picture included in an output video sequence.

[0413] Alternatives to the neural networks described above may be any combination of the neural networks described above, other types of generative models, etc. may also be used.

[0414] Furthermore, the generative model may be used to detect or derive geometric attributes, or may be used to detect or derive basic attributes.

[0415] [Configuration Example for Encoding and Decoding Video] Figure 51 is a block diagram showing a configuration example for encoding a video by the encoding device 100 in this embodiment. For example, the encoding device 100 may include the multiple components shown in Figure 51 as multiple components for encoding images included in a video block by block in accordance with VVC. The encoding device 100 may include the multiple components shown in Figure 51 in addition to the multiple components described above, or at least some of the multiple components described above may be integrated into the multiple components shown in Figure 51.

[0416] 51 , the encoding device 100 includes a dividing unit 102, a subtraction unit 104, a transform unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transform unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, a prediction control unit 128, and a prediction parameter generation unit 130. Note that the intra prediction unit 124 and the inter prediction unit 126 are each configured as part of a prediction processing unit.

[0417] The partitioning unit 102 partitions an image into multiple blocks and provides parameters related to the partitioning to the entropy coding unit 110. The subtraction unit 104 subtracts a predicted image block from a current block to obtain a predicted residual block. The transformation unit 106 performs transformation on the predicted residual block to obtain a transform coefficient block. The quantization unit 108 performs quantization on the transform coefficient block to obtain a quantized coefficient block. The entropy coding unit 110 performs entropy coding on the quantized coefficient block and the parameters to generate a bitstream.

[0418] The inverse quantization unit 112 performs inverse quantization on the quantized coefficient block to obtain a transform coefficient block. The inverse transform unit 114 performs inverse transform on the transform coefficient block to obtain a prediction residual block. The adder 116 adds the prediction image block to the prediction residual block to obtain a reconstructed image block. The block memory 118 stores the reconstructed image block. The loop filter unit 120 applies a loop filter to the reconstructed image block. The frame memory 122 stores the reconstructed image block to which the loop filter has been applied.

[0419] The intra prediction unit 124 performs intra prediction with reference to the block memory 118 to generate a predicted image block. The inter prediction unit 126 performs inter prediction with reference to the frame memory 122 to generate a predicted image block. The prediction control unit 128 provides the predicted image block generated by the intra prediction unit 124 or the predicted image block generated by the inter prediction unit 126 to the subtraction unit 104 and the addition unit 116. The prediction parameter generation unit 130 provides parameters related to intra prediction or inter prediction to the entropy coding unit 110.

[0420] Fig. 52 is a block diagram showing an example configuration for the decoding device 200 in this embodiment to decode a moving image. For example, the decoding device 200 may include the multiple components shown in Fig. 52 as multiple components for decoding images included in a moving image on a block-by-block basis in accordance with VVC. The decoding device 200 may include the multiple components shown in Fig. 52 in addition to the multiple components described above, or at least some of the multiple components described above may be integrated into the multiple components shown in Fig. 52.

[0421] 52, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an adder 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra prediction unit 216, an inter prediction unit 218, a prediction control unit 220, a prediction parameter generation unit 222, and a partition determination unit 224. Note that the intra prediction unit 216 and the inter prediction unit 218 are each configured as part of a prediction processing unit.

[0422] The entropy decoding unit 202 performs entropy decoding on the bitstream to obtain a quantized coefficient block and parameters. The inverse quantization unit 204 performs inverse quantization on the quantized coefficient block to obtain a transform coefficient block. The inverse transform unit 206 performs inverse transform on the transform coefficient block to obtain a prediction residual block. The adder 208 adds the prediction image block to the prediction residual block to obtain a reconstructed image block. The loop filter unit 212 applies a loop filter to the reconstructed image block.

[0423] The block memory 210 stores reconstructed image blocks, and the frame memory 214 stores reconstructed image blocks to which the loop filter has been applied.

[0424] The intra prediction unit 216 performs intra prediction by referring to the block memory 210 to generate a predicted image block. The inter prediction unit 218 performs inter prediction by referring to the frame memory 214 to generate a predicted image block. The prediction control unit 220 provides the predicted image block generated by the intra prediction unit 216 or the predicted image block generated by the inter prediction unit 218 to the adder 208. The prediction parameter generation unit 222 provides parameters related to intra prediction or inter prediction to the prediction control unit 220, etc.

[0425] The division determination unit 224 determines blocks for decoding the image block by block in accordance with the division parameters.

[0426] [Combination] Multiple configuration examples in the present disclosure may be combined in any manner. Furthermore, multiple operation examples in the present disclosure may be combined in any manner. Furthermore, overlapping descriptions may be omitted in multiple examples in the present disclosure. Furthermore, configurations and processes corresponding to encoding configurations and processes may be applied to decoding, and configurations and processes corresponding to decoding configurations and processes may be applied to encoding. Furthermore, only a portion of the examples included in multiple examples in the present disclosure may be implemented.

[0427] 53 is a block diagram showing an implementation example of the encoding device 100. The encoding device 100 includes a circuit 151 and a memory 152. For example, several components of the encoding device 100 described above are implemented by the circuit 151 and the memory 152.

[0428] The circuit 151 is an electric circuit that performs information processing and can access the memory 152. For example, the circuit 151 may be a dedicated circuit that executes the encoding method of the present disclosure, or may be a general-purpose circuit that executes a program corresponding to the encoding method of the present disclosure. The circuit 151 may also be a processor such as a CPU. Furthermore, the circuit 151 may be a collection of multiple circuits.

[0429] The memory 152 is a dedicated or general-purpose memory that stores information used by the circuit 151 to encode an image. The memory 152 may be an electric circuit and may be connected to the circuit 151. The memory 152 may also be included in the circuit 151. The memory 152 may also be a collection of multiple circuits. The memory 152 may also be a magnetic disk, an optical disk, or the like, and may also be expressed as a storage, a recording medium, or the like. The memory 152 may also be a non-volatile memory or a volatile memory.

[0430] For example, the memory 152 may store data to be coded, such as an image, or may store coded data, such as a bit stream. The memory 152 may also store a program for causing the circuit 151 to perform image processing. The memory 152 may also store a generative model for the circuit 151.

[0431] Also, for example, the memory 152 may correspond to the above-described buffer 135, block memory 118, frame memory 122, etc. And the circuit 151 may correspond to multiple other components in the encoding device 100.

[0432] 54 is a flowchart showing a first basic operation example of the encoding device 100. In the operation of this example, the circuit 151 of the encoding device 100 uses the memory 152 to perform the following.

[0433] Specifically, the circuit 151 encodes, into a bitstream, base data of facial images related to the facial video sequence and geometric information, which corresponds to each of multiple frames of the facial video sequence and indicates geometric attributes within an area including a person's face (S601). The circuit 151 also determines whether the geometric information has been properly acquired (S602). The circuit 151 then further encodes, into the bitstream, concealment parameters related to error recovery control in the decoding device 200 when the geometric information has not been properly acquired (S603).

[0434] This may enable the encoding device 100 to notify the decoding device 200 of information regarding error recovery control when the geometry information is not properly acquired in the encoding device 100. Therefore, when the geometry information is not properly acquired in the encoding device 100, it may become possible for the decoding device 200 to properly perform error recovery control.

[0435] Here, the base data corresponds to a reference image or a reference attribute, and the geometric information corresponds to a geometric attribute set. Whether the geometric information has been properly acquired corresponds, for example, to whether the geometric information has been acquired according to a predetermined criterion corresponding to proper acquisition of the geometric information, more specifically, whether the geometric information conforms to a predetermined condition.

[0436] For example, the circuit 151 may encode the concealment parameters into a header region of the bitstream. This may allow the encoding device 100 to notify the decoding device 200 of information regarding error recovery control when the encoding device 100 has not properly acquired the geometric information via the header region of the bitstream. Therefore, it may be possible to properly perform error recovery control based on the information in the header region.

[0437] Furthermore, for example, the concealment parameter may indicate, depending on the value of the concealment parameter, that the reliability of the geometric information corresponding to the current frame obtained by the encoding device 100 is low. This may enable the encoding device 100 to notify the decoding device 200 that the reliability of the geometric information obtained by the encoding device 100 is low. Therefore, it may be possible to appropriately perform error recovery control in the decoding device 200.

[0438] Here, a low reliability of the geometric information may correspond to a reliability of the geometric information being lower than a threshold value, for example, which may be an average reliability or an arbitrarily determined reliability.

[0439] Alternatively, a low reliability of the geometric information may correspond to the geometric information not conforming to a predetermined condition, or the geometric information not being acquired in accordance with a predetermined criterion, etc. Furthermore, if the geometric information does not conform to a predetermined condition, or the geometric information not being acquired in accordance with a predetermined criterion, the reliability of the geometric information may be considered to be lower than a threshold value.

[0440] Furthermore, for example, the concealment parameter may indicate, depending on the value of the concealment parameter, that the encoding device 100 has not acquired geometric information corresponding to the current frame. This may enable the encoding device 100 to notify the decoding device 200 that the encoding device 100 has not acquired geometric information. This may enable the decoding device 200 to appropriately perform error recovery control.

[0441] Furthermore, for example, the concealment parameter may indicate, depending on the value of the concealment parameter, that the number of pieces of geometric information corresponding to the current frame acquired by the encoding device 100 is different from the original number. This may enable the encoding device 100 to notify the decoding device 200 that the number of attributes in the geometric information acquired by the encoding device 100 is inappropriate. This may enable the decoding device 200 to perform appropriate error recovery control.

[0442] Furthermore, for example, the concealment parameter may indicate, depending on the value of the concealment parameter, that stored geometric information is to be applied to the current frame of the facial video image instead of geometric information decoded from the bitstream. This may enable the encoding device 100 to notify the decoding device 200 that the stored geometric information is to be applied to the current frame. Therefore, it may be possible to appropriately perform error recovery control in the decoding device 200.

[0443] Furthermore, for example, the concealment parameter may indicate that geometric information corresponding to the current frame is to be stored for a subsequent frame depending on the value of the concealment parameter, which may allow the encoding device 100 to notify the decoding device 200 that geometric information corresponding to the current frame is to be stored, thereby making it possible to appropriately perform error recovery control for the subsequent frame.

[0444] Furthermore, for example, there may be cases where the reliability of the geometric information corresponding to the current frame obtained by the encoding device 100 is lower than a threshold. In this case, the circuit 151 may set the concealment parameter to a value indicating that the reliability of the geometric information is low. This may allow the encoding device 100 to notify the decoding device 200 that the reliability of the geometric information is low when the reliability of the geometric information is lower than the threshold. Therefore, there may be cases where it is possible to appropriately perform error recovery control in the decoding device 200.

[0445] Furthermore, for example, there may be cases where the reliability of the geometric information corresponding to the current frame acquired by the encoding device 100 is lower than a threshold, or where the current frame acquired by the encoding device 100 does not include a face. In this case, the circuit 151 may set the concealment parameter to a value indicating that no geometric information was acquired, without encoding the geometric information corresponding to the current frame into the bitstream.

[0446] This may enable the encoding device 100 to notify the decoding device 200 that the geometric information was not acquired when the reliability of the geometric information is lower than a threshold or when a face is not included. Therefore, it may become possible for the decoding device 200 to appropriately perform error recovery control.

[0447] Furthermore, for example, there may be cases where the number of attributes in the geometric information corresponding to the current frame, acquired by the encoding device 100, differs from the predefined number of attributes. In this case, the circuit 151 may set the concealment parameter to a value indicating that the number of attributes in the geometric information differs from the original number.

[0448] This may enable the encoding device 100 to notify the decoding device 200 that the number of attributes in the geometric information is different from the original number when the number of attributes in the geometric information is different from the predefined number of attributes. This may enable the decoding device 200 to appropriately perform error recovery control.

[0449] Furthermore, for example, there may be cases where the reliability of the geometric information corresponding to the current frame acquired by the encoding device 100 is lower than a threshold, or where the current frame acquired by the encoding device 100 does not include a face. In this case, the circuit 151 may set the concealment parameter to a value indicating that the stored geometric information is to be applied to the geometric information corresponding to the current frame.

[0450] This may enable the encoding device 100 to notify the decoding device 200 that the stored geometric information should be applied to the current frame when the reliability of the geometric information is lower than a threshold or when no face is included, which may enable the decoding device 200 to appropriately perform error recovery control.

[0451] Also, for example, the reliability of the geometric information corresponding to the current frame obtained by the encoding device 100 may be equal to or greater than a certain threshold. In this case, the circuit 151 may set the concealment parameter to a value indicating that the geometric information is to be stored for the subsequent frame.

[0452] This may enable the encoding device 100 to notify the decoding device 200 that the geometric information will be stored when the reliability of the geometric information is equal to or greater than a threshold value. This may enable the decoding device 200 to appropriately perform error recovery control for subsequent frames.

[0453] Also, for example, there may be cases where the reliability of the geometric information corresponding to the current frame obtained by the encoding device 100 is lower than a threshold. In this case, the circuit 151 may correct the obtained geometric information using the stored geometric information. Then, the circuit 151 may encode the corrected geometric information into a bitstream.

[0454] As a result, when the encoding device 100 does not properly acquire geometric information, it may be possible to correct the improperly acquired geometric information using the stored geometric information. Therefore, it may be possible to properly perform error recovery control in the encoding device 100.

[0455] Furthermore, for example, there may be cases where the reliability of the geometric information corresponding to the current frame obtained by the encoding device 100 is lower than a threshold value. In this case, the circuit 151 may encode the stored geometric information into a bitstream as the geometric information corresponding to the current frame.

[0456] This may allow the encoding device 100 to use stored geometric information in place of the geometric information that was not properly acquired when the geometric information was not properly acquired, thereby enabling the encoding device 100 to perform error recovery control appropriately.

[0457] Furthermore, for example, the stored geometric information may be geometric information corresponding to a past frame acquired by the encoding device 100. This may make it possible to appropriately perform error recovery control using the geometric information corresponding to the past frame.

[0458] Furthermore, for example, the stored geometric information may be geometric information of a predefined standard, which may enable appropriate error recovery control using the predefined standard geometric information.

[0459] 55 is a flowchart showing a second basic operation example of the encoding device 100. In the operation of this example, the circuit 151 of the encoding device 100 uses the memory 152 to perform the following.

[0460] Specifically, the circuit 151 encodes base data of an image included in the video (S611). The circuit 151 also encodes face attribute parameters, which indicate faces included in the image, into a bitstream (S612). The circuit 151 also encodes reliability parameters, which relate to the reliability of the face attribute parameters, into the bitstream (S613). Here, the reliability parameters indicate low reliability depending on the value of the reliability parameters.

[0461] For example, if the reliability of a face attribute parameter is unknown in the decoding device 200, the decoding device 200 may estimate the reliability and perform processing, which may result in the generation of an inappropriate image. With the above configuration, if the reliability of a face attribute parameter is low, it may be possible for the encoding device 100 to notify the decoding device 200 that the reliability of the face attribute parameter is low.

[0462] Therefore, it may be possible to appropriately determine whether the reliability of the face attribute parameter is low in the decoding device 200. Furthermore, it may be possible for the encoding device 100 to control the determination by the decoding device 200 as to whether the reliability of the face attribute parameter is low or not.

[0463] Here, the base data corresponds to a reference image or a reference attribute, the facial attribute parameter corresponds to a geometric attribute set, and the reliability parameter corresponds to an occlusion parameter. If the reliability parameter does not indicate that the facial attribute parameter has low reliability, the reliability of the facial attribute parameter may be high or indefinite.

[0464] For example, the image may correspond to each of a plurality of pictures included in a video, the base data may be data common to a plurality of pictures, and the bitstream may include a reliability parameter and a face attribute parameter for each of the plurality of pictures.

[0465] This may enable the encoding device 100 to notify the decoding device 200 that the reliability of the face attribute parameter is low for each picture, if the reliability of the face attribute parameter is low. Therefore, the decoding device 200 may be able to appropriately determine that the reliability of the face attribute parameter is low for each picture.

[0466] Furthermore, for example, the bitstream may include a reliability parameter before the facial attribute parameter for each of a plurality of pictures. This may enable the encoding device 100 to notify the decoding device 200 of the reliability of the facial attribute parameter before notifying the facial attribute parameter. Therefore, the decoding device 200 may be able to process the facial attribute parameter after appropriately determining that the reliability of the facial attribute parameter is low.

[0467] Furthermore, for example, if the reliability parameter does not indicate that the reliability of the face attribute parameter is low, the face attribute parameter may be stored. This may make it possible to suppress the storage of face attribute parameters with low reliability. Therefore, it may make it possible to suppress the reuse of face attribute parameters with low reliability.

[0468] Furthermore, the encoding device 100 may include an input terminal, an entropy encoder, and an output terminal. The operations performed by the circuit 151 may be performed by the entropy encoder. Data used in the operation of the entropy encoder may be input to the input terminal. Data obtained by the operation of the entropy encoder may be output from the output terminal.

[0469] 56 is a block diagram showing an example implementation of the decoding device 200. The decoding device 200 includes a circuit 251 and a memory 252. For example, several components of the decoding device 200 described above are implemented by the circuit 251 and the memory 252.

[0470] The circuit 251 is an electric circuit that performs information processing and can access the memory 252. For example, the circuit 251 may be a dedicated circuit that executes the decoding method of the present disclosure, or may be a general-purpose circuit that executes a program corresponding to the decoding method of the present disclosure. The circuit 251 may also be a processor such as a CPU. Furthermore, the circuit 251 may be a collection of multiple circuits.

[0471] The memory 252 is a dedicated or general-purpose memory that stores information for the circuit 251 to decode images. The memory 252 may be an electric circuit and may be connected to the circuit 251. The memory 252 may also be included in the circuit 251. The memory 252 may also be a collection of multiple circuits. The memory 252 may also be a magnetic disk, an optical disk, or the like, and may also be expressed as a storage, a recording medium, or the like. The memory 252 may also be a non-volatile memory or a volatile memory.

[0472] For example, the memory 252 may store data to be decoded, such as a bitstream, or may store decoded data, such as an image. The memory 252 may also store a program for causing the circuit 251 to perform image processing. The memory 252 may also store a generative model for the circuit 251.

[0473] Also, for example, the memory 252 may correspond to the above-described buffer 237, block memory 210, frame memory 214, etc. And the circuit 251 may correspond to multiple other components in the decoding device 200.

[0474] 57 is a flowchart showing a first basic operation example of the decoding device 200. In the operation of this example, the circuit 251 of the decoding device 200 uses the memory 252 to perform the following.

[0475] Specifically, the circuit 251 decodes, from the bitstream, base data of facial images related to the facial video and geometric information, which corresponds to each of a plurality of frames of the facial video and indicates geometric attributes within an area including a person's face (S701).The circuit 251 also decodes, from the bitstream, concealment parameters related to error recovery control in the case where the geometric information is not properly acquired by the encoding device 100 (S702).

[0476] Then, the circuit 251 generates a facial motion image from the base data, geometric information, and occlusion parameters using a generative model (S703).

[0477] This may enable the encoding device 100 to notify the decoding device 200 of information regarding error recovery control when the geometry information is not properly acquired in the encoding device 100. Therefore, when the geometry information is not properly acquired in the encoding device 100, it may become possible for the decoding device 200 to properly perform error recovery control.

[0478] Here, the base data corresponds to a reference image or a reference attribute, and the geometric information corresponds to a geometric attribute set. Whether the geometric information has been properly acquired corresponds, for example, to whether the geometric information has been acquired according to a predetermined criterion corresponding to proper acquisition of the geometric information, more specifically, whether the geometric information conforms to a predetermined condition.

[0479] For example, the circuit 251 may decode the concealment parameters from a header region in the bitstream. This may enable the encoding device 100 to notify the decoding device 200 of information regarding error recovery control when the encoding device 100 has not properly acquired geometry information via the header region in the bitstream. Therefore, it may be possible for the decoding device 200 to properly perform error recovery control based on the information in the header region.

[0480] Furthermore, for example, the concealment parameter may indicate, depending on the value of the concealment parameter, that the reliability of the geometric information corresponding to the current frame obtained by the encoding device 100 is low. This may enable the encoding device 100 to notify the decoding device 200 that the reliability of the geometric information obtained by the encoding device 100 is low. Therefore, it may be possible to appropriately perform error recovery control in the decoding device 200.

[0481] Here, a low reliability of the geometric information may correspond to a reliability of the geometric information being lower than a threshold value, for example, which may be an average reliability or an arbitrarily determined reliability.

[0482] Alternatively, a low reliability of the geometric information may correspond to the geometric information not conforming to a predetermined condition, or the geometric information not being acquired in accordance with a predetermined criterion, etc. Furthermore, if the geometric information does not conform to a predetermined condition, or the geometric information not being acquired in accordance with a predetermined criterion, the reliability of the geometric information may be considered to be lower than a threshold value.

[0483] Furthermore, for example, the concealment parameter may indicate, depending on the value of the concealment parameter, that the encoding device 100 has not acquired geometric information corresponding to the current frame. This may enable the encoding device 100 to notify the decoding device 200 that the encoding device 100 has not acquired geometric information. This may enable the decoding device 200 to appropriately perform error recovery control.

[0484] Furthermore, for example, the concealment parameter may indicate, depending on the value of the concealment parameter, that the number of attributes in the geometric information corresponding to the current frame acquired by the encoding device 100 is different from the original number of attributes. This may enable the encoding device 100 to notify the decoding device 200 that the number of attributes in the geometric information acquired by the encoding device 100 is inappropriate. This may enable the decoding device 200 to perform appropriate error recovery control.

[0485] Furthermore, for example, the concealment parameter may indicate, depending on the value of the concealment parameter, that stored geometric information is to be applied to the current frame of the facial video image instead of geometric information decoded from the bitstream. This may enable the encoding device 100 to notify the decoding device 200 that the stored geometric information is to be applied to the current frame. Therefore, it may be possible to appropriately perform error recovery control in the decoding device 200.

[0486] Furthermore, for example, the concealment parameter may indicate that geometric information corresponding to the current frame is to be stored for a subsequent frame depending on the value of the concealment parameter, which may allow the encoding device 100 to notify the decoding device 200 that geometric information corresponding to the current frame is to be stored, thereby enabling the decoding device 200 to appropriately perform error recovery control for the subsequent frame.

[0487] Also, for example, the concealment parameters may indicate that the geometric information corresponding to the current frame obtained by the encoding device 100 is of low reliability, or that the number of attributes in the geometric information is different from the original number of attributes.

[0488] In this case, the circuit 251 may correct geometric information corresponding to the current frame decoded from the bitstream using the stored geometric information, and then apply the corrected geometric information to geometric information for generating a facial image corresponding to the current frame in the facial video sequence.

[0489] As a result, if the encoding device 100 does not properly acquire the geometric information, it may be possible to correct the improperly acquired geometric information using the stored geometric information, which may enable appropriate error recovery control.

[0490] Furthermore, for example, the concealment parameters may indicate first information, second information, third information, or fourth information. Here, the first information indicates that the geometric information corresponding to the current frame, acquired by the encoding device 100, has low reliability. The second information indicates that the geometric information was not acquired. The third information indicates that the number of attributes in the geometric information is different from the original number of attributes. The fourth information indicates that stored geometric information is applied to the geometric information corresponding to the current frame.

[0491] In this case, the circuit 251 may apply the stored geometric information to the geometric information for generating the facial image corresponding to the current frame in the facial motion picture.

[0492] As a result, when the geometry information is not properly acquired in the encoding device 100, it may be possible to use the stored geometry information instead of the geometry information that was not properly acquired, which may make it possible to perform error recovery control appropriately.

[0493] Furthermore, for example, the concealment parameters may indicate first information, second information, or third information. Here, the first information indicates that the geometric information corresponding to the current frame, which is acquired by the encoding device 100, has low reliability. The second information indicates that the geometric information has not been acquired. The third information indicates that the number of attributes in the geometric information is different from the original number of attributes.

[0494] In this case, instead of using a generative model to generate a facial image corresponding to the current frame in the facial video sequence, an already generated facial image may be applied to the facial image corresponding to the current frame.

[0495] This may enable the encoding device 100 to use an already generated facial image for the current frame instead of generating a new one when the geometric information is not properly acquired. This may enable the encoding device 100 to properly perform error recovery control.

[0496] Furthermore, for example, the stored geometric information may be geometric information corresponding to a past frame decoded from a bitstream, which may enable appropriate error recovery control using the geometric information corresponding to the past frame.

[0497] Furthermore, for example, the stored geometric information may be geometric information of a predefined standard, which may enable appropriate error recovery control using the predefined standard geometric information.

[0498] 58 is a flowchart showing a second basic operation example of the decoding device 200. In the operation of this example, the circuit 251 of the decoding device 200 uses the memory 252 to perform the following.

[0499] Specifically, the circuit 251 acquires base data of an image included in the video (S711). The circuit 251 also decodes facial attribute parameters indicating faces included in the image from the bitstream (S712). The base data and facial attribute parameters are input to a generative model to generate an output image corresponding to the image. The bitstream includes a reliability parameter related to the reliability of the facial attribute parameter. The reliability parameter indicates low reliability depending on the value of the reliability parameter.

[0500] This may enable the encoding device 100 to notify the decoding device 200 that the reliability of the face attribute parameter is low when the reliability of the face attribute parameter is low. Therefore, the decoding device 200 may be able to appropriately determine that the reliability of the face attribute parameter is low.

[0501] Here, the base data corresponds to a reference image or a reference attribute, the facial attribute parameter corresponds to a geometric attribute set, and the reliability parameter corresponds to an occlusion parameter. If the reliability parameter does not indicate that the facial attribute parameter has low reliability, the reliability of the facial attribute parameter may be high or indefinite.

[0502] For example, the image may correspond to each of a plurality of pictures included in a video, the base data may be data common to a plurality of pictures, and the bitstream may include a reliability parameter and a face attribute parameter for each of the plurality of pictures.

[0503] This may enable the encoding device 100 to notify the decoding device 200 that the reliability of the face attribute parameter is low for each picture, if the reliability of the face attribute parameter is low. Therefore, the decoding device 200 may be able to appropriately determine that the reliability of the face attribute parameter is low for each picture.

[0504] Furthermore, for example, the bitstream may include a reliability parameter before the facial attribute parameter for each of a plurality of pictures. This may enable the encoding device 100 to notify the decoding device 200 of the reliability of the facial attribute parameter before notifying the facial attribute parameter. Therefore, the decoding device 200 may be able to process the facial attribute parameter after appropriately determining that the reliability of the facial attribute parameter is low.

[0505] Furthermore, for example, if the reliability parameter does not indicate that the reliability of the face attribute parameter is low, the face attribute parameter may be stored. This may make it possible to suppress the storage of face attribute parameters with low reliability. Therefore, it may make it possible to suppress the reuse of face attribute parameters with low reliability.

[0506] Furthermore, for example, the decoding device 200 may include an input terminal, an entropy decoder, and an output terminal. The operations performed by the circuit 251 may be performed by the entropy decoder. Data used in the operation of the entropy decoder may be input to the input terminal. Data obtained by the operation of the entropy decoder may be output from the output terminal.

[0507] [Other Examples] The encoding device 100 and the decoding device 200 in each of the above-described examples may be used as an image encoding device and an image decoding device, or as a video encoding device and a video decoding device, respectively. Furthermore, multiple components included in the encoding device 100 and multiple components included in the decoding device 200 may perform corresponding operations.

[0508] Additionally, the term "encode" may be substituted with terms such as "store," "include," "write," "write," "signal," "send," "notify," or "preserve," and these terms may be interchangeable. For example, encoding information may mean including the information in a bitstream. Also, encoding information into a bitstream may mean encoding the information to generate a bitstream that includes the encoded information.

[0509] Furthermore, the term "decode" may be replaced with terms such as "read," "decode," "read," "load," "derive," "obtain," "receive," "extract," or "restore," and these terms may be interchangeable. For example, decoding information may mean obtaining information from a bitstream. Decoding information from a bitstream may mean decoding the bitstream to obtain information contained in the bitstream.

[0510] Furthermore, for example, the coding information and compression information contained in the bitstream may be simply referred to as information.

[0511] Furthermore, at least some of the examples described above may be used as an encoding method, a decoding method, an entropy encoding method, an entropy decoding method, or some other method.

[0512] Each component may be configured with dedicated hardware, or may be realized by executing a software program suitable for that component. Each component may be realized by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.

[0513] Specifically, each of the encoding device 100 and the decoding device 200 may include a processing circuit and a storage device electrically connected to and accessible from the processing circuit. For example, the processing circuit corresponds to the circuit 151 or 251, and the storage device corresponds to the memory 152 or 252.

[0514] The processing circuit includes at least one of dedicated hardware and a program execution unit, and executes processing using a storage device. If the processing circuit includes a program execution unit, the storage device stores the software program executed by the program execution unit.

[0515] An example of the above-mentioned software program is a bitstream. The bitstream includes an encoded image and a syntax for performing a decoding process to decode the image. The bitstream causes the decoding device 200 to decode the image by executing a process based on the syntax. Furthermore, for example, software for realizing the above-mentioned encoding device 100 or decoding device 200 is a program such as the following.

[0516] For example, this program may cause a computer to execute an encoding method that encodes base data of a facial image related to a facial moving image and geometric information, which corresponds to each of multiple frames of the facial moving image and indicates geometric attributes within an area including a person's face, into a bit stream, determines whether the geometric information has been properly acquired, and further encodes into the bit stream concealment parameters related to error recovery control in a decoding device in the event that the geometric information has not been properly acquired.

[0517] Furthermore, for example, the program may cause a computer to execute an encoding method that encodes base data of an image included in a video, encodes facial attribute parameters indicating faces included in the image into a bitstream, and further encodes reliability parameters relating to the reliability of the facial attribute parameters into the bitstream, the reliability parameters indicating that the reliability is low depending on the value of the reliability parameter.

[0518] Furthermore, for example, this program may cause a computer to execute a decoding method in which base data of a facial image relating to a facial moving image and geometric information, which corresponds to each of multiple frames of the facial moving image and indicates geometric attributes within an area including a person's face, are decoded from a bit stream, and concealment parameters relating to error recovery control in the event that the geometric information is not properly acquired in the encoding device are further decoded from the bit stream, and the facial moving image is generated using a generative model from the base data, the geometric information, and the concealment parameters.

[0519] Furthermore, for example, the program may cause a computer to execute a decoding method in which base data of an image included in a video is obtained, facial attribute parameters indicating a face included in the image are decoded from a bitstream, and an output image corresponding to the image is generated by inputting the base data and the facial attribute parameters into a generative model, the bitstream including a reliability parameter relating to the reliability of the facial attribute parameter, and the reliability parameter indicating that the reliability is low depending on the value of the reliability parameter.

[0520] Furthermore, each of the above-described components may be a circuit. These circuits may form a single circuit as a whole, or each may be a separate circuit. Furthermore, each component may be realized by a general-purpose processor or a dedicated processor.

[0521] Furthermore, a process performed by a specific component may be performed by another component. The order in which the processes are performed may be changed, or multiple processes may be performed in parallel. Any two or more of the multiple examples of the present disclosure may be appropriately combined and implemented. The encoding / decoding device may include the encoding device 100 and the decoding device 200.

[0522] Furthermore, not all of the components in the present disclosure may be implemented, and only some of the components in the present disclosure may be implemented. Similarly, not all of the processes in the present disclosure may be executed, and only some of the processes in the present disclosure may be executed.

[0523] Furthermore, ordinal numbers such as "first" and "second" used in the description may be changed as appropriate. Furthermore, new ordinal numbers may be assigned to components or removed. Furthermore, these ordinal numbers may be assigned to elements in order to identify them, and may not correspond to a meaningful order.

[0524] Also, for example, a phrase "at least one of" a first element, a second element, and a third element corresponds to the first element, the second element, the third element, or any combination thereof.

[0525] Although the aspects of the encoding device 100 and the decoding device 200 have been described above based on a number of examples, the aspects of the encoding device 100 and the decoding device 200 are not limited to these examples. As long as they do not deviate from the spirit of the present disclosure, various modifications conceivable by those skilled in the art to each example, or configurations constructed by combining components of different examples, may also be included within the scope of the aspects of the encoding device 100 and the decoding device 200.

[0526] One or more aspects disclosed herein may be implemented in combination with at least a part of other aspects of the present disclosure. Also, some processes shown in the flowcharts of one or more aspects disclosed herein, some configurations of devices, some syntax, etc. may be implemented in combination with other aspects.

[0527] [Implementation and Application] In each of the above embodiments, each of the functional or operational blocks can typically be realized by an MPU (micro processing unit), memory, etc. Furthermore, the processing by each of the functional blocks may be realized as a program execution unit such as a processor that reads and executes software (programs) recorded on a recording medium such as a ROM. The software may be distributed. The software may be recorded on various recording media such as semiconductor memory. It is also possible to realize each functional block by hardware (dedicated circuitry).

[0528] The processing described in each embodiment may be realized by centralized processing using a single device (system), or may be realized by distributed processing using multiple devices. Furthermore, the processor that executes the program may be a single processor or multiple processors. That is, centralized processing or distributed processing may be performed.

[0529] The aspects of the present disclosure are not limited to the above examples, and various modifications are possible, and these modifications are also included within the scope of the aspects of the present disclosure.

[0530] Furthermore, application examples of the video coding method (image coding method) or video decoding method (image decoding method) shown in each of the above embodiments and various systems for implementing the application examples will be described below. Such systems may be characterized by having an image coding device using the image coding method, an image decoding device using the image decoding method, or an image coding / decoding device that includes both. Other configurations of such systems can be appropriately changed depending on the situation.

[0531] [Example of Use] Fig. 59 shows the overall configuration of an appropriate content supply system ex100 that realizes a content distribution service. The area where communication services are provided is divided into cells of a desired size, and base stations ex106, ex107, ex108, ex109, and ex110, which are fixed wireless stations in the illustrated example, are installed in each cell.

[0532] In this content supply system ex100, devices such as a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 are connected to the Internet ex101 via an Internet service provider ex102 or a communication network ex104 and base stations ex106 to ex110. The content supply system ex100 may connect a combination of any of the above devices. In various implementations, the devices may be connected to each other directly or indirectly via a telephone network, short-range wireless communication, or the like, without going through the base stations ex106 to ex110. Furthermore, the streaming server ex103 may be connected to devices such as the computer ex111, the game console ex112, the camera ex113, the home appliance ex114, and the smartphone ex115 via the Internet ex101, etc. The streaming server ex103 may also be connected to a terminal in a hotspot on an airplane ex117 via a satellite ex116.

[0533] Note that wireless access points, hot spots, etc. may be used instead of the base stations ex106 to ex110. Furthermore, the streaming server ex103 may be directly connected to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or may be directly connected to the airplane ex117 without going through the satellite ex116.

[0534] The camera ex113 is a device such as a digital camera that can take still images and videos. The smartphone ex115 is a smartphone, mobile phone, or PHS (Personal Handyphone System) that supports mobile communication systems such as 2G, 3G, 3.9G, 4G, and the upcoming 5G.

[0535] The home appliance ex114 is a refrigerator, or an appliance included in a home fuel cell cogeneration system, or the like.

[0536] In the content supply system ex100, a terminal having a photographing function is connected to a streaming server ex103 via a base station ex106 or the like, thereby enabling live streaming and the like. In live streaming, a terminal (such as a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, a smartphone ex115, or a terminal on an airplane ex117) may perform the encoding process described in each of the above embodiments on still image or video content captured by a user using the terminal, may multiplex the video data obtained by encoding with audio data obtained by encoding audio corresponding to the video, and may transmit the obtained data to the streaming server ex103. In other words, each terminal functions as an image encoding device according to one aspect of the present disclosure.

[0537] Meanwhile, the streaming server ex103 streams the transmitted content data to the requesting client. The client is a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, a smartphone ex115, a terminal on an airplane ex117, or the like, which is capable of decoding the encoded data. Each device that receives the distributed data decodes and plays back the received data. That is, each device may function as an image decoding device according to one aspect of the present disclosure.

[0538] [Distributed Processing] The streaming server ex103 may also be multiple servers or multiple computers that process, record, and distribute data in a distributed manner. For example, the streaming server ex103 may be implemented using a CDN (Content Delivery Network), where content distribution is achieved through a network connecting numerous edge servers distributed around the world. In a CDN, a physically nearby edge server is dynamically assigned depending on the client. Content is then cached and distributed to that edge server, thereby reducing delays. Furthermore, when certain types of errors occur or communication conditions change due to increased traffic, processing can be distributed among multiple edge servers, the distribution entity can be switched to another edge server, or distribution can be continued by bypassing the failed portion of the network, thereby achieving high-speed and stable distribution.

[0539] In addition to the distributed processing of the distribution itself, the encoding of captured data may be performed by each terminal, by the server, or by multiple terminals. For example, encoding generally involves two processing loops. The first loop detects the complexity of the image or the amount of code for each frame or scene. The second loop maintains image quality while improving encoding efficiency. For example, a terminal may perform the first encoding process, and the server that receives the content may perform the second encoding process, thereby improving content quality and efficiency while reducing the processing load on each terminal. In this case, if there is a request to receive and decode the data in near real time, the data encoded by a terminal can be received and played back by another terminal, enabling more flexible real-time distribution.

[0540] As another example, the camera ex113 or the like extracts features from an image, compresses the data related to the features as metadata, and transmits the compressed data to the server. The server performs compression according to the meaning (or importance of the content) of the image, for example, by determining the importance of an object from the features and switching the quantization precision accordingly. The feature data is particularly effective in improving the accuracy and efficiency of motion vector prediction when the server re-compresses the image. Alternatively, the terminal may perform simple encoding such as VLC (variable length coding), and the server may perform encoding with a high processing load such as CABAC (context-adaptive binary arithmetic coding).

[0541] As another example, in a stadium, shopping mall, factory, or the like, there may be multiple pieces of video data that have been shot by multiple terminals of almost the same scene. In this case, using the multiple terminals that shot the video and, as necessary, other terminals and servers that did not shoot the video, encoding processes are assigned to each of them, for example, in units of GOPs (Group of Pictures), pictures, or tiles obtained by dividing a picture, for distributed processing. This reduces delays and achieves better real-time performance.

[0542] Since multiple video data are of almost the same scene, the server may manage and / or instruct the video data shot by each terminal to be mutually referential. The server may also receive encoded data from each terminal and change the reference relationships between multiple data, or correct or replace the pictures themselves and re-encode them. This allows for the generation of streams with improved quality and efficiency for each piece of data.

[0543] Furthermore, the server may perform transcoding to change the encoding method of the video data before distributing it. For example, the server may convert an MPEG-based encoding method into a VP-based encoding method (e.g., VP9), or convert H.264 to H.265.

[0544] In this way, the encoding process can be performed by a terminal or one or more servers. Therefore, although the following uses terms such as "server" or "terminal" to refer to the entity performing the process, some or all of the processing performed by the server may be performed by the terminal, and some or all of the processing performed by the terminal may be performed by the server. The same applies to the decoding process.

[0545] [3D, Multi-Angle] Images or videos of different scenes or the same scene taken from different angles using multiple devices such as cameras ex113 and / or smartphones ex115 that are approximately synchronized with each other are increasingly being integrated and used. The videos taken by each device are integrated based on the relative positional relationship between the devices obtained separately, or on areas where feature points in the videos match.

[0546] The server may not only encode two-dimensional video, but also encode still images automatically or at a time specified by the user based on scene analysis of the video and transmit them to the receiving terminal. Furthermore, if the server can acquire the relative positional relationship between the capturing terminals, it can generate a three-dimensional shape of the scene based on not only two-dimensional video but also video of the same scene captured from different angles. The server may separately encode three-dimensional data generated by a point cloud or the like, or may select or reconstruct the video to be transmitted to the receiving terminal from video captured by multiple terminals based on the results of recognizing or tracking people or objects using the three-dimensional data.

[0547] In this way, the user can enjoy a scene by arbitrarily selecting each video corresponding to each shooting terminal, or can enjoy content in which a video from a selected viewpoint is cut out from 3D data reconstructed using multiple images or videos. Furthermore, together with the video, sound may also be collected from multiple different angles, and the server may multiplex the sound from a specific angle or space with the corresponding video and transmit the multiplexed video and sound.

[0548] In recent years, content that associates the real world with a virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has also become popular. In the case of VR images, the server creates viewpoint images for the right eye and left eye, respectively, and may perform encoding that allows reference between each viewpoint video using Multi-View Coding (MVC) or the like, or may encode them as separate streams without referencing each other. When decoding the separate streams, it is preferable to play them in synchronization with each other so that a virtual three-dimensional space is reproduced according to the user's viewpoint.

[0549] In the case of AR images, the server superimposes virtual object information in the virtual space onto camera information in the real space based on the three-dimensional position or the movement of the user's viewpoint. The decoding device may acquire or store virtual object information and three-dimensional data, generate a two-dimensional image according to the movement of the user's viewpoint, and smoothly connect the two-dimensional image to create superimposed data. Alternatively, the decoding device may send the movement of the user's viewpoint to the server in addition to a request for virtual object information. The server may create superimposed data according to the movement of the viewpoint received from the three-dimensional data stored on the server, encode the superimposed data, and distribute it to the decoding device. Note that the superimposed data may have an α value indicating transparency in addition to RGB, and the server may set the α value of parts other than the object created from the three-dimensional data to 0, etc., to encode the parts in a transparent state. Alternatively, the server may generate data by setting a predetermined RGB value as the background, like a chromakey, and using the background color for parts other than the object.

[0550] Similarly, the decoding of distributed data may be performed by each client terminal, by the server, or by multiple terminals. For example, one terminal may first send a reception request to the server, and then other terminals may receive and decode content according to the request, after which the decoded signal is transmitted to a device having a display. By distributing the processing and selecting appropriate content regardless of the capabilities of the communication terminals themselves, high-quality data can be reproduced. As another example, large-sized image data may be received on a TV or other device, and only a portion of the picture, such as a tile into which the picture is divided, may be decoded and displayed on the viewer's personal device. This allows the viewer to share the overall picture while checking their own area of ​​responsibility or an area of ​​interest in more detail.

[0551] In situations where multiple short-, medium-, or long-range wireless communications are available indoors and outdoors, seamless content reception may be possible using distribution system standards such as MPEG-DASH (Dynamic Adaptive Streaming over HTTP). Users may freely select and switch in real time between decoding devices or display devices, such as their own terminals and indoor / outdoor displays. Furthermore, decoding can be performed while switching between decoding and display devices using their own location information, etc. This allows information to be mapped and displayed on a part of the wall or ground of a neighboring building with an embedded display device while the user is traveling to their destination. It is also possible to switch the bit rate of received data based on the accessibility of the encoded data on the network, such as when the encoded data is cached on a server that can be quickly accessed from the receiving terminal or copied to an edge server in a content delivery service.

[0552] [Web Page Optimization] FIG. 60 is a diagram showing an example of a display screen of a web page on a computer ex111 or the like. FIG. 61 is a diagram showing an example of a display screen of a web page on a smartphone ex115 or the like. As shown in FIGS. 60 and 61 , a web page may include multiple link images that are links to image content, and the appearance of the link images may differ depending on the device used to view the page. When multiple link images are visible on the screen, the display device (decoding device) may display a still image or I-picture contained in each content as a link image until the user explicitly selects the link image, or until the link image approaches the center of the screen or the entire link image is within the screen, or may display a video such as a GIF animation using multiple still images or I-pictures, or may receive only the base layer and decode and display the video.

[0553] When a link image is selected by a user, the display device performs decoding while giving top priority to the base layer. Note that if the HTML (HyperText Markup Language) constituting the web page contains information indicating that the content is scalable, the display device may decode up to the enhancement layer. Furthermore, to ensure real-time performance, before selection or when the communication bandwidth is very limited, the display device decodes and displays only forward-reference pictures (I pictures, P pictures, and forward-reference-only B pictures), thereby reducing the delay between the decoding time of the first picture and the display time (the delay from the start of content decoding to the start of display). Furthermore, the display device may intentionally ignore the picture reference relationships and roughly decode all B pictures and P pictures using forward reference, and then perform normal decoding as the number of received pictures increases over time.

[0554] [Autonomous Driving] When transmitting and receiving still image or video data such as two-dimensional or three-dimensional map information for automatic driving or driving assistance of a vehicle, the receiving terminal may receive weather or construction information as meta information in addition to image data belonging to one or more layers, and may associate and decode these. Note that the meta information may belong to a layer, or may simply be multiplexed with the image data.

[0555] In this case, since a vehicle, drone, airplane, or the like including the receiving terminal is moving, the receiving terminal can transmit location information of the receiving terminal, thereby realizing seamless reception and decoding while switching between base stations ex106 to ex110. Furthermore, the receiving terminal can dynamically switch how much meta information to receive or how much to update map information depending on the user's selection, the user's situation, and / or the state of the communication bandwidth.

[0556] In the content supply system ex100, the client can receive, decode, and play back encoded information sent by a user in real time.

[0557] [Distribution of Personal Content] The content supply system ex100 also allows for unicast or multicast distribution of not only high-quality, long-duration content from video distribution companies, but also low-quality, short-duration content from individuals. It is expected that such personal content will continue to increase in the future. To improve the quality of personal content, the server may perform editing before encoding. This can be achieved, for example, using the following configuration.

[0558] During shooting, either in real time or after accumulating and shooting, the server performs recognition processing such as detecting shooting errors, scene search, semantic analysis, and object detection from the original image data or encoded data. Based on the recognition results, the server manually or automatically corrects out-of-focus or camera shake, deletes less important scenes such as scenes with lower brightness or out-of-focus compared to other pictures, emphasizes object edges, changes color, and performs other editing. The server then encodes the edited data based on the editing results. It is also known that viewing rates decrease if the shooting time is too long. Therefore, the server may automatically clip not only less important scenes as described above but also scenes with little movement, based on the image processing results, so that the content falls within a specific time range depending on the shooting time. Alternatively, the server may generate and encode a digest based on the results of the semantic analysis of the scenes.

[0559] Personal content may contain content that, if left as is, violates copyright, moral rights, or portrait rights, and may cause the scope of sharing to exceed the intended scope, resulting in inconvenience to individuals. Therefore, for example, the server may intentionally defocus images of people's faces on the periphery of the screen or the interior of a house before encoding. Furthermore, the server may recognize whether the image to be encoded contains the face of a person other than a pre-registered person, and if so, perform processing such as blurring the face. Alternatively, as pre- or post-processing before encoding, the user may specify a person or background area they wish to modify in the image for copyright or other reasons. The server may replace the specified area with another image or blur the focus. If the image contains a person, the server may track the person in the video and replace the image of the person's face.

[0560] Because viewing personal content with small data volumes requires high real-time performance, the decoding device first receives the base layer as a top priority, and then decodes and plays it back, depending on the bandwidth. The decoding device may also receive an enhancement layer during this time, and if the content is played back more than once, such as when playback is looped, it may play back high-quality video including the enhancement layer. A stream that has undergone scalable encoding in this way can provide an experience in which the video appears rough when not selected or when viewing begins, but gradually becomes smarter and the image quality improves. In addition to scalable encoding, a similar experience can also be provided when a rough stream played the first time and a second stream that is encoded with reference to the first video are configured as a single stream.

[0561] [Other Application Examples] Furthermore, these encoding or decoding processes are generally performed by the LSIex500 possessed by each terminal. The LSI (large scale integration circuitry) ex500 (see FIG. 59) may be a single-chip or multi-chip configuration. Furthermore, video encoding or decoding software may be embedded in some kind of recording medium (such as a CD-ROM, flexible disk, or hard disk) readable by the computer ex111, and the encoding or decoding process may be performed using that software. Furthermore, if the smartphone ex115 is equipped with a camera, video data captured by the camera may be transmitted. This video data is data encoded and processed by the LSIex500 possessed by the smartphone ex115.

[0562] The LSIex500 may be configured to download and activate application software. In this case, the terminal first determines whether it supports the content encoding method or has the capability to execute a specific service. If the terminal does not support the content encoding method or does not have the capability to execute a specific service, the terminal downloads a codec or application software, and then acquires and plays the content.

[0563] Furthermore, at least one of the moving image encoding device (image encoding device) or moving image decoding device (image decoding device) of each of the above embodiments can be incorporated into a digital broadcasting system, not limited to the content supply system ex100 via the Internet ex101. Since multiplexed data in which video and audio are multiplexed is transmitted and received over broadcast radio waves using a satellite or the like, there is a difference in that it is more suited to multicast than the content supply system ex100, which has a configuration that is easy to use for unicast, but similar applications are possible with regard to encoding and decoding processes.

[0564] [Hardware Configuration] Fig. 62 is a diagram showing further details of the smartphone ex115 shown in Fig. 59. Fig. 63 is a diagram showing an example configuration of the smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves to and from the base station ex110, a camera unit ex465 capable of capturing video and still images, and a display unit ex458 for displaying video captured by the camera unit ex465 and decoded data of the video and the like received by the antenna ex450. The smartphone ex115 further includes an operation unit ex466 such as a touch panel, an audio output unit ex457 such as a speaker for outputting voice or sound, an audio input unit ex456 such as a microphone for inputting voice, a memory unit ex467 capable of storing captured video or still images, recorded voice, received video or still images, encoded data such as email, or decoded data, and a slot unit ex464 that is an interface with a SIM (Subscriber Identity Module) ex468 for identifying a user and authenticating access to various data including the network. Note that an external memory may be used instead of the memory unit ex467.

[0565] A main control unit ex460 that comprehensively controls the display unit ex458 and operation unit ex466, etc., is connected to a power supply circuit unit ex461, an operation input control unit ex462, a video signal processing unit ex455, a camera interface unit ex463, a display control unit ex459, a modulation / demodulation unit ex452, a multiplexing / separation unit ex453, an audio signal processing unit ex454, a slot unit ex464, and a memory unit ex467 via a synchronization bus ex470.

[0566] When the power key is turned on by a user's operation, the power supply circuit unit ex461 starts up the smartphone ex115 to an operable state and supplies power to each unit from the battery pack.

[0567] The smartphone ex115 performs processes such as telephone calls and data communications under the control of a main control unit ex460 having a CPU, ROM, RAM, etc. During a call, an audio signal collected by an audio input unit ex456 is converted into a digital audio signal by an audio signal processing unit ex454, subjected to spectrum spread processing by a modulation / demodulation unit ex452, subjected to digital-to-analog conversion and frequency conversion processing by a transmission / reception unit ex451, and the resulting signal is transmitted via an antenna ex450. The received data is also amplified and subjected to frequency conversion and analog-to-digital conversion processing, subjected to spectrum despreading processing by a modulation / demodulation unit ex452, and converted into an analog audio signal by an audio signal processing unit ex454, which is then output from an audio output unit ex457. During data communication mode, text, still images, or video data is sent to the main control unit ex460 via an operation input control unit ex462 based on operations on the main unit's operation unit ex466, etc. Similar transmission and reception processing is performed. When transmitting video, still images, or video and audio in the data communication mode, the video signal processing unit ex455 compresses and encodes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 using the moving image encoding method described in each of the above embodiments, and sends the encoded video data to the multiplexing / demultiplexing unit ex453. The audio signal processing unit ex454 encodes the audio signal collected by the audio input unit ex456 while the camera unit ex465 is capturing video or still images, and sends the encoded audio data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded video data and encoded audio data using a predetermined method, and modulates and converts the multiplexed video data and audio data in the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, before transmitting the multiplexed video data and audio data via the antenna ex450.

[0568] In order to decode the multiplexed data received via the antenna ex450, such as when receiving video attached to an email or chat, or video linked to a web page, the multiplexing / separation unit ex453 separates the multiplexed data into a video data bit stream and an audio data bit stream, and supplies the encoded video data to the video signal processing unit ex455 and the encoded audio data to the audio signal processing unit ex454 via the synchronization bus ex470. The video signal processing unit ex455 decodes the video signal using a video decoding method corresponding to the video encoding method described in each of the above embodiments, and the video or still image contained in the linked video file is displayed on the display unit ex458 via the display control unit ex459. The audio signal processing unit ex454 decodes the audio signal, and audio is output from the audio output unit ex457. As real-time streaming becomes increasingly common, audio playback may be socially inappropriate depending on the user's situation. Therefore, it is preferable that the initial setting be a configuration in which only the video data is played without playing the audio signal, and audio may be played in sync only when the user performs an operation such as clicking on the video data.

[0569] Although the smartphone ex115 has been used as an example, three other implementation formats are possible: a transmitting / receiving terminal having both an encoder and a decoder, a transmitting terminal having only an encoder, and a receiving terminal having only a decoder. In the digital broadcasting system, multiplexed data in which audio data is multiplexed with video data is received or transmitted. However, in addition to audio data, text data related to the video may also be multiplexed into the multiplexed data. Furthermore, the video data itself may be received or transmitted instead of the multiplexed data.

[0570] Although the main control unit ex460 including a CPU has been described as controlling the encoding or decoding process, various terminals often include a GPU (Graphics Processing Unit). Therefore, a configuration may be adopted in which a memory shared by the CPU and GPU, or a memory whose addresses are managed for common use, is used to take advantage of the GPU's performance to process a large area in a batch. This shortens the encoding time, ensures real-time performance, and achieves low latency. It is particularly efficient to perform motion estimation, deblocking filtering, SAO (Sample Adaptive Offset), and transformation / quantization processes in a batch, such as by picture, by the GPU rather than by the CPU.

[0571] The present disclosure is applicable to encoding devices that encode moving images, and is applicable to video conferencing systems, etc.

[0572] 100, 700 Encoding device 102 Division unit 104 Subtraction unit 106 Transform unit 108 Quantization unit 110 Entropy coding unit 112, 204 Inverse quantization unit 114, 206 Inverse transformation unit 116, 208 Addition unit 118, 210 Block memory 120, 212 Loop filter unit 122, 214 Frame memory 124, 216 Intra prediction unit 126, 218 Inter prediction unit 128, 220 Prediction control unit 130, 222 Prediction parameter generation unit 131, 133, 134, 731, 733 Compressor 132, 232, 236, 732, 832 Derivation unit 135, 237 Buffer 151, 251 Circuit 152, 252 Memory 200, 800 Decoding device 202 Entropy decoding unit 224 Partition determination unit 231, 233, 235, 831, 833 Decompressor 234, 406, 834 Generator 401 Video decoder 402 Entropy decoder 403 Inverse affine transformer 404 Decision unit 405 Geometry attribute buffer

Claims

1. A decoding device comprising a circuit and a memory connected to the circuit, wherein in operation, the circuit decodes from a bitstream base data of a face image related to a facial motion image and geometric information which is information corresponding to each of a plurality of frames of the facial motion image and indicates geometric attributes within a region including a person's face, further decodes from the bitstream a concealment parameter related to error recovery control when the geometric information is not properly acquired by an encoding device, and generates the facial motion image from the base data, the geometric information, and the concealment parameter using a generation model.

2. The decoding device according to claim 1, wherein the circuit decodes the concealment parameter from a header region in the bitstream.

3. The decoding device according to claim 1 or 2, wherein the concealment parameter indicates, according to a value of the concealment parameter, that the reliability of the geometric information corresponding to a current frame acquired by the encoding device is low.

4. The decoding device according to claim 1 or 2, wherein the concealment parameter indicates, according to a value of the concealment parameter, that the geometric information corresponding to a current frame is not acquired by the encoding device.

5. The decoding device according to claim 1 or 2, wherein the concealment parameter indicates, according to a value of the concealment parameter, that the number of attributes in the geometric information corresponding to a current frame acquired by the encoding device is different from the original number of attributes.

6. The decoding device according to claim 1 or 2, wherein the concealment parameter indicates, according to a value of the concealment parameter, that stored geometric information is applied instead of the geometric information decoded from the bitstream for a current frame of the facial motion image.

7. The decoding device according to claim 1 or 2, wherein the concealment parameter indicates, according to a value of the concealment parameter, that the geometric information corresponding to a current frame is stored for a subsequent frame.

8. When the concealment parameter indicates that the reliability of the geometric information corresponding to the current frame obtained in the encoding device is low, or that the number of attributes in the geometric information is different from the original number of attributes, the circuit corrects the geometric information corresponding to the current frame decoded from the bitstream using the stored geometric information, and applies the corrected geometric information to the geometric information for generating the face image corresponding to the current frame in the face motion image. The decoding device according to claim 1 or 2.

9. When the concealment parameter indicates that the reliability of the geometric information corresponding to the current frame obtained in the encoding device is low, that the geometric information has not been obtained, that the number of attributes in the geometric information is different from the original number of attributes, or that the stored geometric information is applied to the geometric information corresponding to the current frame, the circuit applies the stored geometric information to the geometric information for generating the face image corresponding to the current frame in the face motion image. The decoding device according to claim 1 or 2.

10. When the concealment parameter indicates that the reliability of the geometric information corresponding to the current frame obtained in the encoding device is low, that the geometric information has not been obtained, or that the number of attributes in the geometric information is different from the original number of attributes, the circuit applies the already generated face image to the face image corresponding to the current frame without generating the face image corresponding to the current frame in the face motion image using the generation model. The decoding device according to claim 1 or 2.

11. The stored geometric information is geometric information corresponding to a past frame decoded from the bitstream. The decoding device according to claim 6.

12. The stored geometric information is predefined reference geometric information. The decoding device according to claim 6.

13. A decoding device comprising a circuit and a memory connected to the circuit, wherein the circuit, in operation, acquires base data of an image included in a moving image, decodes face attribute parameters indicating a face included in the image from a bit stream, and generates an output image corresponding to the image by inputting the base data and the face attribute parameters into a generation model, the bit stream includes a reliability parameter regarding the reliability of the face attribute parameters, and the reliability parameter indicates that the reliability is low according to the value of the reliability parameter.

14. The decoding device according to claim 13, wherein the image corresponds to each of a plurality of pictures included in the moving image, the base data is data common to the plurality of pictures, and the bit stream includes the reliability parameter and the face attribute parameters for each of the plurality of pictures.

15. The decoding device according to claim 14, wherein the bit stream includes the reliability parameter before the face attribute parameters for each of the plurality of pictures.

16. The decoding device according to any one of claims 13 to 15, wherein when the reliability parameter does not indicate that the reliability of the face attribute parameters is low, the face attribute parameters are stored.

17. An encoding device comprising a circuit and a memory connected to the circuit, wherein the circuit, in operation, encodes base data of a face image regarding a face moving image and geometric information, which is information corresponding to each of a plurality of frames of the face moving image and indicates geometric attributes within a region including a person's face, into a bit stream, determines whether the geometric information is appropriately acquired, and further encodes a concealment parameter regarding error recovery control in the decoding device when the geometric information is not appropriately acquired into the bit stream.

18. The encoding device according to claim 17, wherein the circuit encodes the concealment parameter into a header region in the bit stream.

19. The encoding device according to claim 17 or 18, wherein the concealment parameter indicates that the reliability of the geometric information corresponding to the current frame acquired in the encoding device is low according to the value of the concealment parameter.

20. The encoding device according to claim 17 or 18, wherein the concealment parameter indicates, according to the value of the concealment parameter, that the geometric information corresponding to the current frame has not been acquired in the encoding device.

21. The encoding device according to claim 17 or 18, wherein the concealment parameter indicates, according to the value of the concealment parameter, that the number of the geometric information corresponding to the current frame acquired in the encoding device is different from the original number.

22. The encoding device according to claim 17 or 18, wherein the concealment parameter indicates, according to the value of the concealment parameter, that stored geometric information is applied instead of the geometric information decoded from the bit stream for the current frame of the face motion image.

23. The encoding device according to claim 17 or 18, wherein the concealment parameter indicates, according to the value of the concealment parameter, that the geometric information corresponding to the current frame is stored for subsequent frames.

24. The encoding device according to claim 17 or 18, wherein when the reliability of the geometric information corresponding to the current frame acquired in the encoding device is lower than a threshold value, the circuit sets a value indicating that the reliability of the geometric information is low in the concealment parameter.

25. The encoding device according to claim 17 or 18, wherein when the reliability of the geometric information corresponding to the current frame acquired in the encoding device is lower than a threshold value, or when no face is included in the current frame acquired in the encoding device, the circuit sets, in the concealment parameter, a value indicating that the geometric information has not been acquired without encoding the geometric information corresponding to the current frame into the bit stream.

26. The encoding device according to claim 17 or 18, wherein when the number of attributes in the geometric information corresponding to the current frame acquired in the encoding device is different from a predefined number of attributes, the circuit sets, in the concealment parameter, a value indicating that the number of the attributes in the geometric information is different from the original number.

27. The circuit sets, in the concealment parameter, a value indicating that when the reliability of the geometric information corresponding to the current frame, which is obtained in the encoding device, is lower than a threshold value, or when no face is included in the current frame obtained in the encoding device, the stored geometric information is applied to the geometric information corresponding to the current frame. The encoding device according to claim 17 or 18.

28. The circuit sets, in the concealment parameter, a value indicating that when the reliability of the geometric information corresponding to the current frame, which is obtained in the encoding device, is equal to or higher than a specific threshold value, the geometric information is stored for a subsequent frame. The encoding device according to claim 17 or 18.

29. When the reliability of the geometric information corresponding to the current frame, which is obtained in the encoding device, is lower than the threshold value, the circuit corrects the obtained geometric information using the stored geometric information, and encodes the corrected geometric information into the bitstream. The encoding device according to claim 17 or 18.

30. When the reliability of the geometric information corresponding to the current frame, which is obtained in the encoding device, is lower than the threshold value, the circuit encodes the stored geometric information as the geometric information corresponding to the current frame into the bitstream. The encoding device according to claim 17 or 18.

31. The stored geometric information is geometric information corresponding to a past frame, which is obtained in the encoding device. The encoding device according to claim 22.

32. The stored geometric information is predefined reference geometric information. The encoding device according to claim 22.

33. An encoding device comprising a circuit and a memory connected to the circuit, wherein the circuit, in operation, encodes base data of an image included in a moving image, encodes face attribute parameters indicating a face included in the image into a bitstream, and further encodes a reliability parameter regarding the reliability of the face attribute parameters into the bitstream, and the reliability parameter indicates that the reliability is low according to the value of the reliability parameter.

34. The image corresponds to each of a plurality of pictures included in the moving image, the base data is data common to the plurality of pictures, and the bit stream includes the reliability parameter and the face attribute parameter for each of the plurality of pictures. The encoding device according to claim 33.

35. The bit stream includes the reliability parameter before the face attribute parameter for each of the plurality of pictures. The encoding device according to claim 34.

36. When the reliability parameter does not indicate that the reliability of the face attribute parameter is low, the face attribute parameter is stored. The encoding device according to any one of claims 33 to 35.

37. Decoding from a bit stream base data of a face image related to a face moving image and geometric information which is information corresponding to each of a plurality of frames of the face moving image and which indicates geometric attributes within a region including a person's face, and further decoding from the bit stream a concealment parameter related to error recovery control when the geometric information is not appropriately acquired by the encoding device, and generating the face moving image using a generation model from the base data, the geometric information, and the concealment parameter. A decoding method.

38. Acquiring base data of an image included in a moving image, decoding from a bit stream a face attribute parameter indicating a face included in the image, and generating an output image corresponding to the image by inputting the base data and the face attribute parameter into a generation model. The bit stream includes a reliability parameter related to the reliability of the face attribute parameter, and the reliability parameter indicates that the reliability is low according to the value of the reliability parameter. A decoding method.

39. Encoding in a bit stream base data of a face image related to a face moving image and geometric information which is information corresponding to each of a plurality of frames of the face moving image and which indicates geometric attributes within a region including a person's face, determining whether the geometric information is appropriately acquired, and further encoding in the bit stream a concealment parameter related to error recovery control in a decoding device when the geometric information is not appropriately acquired. An encoding method.

40. A coding method for coding base data of an image included in a moving image, coding face attribute parameters indicating a face included in the image into a bit stream, further coding a reliability parameter regarding the reliability of the face attribute parameters into the bit stream, and the reliability parameter indicating that the reliability is low according to a value of the reliability parameter.

Citation Information

Patent Citations

  • Image display device, image transmission system, transmission device and reception device

    JP2002008051A

  • video transcoder

    JP2005526457A

  • Image encoding device

    JP2012191450A

  • Scalable system and method using logical entities for production of programs that use multi-media signals

    WO2023022697A1