Decoding device, encoding device, bitstream generation device, decoding method, and encoding method
By generating synthetic facial motion images by combining a generative model with a reference image, geometric information, and background information, the problems of low encoding efficiency, poor image quality, and large processing volume in existing video encoding and decoding technologies are solved, achieving efficient video encoding and decoding.
Patent Information
- Application Number
- CN202480027424.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-28
- Filing Date
- 2024-04-03
- Publication Date
- 2025-11-21
AI Technical Summary
Existing video encoding and decoding technologies suffer from problems such as low encoding efficiency, poor image quality, large processing volume, large circuit scale, and slow processing speed when processing increasing amounts of video data. In particular, it is difficult to achieve appropriate element or motion selection in face reproduction processing.
By using a generative model, synthetic facial motion images are generated using a reference image, geometric information, and background information, suppressing background distortion and optimizing the encoding and decoding process.
It improves encoding efficiency, enhances image quality, reduces processing load and circuit size, increases processing speed, and enables appropriate selection of elements or actions.
Smart Images

Figure CN121002886A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to a decoding apparatus or the like. BACKGROUND
[0002] Video coding techniques have evolved from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Codec). With this progress, in order to handle the ever-increasing amount of digital video data in various uses, there is always a need to provide improvements and optimizations of video coding techniques. The present disclosure relates to further progress, improvements, and optimizations in video coding.
[0003] In addition, Non-Patent Literature 1 relates to an example of an existing standard related to the above-described video coding techniques. In addition, Non-Patent Literature 2 relates to a new proposal related to the above-described video coding techniques.
[0004] Prior Art Documents Non-Patent Literature Non-Patent Literature 1: H.265 (ISO / IEC 23008-2 HEVC) / HEVC (High Efficiency Video Coding) Non-Patent Literature 2: “AHG9 / AHG16: Common text for proposed generative face video SEI message”, JVET-AG0203-v1 SUMMARY
[0005] PROBLEMS TO BE SOLVED BY THE INVENTION With regard to the above-described coding mode, in order to improve coding efficiency, improve picture quality, reduce processing amount, reduce circuit size, or appropriately select elements or actions such as filters, blocks, sizes, motion vectors, reference pictures, or reference blocks, or the like, it is desirable to propose a new mode.
[0006] The present disclosure provides a structure or a method that can contribute to one or more of, for example, improvement of coding efficiency, improvement of picture quality, reduction of processing amount, reduction of circuit size, improvement of processing speed, and appropriate selection of elements or actions. In addition, the present disclosure can include a structure or a method that can contribute to benefits other than the above.
[0007] MEANS FOR SOLVING THE PROBLEMS For example, a decoding apparatus according to an aspect of the present disclosure includes a memory and a circuit connected to the memory, and in operation, decodes (i) a reference image that is an image including a face, (ii) geometry information that corresponds to a plurality of frames of a captured moving image obtained by a camera and represents a geometric property of a subject, and (iii) background information related to a background image, from one or more streams, and generates a synthesized face moving image using a generative model based on the reference image, the geometry information, and the background information, the synthesized face moving image being a moving image including the face and obtained by synthesizing the background image.
[0008] Each of the embodiments in the present disclosure or a part of the structure or the method thereof can achieve at least one of, for example, improvement of coding efficiency, improvement of image quality, reduction of processing amount of encoding / decoding, reduction of circuit size, or improvement of processing speed of encoding / decoding. Alternatively, each of the embodiments in the present disclosure or a part of the structure or the method thereof can achieve appropriate selection of constituent elements / actions such as filters, blocks, sizes, motion vectors, reference pictures, reference blocks, and the like, in encoding and decoding, respectively. In addition, the present disclosure also includes disclosure of a structure or a method that can provide benefits other than the above, such as a structure or a method that improves coding efficiency while suppressing an increase in processing amount.
[0009] Further advantages and effects of the aspect of the present disclosure will be clear from the description and the drawings. The advantages and / or effects can be obtained by several embodiments and the features described in the description and the drawings, but it is not necessary to provide all of them in order to obtain one or more advantages and / or effects.
[0010] In addition, these general or specific aspects can be implemented by a system, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, and can also be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.
[0011] Effects of the Invention The structure or the method according to an aspect of the present disclosure can provide one or more of, for example, improvement of coding efficiency, improvement of image quality, reduction of processing amount, reduction of circuit size, improvement of processing speed, and appropriate selection of elements or actions. In addition, the structure or the method according to an aspect of the present disclosure can also contribute to benefits other than the above. BRIEF DESCRIPTION OF DRAWINGS
[0012] Figure 1 is a block diagram showing the structure of an encoding and decoding system in a reference example.
[0013] Figure 2 is a block diagram showing the structure of an encoding apparatus in a reference example. is a block diagram showing the structure of an encoding and decoding system in a reference example.
[0014] Figure 3 is a block diagram showing the structure of a decoding device in a reference example.
[0015] Figure 4 is a conceptual diagram showing an example of a reference image.
[0016] Figure 5 is a conceptual diagram showing an example of a geometric property.
[0017] Figure 6 is a conceptual diagram showing an example of a face motion image.
[0018] Figure 7 is a block diagram showing a structure example of an encoding decoding system in an embodiment.
[0019] Figure 8 is a diagram showing an example of a hierarchical structure of data in a stream.
[0020] Figure 9 is a block diagram showing a structure example of an encoding device in an embodiment.
[0021] Figure 10 is a flowchart showing an example of an action of an encoding device in an embodiment.
[0022] Figure 11 is a block diagram showing a structure example of a decoding device in an embodiment.
[0023] Figure 12 is a flowchart showing an example of an action of a decoding device in an embodiment.
[0024] Figure 13 is a conceptual diagram showing an example of a decoding process in each time instance.
[0025] Figure 14 is a conceptual diagram showing another example of a decoding process in each time instance.
[0026] Figure 15 is a block diagram showing another structure example of a decoding device in an embodiment.
[0027] Figure 16 is a block diagram showing still another structure example of a decoding device in an embodiment.
[0028] Figure 17 is a block diagram showing another structure example of an encoding device in an embodiment.
[0029] Figure 18 is a block diagram showing still another structure example of a decoding device in an embodiment.
[0030] Figure 19is a block diagram showing another configuration example of a decoding apparatus in the embodiment.
[0031] Figure 20 is a block diagram showing a configuration example of an encoding apparatus for encoding a moving image in the embodiment.
[0032] Figure 21 is a block diagram showing a configuration example of a decoding apparatus for decoding a moving image in the embodiment.
[0033] Figure 22 is a conceptual diagram showing a configuration example of a bitstream.
[0034] Figure 23 is a conceptual diagram showing another configuration example of a bitstream.
[0035] Figure 24 is a conceptual diagram showing another configuration example of a bitstream.
[0036] Figure 25 is a conceptual diagram showing another configuration example of a bitstream.
[0037] Figure 26 is a conceptual diagram showing another configuration example of a bitstream.
[0038] Figure 27 is a conceptual diagram showing a configuration example of a bitstream conforming to VVC.
[0039] Figure 28 is a conceptual diagram showing another configuration example of a bitstream conforming to VVC.
[0040] Figure 29 is a conceptual diagram showing another configuration example of a bitstream conforming to VVC.
[0041] Figure 30 is a diagram showing examples of various neural networks that can be used as a generative model.
[0042] Figure 31 is a block diagram showing an implementation example of an encoding apparatus in the embodiment.
[0043] Figure 32 is a flowchart showing a basic first action example of an encoding apparatus in the embodiment.
[0044] Figure 33 is a flowchart showing a basic second action example of an encoding apparatus in the embodiment.
[0045] Figure 34 is a block diagram showing an implementation example of a bitstream generation apparatus in the embodiment.
[0046] Figure 35is a flowchart showing a basic first example of actions of a bitstream generation apparatus in an embodiment.
[0047] Figure 36 is a flowchart showing a basic second example of actions of a bitstream generation apparatus in an embodiment.
[0048] Figure 37 is a block diagram showing an example of a decoding apparatus in an embodiment.
[0049] Figure 38 is a flowchart showing a basic first example of actions of a decoding apparatus in an embodiment.
[0050] Figure 39 is a flowchart showing a basic second example of actions of a decoding apparatus in an embodiment.
[0051] Figure 40 is a diagram showing an overall structure of a content supply system that implements a content distribution service.
[0052] Figure 41 is a diagram showing an example of a display screen of a web page.
[0053] Figure 42 is a diagram showing an example of a display screen of a web page.
[0054] Figure 43 is a diagram showing an example of a smart phone.
[0055] Figure 44 is a block diagram showing an example of a structure of a smart phone. DETAILED DESCRIPTION
[0056] [INTRODUCTION] Facial reenactment refers to a process of mapping a person's expression and pose to a face image while maintaining the person's identity. Currently, facial reenactment can be used for a wide range of applications including video conferencing, entertainment industry, and social media. The present disclosure can be used in encoding of data regarding facial reenactment.
[0057] Figure 1 is a block diagram showing a structure of an encoding and decoding system in a reference example. For example, the encoding and decoding system is provided with an encoding apparatus 700 and a decoding apparatus 800. First, the encoding apparatus 700 receives a reference image and a driving motion image, and generates a bitstream. Next, the encoding apparatus 700 transmits the bitstream to the decoding apparatus 800 using a transmission channel. Finally, the decoding apparatus 800 reconstructs a face motion image from the received bitstream.
[0058] For example, the reference image is an image containing a face, and can also be expressed as a face image or an identity image. The reference image indicates a static visual feature used for reconstructing a face image. The driving motion image is a moving image containing a face, and is a captured moving image obtained by a camera. The driving motion image functions to impart motion to the reference image.
[0059] Figure 2 is a block diagram indicating the structure of the encoding apparatus 700 in the reference example. In this example, the encoding apparatus 700 is provided with a compressor 731, a derivator 732, and a compressor 733.
[0060] The compressor 731 encodes at least one reference image using a video compression technique. The reference image can be a frame of a driving motion image, a pre-obtained image including a face of a person, or a virtual character.
[0061] The derivator 732 derives geometry information indicating a geometric property corresponding to each frame of the driving motion image. The geometry information indicating the geometric property is also simply expressed as the geometric property. Here, the geometric property, for example, corresponding to a dynamic property, can be expressed by a point group like a face landmark, or can be expressed by a polygon model for expressing the shape of an object by the combination of a plurality of polygons. Further, the geometric property can be expressed by other geometry models. Further, the geometric property can be expressed by the position of a part of a face.
[0062] The compressor 733 compresses the geometric property into a bitstream using an entropy coding or the like. The compressor 731 and the compressor 733 can be the same constituent element, or can be different constituent elements.
[0063] Finally, the bitstream is transmitted to the decoding apparatus 800 via a transmission channel. For example, the compressed geometric property is transmitted from the encoding apparatus 700 to the decoding apparatus 800 per frame of the driving motion image, that is, per time instance.
[0064] Figure 3 is a block diagram indicating the structure of the decoding apparatus 800 in the reference example. In this example, the decoding apparatus 800 is provided with a decompressor 831, a derivator 832, a decompressor 833, and a generator 834.
[0065] The decompressor 831 decodes and reconstructs at least one reference image from the bitstream. Then, the decompressor 831 supplies the reference image to the derivator 832. The derivator 832 derives a reference property from the reference image. Here, the reference property is a static visual property, and can also be expressed as an identity. The decompressor 833 decodes and reconstructs the geometric property per frame.
[0066] The generator 834 generates a face motion image from the reference attribute and the geometry attribute using a neural network. The generator 834 can also use the reference image itself instead of the reference attribute, or use the reference image together with the reference attribute to render the face motion image.
[0067] Figure 4 is a conceptual diagram representing an example of a reference image. As shown in the example, the reference image is an image containing a face.
[0068] Figure 5 is a conceptual diagram representing an example of a geometry attribute. In this example, the geometry attribute is a landmark of a face. For example, the geometry attribute is derived for each frame in which a motion image is captured.
[0069] Figure 6 is a conceptual diagram representing an example of a face motion image. As shown in the example, the face motion image is a motion image containing a face. In the face motion image, for each of a plurality of frames, a geometry attribute of the frame is reflected in a reference image. Thus, the reference image is given a motion.
[0070] The amount of information of the plurality of geometry attributes corresponding to the reference image and the plurality of frames is less than the amount of information of the plurality of frames included in the captured motion image. Therefore, by encoding the plurality of geometry attributes corresponding to the reference image and the plurality of frames, the amount of encoding is reduced compared to encoding the plurality of frames included in the captured motion image. In addition, the reference image is given a motion by each geometry attribute. Thus, the face of the display object is given a motion, and rich expression can be performed.
[0071] However, by giving a motion to the face included in the reference image, the background region included in the reference image can be distorted. For example, since the position of the face is shifted from Figure 4 to Figure 6 , the background can be stretched or missing. Thus, the image quality can be degraded.
[0072] Therefore, the decoding device of Example 1 has a memory and a circuit connected to the memory, which, in operation, decodes (i) a reference image that is an image containing a face, (ii) geometry information that corresponds to a plurality of frames of a captured motion image obtained by a camera and represents a geometry attribute of a subject, and (iii) background information related to a background image, from one or more streams, and generates a synthetic face motion image using a generative model based on the reference image, the geometry information, and the background information, the synthetic face motion image being a motion image containing the face and obtained by synthesizing the background image.
[0073] Thus, it is possible to apply the reference image, the geometric attribute, and the background image to generation of the synthesized face motion image. Therefore, it is possible to impart motion to the face of the reference image by the geometric attribute corresponding to each frame while suppressing distortion of the background of the reference image by the background image. Thus, it is possible to suppress degradation of image quality when generating the synthesized face motion image.
[0074] Further, the decoding apparatus of Example 2 can be the decoding apparatus of Example 1, wherein the circuit inputs the reference image, the geometric information, and the background image to the generation model, and acquires the synthesized face motion image from the generation model.
[0075] Thus, it is possible to easily acquire the synthesized face motion image from the generation model. Further, in the generation model, it is possible to apply the reference image, the geometric attribute, and the background image to generation of the synthesized face motion image. Therefore, it is possible to impart motion to the face of the reference image by the geometric attribute corresponding to each frame while suppressing distortion of the background of the reference image by the background image.
[0076] Further, the decoding apparatus of Example 3 can be the decoding apparatus of Example 1, wherein the circuit inputs the reference image and the geometric information to the generation model, acquires an intermediate face motion image from the generation model, the intermediate face motion image being a motion image containing the face and being a motion image in which the background image is not synthesized, and generates the synthesized face motion image by embedding a corresponding region in the background image in a background region in the intermediate face motion image.
[0077] Thus, it is possible to acquire the intermediate face motion image from the generation model, the intermediate face motion image being a motion image in which motion is imparted to the face of the reference image by the geometric attribute corresponding to each frame. Then, it is possible to apply the background image to the intermediate face motion image. Therefore, it is possible to impart motion to the face of the reference image by the geometric attribute corresponding to each frame while suppressing distortion of the background of the reference image by the background image.
[0078] Further, the decoding apparatus of Example 4 can be the decoding apparatus of Example 3, wherein the circuit performs segmentation processing on the intermediate face motion image, acquires intermediate face motion image segmentation information indicating a foreground region and the background region in the intermediate face motion image, and determines the background region in the intermediate face motion image using the intermediate face motion image segmentation information.
[0079] Thus, it is possible to appropriately determine the background region in the intermediate face motion image based on the intermediate face motion image segmentation information obtained as a result of the segmentation processing on the intermediate face motion image. Therefore, it is possible to appropriately apply the corresponding region in the background image to the background region in the intermediate face motion image.
[0080] In addition, the decoding apparatus of Example 5 can be the decoding apparatus of Example 3, wherein the circuit decodes, from the one or more streams, photographed motion image segmentation information indicating the foreground region and the background region in the photographed motion image, and determines the background region in the intermediate face motion image using the photographed motion image segmentation information.
[0081] Thus, it is possible to appropriately determine the background region in the intermediate face motion image based on the photographed motion image segmentation information obtained from the one or more streams. Therefore, it is possible to appropriately apply the corresponding region in the background image to the background region in the intermediate face motion image.
[0082] In addition, the decoding apparatus of Example 6 can be the decoding apparatus of Example 3, wherein the background region in the reference image has a prescribed background color code embedded therein.
[0083] Thus, it is possible to efficiently determine the background region in the reference image based on the prescribed background color code. Furthermore, it is possible to suppress distortion occurring in the background of the reference image even if motion is imparted to the face in the reference image.
[0084] In addition, the decoding apparatus of Example 7 can be the decoding apparatus of Example 6, wherein the circuit determines a region having the prescribed background color code embedded therein in the intermediate face motion image as the background region in the intermediate face motion image.
[0085] Thus, it is possible to efficiently determine the background region in the intermediate face motion image based on the prescribed background color code. Specifically, assuming that the intermediate face motion image obtained by imparting motion to the face of the reference image has a background region with a prescribed background color code embedded therein, the reference image has a background region with a prescribed background color code embedded therein. Therefore, it is possible to efficiently determine the background region in the intermediate face motion image based on the prescribed background color code.
[0086] In addition, the decoding apparatus of Example 8 can be the decoding apparatus of Example 6 or 7, wherein the circuit decodes, from the one or more streams, background color code information indicating the prescribed background color code.
[0087] Thus, it is possible to efficiently determine the background region in the reference image in accordance with the prescribed background color code obtained from the one or more streams. Also, it is possible to change the prescribed background color code in accordance with the reference image.
[0088] Further, the decoding apparatus of Example 9 can be the decoding apparatus of Example 8, wherein the background color code information indicates a range including a plurality of values as the prescribed background color code, the prescribed background color code being prescribed within the range indicated by the background color code information.
[0089] Thus, it is possible to flexibly prescribe the prescribed background color code. Also, it is possible to flexibly apply the prescribed background color code to the background region.
[0090] Further, the decoding apparatus of Example 10 can be the decoding apparatus of any one of Examples 6 to 9, wherein the prescribed background color code is prescribed by a color code having a frequency of occurrence in the foreground region in the reference image that is below a threshold value.
[0091] Thus, it is possible to suppress a part of the foreground region from being erroneously determined as a part of the background region. Therefore, it is possible to appropriately determine the background region.
[0092] Further, the decoding apparatus of Example 11 can be the decoding apparatus of Example 3, wherein the circuit decodes, from the one or more streams, reference image segmentation information indicating the foreground region and the background region in the reference image, and embeds the prescribed background color code in the background region in the reference image using the reference image segmentation information.
[0093] Thus, it is possible to efficiently determine the background region in the reference image in accordance with the reference image segmentation information obtained from the one or more streams. Further, since the prescribed background color code is embedded in the background region in the reference image, it is possible to suppress distortion occurring in the background of the reference image even if the face in the reference image is given a movement.
[0094] Further, the decoding apparatus of Example 12 can be the decoding apparatus of any one of Examples 1 to 11, wherein the background image is an image prepared independently of the reference image and the captured moving image.
[0095] Thus, it is possible to apply the background image prepared separately from the reference image and the captured moving image to the synthesized face moving image. Therefore, it is possible to suppress the influence of the foreground region in the background image and the like.
[0096] Further, the decoding device of Example 13 can be the decoding device of any one of Examples 1 to 11, wherein the circuit decodes the identifier of the background image as the background information, and selects the background image from a plurality of background image candidates using the identifier.
[0097] Thus, it is possible to flexibly select a background image from a plurality of background image candidates. Therefore, it is possible to apply an appropriate background image to the synthesized face moving image in accordance with the use of the synthesized face moving image.
[0098] Further, the decoding device of Example 14 can also be the decoding device of any one of Examples 1 to 11, wherein the background image is an image included in the captured moving image, or an image obtained by synthesizing a plurality of images included in the captured moving image.
[0099] Thus, it is possible to apply a background image obtained from a captured moving image to the synthesized face moving image. Therefore, it is possible to apply a background image corresponding to a capturing situation to the synthesized face moving image.
[0100] Further, the decoding device of Example 15 can also be the decoding device of any one of Examples 1 to 11, wherein the circuit decodes, as the background information, a reference image applied to the background image.
[0101] Thus, it is possible to use a reference image as a background image. Furthermore, it is possible to suppress distortion of the background of the reference image by using the original reference image as the background image while imparting motion to the face of the reference image by the geometric attribute corresponding to each frame.
[0102] Further, the decoding device of Example 16 can be the decoding device of Example 14 or 15, wherein the circuit interpolates a deficiency of a background region in the background image using a peripheral region of the foreground region in the background image or a background region in a past synthesized face moving image, in a case where the background image includes the foreground region.
[0103] Thus, it is possible to appropriately interpolate a deficiency of a background region even if the background image includes a foreground region. Therefore, it is possible to suppress a deficiency of a background region in the synthesized face moving image.
[0104] Further, the decoding device of Example 17 can be the decoding device of Example 16, wherein the circuit performs segmentation processing on the background image, acquires background image segmentation information indicating the foreground region and the background region in the background image, and determines the foreground region and the background region in the background image using the background image segmentation information.
[0105] Thus, it is possible to appropriately determine the foreground region and the background region in the background image on the basis of the background image segmentation information obtained as a result of the segmentation processing on the background image. Therefore, it is possible to appropriately interpolate the absence of the background region in the background image.
[0106] Further, the decoding apparatus of Example 18 can be the decoding apparatus of Example 16, in which the prescribed foreground color code is embedded in the foreground region in the background image.
[0107] Thus, it is possible to efficiently determine the foreground region in the background image on the basis of the prescribed foreground color code. In addition, it is possible to suppress the foreground reflection of the face and the like in the background region in the synthesized face moving image.
[0108] Further, the decoding apparatus of Example 19 can be the decoding apparatus of Example 18, in which the circuit determines, as the foreground region in the background image, a region having the prescribed foreground color code in the background image from foreground color code information indicating the prescribed foreground color code decoded from the one or more streams.
[0109] Thus, it is possible to efficiently determine the foreground region in the background image on the basis of the prescribed foreground color code obtained from the one or more streams. Further, it is possible to change the prescribed foreground color code on the basis of the background image.
[0110] Further, the decoding apparatus of Example 20 can be the decoding apparatus of Example 19, in which the foreground color code information indicates a range including a plurality of consecutive values as the prescribed foreground color code, and the prescribed foreground color code is prescribed within the range indicated by the foreground color code information.
[0111] Thus, it is possible to flexibly prescribe the prescribed foreground color code. Further, it is possible to flexibly apply the prescribed foreground color code to the foreground region.
[0112] Further, the decoding apparatus of Example 21 can be the decoding apparatus of any one of Examples 18 to 20, in which the prescribed foreground color code is prescribed by a color code having a frequency of occurrence in the background region in the background image that is below a threshold value.
[0113] Thus, it is possible to suppress a part of the background region from being erroneously determined as a part of the foreground region. Therefore, it is possible to appropriately determine the foreground region.
[0114] Further, the decoding apparatus of Example 22 can be the decoding apparatus of any one of Examples 1 to 21, in which the stream in which the background information is decoded is the same as the stream in which the reference image is decoded or the stream in which the geometry information is decoded in the one or more streams.
[0115] Thus, it is possible to decode the background information from the same stream as the stream of the reference image or the stream of the geometry information, rather than from another stream. Therefore, it is possible to efficiently decode the background information together with the reference image or the geometry information.
[0116] Further, the decoding apparatus of Example 23 can be the decoding apparatus of any one of Examples 1 to 14, in which, in the one or more streams, the stream in which the background information is decoded is different from both the stream in which the reference image is decoded and the stream in which the geometry information is decoded.
[0117] Thus, it is possible to decode the background information from a different stream than the stream of the reference image and the stream of the geometry information, rather than from the same stream. Therefore, it is possible to decode the background information separately from the reference image or the geometry information at an arbitrary timing.
[0118] Further, the decoding apparatus of Example 24 can be the decoding apparatus of any one of Examples 1 to 22, in which the background image is decoded as a picture at the beginning of a sequence including a plurality of pictures or a picture at the beginning of a GOP (Group of Picture).
[0119] Thus, it is possible to obtain the background image early. Therefore, it is possible to apply the background image to a synthesized face motion image early.
[0120] Further, the decoding apparatus of Example 25 can be the decoding apparatus of any one of Examples 1 to 14, in which the background image is decoded as a picture from an access unit in the one or more streams.
[0121] Thus, it is possible to process the background image as a picture within an access unit. That is, it is possible to process the background image as a normal picture.
[0122] Further, the decoding apparatus of Example 26 can be the decoding apparatus of Example 25, in which the access unit in which the background image is decoded is the same as the access unit in which the reference image is decoded.
[0123] Thus, it is possible to decode the background image from the same access unit as the access unit of the reference image, rather than from another access unit. Therefore, it is possible to efficiently decode the background image together with the reference image.
[0124] Further, the decoding apparatus of Example 27 can be the decoding apparatus of Example 25, in which the access unit in which the background image is decoded is different from the access unit in which the reference image is decoded.
[0125] Thus, it is possible to decode the background image from an access unit different from the access unit of the reference image. Therefore, it is possible to decode the background image separately from the reference image at an arbitrary timing.
[0126] Further, the decoding apparatus of Example 28 can be the decoding apparatus of any one of Examples 25 to 27, wherein a signal indicating that the background image exists in the access unit is decoded from SEI (Supplemental Enhancement Information) corresponding to the access unit containing the background image.
[0127] Thus, it is possible to recognize the existence of the background image in the access unit according to the signal obtained from the SEI in the access unit. Therefore, it is possible to appropriately deliver the background image.
[0128] Further, the decoding apparatus of Example 29 can be the decoding apparatus of any one of Examples 1 to 14 and 25 to 28, wherein the background image is decoded as a picture from the access unit in the one or more streams, and a signal indicating that the background image exists in the access unit is decoded from SEI (Supplemental Enhancement Information) corresponding to the access unit containing the background image.
[0129] Thus, it is possible to process the background image as a picture within the access unit. That is, it is possible to process the background image as well as a normal picture. In addition, it is possible to recognize the existence of the background image in the access unit according to the signal obtained from the SEI in the access unit. Therefore, it is possible to appropriately deliver the background image.
[0130] Further, the decoding apparatus of Example 30 can be the decoding apparatus of any one of Examples 1 to 29, wherein the background image is decoded as an intra-picture.
[0131] Thus, it is possible to process the background image as an intra-picture. That is, it is possible to process the background image independently of other pictures.
[0132] In addition, the decoding apparatus of Example 31 can be the decoding apparatus of any one of Examples 1 to 30, wherein the background image is commonly applied to a plurality of frames of the synthesized face moving image.
[0133] Thus, it is possible to suppress the amount of encoding of the entire synthesized face moving image. In addition, it is possible to suppress the amount of processing for decoding the background image.
[0134] Further, the decoding apparatus of Example 32 can be the decoding apparatus of Example 8, wherein the circuit decodes the background color code information from SEI (Supplemental Enhancement Information) in the one or more streams.
[0135] Thus, it is possible to efficiently determine the background region in the reference image in accordance with the prescribed background color code obtained from the SEI. Further, it is possible to change the prescribed background color code in accordance with the reference image.
[0136] Further, the decoding apparatus of Example 33 can be the decoding apparatus of Example 19, wherein the circuit decodes the foreground color code information from SEI (Supplemental Enhancement Information) in the one or more streams.
[0137] Thus, it is possible to efficiently determine the foreground region in the background image in accordance with the prescribed foreground color code obtained from the SEI. Then, it is possible to change the prescribed foreground color code in accordance with the background image.
[0138] Further, the decoding apparatus of Example 34 is the decoding apparatus of any one of Examples 1 to 33, wherein the circuit decodes at least one of background color code information indicating a prescribed background color code and foreground color code information indicating a prescribed foreground color code from SEI (Supplemental Enhancement Information) in the one or more streams.
[0139] Thus, it is possible to efficiently determine the background region in accordance with the prescribed background color code obtained from the SEI.
[0140] Further, the decoding apparatus of Example 35 can be the decoding apparatus of Example 5, wherein the circuit decodes the photographed moving image segmentation information from SEI (Supplemental Enhancement Information) in the one or more streams.
[0141] Thus, it is possible to appropriately determine the background region in the intermediate facial moving image in accordance with the photographed moving image segmentation information obtained from the SEI. Therefore, it is possible to appropriately apply the corresponding region in the background image to the background region in the intermediate facial moving image.
[0142] Further, the decoding apparatus of Example 36 can be the decoding apparatus of Example 11, wherein the circuit decodes the reference image segmentation information from SEI (Supplemental Enhancement Information) in the one or more streams.
[0143] Thus, it is possible to efficiently determine the background region in the reference image from the reference image segmentation information obtained from the SEI. Also, it is possible to appropriately embed the prescribed background color code in the background region in the reference image.
[0144] Further, the decoding apparatus of Example 37 can be the decoding apparatus of any one of Examples 1 to 36, wherein the circuit decodes at least one of the captured moving image segmentation information indicating the foreground region and the background region in the captured moving image and the reference image segmentation information indicating the foreground region and the background region in the reference image from SEI (Supplemental Enhancement Information) in the one or more streams.
[0145] Thus, it is possible to efficiently determine the background region from the segmentation information obtained from the SEI.
[0146] Further, the decoding apparatus of Example 38 can include a memory and a circuit connected to the memory, wherein the circuit, in operation, decodes (i) a reference image that is an image including a face and (ii) geometry information that corresponds to a plurality of frames of a captured moving image obtained by a camera and indicates a geometry property of a subject, from one or more streams, inputs the reference image and the geometry information into a generative model, acquires an intermediate face moving image that is a moving image including the face from the generative model, and generates a synthetic face moving image by embedding a corresponding region in the reference image in a background region in the intermediate face moving image.
[0147] Thus, it is possible to acquire an intermediate face moving image from the generative model, the intermediate face moving image being obtained by imparting motion to the face of the reference image with the geometry property corresponding to each frame. Also, it is possible to apply the corresponding region of the original reference image to the background region of the intermediate face moving image. Therefore, it is possible to suppress distortion of the background of the reference image by the original reference image while imparting motion to the face of the reference image with the geometry property corresponding to each frame. Thus, it is possible to suppress degradation of the image quality when generating the synthetic face moving image.
[0148] Further, the encoding device of Example 39 includes a storage and a circuit connected to the storage, and the circuit, in operation, encodes (i) a reference image, (ii) geometry information, and (iii) background information, which are used to generate a synthesized face moving image, into one or more streams, the synthesized face moving image being a moving image including a face and being a moving image obtained by synthesizing a background image, the (i) reference image being an image including the face, the (ii) geometry information corresponding to a plurality of frames of a captured moving image obtained by a camera respectively and indicating a geometric property of a subject, and the (iii) background information being related to the background image.
[0149] Thus, it is possible to provide a reference image, a geometric property, and a background image used to generate a synthesized face moving image. Therefore, when generating the synthesized face moving image, it is possible to suppress distortion of a background of the reference image by the background image while giving motion to the face of the reference image by the geometric property corresponding to each frame. Thus, it is possible to contribute to suppression of degradation of image quality.
[0150] Further, the encoding device of Example 40 can be the encoding device of Example 39, wherein the circuit performs segmentation processing on the captured moving image, acquires captured moving image segmentation information indicating a foreground region and a background region in the captured moving image, and encodes the captured moving image segmentation information into the one or more streams.
[0151] Thus, it is possible to provide captured moving image segmentation information used to determine a foreground region and a background region in a moving image of the same kind as the captured moving image via the one or more streams. Therefore, it is possible to contribute to determination of a foreground region and a background region in an intermediate face moving image obtained by giving motion to the face of the reference image by the geometric property corresponding to each frame.
[0152] Further, the encoding device of Example 41 can be the encoding device of Example 39, wherein the circuit performs segmentation processing on the reference image, acquires reference image segmentation information indicating a foreground region and a background region in the reference image, embeds a prescribed background color code in the background region in the reference image using the reference image segmentation information, and encodes the reference image in which the prescribed background color code is embedded in the background region.
[0153] Thus, it is possible to appropriately determine a background region in the reference image from reference image segmentation information obtained as a result of segmentation processing on the reference image. Further, since the prescribed background color code is embedded in the background region in the reference image, it is possible to suppress distortion occurring in the background of the reference image even when motion is given to the face of the reference image.
[0154] Further, the encoding apparatus of Example 42 can be the encoding apparatus of Example 41, in which the circuitry encodes background color code information indicating the prescribed background color code into the one or more streams.
[0155] Thus, it is possible to provide the prescribed background color code for efficiently determining a background region in a reference image via one or more streams. Also, it is possible to change the prescribed background color code depending on the reference image.
[0156] Further, the encoding apparatus of Example 43 can be the encoding apparatus of Example 42, in which the background color code information indicates a range including a plurality of consecutive values as the prescribed background color code, the prescribed background color code being prescribed within the range indicated by the background color code information.
[0157] Thus, it is possible to flexibly prescribe the prescribed background color code. Also, it is possible to flexibly apply the prescribed background color code to a background region.
[0158] Further, the encoding apparatus of Example 44 can be the encoding apparatus of any one of Examples 41 to 43, in which the prescribed background color code is prescribed by a color code having a frequency of occurrence in the foreground region in the reference image that is below a threshold value.
[0159] Thus, it is possible to suppress a portion of the foreground region from being erroneously determined as a portion of the background region. Therefore, it is possible to appropriately determine the background region.
[0160] Further, the encoding apparatus of Example 45 can be the encoding apparatus of Example 39, in which the circuitry performs segmentation processing on the reference image, acquires reference image segmentation information indicating a foreground region and a background region in the reference image, and encodes the reference image segmentation information into the one or more streams.
[0161] Thus, it is possible to provide reference image segmentation information for determining a foreground region and a background region in a reference image via one or more streams. Therefore, it is possible to facilitate determination of a background region in a reference image.
[0162] Further, the encoding apparatus of Example 46 can be the encoding apparatus of any one of Examples 39 to 45, in which the background image is an image prepared independently of the reference image and the captured moving image.
[0163] Thus, it is possible to apply a background image prepared separately from a reference image and a captured moving image to a synthesized face moving image. Therefore, it is possible to suppress the influence of a foreground region in a background image, and the like.
[0164] Further, the encoding apparatus of Example 47 can be the encoding apparatus of any one of Examples 39 to 45, in which the circuitry selects the background image from among a plurality of background image candidates, and encodes an identifier of the background image as the background information.
[0165] Thus, it becomes possible to flexibly select a background image from among a plurality of background image candidates. Therefore, it becomes possible to apply an appropriate background image to the synthesized face moving image in accordance with the use of the synthesized face moving image.
[0166] Further, the encoding apparatus of Example 48 can be the encoding apparatus of any one of Examples 39 to 45, in which the background image is an image included in the captured moving image, or an image obtained by synthesizing a plurality of images included in the captured moving image.
[0167] Thus, it becomes possible to apply a background image obtained from a captured moving image to the synthesized face moving image. Therefore, it becomes possible to apply a background image corresponding to a capturing situation to the synthesized face moving image.
[0168] Further, the encoding apparatus of Example 49 can be the encoding apparatus of any one of Examples 39 to 45, in which the circuitry encodes, as the background information, a reference image that is applied to the background image.
[0169] Thus, it becomes possible to use a reference image as a background image. Also, it becomes possible to suppress distortion of the background of the reference image by using the original reference image as the background image while imparting motion to the face of the reference image by the geometric attribute corresponding to each frame.
[0170] Further, the encoding apparatus of Example 50 can be the encoding apparatus of Example 48, in which the circuitry interpolates a deficiency of a background region in the background image using a peripheral region of the foreground region in the background image or a background region in another image included in the captured moving image, in a case where the background image includes a foreground region.
[0171] Thus, it becomes possible to appropriately interpolate a deficiency of a background region even if the background image includes a foreground region. Therefore, it becomes possible to suppress a deficiency of a background region in the synthesized face moving image.
[0172] Further, the encoding apparatus of Example 51 can be the encoding apparatus of Example 50, in which the circuitry performs segmentation processing on the background image, acquires background image segmentation information indicating the foreground region and the background region in the background image, and determines the foreground region and the background region in the background image using the background image segmentation information.
[0173] Thus, it is possible to appropriately determine the foreground region and the background region in the background image on the basis of the background image segmentation information obtained as a result of the segmentation processing on the background image. Therefore, it is possible to appropriately interpolate the absence of the background region in the background image.
[0174] Further, the encoding apparatus of Example 52 can be the encoding apparatus of Example 48, in which the circuitry segments the background image, obtains background image segmentation information indicating a foreground region and a background region in the background image, and embeds a prescribed foreground color code in the foreground region in the background image using the background image segmentation information.
[0175] Thus, it is possible to efficiently determine the foreground region in the background image on the basis of the prescribed foreground color code. In addition, it is possible to suppress the foreground reflection of the face and the like in the background region in the synthesized face moving image.
[0176] Further, the encoding apparatus of Example 53 can be the encoding apparatus of Example 52, in which the circuitry encodes foreground color code information indicating the prescribed foreground color code to the one or more streams.
[0177] Thus, it is possible to provide the prescribed foreground color code for efficiently determining the foreground region in the background image via the one or more streams. Furthermore, it is possible to change the prescribed foreground color code in accordance with the background image.
[0178] Further, the encoding apparatus of Example 54 can be the encoding apparatus of Example 53, in which the foreground color code information indicates a range including a plurality of consecutive values as the prescribed foreground color code, and the prescribed foreground color code is prescribed within the range indicated by the foreground color code information.
[0179] Thus, it is possible to flexibly prescribe the prescribed foreground color code. Furthermore, it is possible to flexibly apply the prescribed foreground color code to the foreground region.
[0180] Further, the encoding apparatus of Example 55 can be the encoding apparatus of any one of Examples 52 to 54, in which the prescribed foreground color code is prescribed by a color code having a frequency of occurrence in the background region in the background image that is below a threshold value.
[0181] Thus, it is possible to suppress a part of the background region from being erroneously determined as a part of the foreground region. Therefore, it is possible to appropriately determine the foreground region.
[0182] Further, the encoding apparatus of Example 56 can be the encoding apparatus of any one of Examples 39 to 55, in which, among the one or more streams, the stream in which the background information is encoded is the same as the stream in which the reference image is encoded or the stream in which the geometry information is encoded.
[0183] Thus, it is possible to encode the background information into the same stream as the stream of the reference image or the stream of the geometry information, instead of into another stream. Therefore, it is possible to efficiently encode the background information together with the reference image or the geometry information.
[0184] Further, the encoding apparatus of Example 57 can be the encoding apparatus of any one of Examples 39 to 48, in which, among the one or more streams, the stream in which the background information is encoded is different from both the stream in which the reference image is encoded and the stream in which the geometry information is encoded.
[0185] Thus, it is possible to encode the background information from a stream different from the stream of the reference image and the stream of the geometry information, instead of from the same stream. Therefore, it is possible to encode the background information separately from the reference image or the geometry information at an arbitrary timing.
[0186] Further, the encoding apparatus of Example 58 can be the encoding apparatus of any one of Examples 39 to 56, in which the background image is encoded as a picture at the beginning of a sequence including a plurality of pictures or a picture at the beginning of a GOP (Group of Picture).
[0187] Thus, it is possible to provide the background image early. Therefore, it is possible to apply the background image to a synthesized face motion image early.
[0188] Further, the encoding apparatus of Example 59 can be the encoding apparatus of any one of Examples 39 to 48, in which the background image is encoded as a picture in an access unit in the one or more streams.
[0189] Thus, it is possible to process the background image as a picture within an access unit. That is, it is possible to process the background image as a normal picture.
[0190] Further, the encoding apparatus of Example 60 can be the encoding apparatus of Example 59, in which the access unit in which the background image is encoded is the same as the access unit in which the reference image is encoded.
[0191] Thus, it is possible to encode the background image into the same access unit as the access unit of the reference image, instead of into another access unit. Therefore, it is possible to efficiently encode the background image together with the reference image.
[0192] Further, the encoding apparatus of Example 61 can be the encoding apparatus of Example 59, wherein the access unit in which the background image is encoded is different from the access unit in which the reference image is encoded.
[0193] Thus, it is possible to encode the background image into an access unit different from the access unit of the reference image, instead of the same access unit. Therefore, it is possible to encode the background image separately from the reference image at an arbitrary timing.
[0194] Further, the encoding apparatus of Example 62 can be the encoding apparatus of any one of Examples 59 to 61, wherein a signal indicating that the background image exists in the access unit is encoded into SEI (Supplemental Enhancement Information) corresponding to the access unit in which the background image is encoded.
[0195] Thus, it is possible to notify the existence of the background image in the access unit by the SEI signal in the access unit. Therefore, it is possible to appropriately deliver the background image.
[0196] Further, the encoding apparatus of Example 63 can be the encoding apparatus of any one of Examples 39 to 48 and 59 to 62, wherein the background image is encoded as a picture into the access unit in the one or more streams, and a signal indicating that the background image exists in the access unit is encoded into SEI (Supplemental Enhancement Information) corresponding to the access unit in which the background image is encoded.
[0197] Thus, it is possible to handle the background image as a picture within the access unit. That is, it is possible to handle the background image as a normal picture. Thus, it is possible to notify the existence of the background image in the access unit by the SEI signal in the access unit. Therefore, it is possible to appropriately deliver the background image.
[0198] Further, the encoding apparatus of Example 64 can be the encoding apparatus of any one of Examples 39 to 63, wherein the background image is encoded as an intra picture.
[0199] Thus, it is possible to handle the background image as an intra picture. That is, it is possible to handle the background image independently of other pictures.
[0200] Further, the encoding apparatus of Example 65 can be the encoding apparatus of Example 42, wherein the circuitry encodes the background color code information into SEI (Supplemental Enhancement Information) in the one or more streams.
[0201] Thus, it is possible to provide the prescribed foreground color code for efficiently determining the foreground region in the background image via the SEI. Also, it is possible to vary the prescribed foreground color code depending on the background image.
[0202] Further, the encoding apparatus of Example 66 can be the encoding apparatus of Example 53, wherein the circuitry encodes the foreground color code information to SEI (Supplemental Enhancement Information) in the one or more streams.
[0203] Thus, it is possible to provide the prescribed foreground color code for efficiently determining the foreground region in the background image via the SEI. Also, it is possible to vary the prescribed foreground color code depending on the background image.
[0204] Further, the encoding apparatus of Example 67 can be the encoding apparatus of any one of Examples 39 to 66, wherein the circuitry encodes at least one of background color code information indicating the prescribed background color code and foreground color code information indicating the prescribed foreground color code to SEI (Supplemental Enhancement Information) in the one or more streams.
[0205] Thus, it is possible to provide the prescribed foreground color code for efficiently determining the foreground region via the SEI.
[0206] Further, the encoding apparatus of Example 68 can be the encoding apparatus of Example 40, wherein the circuitry encodes the photographed moving image segmentation information to SEI (Supplemental Enhancement Information) in the one or more streams.
[0207] Thus, it is possible to provide the photographed moving image segmentation information for determining the foreground region and the background region in a moving image of the same kind as the photographed moving image via the SEI. Therefore, it is possible to contribute to determination of the foreground region and the background region in an intermediate face moving image obtained by imparting a motion to a face of a reference image by a geometric attribute corresponding to each frame.
[0208] Further, the encoding apparatus of Example 69 can be the encoding apparatus of Example 45, wherein the circuitry encodes the reference image segmentation information to SEI (Supplemental Enhancement Information) in the one or more streams.
[0209] Thus, it is possible to provide the reference image segmentation information for determining the foreground region and the background region in the reference image via the SEI. Therefore, it is possible to facilitate the determination of the background region in the reference image.
[0210] Further, the encoding apparatus of Example 70 is the encoding apparatus of any one of Examples 39 to 69, wherein the circuitry encodes at least one of the captured moving image segmentation information indicating the foreground region and the background region in the captured moving image, and the reference image segmentation information indicating the foreground region and the background region in the reference image into the SEI (Supplemental Enhancement Information) in the one or more streams.
[0211] Thus, it is possible to provide the segmentation information for determining the foreground region and the background region via the SEI.
[0212] Further, the encoding apparatus of Example 71 includes a memory and a circuitry connected to the memory, the circuitry, in operation, encodes (i) a reference image that is an image containing a face, and (ii) geometry information that corresponds to a plurality of frames of a captured moving image obtained by a camera respectively and indicates a geometric property of a subject, into one or more streams, the reference image and the geometry information being used to input the reference image and the geometry information into a generation model when generating a synthesized face moving image, to obtain an intermediate face moving image that is a moving image containing the face from the generation model, the reference image being further used to generate the synthesized face moving image by embedding a corresponding region in the reference image in a background region in the intermediate face moving image.
[0213] Thus, it is possible to provide the reference image and the geometric property for generating the synthesized face moving image. Also, when generating the synthesized face moving image, it is possible to suppress distortion of a background of the reference image by the original reference image while imparting motion to the face of the reference image by the geometric property corresponding to each frame. Thus, it is possible to facilitate suppression of degradation of image quality.
[0214] Further, the bitstream generation apparatus of Example 72 includes a storage and a circuit connected to the storage, the circuit, in operation, generates a stream including (i) a reference image that is an image including a face, (ii) geometry information that corresponds to a plurality of frames of a captured moving image obtained by a camera respectively and represents a geometry property of a subject, and (iii) background information that is related to a background image, and generates a synthesized face moving image that is a moving image including the face and is obtained by synthesizing the background image.
[0215] Thus, it is possible to provide a reference image, a geometry property, and a background image for generating a synthesized face moving image. Therefore, when generating the synthesized face moving image, it is possible to suppress distortion of a background of the reference image by the background image while imparting motion to the face of the reference image by the geometry property corresponding to each frame. Thus, it is possible to contribute to suppression of degradation of image quality.
[0216] Further, the decoding method of Example 73 decodes (i) a reference image that is an image including a face, (ii) geometry information that corresponds to a plurality of frames of a captured moving image obtained by a camera respectively and represents a geometry property of a subject, and (iii) background information that is related to a background image, from one or more streams, and generates a synthesized face moving image that is a moving image including the face and is obtained by synthesizing the background image, using a generation model, from the reference image, the geometry information, and the background information.
[0217] Thus, it is possible to apply a reference image, a geometry property, and a background image to generation of a synthesized face moving image. Therefore, it is possible to suppress distortion of a background of the reference image by the background image while imparting motion to the face of the reference image by the geometry property corresponding to each frame. Thus, it is possible to suppress degradation of image quality when generating the synthesized face moving image.
[0218] Further, the encoding method of Example 74 encodes (i) a reference image that is an image including a face, (ii) geometry information that corresponds to a plurality of frames of a captured moving image obtained by a camera respectively and represents a geometry property of a subject, and (iii) background information that is related to a background image, into one or more streams, for generating a synthesized face moving image that is a moving image including the face and is obtained by synthesizing the background image.
[0219] Thus, it is possible to provide a reference image, a geometry attribute, and a background image used for generating a synthesized face moving image. Therefore, when generating a synthesized face moving image, it is possible to suppress distortion of a background of a reference image by a background image while giving motion to a face of the reference image by a geometry attribute corresponding to each frame. Thus, it is possible to contribute to suppression of degradation of image quality.
[0220] Further, these general or specific aspects can be implemented by a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a CD-ROM, or any combination of the system, the apparatus, the method, the integrated circuit, the computer program, and the recording medium.
[0221] [Definition of Terms] As an example, each term can be defined as follows.
[0222] (1) Image A unit of data composed of a set of pixels, composed of a picture, a block smaller than a picture, and including a still image in addition to a moving image.
[0223] (2) Picture A processing unit of an image composed of a set of pixels, sometimes also referred to as a frame, a field.
[0224] (3) Block A processing unit including a set of a specific number of pixels, the name is not limited as exemplified by the following examples. In addition, the shape is not limited, for example, not only including a rectangle composed of MxN pixels, a square composed of MxM pixels, but also a triangle, a circle, and other shapes.
[0225] (Examples of Blocks) • Slice / Tile / Brick • CTU / Superblock / Basic Partition Unit • VPDU / Hardware Processing Partition Unit • CU / Processing Block Unit / Prediction Block Unit (PU) / Orthogonal Transformation Block Unit (TU) / Cell • Subblock (4) Pixel / Sample As a point constituting the smallest unit of an image, not only an integer position pixel, but also a fractional position pixel generated based on an integer position pixel is included.
[0226] (5) Pixel Value / Sample Value As an inherent value possessed by a pixel, not only a luminance value, a color difference value, a gray scale of RGB, but also a depth value or two values of 0 and 1 are included.
[0227] (6) Mark In addition to 1 bit, there are cases where multiple bits are included, such as parameters or indexes of 2 bits or more. Furthermore, not only binary numbers but also multivalues using other base numbers can be used.
[0228] (7) Signal In order to convey information, a symbolization or encoding is performed, and in addition to a digital signal after discretization, an analog signal that takes a continuous value is also included.
[0229] (8) Stream / bit stream It refers to a data string of digital data or a stream of digital data. In addition to 1 stream, a stream / bit stream can be composed of multiple streams divided into multiple levels. Furthermore, in addition to the case where it is transmitted in a serial communication manner through a single transmission path, a case where it is transmitted in a packet communication manner through multiple transmission paths is also included.
[0230] (9) Difference / differential In the case of a scalar, in addition to a simple difference (x-y), as long as an operation including a difference is included, an absolute value of a difference (|x-y|), a squared difference (x^2-y^2), a square root of a difference (sqrt(x-y)), a weighted difference (ax-by: a, b are constants), and an offset difference (x-y+a: a is an offset) are included.
[0231] (10) Sum In the case of a scalar, in addition to a simple sum (x+y), as long as an operation including a sum is included, an absolute value of a sum (|x+y|), a squared sum (x^2+y^2), a square root of a sum (sqrt(x+y)), a weighted sum (ax+by: a, b are constants), and an offset sum (x+y+a: a is an offset) are included.
[0232] (11) Based on Cases where elements other than elements to be considered as the basis are also included are also included. Furthermore, in addition to the case where the result is directly obtained, a case where the result is obtained via an intermediate result is also included.
[0233] (12) Used, using Cases where elements other than elements to be considered as the basis are also included are also included. Furthermore, in addition to the case where the result is directly obtained, a case where the result is obtained via an intermediate result is also included.
[0234] (13) Prohibit, forbid In other words, not allowed. Furthermore, not prohibiting or allowing does not necessarily mean an obligation.
[0235] (14) Limit, restriction, restrict, restricted In other words, not allowed. Furthermore, not prohibited or allowed does not necessarily mean an obligation. Also, as long as a part is prohibited in terms of quantity or quality, cases where it is completely prohibited are also included.
[0236] (15) Chroma An adjective denoted by the symbols Cb and Cr that specifies a sample array or a single sample representation of one of the two color difference signals associated with the primary colors. The term chrominance can also be used instead of the term chroma.
[0237] (16) Luma An adjective denoted by the symbol or subscript Y or L that specifies a sample array or a single sample representation of a monochrome signal associated with the primary colors. The term luminance can also be used instead of the term luma.
[0238] [Explanation of Descriptions] In the drawings, the same reference numbers denote the same or similar constituent elements. Furthermore, the sizes and relative positions of the constituent elements in the drawings are not necessarily to scale.
[0239] Hereinafter, the embodiments will be specifically described with reference to the drawings. In addition, the embodiments described below each represent an inclusive or specific example. The numerical values, shapes, materials, arrangement positions and connection modes of the constituent elements, steps, relationships and orders of the steps, and the like shown in the embodiments below are one example, and are not intended to limit the scope of the request.
[0240] Hereinafter, embodiments of an encoding apparatus and a decoding apparatus will be described. The embodiments are examples of an encoding apparatus and a decoding apparatus that can apply the processing and / or structure described in each aspect of the present disclosure. The processing and / or structure can also be implemented in an encoding apparatus and a decoding apparatus different from the embodiments. For example, with respect to the processing and / or structure applied to the embodiments, for example, any one of the following can also be implemented.
[0241] (1) Any one of the plurality of constituent elements of the encoding apparatus or the decoding apparatus of the embodiments described in each aspect of the present disclosure can be replaced with or combined with other constituent elements described in any one of the aspects of the present disclosure.
[0242] (2) In the encoding apparatus or the decoding apparatus of the embodiments, arbitrary changes such as addition, substitution, deletion, and the like of functions or processes performed by some of the constituent elements of the encoding apparatus or the decoding apparatus can be made. For example, any of the functions or processes can be substituted with or combined with other functions or processes described in any of the aspects of the present disclosure.
[0243] (3) In the method implemented by the encoding apparatus or the decoding apparatus of the embodiments, arbitrary changes such as addition, substitution, deletion, and the like of some of the processes included in the method can be made. For example, any of the processes in the method can be substituted with or combined with other processes described in any of the aspects of the present disclosure.
[0244] (4) Some of the constituent elements constituting the encoding apparatus or the decoding apparatus of the embodiments can be combined with the constituent elements described in any of the aspects of the present disclosure, can be combined with the constituent elements having some of the functions described in any of the aspects of the present disclosure, and can be combined with the constituent elements performing some of the processes implemented by the constituent elements described in any of the aspects of the present disclosure.
[0245] (5) The constituent elements having some of the functions of the encoding apparatus or the decoding apparatus of the embodiments, or the constituent elements performing some of the processes implemented by the encoding apparatus or the decoding apparatus of the embodiments can be combined with or substituted with the constituent elements described in any of the aspects of the present disclosure, the constituent elements having some of the functions described in any of the aspects of the present disclosure, or the constituent elements performing some of the processes described in any of the aspects of the present disclosure.
[0246] (6) In the method implemented by the encoding apparatus or the decoding apparatus of the embodiments, any of the processes included in the method can be substituted with or combined with the processes described in any of the aspects of the present disclosure, or the same any of the processes.
[0247] (7) Some of the processes included in the method implemented by the encoding apparatus or the decoding apparatus of the embodiments can be combined with the processes described in any of the aspects of the present disclosure.
[0248] (8) The embodiments of the processes and / or structures described in the aspects of the present disclosure are not limited to the encoding apparatus or the decoding apparatus of the embodiments. For example, the processes and / or structures can be implemented in an apparatus utilized for a purpose different from the motion picture encoding or the motion picture decoding disclosed in the embodiments.
[0249] [Structure of an encoding / decoding system] Figure 4is a block diagram showing an example of a configuration of an encoding and decoding system in the present embodiment. The encoding and decoding system includes, for example, an encoding apparatus 100 and a decoding apparatus 200. Figure 4 The example of Figure 1 is similar to that of Figure 4 , but the specific configuration and processing of the encoding apparatus 100, the specific configuration and processing of the decoding apparatus 200, and the bitstream differ from those of Figure 1
[0250] The encoding apparatus 100 receives a reference image, a driving motion image, and a background image, and generates a bitstream. Next, the encoding apparatus 100 transmits the bitstream to the decoding apparatus 200 using a transmission channel. Finally, the decoding apparatus 200 reconstructs a synthesized face motion image from the bitstream.
[0251] The reference image is an image including a face, and can be expressed as a face image or an identity image, for example. The reference image indicates a static visual feature used for reconstructing a synthesized face motion image. The driving motion image is a motion image including a face, and is a captured motion image obtained by a camera. The driving motion image functions to impart motion to the reference image. The bitstream is simply expressed as a stream. Furthermore, the use of one bitstream is not limited, and multiple bitstreams can be used.
[0252] The person included in the reference image and the person included in the driving motion image can be the same person or can not be the same person.
[0253] In addition, the encoding and decoding system in the present embodiment can be applied to video conferencing, video production and editing in the entertainment industry, social media and e-commerce industry, and the like. However, the range of application is not limited thereto.
[0254] [Data Structure] Figure 8 is a diagram showing an example of a hierarchical structure of data in a stream. The stream includes, for example, a video sequence. For example, as shown in (a) of Figure 8 , the video sequence includes a VPS (Video Parameter Set), an SPS (Sequence Parameter Set), a PPS (Picture Parameter Set), an SEI (Supplemental Enhancement Information), and a plurality of pictures.
[0255] The VPS includes, in a motion image composed of a plurality of layers, encoding parameters common to the plurality of layers, and encoding parameters associated with the plurality of layers or each layer included in the motion image.
[0256] The SPS includes parameters used for the sequence, i.e., encoding parameters referenced by the decoding device 200 for decoding the sequence. For example, these encoding parameters can represent the width or height of the image. Furthermore, multiple SPSs can exist.
[0257] The PPS includes parameters used for the images, i.e., encoding parameters referenced by the decoding device 200 for decoding each image in the sequence. For example, these encoding parameters may include a reference value for the quantization width used for image decoding and a flag indicating the application of weighted prediction. Furthermore, multiple PPSs may exist. Additionally, SPS and PPS are sometimes simply referred to as parameter sets.
[0258] like Figure 8 As shown in (b), the image may include an image header and one or more slices. The image header includes encoding parameters referenced by the decoding device 200 for decoding the one or more slices.
[0259] like Figure 8 As shown in (c), the slice includes a slice header and one or more bricks. The slice header includes encoding parameters referenced by the decoding device 200 for decoding the one or more bricks.
[0260] like Figure 8 As shown in (d), the brick contains more than one CTU (Coding Tree Unit).
[0261] Alternatively, the image may exclude the slice and instead include a set of tiles. In this case, the set of tiles includes more than one tile. Furthermore, the brick may include a slice.
[0262] CTU is also known as a superblock or basic partitioning unit. For example... Figure 8 As shown in (e), this CTU includes a CTU header and one or more CUs (Coding Units). The CTU header includes encoding parameters referenced by the decoding device 200 for decoding the one or more CUs.
[0263] A CU can be divided into multiple smaller CUs. Furthermore, as... Figure 8The CU includes a CU header, prediction information, and residual coefficient information as shown in (f). The prediction information is information used for predicting the CU, and the residual coefficient information is information indicating a prediction residual described later. In addition, the CU is basically the same as a PU (Prediction Unit) and a TU (Transform Unit), but can include a plurality of TUs smaller than the CU, for example, in SBT described later. Furthermore, the CU can be processed per each VPDU (Virtual Pipeline Decoding Unit) constituting the CU. The VPDU is, for example, a fixed unit capable of being processed in one stage when pipeline processing is performed in hardware.
[0264] In addition, the stream can not have Figure 8 In addition, the order of these levels can be exchanged, and any level can be replaced with another level. Furthermore, a picture that is an object of processing performed by the encoding apparatus 100 or the decoding apparatus 200 or the like at a current time point is referred to as a current picture. If the processing is encoding, the current picture is synonymous with an encoding target picture, and if the processing is decoding, the current picture is synonymous with a decoding target picture. Furthermore, a block such as a CU or the like that is an object of processing performed by the encoding apparatus 100 or the decoding apparatus 200 or the like at a current time point is referred to as a current block. If the processing is encoding, the current block is synonymous with an encoding target block, and if the processing is decoding, the current block is synonymous with a decoding target block.
[0265] Here, a region in which a parameter used for encoding and decoding is described can be expressed as a header region. For example, the header region is a region including SEI. The header region can also include VPS, SPS, PPS, SEI, a picture header, a slice header, a CTU header, and a CU header.
[0266] Furthermore, for example, a picture can be classified into any one of a plurality of kinds including an I picture, a P picture, and a B picture. The I picture is an intra prediction picture, also referred to as an intra picture, and is a picture that is encoded and decoded without referring to other pictures. The P picture is a unidirectional prediction picture and is a picture that can be encoded and decoded by referring to one other picture. The B picture is a bidirectional prediction picture and is a picture that can be encoded and decoded by referring to two other pictures.
[0267] Further, a moving image can be constituted of a plurality of GOPs (Group of Pictures). The GOP refers to a set of pictures. The GOP includes one or more I pictures. The GOP can include one or more P pictures, and can include one or more B pictures. The GOP can be a unit capable of editing and random access of a video, and the like. The GOP can have a certain number of pictures, and can have a certain configuration order related to I pictures, P pictures, and B pictures as a GOP structure.
[0268] [Structure and processing of encoding] Figure 9 is a block diagram showing a configuration example of the encoding apparatus 100 in the present embodiment. The encoding apparatus 100 generates a bitstream from a reference image, a driving moving image, and a background image. In this example, the encoding apparatus 100 includes a compressor 131, a deriver 132, a compressor 133, and a compressor 134. Each constituent element is, for example, a circuit that performs information processing. Two or more of the compressor 131, the compressor 133, and the compressor 134 can be combined.
[0269] Figure 10 is a flowchart showing an action example of the encoding apparatus 100 in the present embodiment. For example, Figure 9 the plurality of constituent elements of the encoding apparatus 100 shown in Figure 10 act in accordance with the flowchart.
[0270] In this example, first, the compressor 131 compresses at least one reference image by encoding it into a bitstream (S101). The reference image can be encoded in accordance with a video coding method such as VVC. The reference image can be a frame of a driving moving image, can be a pre-prepared image including a face of a person, or can be a virtual avatar.
[0271] In addition, the deriver 132 derives geometry information indicating a geometric property corresponding to each frame of the driving moving image (S102). The geometry information is also simply referred to as a geometric property. Specifically, the deriver 132 inputs each frame of the driving moving image into a recognition model such as a neural network, and acquires a geometric property corresponding to each frame from the recognition model. The geometric property corresponds to a temporal instance of each frame of the driving moving image.
[0272] Here, the geometric property, for example, corresponds to a dynamic property, and can be represented by a point group such as a face landmark, or can be represented by a polygon model for representing a shape of an object by a combination of a plurality of polygons. Further, the geometric property can be represented by another geometric model. Further, the geometric property can be represented by a position of a part of a face. Further, the geometric property can be represented as a face property. Further, the geometric property can be processed as a set of geometric properties.
[0273] For example, the face landmark used as the geometric attribute indicates positions of points on main areas of the face including the outline of the face, the eyes, the eyebrows, the nose, the mouth, the lips, and the chin. Since such a geometric attribute can be interpreted by other people or other devices, the attribute can be corrected, and the processing of the attribute can be improved.
[0274] The compressor 133 compresses by encoding the geometric attribute to a bitstream using a method such as entropy coding (S103).
[0275] The compressor 134 compresses by encoding at least 1 background image to a bitstream (S104). The background image can be encoded in a video coding manner such as VVC.
[0276] The background image is used for synthesizing a background area in the face motion image. That is, the background image represents a background superimposed on the face motion image containing the face. The background image can be an image corresponding to a frame contained in the driving motion image, or can be generated from a plurality of images corresponding to a plurality of frames contained in the driving motion image. The background image can be a background image prepared in advance separately from the reference image and the driving motion image.
[0277] Further, the background image can be selected from a plurality of background image candidates. In addition to the background image, a selection parameter for selecting the background image from the plurality of background image candidates can be encoded. The selection parameter can be an identifier of the background image corresponding to any one of the plurality of background image candidates.
[0278] Further, the background image can be a solid color style, a texture style, a gradient style, a pattern style, a blur style, an illustration style, a high-contrast style, a real-world scene, or a synthetic scene, or the like. Further, the background image can use any combination thereof, or can use an image different from them.
[0279] The encoding apparatus 100 can generate a synthesized face motion image based on the reference image, the geometric attribute, and the background image, as well as the action of the decoding apparatus 200 to be described later (S105). To generate the synthesized face motion image, the encoding apparatus 100 can have a plurality of constituent elements similar to the decoding apparatus 200. Thereby, the synthesized face motion image generated by the decoding apparatus 200 can be confirmed by the encoding apparatus 100. In addition, this processing can be omitted.
[0280] The encoding apparatus 100 transmits the bitstream to the decoding apparatus 200 via a transmission channel after encoding the reference image, the geometry attribute, and the background image into the bitstream. For example, the geometry attribute is transmitted as a bitstream from the encoding apparatus 100 to the decoding apparatus 200 in accordance with each frame of the driving moving image, that is, each time instance. The geometry attribute can be transmitted as SEI (Supplemental Enhancement Information).
[0281] For the generation of the synthesized face moving image, the reference image can be transmitted only once. Also, in the decoding apparatus 200, the same reference image can be used when generating each frame of the synthesized face moving image. Likewise, for the generation of the synthesized face moving image, the background image can be transmitted only once. Also, in the decoding apparatus 200, the same background image can be used when generating each frame of the synthesized face moving image.
[0282] Alternatively, the background image can be transmitted in accordance with each frame of the driving moving image, that is, each time instance, as with the geometry attribute. Alternatively, the background image can be transmitted as a plurality of key frames respectively, the refresh rate of the background image corresponding to the key frame interval.
[0283] Alternatively, the encoding apparatus 100 can track the position of the face in the driving moving image and transmit the background image in a case where the motion of the face such as parallel movement or rotation exceeds a threshold value. In this way, the refresh rate of the background image can depend on the motion of the face in the driving moving image.
[0284] The reference image, the geometry attribute, and the background image can be encoded into the same one bitstream and transmitted, or can be encoded into different bitstreams respectively and transmitted. Alternatively, any two of the reference image, the geometry attribute, and the background image can be encoded into the same one bitstream and transmitted, and the remaining one can be encoded into another bitstream and transmitted.
[0285] [Structure and processing of decoding apparatus] Figure 11 is a block diagram showing an example of the structure of the decoding apparatus 200 in the present embodiment. The decoding apparatus 200 generates a synthesized face moving image from a bitstream. In this example, the decoding apparatus 200 is provided with a decompressor 231, an extractor 232, a decompressor 233, a generator 234, a decompressor 235, and a synthesizer 236. Each constituent element is, for example, a circuit that performs information processing. Two or more of the decompressor 231, the decompressor 233, and the decompressor 235 can be unified.
[0286] Figure 12 is a flowchart showing an example of the operation of the decoding apparatus 200 in the present embodiment. For example, Figure 11The plurality of constituent elements of the decoding apparatus 200 shown act in accordance with the flowchart of Figure 12 Hereinafter, it is possible to omit the same explanation as that of the encoding.
[0287] The decompressor 231 decompresses by decoding at least one reference image from the bitstream (S201). The reference image can be decoded in accordance with a video coding scheme such as VVC. Then, the decompressor 231 supplies the reference image to the extractor 232.
[0288] The extractor 232 extracts reference information indicating a reference attribute from the reference image (S202). Here, the reference information indicating the reference attribute can simply be expressed as the reference attribute. The reference image is a static visual attribute, and can also be expressed as an identity. The reference attribute can include information on at least one of hair, glasses, a beard, eyebrows, eyes, a mouth, a nose, skin, a face contour, clothing, and accessories.
[0289] The decompressor 233 decompresses by decoding the geometry attribute per frame from the bitstream using an entropy decoding or the like (S203).
[0290] The generator 234 generates an intermediate facial motion image from the reference attribute and the geometry attribute using a generative model of a neural network or the like (S204).
[0291] The generative model can be a GAN (Generative Adversarial Network), a VAE (Variational Autoencoder), an autoregressive model, or a diffusion model, or the like. For example, the generative model can be a machine learning framework that generates new data based on a provided data set, and in the generative model, analysis and learning of a basic distribution of the data set can be performed.
[0292] For example, the generator 234 inputs the reference attribute and the geometry attribute per frame to the generative model, and acquires the intermediate facial motion image and the segmentation mask from the generative model. More specifically, the generator 234 inputs the reference attribute and the geometry attribute per frame to the generative model, and acquires a frame of the intermediate facial motion image and a segmentation mask of the frame from the generative model. The segmentation mask indicates a foreground region and a background region in the intermediate facial motion image (specifically, the frame of the intermediate facial motion image).
[0293] A segmented mask can be represented as a two-dimensional map where all pixels in the foreground region are 1 and all pixels in the background region are 0, or it can be represented as a two-dimensional map where all pixels in the foreground region are 0 and all pixels in the background region are 1. For example, the foreground region can be a region containing faces and other features that are motion-dependent, while the background region can be a region not containing faces and other features that are motion-dependent. A segmented mask is also called segmented information.
[0294] Generator 234 can render the intermediate face motion image using the reference image itself instead of the reference attributes, or it can render the intermediate face motion image using the reference image itself in addition to the reference attributes. Furthermore, the exporter 232 may or may not be included in generator 234. Moreover, the recognition model used to export the reference attributes in exporter 232 may be included in the generative model used to generate the intermediate face motion image, etc., in generator 234. The same applies to other variations regarding exporter 232 and reference attributes.
[0295] In other words, generator 234 can use a generative model to generate a segmented mask and an intermediate facial motion image based on the reference image and geometric attributes. At this time, generator 234 can input the reference image and geometric attributes into the generative model and obtain the segmented mask and the intermediate facial motion image from the generative model.
[0296] The decompressor 235 decompresses the data by decoding at least one background image from the bitstream (S205). The background image can be decoded using a video codec such as VVC. In addition to the background image, selection parameters for choosing a background image from multiple background image candidates can also be decoded. The selection parameters can be identifiers of the background image corresponding to any one of the multiple background image candidates.
[0297] The synthesizer 236 uses an intermediate face motion image, a segmented mask, and a background image to generate a synthesized face motion image by embedding the corresponding region in the background image into the background region in the intermediate face motion image (S206).
[0298] Figure 13 This is a conceptual diagram illustrating an example of decoding processing in each time instance. In this example, the decoding device 200 receives a compressed reference image, compressed geometric properties, and a compressed background image at the initial time instance (t=0). The decoding device 200 then stores the compressed reference image and the compressed background image in its memory 252.
[0299] In addition, the decoding device 200 also decodes the compressed reference image, the compressed geometric properties of the initial time instance (t=0), and the compressed background image. Then, the decoding device 200 uses a generative model to generate an image of the initial time instance (t=0) in the synthesized face motion image based on the reference image, geometric properties, and background image.
[0300] Furthermore, the decoding device 200 receives compressed geometric attributes at a subsequent time instance (t=T) and retrieves and obtains a compressed reference image and a compressed background image from its memory 252. Then, the decoding device 200 performs decoding processing on the compressed reference image, the compressed geometric attributes of that time instance (t=T), and the compressed background image. Finally, the decoding device 200 uses a generative model to generate an image of that time instance (t=T) in the synthesized facial motion image based on the reference image, geometric attributes, and background image.
[0301] The decoding device 200 can store the reference image and background image obtained by decoding the compressed reference image and the compressed background image in its memory 252. Then, in a subsequent time instance (t=T), the decoding device 200 can retrieve the reference image and background image that have been decoded from the memory 252 and apply the reference image and background image retrieved from the memory 252 to the generation of the image in the synthesized facial motion image.
[0302] As this example illustrates, compressed geometric attributes are received in each temporal instance. The geometric attributes of a temporal instance are derived from a frame in the driving motion picture. On the other hand, the compressed reference image and the compressed background image can be received only in the initial temporal instance. This makes it possible to reduce the amount of coding.
[0303] Figure 14 This is a conceptual diagram illustrating another example of the decoding process in each time instance. In this example, the decoding device 200 receives a compressed reference image, compressed geometric properties, and a compressed background image at the initial time instance (t=0). The decoding device 200 then stores the compressed reference image in its memory 252.
[0304] In addition, the decoding device 200 also decodes the compressed reference image, the compressed geometric properties of the initial time instance (t=0), and the compressed background image. Then, the decoding device 200 uses a generative model to generate an image of the initial time instance (t=0) in the synthesized face motion image based on the reference image, geometric properties, and background image.
[0305] Furthermore, the decoding device 200 receives compressed geometric attributes and a compressed background image at a subsequent time instance (t=T), and retrieves and obtains a compressed reference image from its memory 252. Then, the decoding device 200 performs decoding processing on the compressed reference image, the compressed geometric attributes of that time instance (t=T), and the compressed background image of that time instance (t=T). Then, the decoding device 200 uses a generative model to generate an image of that time instance (t=T) in the synthesized facial motion image based on the reference image, geometric attributes, and background image.
[0306] The decoding device 200 can store the reference image obtained by decoding the compressed reference image in its memory 252. Then, the decoding device 200 can retrieve the reference image that has been decoded from the memory 252 at a subsequent time instance (t=T) and apply the reference image retrieved from the memory 252 to the generation of the image in the synthesized facial motion image.
[0307] As shown in this example, compressed geometric properties and a compressed background image are received in each temporal instance. The geometric properties of a temporal instance are derived from a frame in the driving motion picture.
[0308] Furthermore, for example, the background image of a time instance can be a background image derived from a frame of the driving motion image. When encoding the background image, the amount of encoding of the background image can be reduced by reducing the resolution, quantizing with a larger quantization step size, or filling the foreground area with foreground color code.
[0309] Furthermore, for example, the compressed reference image can be received only in the initial time instance. This makes it possible to reduce the amount of encoding.
[0310] [Variation Example] Figure 15 This is a block diagram illustrating another structural example of the decoding device 200 in this embodiment. In the example above, namely... Figure 11 In the example, generator 234 inputs baseline and geometric attributes into the generative model and obtains the intermediate facial motion image and segmented mask from the generative model.
[0311] In contrast, in this example, Figure 15 In the example, generator 234 inputs baseline and geometric attributes into the generative model and obtains an intermediate facial motion image from the generative model. Then, generator 234 segments the intermediate facial motion image to obtain a segmented mask.
[0312] Specifically, generator 234 inputs the baseline and geometric attributes into the generation model for each frame and obtains the frames of the intermediate facial motion image from the generation model. Then, generator 234 performs segmentation processing on each frame of the intermediate facial motion image and obtains the segmentation mask of that frame.
[0313] This allows for segmented processing, potentially enabling smoother processing. A segmentation processor (not shown) can replace generator 234 for segmentation processing.
[0314] Segmentation can be performed using machine learning models such as neural networks. Other segmentation processes disclosed herein are similar.
[0315] The foreground and background regions in the intermediate facial motion image and the synthesized facial motion image generated in the decoding device 200 correspond to the foreground and background regions in the driving motion image. Therefore, the encoding device 100 can segment the driving motion image and encode a segmentation mask for the driving motion image. Then, the decoding device 200 can decode the segmentation mask and use the segmentation mask to generate a synthesized facial motion image.
[0316] Specifically, in the encoding apparatus 100, the extractor 132 can segment the driving motion image for each frame to generate segmented masks representing the foreground and background regions in the driving motion image. Additionally, the compressor 133 can perform compression by encoding the segmented masks into a bitstream.
[0317] Then, in the decoding device 200, the decompressor 233 can decompress the data by decoding the segmented mask from the bitstream. Furthermore, the synthesizer 236 can generate a synthesized facial motion image using the segmented mask, etc. This makes it possible to reduce the processing load in the decoding device 200.
[0318] Furthermore, in the encoding device 100, a segmentation processor (not shown), different from the extractor 132, can perform segmentation processing. Additionally, a compressor (not shown), different from the compressor 133, can encode the segmented mask into the bitstream. Furthermore, in the decoding device 200, a decompressor (not shown), different from the decompressor 233, can decode the segmented mask from the bitstream.
[0319] Furthermore, the segmented mask can be transmitted from the encoding device 100 to the decoding device 200 in the SEI for each frame.
[0320] Figure 16This is a block diagram illustrating another structural example of the decoding apparatus 200 in this embodiment. In this example, the generator 234 inputs reference attributes and geometric attributes into the generation model and obtains an intermediate face motion image from the generation model in which a predetermined background color code is embedded in the background region. That is, in the intermediate face motion image obtained from the generation model, the background region is filled with a predetermined background color. In other words, all pixels in the background region of the intermediate face motion image have the same predetermined background color code as pixel values.
[0321] Therefore, it is possible to determine the background region in a mid-motion image of a face without using segmented masks.
[0322] The synthesizer 236 embeds a corresponding region from the background image into the background region of the intermediate face motion image. Specifically, in the intermediate face motion image, pixels with a specified background color code are replaced with corresponding pixels in the background image, while pixels without a specified background color code remain unchanged. Thus, the synthesizer 236 generates a synthesized face motion image based on the intermediate face motion image with the specified background color code embedded in the background region and the background image.
[0323] Figure 17 This is a block diagram illustrating another structural example of the encoding device 100 in this embodiment. In this example, a reference image with a predetermined background color code embedded in its background region is used. That is, a reference image with its background region filled with a predetermined background color is used. In other words, all pixels in the background region of the reference image have the same predetermined background color code as their pixel value. For example, the predetermined background color code may already be embedded in the background region of the reference image.
[0324] Alternatively, in the encoding device 100, the compressor 131 can embed a predetermined background color code in the background region of the reference image. Specifically, the compressor 131 can segment the reference image to obtain a segmented mask representing the foreground and background regions in the reference image. Then, the compressor 131 can determine the background region in the reference image based on the segmented mask and embed the predetermined background color code in the background region of the reference image.
[0325] A preprocessor (not shown) can replace compressor 131 to segment the reference image and embed a specified background color code in the background area of the reference image.
[0326] Compressor 131 encodes a reference image into a bitstream by embedding a specified background color code in the background region. Compressor 131 can also encode background color code information representing the specified background color code into the bitstream. For example, the background color code information can be transmitted in SEI.
[0327] Figure 18 This is a block diagram illustrating another structural example of the decoding device 200 in this embodiment. Figure 18 The decoding device 200 is based on the Figure 17 The encoding device 100 generates a bitstream to produce a synthetic facial motion image. Specifically, in this example, a reference image with a specified background color code embedded in the background region is used. In other words, a reference image with the background region filled with a specified background color is used.
[0328] In the decoding apparatus 200, the decompressor 231 decodes a reference image from the bitstream in which a specified background color code is embedded in the background region. The exporter 232 exports reference attributes from the reference image in which the specified background color code is embedded in the background region. For example, since the specified background color code is embedded in the background region of the reference image, the reference attributes can be appropriately exported from the foreground region of the reference image alone.
[0329] Then, the reference attributes and geometric attributes are input into the generator model through generator 234, and the intermediate face motion image is obtained from the generator model to generate the intermediate face motion image.
[0330] The intermediate facial motion image generated by generator 234 corresponds to a reference image to which motion has been assigned through geometric attributes. Furthermore, a prescribed background color code is embedded in the background region of the reference image. Therefore, generator 234 generates an intermediate facial motion image with the prescribed background color code embedded in the background region.
[0331] Therefore, with Figure 16 Similarly, it's possible to determine the background region in a mid-motion image of a face without using segmented masks. And, with... Figure 16 Similarly, it is possible to generate a synthetic facial motion image based on an intermediate facial motion image and a background image with a specified background color code embedded in the background region.
[0332] The decompressor 231 can decode the background color code information representing a specified background color code. Then, the synthesizer 236 can determine the region in the intermediate facial motion image that has the specified background color code represented by the decoded background color code information as the background region.
[0333] exist Figure 17 In the example encoding device 100, the compressor 131 may encode the segmented mask into the bitstream (e.g., the SEI in the bitstream) instead of embedding a prescribed background color code in the background region of the reference image. Then, in Figure 18In the example decoding device 200, the decompressor 231 can decode the segmented mask and the reference image from the bitstream, and embed a specified background color code in the background area of the reference image according to the segmented mask.
[0334] In addition, not limited to Figure 17 and Figure 18 Examples, in the corresponding Figure 16 In the example encoding device 100 and decoding device 200, background color code information can be transmitted. Alternatively, without transmitting background color code information, the specified background color code in the encoding device 100 and decoding device 200 can be specified using a neural network based on a reference image, or it can be specified independently of the reference image.
[0335] For the generation of synthetic facial motion images, the background color code information only needs to be transmitted once. Moreover, in the decoding device 200, the same background color code information can be used when generating each frame of the synthetic facial motion image.
[0336] For example, the specified background color code is a color code that differs from the multiple colors (i.e., multiple pixel values) of foreground areas such as faces. This allows for the appropriate identification of background areas in moving images of the middle face.
[0337] Specifically, the prescribed background color code can be a color that is not commonly seen in all typical facial or body areas. For example, green or blue, which are colors furthest from the human body color, can be chosen as the prescribed background color code because they are used in keying techniques.
[0338] Alternatively, you can first construct a list of all possible colors. Or, you can extract all colors from the entire foreground area. Then, you can remove all extracted colors from the list. Finally, you can select the remaining colors in the list as the specified background color code.
[0339] Alternatively, all colors within the foreground area can be entered into a frequency table. The color with the highest frequency can then be determined from the frequency table. The color on the opposite side of the color wheel from the determined color can then be designated as the background color code.
[0340] Furthermore, when generating synthetic facial motion images, the pixel values of the foreground region and the background region with embedded background color codes may vary. Additionally, if the specified background color code matches the pixel values contained in the foreground region, it becomes difficult to properly determine the background region.
[0341] Therefore, the specified background color code can be defined by a range of consecutive values that are inconsistent with the pixel values contained in the foreground area. Alternatively, the encoding device 100 can encode the range of consecutive values into the specified background color code, and the decoding device 200 can decode the range into the specified background color code.
[0342] For example, the minimum value (y) of multiple consecutive ranges. min u min v min ) and the maximum value of multiple consecutive ranges (y max u max v max ( ) can be transmitted as a background color code via a bitstream. Alternatively, the intermediate value (y) of a consecutive range of values. mean u mean v mean ) and the difference (y) between the median and minimum values of multiple consecutive ranges. delta u delta v delta () can be transmitted via bitstream as a background color code.
[0343] Figure 19 This is a block diagram illustrating another structural example of the decoding apparatus 200 in this embodiment. In this example, the generator 234 generates a synthetic facial motion image using a generative model based on reference attributes, geometric attributes, and a background image. Specifically, the generator 234 generates a synthetic facial motion image by inputting the reference attributes, geometric attributes, and a background image into the generative model and obtaining the synthetic facial motion image from the generative model. This simplifies the process. In this case, the decoding apparatus 200 may not require a separate synthesizer 236.
[0344] In the examples above, the background image can be an image that does not contain foreground areas such as faces, or an image that does contain foreground areas such as faces. The foreground areas in the background image can be embedded with a specified foreground color code. That is, the foreground areas in the background image can be filled with a specified foreground color. This suppresses the appearance of foreground areas such as faces in the background image in the synthesized face image.
[0345] For example, in the encoding apparatus 100, the compressor 134 segments the background image to obtain a segment mask representing the foreground and background regions in the background image. Then, the compressor 134 can use the segment mask to embed a specified foreground color code into the foreground region of the background image. The compressor 134 can then encode the background image with the specified foreground color code embedded in the foreground region into a bitstream.
[0346] Furthermore, in the decoding device 200, the decompressor 235 can decode a background image with a specified foreground color code embedded in the foreground region from the bitstream. Thus, the decoding device 200 is able to determine the foreground region and the background region in the background image.
[0347] When a background image contains a foreground region, the missing portion or all of the background region in the background image can be interpolated through restoration or other methods. Specifically, the surrounding area of the foreground region in the background image (i.e., the background region) can be used to interpolate the missing background region in the background image. Alternatively, a background region from another image can be used to interpolate the missing background region in the background image.
[0348] Interpolation of missing background regions in the background image can be performed in the encoding device 100 or the decoding device 200.
[0349] For example, the background image can be an image contained within the driving motion image, or it can be a composite image of multiple images contained within the driving motion image. In this case, in the encoding device 100, the compressor 134 can interpolate the missing background region in the background image using the surrounding region of the foreground region in the background image or the background region in another image contained within the driving motion image. Then, the compressor 134 can encode the background image with the interpolated missing background region.
[0350] Furthermore, in the decoding device 200, the decompressor 235 or the synthesizer 236 can interpolate for missing background regions in the background image using the surrounding region of the foreground region in the background image or the background region in a previous synthesized facial motion image. Additionally, when timing the embedding of a corresponding region in the background image into the background region of an intermediate motion image, if the corresponding region contains a foreground region, the synthesizer 236 can interpolate for missing background regions in that corresponding region.
[0351] As mentioned above, selection parameters can be transmitted via bitstream instead of transmitting a background image. These selection parameters provide information related to the background.
[0352] The selection parameters can include information related to the content rating (content assessment) of the background image. Specifically, the selection parameters can represent the age group suitable for viewing the background image as a content rating.
[0353] For example, the selection parameters could indicate that the background image is classified as NC16 (not available for children under 16 years old). Furthermore, if the viewer falls into a smaller category, further processing such as blurring the background image can be performed based on the selection parameters.
[0354] This allows for the protection of minors and other groups from inappropriate background images on topics such as violence, similar to media content rating systems.
[0355] Furthermore, for example, selection parameters may include information related to the customization of the background image. Specifically, selection parameters may include information for further customizing the background image based on the viewer's profile. For example, if the viewer is female, a background image frequently chosen by women can be selected (specifically, a pink background image, etc.). In another example, if the viewer is a child, a background image frequently chosen by children can be selected (specifically, a colorful background image, etc.).
[0356] The final background image applied by each of the more than one decoding devices 200 can be changed by selecting parameters in this way.
[0357] In the examples above, the background image or background information associated with the background image is transmitted separately from the reference image and geometric attributes. The background information may be, for example, information used to obtain the background image and to select a background image from multiple background image candidates, or information about the destination of the background image. However, in other examples, the background image or background information may not be transmitted, or the background image may not be predetermined.
[0358] Specifically, for example, a reference image can be used as a background image instead of a background image transmitted separately from the reference image. Then, for example, in the decoding device 200, the synthesizer 236 can generate a synthesized facial motion image by embedding a corresponding region in the reference image into a background region in the intermediate facial motion image.
[0359] Missing parts or all of the background region in a reference image can be interpolated through repair or other methods. Specifically, the surrounding area (i.e., the background area) of the foreground region in the reference image can be used to interpolate the missing background region in the reference image.
[0360] Alternatively, background regions from other images can be used to interpolate for missing background regions in the reference image. In the case of transmitting multiple reference images, this other image can be a different reference image. Alternatively, this other image can be an image contained within an intermediate facial motion image generated using other reference images or a synthetic facial motion image.
[0361] Furthermore, the reference image can be segmented to obtain segmented masks representing the foreground and background regions in the reference image. These segmented masks can then be used to determine the foreground and background regions in the reference image. The segmentation process can be performed by either the encoding device 100 or the decoding device 200. When the segmentation process is performed by the encoding device 100, the segmented masks can be transmitted from the encoding device 100 to the decoding device 200.
[0362] In the examples above, the reference image can be considered as being encoded and decoded as a background image or background information. Alternatively, the reference image can be considered as being applied to the background image. Or, regarding the background image, it can be considered as neither being encoded nor decoded.
[0363] Furthermore, although geometric information representing geometric attributes was used in the examples above, the attribute information corresponding to each frame is not limited to geometric information representing geometric attributes. Dynamic information representing dynamic attributes in a different form than geometric attributes can be used instead of geometric information.
[0364] Figure 20 This is a block diagram illustrating a structural example of the encoding device 100 encoding moving images in this embodiment. For example, the encoding device 100 may include... Figure 20 The multiple components shown are used for encoding images contained in a moving image in block units according to VVC. In addition to the multiple components described above, the encoding device 100 may include... Figure 20 The multiple constituent elements shown can also be combined with at least a portion of the aforementioned multiple constituent elements. Figure 20 Among the various constituent elements shown.
[0365] like Figure 20 As shown, the encoding apparatus 100 includes a segmentation unit 102, a subtraction unit 104, a transform unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transform unit 114, an addition unit 116, a block memory 118, a cyclic filtering unit 120, a frame memory 122, an intra-frame prediction unit 124, an inter-frame prediction unit 126, a prediction control unit 128, and a prediction parameter generation unit 130. Furthermore, the intra-frame prediction unit 124 and the inter-frame prediction unit 126 are each configured as part of the prediction processing unit.
[0366] The segmentation unit 102 segments the image into multiple blocks and provides segmentation-related parameters to the entropy coding unit 110. The subtraction unit 104 subtracts the predicted image block from the current block to obtain a prediction residual block. The transform unit 106 transforms the prediction residual block to obtain a transform coefficient block. The quantization unit 108 quantizes the transform coefficient block to obtain a quantized coefficient block. The entropy coding unit 110 entropy-codes the quantized coefficient block and parameters to generate a bitstream.
[0367] The inverse quantization unit 112 performs inverse quantization on the quantization coefficient block to obtain the transform coefficient block. The inverse transform unit 114 performs inverse transform on the transform coefficient block to obtain the prediction residual block. The addition unit 116 adds the prediction image block and the prediction residual block to obtain the reconstructed image block. The block memory 118 stores the reconstructed image block. The cyclic filtering unit 120 applies a cyclic filter to the reconstructed image block. The frame memory 122 stores the reconstructed image block with the applied cyclic filter.
[0368] Intra-frame prediction unit 124 performs intra-frame prediction with reference to block memory 118 to generate predicted image blocks. Inter-frame prediction unit 126 performs inter-frame prediction with reference to frame memory 122 to generate predicted image blocks. Prediction control unit 128 provides the predicted image blocks generated by intra-frame prediction unit 124 or inter-frame prediction unit 126 to subtraction unit 104 and addition unit 116. Prediction parameter generation unit 130 provides parameters related to intra-frame prediction or inter-frame prediction to entropy coding unit 110.
[0369] Figure 21 This is a block diagram illustrating a structural example of the decoding device 200 in this embodiment for decoding moving images. For example, the decoding device 200 may include... Figure 21 The multiple components shown are used for decoding images contained in a moving image in block units according to VVC. In addition to the multiple components described above, the decoding device 200 may include... Figure 21 The multiple constituent elements shown can also be at least a portion of the aforementioned constituent elements combined into... Figure 21 Among the various constituent elements shown.
[0370] like Figure 21 As shown, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an adder unit 208, a block memory 210, a cyclic filtering unit 212, a frame memory 214, an intra-frame prediction unit 216, an inter-frame prediction unit 218, a prediction control unit 220, a prediction parameter generation unit 222, and a segmentation determination unit 224. Furthermore, the intra-frame prediction unit 216 and the inter-frame prediction unit 218 are each configured as part of the prediction processing unit.
[0371] Entropy decoding unit 202 performs entropy decoding on the bitstream to obtain quantization coefficient blocks and parameters. Inverse quantization unit 204 performs inverse quantization on the quantization coefficient blocks to obtain transform coefficient blocks. Inverse transform unit 206 performs inverse transform on the transform coefficient blocks to obtain prediction residual blocks. Addition unit 208 adds the prediction image block and the prediction residual block to obtain a reconstructed image block. Cyclic filtering unit 212 applies a cyclic filter to the reconstructed image block.
[0372] Block memory 210 stores reconstructed image blocks. Frame memory 214 stores reconstructed image blocks with cyclic filters applied.
[0373] Intra-frame prediction unit 216 performs intra-frame prediction using reference block memory 210 to generate predicted image blocks. Inter-frame prediction unit 218 performs inter-frame prediction using reference frame memory 214 to generate predicted image blocks. Prediction control unit 220 provides the predicted image blocks generated by intra-frame prediction unit 216 or inter-frame prediction unit 218 to addition unit 208. Prediction parameter generation unit 222 provides parameters related to intra-frame prediction or inter-frame prediction to prediction control unit 220, etc.
[0374] The segmentation decision unit 224 determines the blocks used to decode the image in block units based on parameters related to segmentation.
[0375] [Example of a bitstream structure] In this disclosure, three types of information—reference image, geometric attributes, and background image—are transmitted via bit stream.
[0376] Figure 22 , Figure 23 , Figure 24 , Figure 25 and Figure 26 The image shows the bitstream layout. The bitstream layout can be... Figure 22 , Figure 23 , Figure 24 , Figure 25 and Figure 26 The combination of. Furthermore, not limited to. Figure 22 , Figure 23 , Figure 24 , Figure 25 and Figure 26 The bitstream layout shown allows the reference image, geometric attributes, and background image to be encoded in any order.
[0377] Figure 27 This is a conceptual diagram representing a structure example of a bitstream. In this example, the initial GOP contains more than one base image, more than one background image, and multiple geometric attributes, while each of the other GOPs contains more than one background image and multiple geometric attributes.
[0378] Figure 28This is a conceptual diagram representing another structural example of a bitstream. In this example, the initial GOP contains a base image, a background image, and multiple geometric attributes, while other GOPs contain multiple geometric attributes.
[0379] Figure 29 This is a conceptual diagram representing another structural example of a bitstream. In this example, the bitstream includes a first bitstream, a second bitstream, and a third bitstream. The first bitstream contains more than one reference image. The second bitstream contains more than one background image. The third bitstream contains multiple geometric attributes.
[0380] Figure 27 This is a conceptual diagram representing another structural example of a bitstream. In this example, the bitstream includes a first bitstream and a second bitstream. The first bitstream contains more than one reference image. The initial GOP in the second bitstream contains more than one background image and multiple geometric attributes. The other GOPs in the second bitstream each contain multiple geometric attributes.
[0381] Figure 28 This is a conceptual diagram representing another structural example of a bitstream. In this example, each GOP contains more than one base image, more than one background image, and multiple geometric attributes.
[0382] For example, if the bitstream contains multiple background images, a synthetic face image can be generated using one of the background images, or a combination of the background images can be used. Alternatively, the multiple background images in the bitstream can each correspond to multiple frames of the synthetic face image.
[0383] Furthermore, for example, when the bitstream contains multiple reference images, a synthetic face image can be generated using one of the multiple reference images, or a combination of multiple reference images can be used to generate a synthetic face image.
[0384] Furthermore, for example, to identify the content contained within an access unit, header parameters can be used to indicate which of the reference image, background image, or geometric attributes the access unit corresponds to. Header parameters can be encoded as SPS, PPS, PH, VUI, or SEI.
[0385] Furthermore, the synthesized facial motion image can be rendered and displayed by the decoding device 200 after decoding the initial access unit containing geometric attributes.
[0386] In encoding / decoding standards, it's possible to specify that each access unit contains image data corresponding to one image. Therefore, the reference image and background image can be contained in different access units. Alternatively, for example, in hierarchical coding such as multiview coding, it's possible to allow image data corresponding to multiple images to be contained in one access unit. Therefore, based on hierarchical coding such as multiview coding, the reference image and background image can be contained in the same access unit.
[0387] Each access unit may contain other NAL units from the Video Coding Layer (VCL) or non-Video Coding Layer.
[0388] Furthermore, geometric attributes such as facial landmark data can be included in metadata, VSEI (Versatile Supplemental Enhancement Information), or SEI within any motion picture codec or image codec. Additionally, geometric attributes can be included in NAL units with a new nal_unit_type in the VCL.
[0389] Figure 29 , Figure 30 and Figure 30 An example bitstream layout using the VVC codec as the encoding standard is shown. The VVC codec can be replaced with other video or image codecs such as HEVC, AVC, AV1, SVC, EVC, or JPEG.
[0390] Figure 31 This is a conceptual diagram representing a structure example of a bitstream conforming to VVC. In this example, the reference image and the background image are encoded and decoded as intra-frame pictures of VVC, respectively. Geometric attributes are encoded into an SEI represented as a geometric attribute SEI and decoded from that SEI.
[0391] In this example, the background image is transmitted multiple times. The initial access unit contains the image data of the reference image, while subsequent access units contain the image data of the geometric attribute SEI and the background image. That is, the geometric attribute and the background image are contained within the same access unit.
[0392] Figure 32 This is a conceptual diagram representing another structural example of a VVC-compliant bitstream. In this example, the reference image and background image are encoded and decoded as intra-frame pictures of the VVC, respectively. Geometric attributes are encoded into an SEI that represents the geometric attribute SEI and decoded from that SEI.
[0393] In this example, the background image is transmitted only once. The initial access unit contains the image data of the reference image, the next access unit contains the image data of the geometric attribute SEI and the background image. Subsequent access units contain the geometric attribute SEI.
[0394] Figure 33 This is a conceptual diagram representing another structural example of a bitstream conforming to VVC. In this example, the reference image is encoded and decoded as an intra-frame image of VVC. Geometric attributes are encoded into an SEI represented as a geometric attribute SEI and decoded from that SEI. Furthermore, background parameters, such as selection parameters used to choose a background image from multiple background image candidates instead of a background image, are encoded into an SEI represented as a background parameter SEI and decoded from that SEI.
[0395] In addition, the initial access unit contains the image data of the reference image, and the next access unit contains the geometric attribute SEI, background parameter SEI, and image data. Subsequent access units may contain either the geometric attribute SEI or both the geometric attribute SEI and image data.
[0396] For example, a codec standard might specify that each access unit contains image data corresponding to an image. To ensure compatibility with such a codec standard, the access unit used to transmit the SEI may contain not only the SEI but also the image data.
[0397] Specifically, the access unit used to transmit SEI can contain image data corresponding to the image at the minimum allowed resolution as virtual image data. This image can be a constant value such that all pixels have zero values. If the color code of each pixel is encoded into the bitstream, each pixel can be filled with the same color code.
[0398] In addition, the access unit used to transmit SEI may contain slice NAL units representing solid color images as image data, such as pure black, pure white, or pure green used in camera keying techniques.
[0399] Alternatively, the access unit used for transmitting the SEI can contain a copy of the reference image or background image as image data. This access unit can also contain skipped images as image data. Alternatively, this access unit can contain images specifying skip modes for all CUs (Coding Units) as image data. This minimizes the overhead of each access unit.
[0400] Alternatively, the access unit used to transmit SEI may include a parameter indicating that the NAL unit corresponding to the image data should be ignored.
[0401] In another example, the access unit may contain image data of the reference image, and at least one of the geometric attribute SEI and the background parameter SEI. Alternatively, another access unit may contain at least one of the geometric attribute SEI and the background parameter SEI. The geometric attribute SEI can be replaced with a NAL unit with a new nal_unit_type. Or, the background parameter SEI can be replaced with a NAL unit with a new nal_unit_type.
[0402] Furthermore, geometric properties and background parameters do not need to be encoded as separate SEIs, but can be encoded as a single SEI.
[0403] Additionally, as a specific example, the initial identical access units of each GOP can contain image data of a reference image, geometric attribute SEI, and background parameter SEI. Then, when generating a synthetic face motion image, the initial frame of the GOP in the synthetic face motion image can be generated based on the reference image, geometric attributes, and background information obtained from the identical access units. Alternatively, the initial frame of the GOP of the processing object can be generated using the reference image contained in the initial access units of previous GOPs.
[0404] Furthermore, in the initial GOP, the initial access unit may contain image data of the reference image, and subsequent access units may contain geometric attribute SEI and background parameter SEI. Then, in subsequent GOPs, the initial access unit may contain image data of the reference image, geometric attribute SEI, and background parameter SEI, and subsequent access units may contain geometric attribute SEI and background parameter SEI.
[0405] Therefore, when generating frames using geometric attribute SEI and background parameter SEI, a reference image obtained from the processed access unit can be used.
[0406] [Example of a generative model] Figure 34 This is a graph representing various model examples that can be used as generative models. For example, neural networks are used as generative models. Specifically, in Figure 35 The diagram illustrates generative adversarial networks, variational autoencoders, flow-based generative models, and diffusion models.
[0407] In generative adversarial networks (GANs), new data instances similar to the input data are generated by learning the features of the input data. Specifically, the unsupervised task in the generative model is transformed into a supervised task through two sub-models.
[0408] For example, a generator sub-model generates pseudo-samples, and a recognizer sub-model identifies the real input and the pseudo-samples generated by the generator sub-model. Then, an output image is generated through a minimax game, in which the recognition probability of the recognizer sub-model assigning the correct labels to the real input and pseudo-samples is maximized while minimizing the distribution difference between the real input and pseudo-samples.
[0409] In variational autoencoders, the input data is first compressed into a multivariate latent distribution used to reconstruct the data from the latent space as accurately as possible. This enables efficient data compression and dimensionality reduction. In stream-based generative models, the source distribution is transformed into the distribution of the training data via a sequence of more than one invertible transformation. This allows for accurate calculation of the learned data distribution and the likelihood of the final target.
[0410] In the diffusion model, new data instances similar to the training data are also generated. First, in the diffusion model, the structure of the training data is degraded through repeated injections of perturbations and noise before noise removal is initiated in an attempt to recover the original data. As a result, the data is repeatedly mapped to the latent distribution via a Markov chain where the latent states of each step depend only on the latent states of the previous steps. Then, the data is recovered through hierarchical noise removal.
[0411] For example, a neural network could be a face image generation neural network, which can be used to generate an output image using geometric information and an image representing facial parameters in a fixed format. That is, the neural network corresponds to the process of generating multiple sample values, which constitute an output image included as part of the output motion image.
[0412] Alternatively, the neural network can be the generative face video SEI (SEI) neural network discussed by the MPEG (Moving Picture Experts Group). Specifically, for example, it can be the face picture generator neural network labeled "GenerativeNN()" in Non-Patent Document 2.
[0413] Alternative examples of the aforementioned neural networks can be any combination thereof. Alternatively, other types of generative models may also be used.
[0414] Furthermore, machine learning models such as neural networks can be used for segmentation processing. Additionally, they can be used to derive geometric properties and baseline properties.
[0415] [Implementation Example] Figure 36 This is a block diagram illustrating an implementation example of the encoding device 100. The encoding device 100 includes a circuit 151 and a memory 152. For example, multiple components of the encoding device 100 described above are implemented by the circuit 151 and the memory 152.
[0416] Circuit 151 is a circuit that performs information processing and can access memory 152. For example, circuit 151 can be a dedicated circuit that executes the encoding method of this disclosure, or it can be a general-purpose circuit that executes a program corresponding to the encoding method of this disclosure. In addition, circuit 151 can also be a processor such as a CPU. Furthermore, circuit 151 can also be an assembly of multiple circuits.
[0417] Memory 152 is a dedicated or general-purpose memory that stores information used by circuit 151 to encode images. Memory 152 can be a circuit or connected to circuit 151. Alternatively, memory 152 can be included within circuit 151. Furthermore, memory 152 can be an assembly of multiple circuits. Additionally, memory 152 can be a disk or optical disk, or it can take the form of a storage device or recording medium. Furthermore, memory 152 can be non-volatile memory or volatile memory.
[0418] For example, memory 152 can store encoded object data such as images, or encoded data such as bitstreams. Furthermore, memory 152 can also store programs for causing circuit 151 to perform image processing. Additionally, memory 152 can also store a generation model in circuit 151. Furthermore, memory 152 can also store a reference image.
[0419] Figure 37 This is a flowchart illustrating a basic first operation example of the encoding device 100. In this example, the circuit 151 of the encoding device 100 uses the memory 152 to perform the following operations.
[0420] Specifically, circuit 151 encodes the reference image, geometric information, and background information used to generate the synthetic facial motion image into one or more streams (S301). The synthetic facial motion image is a motion image containing a face and is a motion image synthesized with a background image. The reference image is an image containing a face. The geometric information corresponds to each frame in a plurality of frames of the captured motion image obtained by the camera and represents the geometric properties of the subject. The background information is information related to the background image.
[0421] Therefore, it may be possible to provide a reference image, geometric properties, and a background image for generating synthetic facial motion images. Thus, when generating synthetic facial motion images, it may be possible to simultaneously imbue the face in the reference image with motion through geometric properties corresponding to each frame, while suppressing background distortion in the reference image through the background image. This could potentially help suppress image quality degradation.
[0422] For example, circuit 151 can segment the captured motion image to obtain segment information representing the foreground and background regions in the captured motion image. Then, circuit 151 can encode the captured motion image segment information into one or more streams.
[0423] Therefore, it is possible to provide segmentation information for capturing motion images, used to determine foreground and background regions in motion images of the same type as captured motion images, via one or more streams. Thus, it may be possible to help determine foreground and background regions in intermediate face motion images obtained by applying motion to the face of a reference image using geometric properties corresponding to each frame.
[0424] Furthermore, for example, circuit 151 can segment the reference image to obtain reference image segmentation information representing the foreground and background regions in the reference image. Additionally, circuit 151 can use the reference image segmentation information to embed a predetermined background color code in the background region of the reference image. Then, circuit 151 can encode the reference image with the predetermined background color code embedded in the background region.
[0425] Therefore, it is possible to appropriately determine the background region in the reference image based on the reference image segmentation information obtained as a result of segmentation processing of the reference image. Furthermore, since a prescribed background color code is embedded in the background region of the reference image, it is possible to suppress distortion in the background of the reference image even when motion is applied to the face in the reference image.
[0426] Furthermore, for example, circuit 151 can encode background color code information representing a specified background color code into more than one stream. Thus, it is possible to provide a specified background color code for efficiently determining a background region in a reference image via more than one stream. Moreover, it is possible to change the specified background color code according to the reference image.
[0427] Furthermore, for example, background color code information can represent a range containing multiple consecutive values as a defined background color code. Moreover, the defined background color code can be specified within the range represented by the background color code information. Therefore, it is possible to flexibly define the defined background color code. Furthermore, it is possible to flexibly apply the defined background color code to background areas.
[0428] Furthermore, for example, the specified background color code can be defined by a color code that occurs at a frequency below a threshold in the foreground region of the reference image. This may prevent portions of the foreground region from being incorrectly identified as portions of the background region. Therefore, it may be possible to properly determine the background region.
[0429] Furthermore, for example, circuit 151 can segment the reference image to obtain reference image segmentation information representing the foreground and background regions in the reference image. Then, circuit 151 can encode the reference image segmentation information into one or more streams.
[0430] Therefore, it is possible to provide reference image segmentation information for determining foreground and background regions in a reference image via more than one stream. This could potentially help in determining background regions in a reference image.
[0431] Furthermore, for example, the background image can be an image prepared independently of the reference image and the captured motion image. Therefore, it is possible to apply a background image prepared separately from the reference image and the captured motion image to a synthesized facial motion image. Consequently, it may be possible to suppress the influence of foreground regions in the background image, etc.
[0432] Furthermore, for example, circuit 151 can select a background image from multiple background image candidates. Then, circuit 151 can encode the identifier of the background image as background information. Thus, it is possible to flexibly select a background image from multiple background image candidates. Therefore, it is possible to apply an appropriate background image to the synthesized face motion image according to its intended use.
[0433] Furthermore, for example, the background image can be an image contained within a captured moving image, or it can be an image obtained by synthesizing multiple images contained within a captured moving image. Therefore, it is possible to apply a background image obtained from a captured moving image to synthesize a facial motion image. Thus, it is possible to apply a background image corresponding to the captured situation to synthesize a facial motion image.
[0434] Furthermore, for example, circuit 151 can encode the reference image into background information. The reference image can then be applied to the background image. Thus, it is possible to use the reference image as a background image. Moreover, it is possible to suppress background distortion of the reference image by using the original reference image as the background image while simultaneously applying motion to the face of the reference image through geometric attributes corresponding to each frame.
[0435] Furthermore, for example, when the background image contains a foreground region, circuit 151 can interpolate the missing background region in the background image using the surrounding area of the foreground region in the background image or the background region of other images included in the captured motion image. Thus, it is possible to appropriately interpolate the missing background region even when the background image contains a foreground region. Therefore, it is possible to suppress the missing background region in the synthesized facial motion image.
[0436] Furthermore, for example, circuit 151 can segment the background image to obtain background image segmentation information representing the foreground and background regions in the background image. Then, circuit 151 can use the background image segmentation information to determine the foreground and background regions in the background image.
[0437] Therefore, it may be possible to appropriately determine the foreground and background regions in a background image based on the background image segmentation information obtained as a result of segmentation processing of the background image. Consequently, it may be possible to appropriately interpolate for missing background regions in the background image.
[0438] Furthermore, for example, circuit 151 can segment the background image to obtain background image segmentation information representing the foreground and background regions in the background image. Then, circuit 151 can use the background image segmentation information to embed a specified foreground color code in the foreground region of the background image.
[0439] Therefore, it may be possible to efficiently determine the foreground region in the background image based on a specified foreground color code. Additionally, it may be possible to suppress the projection of foreground features such as faces onto the background region of a synthesized facial motion image.
[0440] Furthermore, for example, circuit 151 can encode foreground color code information representing a predetermined foreground color code into more than one stream. Thus, it is possible to provide a predetermined foreground color code for efficiently determining a foreground region in a background image via more than one stream. Moreover, it is possible to change the predetermined foreground color code according to the background image.
[0441] Furthermore, for example, the foreground color code information can represent a range containing multiple consecutive values as a defined foreground color code. Moreover, the defined foreground color code can be specified within the range represented by the foreground color code information. Therefore, it is possible to flexibly define the defined foreground color code. Furthermore, it is possible to flexibly apply the defined foreground color code to the foreground area.
[0442] Furthermore, for example, the specified foreground color code can be defined by a color code that occurs in the background region of the background image at a frequency below a threshold. This may prevent parts of the background region from being incorrectly identified as parts of the foreground region. Therefore, it may be possible to properly determine the foreground region.
[0443] Furthermore, for example, in more than one stream, the stream in which background information is encoded can be the same as the stream in which the reference image is encoded, or it can be the same as the stream in which geometric information is encoded. Thus, it is possible to encode background information into the same stream as the reference image stream or the geometric information stream, rather than into a separate stream. Therefore, it is possible to efficiently encode background information together with the reference image or geometric information.
[0444] Furthermore, for example, in more than one stream, the stream in which background information is encoded can be different from both the stream in which the reference image is encoded and the stream in which the geometric information is encoded. Thus, it is possible to encode background information from a stream different from both the reference image stream and the geometric information stream, rather than from the same stream. Therefore, it is possible to encode background information separately from the reference image or geometric information at arbitrary timing.
[0445] Furthermore, for example, the background image can be encoded as the first image in a sequence containing multiple images, or it can be encoded as the first image in a group of pictures (GOPs). This makes it possible to provide the background image earlier. Therefore, it is possible to apply the background image to the synthesized facial motion image earlier.
[0446] Furthermore, for example, a background image can be encoded as an image within one or more access units in a stream. Therefore, it is possible to process the background image as an image within an access unit. In other words, it is possible to process the background image in the same way as a regular image.
[0447] Furthermore, for example, the access unit encoded for the background image can be the same as the access unit encoded for the reference image. Therefore, it is possible to encode the background image to the same access unit as the reference image, instead of encoding it to a different access unit. Thus, it is possible to efficiently encode the background image together with the reference image.
[0448] Furthermore, for example, the access unit encoded for the background image can be different from the access unit encoded for the reference image. Therefore, it is possible to encode the background image to an access unit different from that of the reference image, rather than encoding it to the same access unit. Thus, it is possible to encode the background image separately from the reference image at arbitrary timing.
[0449] Furthermore, for example, a corresponding SEI code can be established for the access unit that is encoded with the background image, representing a signal indicating the presence of the background image in that access unit. Thus, it becomes possible to notify the access unit of the presence of the background image via the SEI signal in the access unit. Therefore, it becomes possible to appropriately transmit the background image.
[0450] Furthermore, for example, the background image can be encoded as an image in one or more access units within a stream. Additionally, a corresponding SEI encoding can be established for the access unit where the background image is encoded, representing the signal indicating the presence of the background image in the access unit.
[0451] Therefore, it is possible to process the background image as an image within the access unit. In other words, it is possible to process the background image in the same way as a regular image. Consequently, it is possible to notify the access unit of the presence of the background image via the SEI signal in the access unit. Therefore, it is possible to appropriately transmit the background image.
[0452] Furthermore, for example, the background image can be encoded as an intra-frame image. Therefore, it is possible to process the background image as an intra-frame image. That is, it is possible to process the background image without relying on other images.
[0453] Furthermore, for example, circuit 151 can encode background color code information into an SEI in one or more streams. Thus, it is possible to provide a specified background color code via the SEI for efficiently determining background regions in a reference image. Moreover, it is possible to change the specified background color code according to the reference image.
[0454] Furthermore, for example, circuit 151 can encode foreground color code information into an SEI in one or more streams. Thus, it becomes possible to provide a specified foreground color code via the SEI for efficiently determining foreground regions in a background image. Moreover, it becomes possible to change the specified foreground color code according to the background image.
[0455] Furthermore, for example, circuit 151 can encode at least one of background color code information representing a specified background color code and foreground color code information representing a specified foreground color code into an SEI in one or more streams. Thus, it may be possible to provide a specified foreground color code for efficiently determining a foreground region via the SEI.
[0456] Furthermore, for example, circuit 151 can encode captured motion image segmentation information into an SEI in one or more streams. Thus, it may be possible to provide captured motion image segmentation information via the SEI for determining foreground and background regions in motion images of the same type as the captured motion images. Therefore, it may be possible to help determine foreground and background regions in intermediate face motion images obtained by applying motion to the face of a reference image using geometric attributes corresponding to each frame.
[0457] Furthermore, for example, circuit 151 can encode reference image segmentation information into an SEI in one or more streams. Thus, it is possible to provide reference image segmentation information via the SEI for determining foreground and background regions in the reference image. Therefore, it is possible to assist in determining background regions in the reference image.
[0458] Furthermore, for example, circuit 151 can encode at least one of the captured motion image segmentation information and the reference image segmentation information into an SEI in one or more streams. The captured motion image segmentation information represents the foreground and background regions in the captured motion image. The reference image segmentation information represents the foreground and background regions in the reference image. Thus, it is possible to provide segmentation information for determining the foreground and background regions via the SEI.
[0459] Figure 38 This is a flowchart illustrating a basic second operation example of the encoding device 100. In this example, the circuit 151 of the encoding device 100 uses the memory 152 to perform the following operations.
[0460] Specifically, circuit 151 encodes the reference image and geometric information used to generate the synthetic facial motion image into one or more streams (S311). The reference image is an image containing the face. The geometric information corresponds to each frame in a plurality of frames of the captured motion image obtained by the camera and represents the geometric properties of the subject.
[0461] In generating synthetic facial motion images, a reference image and geometric information are used as inputs to the generative model, and an intermediate facial motion image is obtained from the generative model. The intermediate facial motion image is a motion image containing the face. The reference image is used to generate the synthetic facial motion image by embedding corresponding regions from the reference image into the background regions of the intermediate facial motion image.
[0462] Therefore, it may be possible to provide a reference image and geometric properties for generating synthetic facial motion images. Furthermore, when generating synthetic facial motion images, it may be possible to simultaneously imbue the face in the reference image with motion through geometric properties corresponding to each frame, while suppressing background distortion in the original reference image. This could potentially help suppress image quality degradation.
[0463] Alternatively, the encoding device 100 may also include an input terminal, an entropy encoder, and an output terminal. Furthermore, the operation performed by the circuit 151 can also be performed by the entropy encoder. Additionally, data for the operation of the entropy encoder can be input to the input terminal. And, data obtained from the operation of the entropy encoder can be output from the output terminal.
[0464] Figure 39 This is a block diagram illustrating an implementation example of the bitstream generation device 300. The bitstream generation device 300 includes circuitry 351 and memory 352. For example, the bitstream generation device 300, circuitry 351, and memory 352 can correspond to encoding device 100, circuitry 151, and memory 152, respectively. Moreover, the bitstream generation device 300, circuitry 351, and memory 352 can each perform the same function as encoding device 100, circuitry 151, and memory 152.
[0465] Circuit 351 is a circuit for information processing and can access memory 352. For example, circuit 351 can be a dedicated circuit for executing the bit stream generation method of this disclosure, or it can be a general-purpose circuit for executing a program corresponding to the bit stream generation method of this disclosure. In addition, circuit 351 can be a processor such as a CPU. Furthermore, circuit 351 can be an assembly of multiple circuits.
[0466] Memory 352 is a dedicated or general-purpose memory for storing information used by circuit 351 to generate a bit stream. Memory 352 can be a circuit or connected to circuit 351. Alternatively, memory 352 can be contained within circuit 351. Alternatively, memory 352 can be an assembly of multiple circuits. Alternatively, memory 352 can be a disk or optical disk, or it can take the form of a storage device or recording medium. Alternatively, memory 352 can be non-volatile memory or volatile memory.
[0467] For example, memory 352 can store data used to generate the bitstream, or it can store the bitstream itself. Additionally, memory 352 can store a program for causing circuit 351 to perform generation processing. Furthermore, memory 352 can store a generation model of circuit 351. Additionally, memory 352 can store a reference image, a background image, and background image candidates.
[0468] Figure 40This is a flowchart illustrating a basic first operation example of the bitstream generation device 300. In this example, the circuit 351 of the bitstream generation device 300 uses the memory 352 to perform the following operations.
[0469] Specifically, circuit 351 generates a bitstream containing a reference image, geometric information, and background information for generating a synthetic facial motion image (S501). The synthetic facial motion image is a motion image containing a face, and is a motion image synthesized with a background image. The reference image is an image containing a face. The geometric information corresponds to each frame in a plurality of frames of the captured motion image obtained by the camera, and represents the geometric properties of the subject. The background information is information related to the background image.
[0470] Therefore, it may be possible to provide a reference image, geometric properties, and a background image for generating synthetic facial motion images. Thus, when generating synthetic facial motion images, it may be possible to simultaneously imbue the face in the reference image with motion through geometric properties corresponding to each frame, while suppressing background distortion in the reference image through the background image. This could potentially help suppress image quality degradation.
[0471] Figure 41 This is a flowchart illustrating a basic second operation example of the bitstream generation device 300. In this example, the circuit 351 of the bitstream generation device 300 uses the memory 352 to perform the following operations.
[0472] Specifically, circuit 351 generates a bitstream containing a reference image and geometric information for generating a synthetic facial motion image (S511). The reference image is an image containing a face. The geometric information corresponds to each frame in a plurality of frames of the captured motion image obtained by the camera and represents the geometric properties of the subject.
[0473] In generating synthetic facial motion images, a reference image and geometric information are used as inputs to the generative model, and an intermediate facial motion image is obtained from the generative model. The intermediate facial motion image is a motion image containing the face. The reference image is used to generate the synthetic facial motion image by embedding corresponding regions from the reference image into the background regions of the intermediate facial motion image.
[0474] Therefore, it may be possible to provide a reference image and geometric properties for generating synthetic facial motion images. Furthermore, when generating synthetic facial motion images, it may be possible to simultaneously imbue the face in the reference image with motion through geometric properties corresponding to each frame, while suppressing background distortion in the original reference image. This could potentially help suppress image quality degradation.
[0475] Figure 42This is a block diagram illustrating an implementation example of the decoding device 200. The decoding device 200 includes circuitry 251 and a memory 252. For example, multiple components of the decoding device 200 described above are implemented using circuitry 251 and memory 252.
[0476] Circuit 251 is an information processing circuit capable of accessing memory 252. For example, circuit 251 can be a dedicated circuit for executing the decoding method of this disclosure, or it can be a general-purpose circuit for executing a program corresponding to the decoding method of this disclosure. Furthermore, circuit 251 can also be a processor such as a CPU. Moreover, circuit 251 can also be an aggregation of multiple circuits.
[0477] Memory 252 is a dedicated or general-purpose memory that stores information used by circuit 251 to decode images. Memory 252 can be a circuit or connected to circuit 251. Alternatively, memory 252 can be included within circuit 251. Furthermore, memory 252 can be an assembly of multiple circuits. Additionally, memory 252 can be a disk or optical disk, or it can take the form of a storage device or recording medium. Furthermore, memory 252 can be non-volatile memory or volatile memory.
[0478] For example, memory 252 can store decoded object data such as bitstreams, or it can store decoded data such as images. Furthermore, memory 252 can also store programs for causing circuit 251 to perform image processing. Additionally, memory 252 can store a generation model in circuit 251. Moreover, memory 252 can store a reference image, a background image, or background image candidates.
[0479] Figure 41 This is a flowchart of a basic first operation example of the decoding device 200. In this example, the circuit 251 of the decoding device 200 uses the memory 252 to perform the following operations.
[0480] Specifically, circuit 251 decodes a reference image, geometric information, and background information from one or more streams (S401). The reference image is an image containing a face. The geometric information corresponds to each frame in a series of frames captured by the camera, representing the geometric properties of the subject. The background information is information related to the background image.
[0481] Then, circuit 251 uses a generative model to generate a synthetic facial motion image based on the reference image, geometric information, and background information (S402). The synthetic facial motion image is a motion image containing the face and is a motion image synthesized with the background image.
[0482] Therefore, it is possible to apply the reference image, geometric attributes, and background image to the generation of synthetic facial motion images. Thus, it is possible to simultaneously imbue the face in the reference image with motion through the geometric attributes corresponding to each frame, while suppressing background distortion in the reference image using the background image. Therefore, it is possible to suppress image quality degradation during the generation of synthetic facial motion images.
[0483] For example, circuit 251 can input a reference image, geometric information and a background image into the generative model, and obtain a synthetic facial motion image from the generative model.
[0484] Therefore, it is possible to easily obtain synthetic facial motion images based on the generative model. Furthermore, in the generative model, it is possible to apply a reference image, geometric attributes, and a background image to the generation of synthetic facial motion images. Thus, it is possible to simultaneously impart motion to the face in the reference image through geometric attributes corresponding to each frame, while suppressing background distortion in the reference image through the background image.
[0485] Furthermore, for example, circuit 251 can input a reference image and geometric information into the generative model and obtain an intermediate facial motion image from the generative model. The intermediate facial motion image is a motion image containing a face and without a synthesized background image. Then, circuit 251 can generate a synthesized facial motion image by embedding the corresponding region from the background image into the background region of the intermediate facial motion image.
[0486] Therefore, it is possible to obtain intermediate facial motion images from the generative model, which are obtained by imbuing the face of a reference image with motion through geometric attributes corresponding to each frame. Then, it is possible to apply a background image to the intermediate facial motion images. Thus, it is possible to suppress background distortion of the reference image while imbuing the face of the reference image with motion through geometric attributes corresponding to each frame.
[0487] Furthermore, for example, circuit 251 can segment the intermediate face motion image to obtain intermediate face motion image segmentation information representing the foreground and background regions in the intermediate face motion image. Then, circuit 251 can use the intermediate face motion image segmentation information to determine the background region in the intermediate face motion image.
[0488] Therefore, it may be possible to appropriately determine the background region in the intermediate face motion image based on the intermediate face motion image segmentation information obtained as a result of segmentation processing of the intermediate face motion image. Consequently, it may be possible to appropriately apply the corresponding region in the background image to the background region in the intermediate face motion image.
[0489] Furthermore, for example, circuit 251 can decode captured motion image segmentation information representing foreground and background regions in captured motion images from one or more streams. Then, circuit 251 can use the captured motion image segmentation information to determine the background region in the intermediate face motion image.
[0490] Therefore, it is possible to appropriately determine the background region in the intermediate face motion image based on the segmentation information of the captured motion images obtained from more than one stream. Consequently, it is possible to appropriately apply the corresponding region in the background image to the background region in the intermediate face motion image.
[0491] Furthermore, for example, a predetermined background color code can be embedded in the background region of the reference image. This makes it possible to efficiently determine the background region in the reference image based on the predetermined background color code. Additionally, it becomes possible to suppress distortion in the background of the reference image even when motion is applied to a face in the reference image.
[0492] In addition, for example, circuit 251 can determine the region with a specified background color code in the intermediate face motion image as the background region in the intermediate face motion image.
[0493] Therefore, it may be possible to efficiently determine the background region in an intermediate face motion image based on a specified background color code. Specifically, suppose that a specified background color code is embedded in the background region of an intermediate face motion image obtained by applying motion to the face of a reference image, which has a background region with the specified background color code embedded. Thus, it may be possible to efficiently determine the background region in an intermediate face motion image based on a specified background color code.
[0494] Furthermore, for example, circuit 251 can decode background color code information representing a specified background color code from one or more streams. Therefore, it is possible to efficiently determine the background region in a reference image based on the specified background color code obtained from one or more streams. Moreover, it is possible to change the specified background color code based on the reference image.
[0495] Furthermore, for example, background color code information can represent a range containing multiple consecutive values as a defined background color code. Then, the defined background color code can be specified within the range represented by the background color code information. Therefore, it is possible to flexibly define the defined background color code. Moreover, it is possible to flexibly apply the defined background color code to background areas.
[0496] Furthermore, for example, the specified background color code can be defined by a color code that occurs at a frequency below a threshold in the foreground region of the reference image. This may prevent portions of the foreground region from being incorrectly identified as portions of the background region. Therefore, it may be possible to properly determine the background region.
[0497] Furthermore, for example, circuit 251 can decode reference image segmentation information representing foreground and background regions in a reference image from one or more streams. Then, circuit 251 can use the reference image segmentation information to embed a specified background color code in the background region of the reference image.
[0498] Therefore, it is possible to efficiently determine the background region in a reference image based on the segmentation information of the reference image obtained from more than one stream. Furthermore, since a specified background color code is embedded in the background region of the reference image, it is possible to suppress distortion in the background of the reference image even when the face in the reference image is moved.
[0499] Furthermore, for example, the background image can be an image prepared independently of the reference image and the captured motion image. Therefore, it is possible to apply a background image prepared separately from the reference image and the captured motion image to a synthesized facial motion image. Consequently, it may be possible to suppress the influence of foreground regions in the background image, etc.
[0500] Furthermore, for example, circuit 251 can decode the identifier of the background image into background information. Then, circuit 251 can use the identifier to select a background image from multiple background image candidates.
[0501] Therefore, it is possible to flexibly select a background image from multiple background image candidates. Consequently, it is possible to apply an appropriate background image to the synthesized face motion image based on its intended use.
[0502] Furthermore, for example, the background image can be an image contained within a captured moving image, or it can be an image obtained by synthesizing multiple images contained within a captured moving image. Therefore, it is possible to apply a background image obtained from a captured moving image to synthesize a facial motion image. Thus, it is possible to apply a background image corresponding to the captured situation to synthesize a facial motion image.
[0503] Furthermore, for example, circuit 251 can decode the reference image into background information. The reference image can then be applied to the background image. Thus, it is possible to use the reference image as a background image. Moreover, it is possible to suppress background distortion of the reference image by using the original reference image as the background image while simultaneously applying motion to the face in the reference image through geometric attributes corresponding to each frame.
[0504] Furthermore, for example, when the background image contains a foreground region, circuit 251 can interpolate the missing background region in the background image using the surrounding area of the foreground region in the background image or the background region in a previous synthesized face motion image. Thus, it is possible to appropriately interpolate the missing background region even when the background image contains a foreground region. Therefore, it is possible to suppress the missing background region in the synthesized face motion image.
[0505] Furthermore, for example, circuit 251 can segment the background image to obtain background image segmentation information representing the foreground and background regions in the background image. Then, circuit 251 can use the background image segmentation information to determine the foreground and background regions in the background image.
[0506] Therefore, it may be possible to appropriately determine the foreground and background regions in a background image based on the background image segmentation information obtained as a result of segmentation processing of the background image. Consequently, it may be possible to appropriately interpolate for missing background regions in the background image.
[0507] Furthermore, for example, a predetermined foreground color code can be embedded in the foreground region of the background image. This makes it possible to efficiently determine the foreground region in the background image based on the predetermined foreground color code. Additionally, it may be possible to suppress the projection of foreground features such as faces into the background region of a synthesized facial motion image.
[0508] Furthermore, for example, circuit 251 can decode foreground color code information representing a specified foreground color code from one or more streams. Then, a region in the background image having the specified foreground color code can be determined as a foreground region in the background image. Thus, it is possible to efficiently determine the foreground region in the background image based on the specified foreground color code obtained from one or more streams. Moreover, it is possible to change the specified foreground color code based on the background image.
[0509] Furthermore, for example, the foreground color code information can represent a range containing multiple consecutive values as a defined foreground color code. Moreover, the defined foreground color code can be specified within the range represented by the foreground color code information. Therefore, it is possible to flexibly define the defined foreground color code. Furthermore, it is possible to flexibly apply the defined foreground color code to the foreground area.
[0510] Furthermore, for example, the specified foreground color code can be defined by a color code that occurs in the background region of the background image at a frequency below a threshold. This may prevent parts of the background region from being incorrectly identified as parts of the foreground region. Therefore, it may be possible to properly determine the foreground region.
[0511] Furthermore, for example, in more than one stream, the stream from which background information is decoded can be the same as the stream from which the reference image is decoded, or it can be the same as the stream from which the geometric information is decoded. Thus, it is possible to decode background information from the same stream as the reference image or the geometric information, rather than from a separate stream. Therefore, it is possible to efficiently decode background information together with the reference image or geometric information.
[0512] Furthermore, for example, in more than one stream, the stream in which the background information is decoded can be different from both the stream in which the reference image is decoded and the stream in which the geometric information is decoded. Therefore, it is possible to decode the background information from a stream different from both the reference image stream and the geometric information stream, rather than from the same stream. Thus, it is possible to decode the background information separately from the reference image or geometric information at arbitrary timing.
[0513] Furthermore, for example, the background image can be decoded as the first image of a sequence containing multiple images, or as the first image of a group of pictures (GOPs). This makes it possible to obtain the background image early. Therefore, it may be possible to apply the background image to the synthesized facial motion image at an earlier stage.
[0514] Furthermore, for example, a background image can be decoded into an image from one or more access units in a stream. Therefore, it is possible to process the background image as an image within an access unit. In other words, it is possible to process the background image in the same way as a regular image.
[0515] Furthermore, for example, the access unit used to decode the background image can be the same as the access unit used to decode the reference image. Therefore, it is possible to decode the background image from the same access unit as the reference image, instead of from a different access unit. Thus, it is possible to efficiently decode the background image together with the reference image.
[0516] Furthermore, for example, the access unit used to decode the background image may be different from the access unit used to decode the reference image. Therefore, it is possible to decode the background image from an access unit different from the reference image's access unit, rather than from the same access unit. Thus, it is possible to decode the background image separately from the reference image at any given time.
[0517] Furthermore, for example, a signal indicating the presence of a background image in an access unit can be decoded from an access unit that contains an SEI corresponding to that access unit. Thus, it becomes possible to identify the presence of a background image in an access unit based on the signal obtained from the SEI in the access unit. Therefore, it becomes possible to appropriately transmit the background image.
[0518] Furthermore, for example, a background image can be decoded into a picture from an access unit in one or more streams. Additionally, a signal indicating the presence of a background image in an access unit can be decoded from an SEI corresponding to the access unit containing the background image.
[0519] Therefore, it becomes possible to process the background image as an image within the access unit. In other words, it becomes possible to process the background image in the same way as a regular image. Furthermore, it becomes possible to identify the presence of a background image within the access unit based on the signal obtained from the SEI within the access unit. Thus, it becomes possible to appropriately transmit the background image.
[0520] Furthermore, for example, the background image can be decoded into an intra-frame image. Therefore, it is possible to process the background image as an intra-frame image. In other words, it may be possible to process the background image without relying on other images.
[0521] Furthermore, for example, background images can be used together in multiple frames of the synthesized facial motion image. This makes it possible to reduce the overall coding complexity of the synthesized facial motion image. Additionally, it may be possible to reduce the processing complexity required for decoding the background image.
[0522] Furthermore, for example, circuit 251 can decode background color code information from SEI in one or more streams. Therefore, it is possible to efficiently determine the background region in the reference image based on the specified background color code obtained from the SEI. Then, it is possible to change the specified background color code based on the reference image.
[0523] Furthermore, for example, circuit 251 can decode foreground color code information from SEIs in one or more streams. Therefore, it is possible to efficiently determine the foreground region in the background image based on the specified foreground color code obtained from the SEI. Then, it is possible to change the specified foreground color code according to the background image.
[0524] Furthermore, for example, circuit 251 can decode at least one of background color code information representing a specified background color code and foreground color code information representing a specified foreground color code from one or more SEI streams. Thus, it is possible to efficiently determine a background region based on the specified background color code obtained from the SEI.
[0525] Furthermore, for example, circuit 251 can decode captured motion image segmentation information from SEI in one or more streams. Therefore, it is possible to appropriately determine the background region in the intermediate face motion image based on the captured motion image segmentation information obtained from the SEI. Consequently, it is possible to appropriately apply the corresponding region in the background image to the background region in the intermediate face motion image.
[0526] Furthermore, for example, circuit 251 can decode reference image segmentation information from SEIs in one or more streams. Therefore, it is possible to efficiently determine background regions in the reference image based on the reference image segmentation information obtained from the SEI. Moreover, it is possible to appropriately embed prescribed background color codes into the background regions of the reference image.
[0527] Figure 42 This is a flowchart illustrating a basic second operation example of the decoding device 200. In this example, the circuit 251 of the decoding device 200 uses the memory 252 to perform the following operations.
[0528] Furthermore, for example, circuit 251 can decode at least one of captured motion image segmentation information and reference image segmentation information from one or more SEI streams. The captured motion image segmentation information represents foreground and background regions in the captured motion image. The reference image segmentation information represents foreground and background regions in the reference image. Thus, it is possible to efficiently determine the background region based on the segmentation information obtained from the SEI.
[0529] Specifically, circuit 251 decodes reference images and geometric information from one or more streams (S411). The reference image is an image containing a face. The geometric information corresponds to each frame in a plurality of frames of a moving image captured by a camera and represents the geometric properties of the subject.
[0530] Next, circuit 251 inputs the reference image and geometric information into the generation model and obtains an intermediate face motion image from the generation model (S412). The intermediate face motion image is a motion image containing the face. Then, circuit 251 generates a synthetic face motion image by embedding the corresponding region in the reference image into the background region in the intermediate face motion image (S413).
[0531] Therefore, it may be possible to obtain intermediate facial motion images from the generative model, which are obtained by imbuing the face of a reference image with motion through geometric attributes corresponding to each frame. Furthermore, it may be possible to apply the corresponding region of the original reference image to the background region of the intermediate facial motion image. Thus, it may be possible to suppress background distortion of the reference image while imbuing the face of the reference image with motion through geometric attributes corresponding to each frame. Therefore, it may be possible to suppress image quality degradation when generating synthetic facial motion images.
[0532] Alternatively, the decoding device 200 may also include an input terminal, an entropy decoder, and an output terminal. Furthermore, the operations performed by the circuit 251 can also be performed by the entropy decoder. Additionally, data for the operation of the entropy decoder can be input to the input terminal. And, data obtained through the operation of the entropy decoder can be output from the output terminal.
[0533] Furthermore, for example, a computer-readable non-transitory recording medium storing more than one bitstream can also be used. The more than one bitstream may also include: at least one reference image for displaying the moving image; and geometric information representing geometric properties within a region including a person's face, as information corresponding to each of a plurality of images of the moving image. Moreover, the more than one bitstream can also enable the decoding device 200 to perform (i) decoding of at least one reference image and (ii) decoding of the geometric information.
[0534] Therefore, it may be possible to realize a recording medium that stores more than one bitstream corresponding to the decoding device and decoding method described above. Thus, it may be possible to obtain the same effect as the decoding device 200 described above using a recording medium.
[0535] [Other examples] The encoding device 100 and decoding device 200 in the above examples can be used as image encoding devices and image decoding devices, respectively, or as moving image encoding devices and moving image decoding devices. Furthermore, the multiple components included in the encoding device 100 and the multiple components included in the decoding device 200 can perform corresponding operations.
[0536] Furthermore, encoding can be replaced by other representations such as storing, containing, writing, describing, signaling, sending, notifying, or saving, and these representations are interchangeable. For example, encoding information can mean including the information in a bitstream. Additionally, encoding information into a bitstream can refer to encoding the information to generate a bitstream that includes the encoded information.
[0537] Furthermore, the term "decoding" can be replaced by terms such as "reading out," "reading through," "reading," "reading in," "exporting," "acquiring," "receiving," "extracting," or "recovering," and these terms are interchangeable. For example, decoding information can mean acquiring information from a bitstream. Additionally, decoding information from a bitstream can refer to decoding the bitstream to obtain the information contained within it.
[0538] In addition, for example, encoded and compressed information contained in a bitstream may simply be represented as information.
[0539] Furthermore, at least some of the above examples can be used as encoding methods, decoding methods, entropy encoding methods, entropy decoding methods, or other methods.
[0540] Furthermore, each component can be constructed using dedicated hardware, or it can be implemented by executing software programs suitable for each component. Each component can also be implemented by a program execution unit such as a CPU or processor reading and executing software programs recorded on recording media such as hard disks or semiconductor memory.
[0541] Specifically, the encoding device 100 and the decoding device 200 may each include a processing circuit and a storage device electrically connected to the processing circuit and accessible from the processing circuit. For example, the processing circuit corresponds to circuit 151 or 251, and the storage device corresponds to memory 152 or 252.
[0542] The processing circuit includes at least one of dedicated hardware and a program execution unit, and performs processing using a storage device. Furthermore, when the processing circuit includes a program execution unit, the storage device stores a software program executed by that program execution unit.
[0543] An example of the software program described above is a bitstream. The bitstream includes an encoded image and a syntax for decoding the image. The bitstream causes the decoding device 200 to decode the image by performing syntax-based processing. Furthermore, for example, the software used to implement the aforementioned encoding device 100 or decoding device 200 is the following program.
[0544] For example, the program can enable a computer to execute an encoding method that encodes (i) a reference image, (ii) geometric information, and (iii) background information for generating a synthetic facial motion image into one or more streams, wherein the synthetic facial motion image is a motion image containing a face and is obtained by synthesizing a background image, wherein the (i) reference image is an image containing the face, the (ii) geometric information corresponds to multiple frames of a captured motion image obtained by a camera and represents the geometric properties of the subject, and the (iii) background information is related to the background image.
[0545] Furthermore, for example, the program can enable a computer to execute a decoding method that decodes from one or more streams (i) a reference image containing a face, (ii) geometric information corresponding to and representing the geometric properties of the subject in multiple frames of a motion image captured by a camera, and (iii) background information related to a background image, and uses a generative model to generate a synthetic face motion image based on the reference image, the geometric information, and the background information, the synthetic face motion image being a motion image containing the face and being a motion image obtained by synthesizing the background image.
[0546] Furthermore, the aforementioned components can be circuits. These circuits can form a single circuit as a whole, or they can be separate circuits. Moreover, each component can be implemented using a general-purpose processor or a dedicated processor.
[0547] Furthermore, the processing performed by a specific component can be performed by other components. Furthermore, the order of execution of the processing can be changed, and multiple processes can be executed in parallel. Furthermore, any two or more of the various examples of this disclosure can be appropriately combined for implementation. Furthermore, the encoding / decoding apparatus can include an encoding device 100 and a decoding device 200.
[0548] Furthermore, it is not necessary to install all of the constituent elements of this disclosure; only a portion of the constituent elements may be installed. Similarly, it is not necessary to install all of the processes in this disclosure; only a portion of the processes may be executed.
[0549] Furthermore, the ordinal numbers such as 1st and 2nd used in the description can be replaced appropriately. Additionally, new ordinal numbers can be assigned to constituent elements, or ordinal numbers can be removed. Moreover, these ordinal numbers are sometimes attached to elements to identify them, and sometimes do not correspond to a meaningful order.
[0550] Furthermore, for example, at least one (or more) of the first, second, and third elements corresponds to the first, second, third elements, or any combination thereof.
[0551] The forms of the encoding device 100 and the decoding device 200 have been described above based on several examples, but the forms of the encoding device 100 and the decoding device 200 are not limited to these examples. As long as they do not depart from the spirit of this disclosure, forms obtained by implementing various modifications to the examples that would be conceived by those skilled in the art, and forms constructed by combining the constituent elements of different examples, may also be included within the scope of the forms of the encoding device 100 and the decoding device 200.
[0552] It is also possible to implement one or more of the forms disclosed herein in combination with at least a portion of other forms disclosed herein. Furthermore, it is also possible to implement a portion of the processing described in the flowcharts of one or more forms disclosed herein, a portion of the structure of the apparatus, a portion of the syntax, etc., in combination with other forms.
[0553] [Implementation and Application] In the above embodiments, each functional block or active block can typically be implemented using an MPU (microprocessor unit) and memory. Alternatively, the processing of each functional block can be implemented by a program execution unit such as a processor that reads and executes software (programs) recorded in a recording medium such as ROM. This software can be distributed. The software can also be recorded in various recording media such as semiconductor memory. Furthermore, each functional block can also be implemented using hardware (dedicated circuitry).
[0554] The processing described in each embodiment can be implemented either centrally using a single device (system) or distributedly using multiple devices. Furthermore, the processor executing the above program can be either single or multiple. That is, it can be either centrally processed or distributed.
[0555] The form of this disclosure is not limited to the above embodiments, and various modifications can be made, which are also included within the scope of this disclosure.
[0556] Furthermore, this section describes application examples of the motion picture encoding method (image encoding method) or motion picture decoding method (image decoding method) shown in the above embodiments, and various systems implementing such application examples. Alternatively, such a system may be characterized by having an image encoding device using the image encoding method, an image decoding device using the image decoding method, or an image encoding / decoding device possessing both. Other structures of such a system can be appropriately modified depending on the circumstances.
[0557] [Usage Examples] Figure 40 This is a diagram showing the overall structure of an appropriate content delivery system ex100 that implements content distribution services. The provision of communication services is divided into desired sizes, and each unit contains base stations ex106, ex107, ex108, ex109, and ex110, which are fixed wireless stations as illustrated in the example.
[0558] In this content delivery system ex100, various devices such as computer ex111, game console ex112, camera ex113, home appliance ex114, and smartphone ex115 are connected to the Internet ex101 via Internet service provider ex102 or communication network ex104, and base stations ex106 to ex110. The content delivery system ex100 can also combine and connect some of the aforementioned devices. In various embodiments, the devices can also be directly or indirectly interconnected via telephone network or short-range wireless without using base stations ex106 to ex110. Furthermore, the streaming media server ex103 can also be connected to the computer ex111, game console ex112, camera ex113, home appliance ex114, and smartphone ex115 via the Internet ex101, etc. Additionally, the streaming media server ex103 can also be connected to terminals within a hotspot on an aircraft ex117 via satellite ex116.
[0559] Alternatively, it can replace base stations ex106 to ex110 by using wireless access points or hotspots. Furthermore, the streaming media server ex103 can connect directly to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or it can connect directly to the aircraft ex117 without going through the satellite ex116.
[0560] The camera ex113 is a digital camera or other device capable of still and moving image photography. Additionally, the smartphone ex115 refers to smartphones, portable phones, or PHS (Personal Handyphone System) phones that correspond to mobile communication systems known as 2G, 3G, 3.9G, 4G, and the future 5G.
[0561] Home appliances EX114 refers to refrigerators or other equipment included in home fuel cell combined heat and power systems.
[0562] In the content supply system ex100, a terminal with photography capabilities connects to the streaming media server ex103 via a base station ex106, thereby enabling on-site distribution. During on-site distribution, terminals (such as computers ex111, game consoles ex112, cameras ex113, home appliances ex114, smartphones ex115, and terminals within airplanes ex117) can perform the encoding processing described in the above embodiments on still or moving images captured by the user using the terminal. They can also multiplex the encoded image data and the encoded audio data, and send the resulting data to the streaming media server ex103. In other words, each terminal functions as an image encoding device according to one aspect of this disclosure.
[0563] On the other hand, the streaming media server ex103 will stream the content data sent by the requesting client. The client can be a terminal within a computer ex111, game console ex112, camera ex113, home appliance ex114, smartphone ex115, or airplane ex117, capable of decoding the encoded data described above. Each device receiving the distributed data can also decode and reproduce the received data. That is, each device can also function as an image decoding apparatus according to one aspect of this disclosure.
[0564] [Decentralized processing] Furthermore, the streaming media server ex103 can also be multiple servers or multiple computers, distributing data through decentralized processing or recording. For example, the streaming media server ex103 can also be implemented using a CDN (Contents Delivery Network), which distributes content by connecting many edge servers scattered around the world. In a CDN, physically nearby edge servers are dynamically allocated based on the client. Furthermore, by caching and distributing content to these edge servers, latency can be reduced. Moreover, in the event of several types of errors or changes in communication status due to increased traffic, processing can be distributed across multiple edge servers, the distribution entity can be switched to other edge servers, or distribution can continue by bypassing parts of the faulty network, thus achieving high-speed and stable distribution.
[0565] Furthermore, beyond the decentralized processing of distribution itself, the encoding processing of the captured data can be performed by each terminal, on the server side, or shared among them. For example, the encoding process typically involves two processing loops. In the first loop, the complexity or encoding amount of the image per frame or scene unit is detected. In the second loop, processing is performed to maintain image quality while improving encoding efficiency. For instance, by having the terminal perform the first encoding processing and the server receiving the content perform the second encoding processing, the processing load on each terminal can be reduced while improving content quality and efficiency. In this case, if there is a request for near real-time reception and decoding, the data encoded by the terminal in the first round can also be received and reproduced by other terminals, thus enabling more flexible real-time distribution.
[0566] As another example, cameras like the EX113 extract features from images, compress the data about these features as metadata, and send it to a server. The server, for instance, adjusts the quantization precision based on the features to determine the importance of the target, performing compression that corresponds to the meaning (or importance) of the image. The feature data is particularly effective in improving the accuracy and efficiency of motion vector prediction during further compression on the server. Alternatively, simple encoding such as VLC (Variable Length Coding) can be performed by the terminal, while more demanding encoding methods like CABAC (Context Adaptive Binary Arithmetic Coding) can be used by the server.
[0567] As another example, in stadiums, shopping malls, or factories, there may be multiple images of roughly the same scene captured by multiple terminals. In such cases, the data is distributed and encoded separately using the terminals that captured the images, as well as other terminals and servers that did not capture images, as needed. This can be done by assigning encoding and processing data to different units, such as GOP (Group of Pictures), image units, or tile units obtained by segmenting images. This reduces latency and improves real-time performance.
[0568] Since multiple image datasets depict roughly the same scene, the server can manage and / or instruct the data to be cross-referenced between images captured by different terminals. Alternatively, the server can receive encoded data from each terminal and change the reference relationships between the multiple datasets, or modify or replace the images themselves for re-encoding. This allows for the generation of streams with improved quality and efficiency for each data point.
[0569] Furthermore, the server can also transcode the image data by changing its encoding method before distributing it. For example, the server can convert MPEG encoding to VP encoding (such as VP9), or H.264 to H.265.
[0570] In this way, encoding processing can be performed by a terminal or one or more servers. Therefore, the terms "server" or "terminal" will be used below to refer to the main body performing the processing, but it is also possible to perform part or all of the processing performed by the server by the terminal, or vice versa. Furthermore, the same applies to decoding processing.
[0571] [3D, Multi-angle] The use of images or videos of different scenes captured by multiple cameras (ex113) and / or smartphones (ex115) that are roughly synchronized with each other, or images or videos of the same scene captured from different angles, is increasing. Images captured by each terminal are merged based on the relative positional relationship between the terminals or regions containing consistent feature points within the images.
[0572] The server not only encodes 2D moving images, but can also automatically or at user-specified times encode still images based on scene analysis of moving images and send them to the receiving terminal. Furthermore, when the server can obtain the relative positional relationships between shooting terminals, it can generate 3D shapes of scenes not only from 2D moving images, but also from images of the same scene captured from different angles. The server can also separately encode 3D data generated from point clouds, or select or reconstruct images from multiple terminals based on the results of identifying or tracking people or targets using 3D data to generate images to be sent to the receiving terminal.
[0573] In this way, users can freely select images corresponding to each shooting terminal to appreciate the scene, and can also appreciate the content of images extracted from the selected viewpoint from 3D data reconstructed using multiple images or videos. Furthermore, along with the images, sound can also be collected from multiple different angles. The server multiplexes the sound from a specific angle or space with the corresponding image and sends the multiplexed image and sound.
[0574] Furthermore, in recent years, content that establishes a correspondence between the real world and the virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has become increasingly popular. In the case of VR images, the server creates separate viewpoint images for the right and left eyes. These images can be encoded using methods such as Multi-View Coding (MVC) to allow reference between the viewpoint images, or they can be encoded as separate streams without reference to each other. During the decoding of these different streams, they can be synchronously reproduced according to the user's viewpoint to recreate the virtual three-dimensional space.
[0575] In the case of AR images, the server can overlay virtual object information in virtual space onto camera information in real space based on 3D position or the user's viewpoint movement. The decoding device acquires or holds the virtual object information and 3D data, generates a 2D image based on the user's viewpoint movement, and creates overlay data by smoothly connecting them. Alternatively, the decoding device can send the user's viewpoint movement to the server in addition to requesting virtual object information. Alternatively, the server can create overlay data by matching the received viewpoint movement with the 3D data held on the server, encode the overlay data, and distribute it to the decoding device. Furthermore, the overlay data has an α value representing transmittance in addition to RGB values. The server sets the α value of the portion outside the target created from the 3D data to 0, etc., and encodes the portion in a state of transmittance. Alternatively, the server can set a predetermined RGB value as the background, similar to a chroma key, to generate data where the portion outside the target is set as the background color.
[0576] Similarly, the decoding of distributed data can be performed by the individual terminals acting as clients, on the server side, or distributed among them. For example, one terminal could first send a receive request to the server, other terminals could receive the content corresponding to that request, decode it, and then send the decoded signal to a device with a display. By distributing the processing independently of the capabilities of the communicating terminals and selecting appropriate content, data with better image quality can be reproduced. Furthermore, as another example, large-format image data can be received by a TV, and the viewer's personal terminal can decode and display a segmented area of the image, such as tiles. This allows for the sharing of the overall image while simultaneously allowing the viewer to identify their own area of responsibility or areas they wish to examine in more detail.
[0577] In situations where multiple short-range, medium-range, or long-range wireless communications are available, both indoors and outdoors, seamless content reception may be possible using distribution system standards such as MPEG-DASH (Dynamic Adaptive Streaming over HTTP). Users can also freely select their terminal, decoding device, or display device (such as an indoor or outdoor monitor) and switch between them in real time. Furthermore, decoding can be performed using location information and other methods to switch between decoding and display terminals. This allows information to be mapped and displayed on a portion of the wall or floor of a building adjacent to a display device while the user is moving towards their destination. Additionally, the bit rate of the received data can be switched based on the ease of access to encoded data on the network, such as data cached on a server accessible only briefly from the receiving terminal or copied to an edge server of the content distribution service.
[0578] [Web page optimization] Figure 43 This is an example of a web page display screen in a computer such as ex111. Figure 40 This is an example image showing the display screen of a web page in a smartphone such as the ex115. Figure 44 and As shown, in cases where a web page contains multiple linked images that serve as links to image content, their visibility can vary depending on the viewing device. When multiple linked images are visible on the screen, before the user explicitly selects a linked image, or before the linked image is near the center of the screen, or before the entire linked image enters the screen, the display device (decoding device) can display still images or I-images of each content as linked images. Alternatively, it can display images such as GIF animations using multiple still images or I-images, or it can decode and display the image by receiving only the basic layer.
[0579] When a user selects a linked image, the display device prioritizes the base layer and decodes it. Additionally, if the HTML (HyperText Markup Language) constituting the web page contains information indicating tiered content, the display device can also decode up to the enhancement layer. Furthermore, in situations where real-time performance is ensured before selection or when communication bandwidth is extremely limited, the display device can reduce the delay between the decoding and display of the first image (the delay from the start of content decoding to the start of display) by decoding and displaying only the preceding reference images (I images, P images, and only B images for preceding reference). Moreover, the display device can also forcibly ignore image reference relationships and coarsely decode all B and P images as preceding references, performing normal decoding as more images are received over time.
[0580] [Autonomous Driving] Furthermore, when receiving still images or video data such as two-dimensional or three-dimensional map information for the purpose of autonomous driving or driving assistance, the receiving terminal can also receive information such as weather or construction information as metadata, in addition to image data belonging to more than one layer, and establish correspondences between them for decoding. Moreover, metadata can belong to a layer or be multiplexed only with image data.
[0581] In this scenario, since the vehicle, drone, or aircraft containing the receiving terminal is moving, the receiving terminal can seamlessly receive and decode data while switching between base stations ex106 and ex110 by transmitting its location information. Furthermore, the receiving terminal can dynamically adjust the level of metadata reception and map information updates based on user selection, user status, and / or the status of the communication frequency band.
[0582] In the content delivery system ex100, the client can receive, decode, and reproduce the encoded information sent by the user in real time.
[0583] [Distribution of Personal Content] Furthermore, the content delivery system ex100 not only handles high-quality, long-duration content provided by video distribution providers, but also enables unicast or multicast distribution of low-quality, short-duration content provided by individuals. It's conceivable that such personal content will increase in the future. To improve the quality of personal content, the server can also encode it after editing. This can be implemented, for example, using the following structure.
[0584] During or after capturing images, the server performs image processing such as identifying shooting errors, scene searching, meaning analysis, and target detection based on the original image data or encoded data. Furthermore, based on the identification results, the server manually or automatically corrects focus deviations or camera shake, deletes scenes of lower importance (e.g., scenes with lower brightness or out-of-focus images), emphasizes target edges, or adjusts tones. The server then encodes the edited data. Additionally, recognizing that longer shooting times reduce audiovisual quality, the server can automatically limit not only low-importance scenes as described above but also scenes with minimal movement to specific timeframes based on image processing results. Alternatively, the server can generate and encode summaries based on the meaning analysis results of the scenes.
[0585] Personal content may contain content that infringes on copyright, the author's moral rights, or portrait rights, or the shared scope may exceed the intended range, causing inconvenience to the individual. Therefore, for example, a server could forcibly encode images such as the faces of people in the periphery of a scene or a home environment as out-of-focus images. Furthermore, the server can identify whether a face different from a pre-registered person is captured within the image to be encoded, and if so, apply processing such as mosaic to the face. Alternatively, as pre- or post-processing for encoding, from a copyright perspective, the user can specify the person or background area to be processed. The server can then replace the specified area with another image or blur the focus. If it involves a person, the server can track the person in a moving image and replace the image of their face.
[0586] For personal content with small data volumes, real-time audiovisual requirements are more stringent. Therefore, although bandwidth also plays a role, the decoding device prioritizes receiving, decoding, and reproducing the base layer. The decoding device can also receive enhancement layers during this process, including them in high-quality image reproduction if the playback is looped or repeated more than twice. In this way, if the stream is scalably encoded, it can provide an experience where the motion images are initially coarse, but gradually become smoother and the image quality improves. Besides scalable encoding, the same experience can be provided when the first, coarser stream and a second stream encoded based on the first motion image are combined into a single stream.
[0587] [Other Implementation Examples] Furthermore, these encoding or decoding processes are typically handled within the LSI ex500 integrated circuit on each terminal. (LSI (large-scale integration circuitry) ex500 (reference)) The structure can be either a single chip or composed of multiple chips. Furthermore, software for motion image encoding or decoding can be installed on a recording medium (CD-ROM, floppy disk, hard disk, etc.) that can be read by a computer such as the ex111, and the software can then be used for encoding and decoding. Moreover, when the ex115 smartphone has a camera, motion image data acquired by that camera can also be transmitted. This motion image data is encoded using the LSIex500 chip found in the ex115 smartphone.
[0588] Alternatively, the LSIex500 can also be a structure that downloads and activates application software. In this case, the terminal first determines whether it corresponds to the content's encoding method or whether it has the capability to execute a specific service. If the terminal does not correspond to the content's encoding method or does not have the capability to execute a specific service, the terminal downloads the codec or application software, and then retrieves and reproduces the content.
[0589] Furthermore, not limited to the content delivery system ex100 via the Internet ex101, at least one of the moving image encoding device (image encoding device) or moving image decoding device (image decoding device) of the above embodiments can also be assembled in a digital broadcasting system. Since multiplexed data that multiplexes images and sound is carried and transmitted and received using radio waves for broadcasting via satellites or the like, it is more suitable for multicasting than the unicast-friendly structure of the content delivery system ex100, but the same applications can be performed for encoding and decoding processing.
[0590] [Hardware Structure] This is a further detailed explanation The image shown is of the EX115 smartphone. Additionally... This diagram illustrates a structural example of a smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves with a base station ex110, a camera unit ex465 for capturing images and still images, and a display unit ex458 for displaying decoded data such as images captured by the camera unit ex465 and images received by the antenna ex450. The smartphone ex115 also includes an operation unit ex466, such as a touch panel; a sound output unit ex457, such as a speaker, for outputting sound or audio; a sound input unit ex456, such as a microphone, for inputting sound; a memory unit ex467 capable of storing encoded or decoded data such as captured images or still images, recorded audio, received images or still images, and emails; and a slot unit ex464 serving as an interface with a SIM (Subscriber Identity Module) ex468, which is used to identify the user and authenticate access to various data sources, including the network. Alternatively, an external memory may be used instead of the memory unit ex467.
[0591] The main control unit ex460, which performs integrated control of the display unit ex458 and the operation unit ex466, is synchronously interconnected with the power supply circuit unit ex461, the operation input control unit ex462, the image signal processing unit ex455, the camera interface unit ex463, the display control unit ex459, the modulation / demodulation unit ex452, the multiplexing / demultiplexing unit ex453, the audio signal processing unit ex454, the slot unit ex464, and the memory unit ex467 via the bus ex470.
[0592] If the power button is turned on by the user, the power circuit section ex461 will activate the smartphone ex115 and supply power to all parts from the battery pack.
[0593] The smart phone ex115 performs call and data communication processing based on the control of the main control unit ex460, which includes a CPU, ROM, and RAM. During a call, the audio signal processing unit ex454 converts the audio signal collected by the audio input unit ex456 into a digital audio signal. The modulation / demodulation unit ex452 performs spectral diffusion processing, and the transmitting / receiving unit ex451 performs digital-to-analog conversion and frequency conversion processing. The resulting signal is transmitted via the antenna ex450. Furthermore, the received data is amplified and subjected to frequency conversion and analog-to-digital conversion processing. The modulation / demodulation unit ex452 performs inverse spectral diffusion processing, and the audio signal processing unit ex454 converts it into an analog audio signal, which is then output from the audio output unit ex457. In data communication mode, text, still images, or video data can be sent to the main control unit ex460 via the operation input control unit ex462 based on the operation of the main unit's operation unit ex466. The same transmission and reception processing is performed. In data communication mode, when transmitting images, still images, or images and sound, the image signal processing unit ex455 compresses and encodes the image signal stored in the memory unit ex467 or the image signal input from the camera unit ex465 using the moving image encoding method described in the above embodiments, and sends the encoded image data to the multiplexing / demultiplexing unit ex453. The sound signal processing unit ex454 encodes the sound signal collected by the sound input unit ex456 during the capture of images and still images by the camera unit ex465, and sends the encoded sound data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded image data and the encoded sound data in a prescribed manner, performs modulation and conversion processing by the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmitting / receiving unit ex451, and transmits it via the antenna ex450.
[0594] When receiving images attached to emails or chat tools, or images linked to web pages, the multiplexing / demultiplexing unit ex453 separates the multiplexed data received via antenna ex450 into bitstreams of image data and bitstreams of audio data to decode the multiplexed data. The encoded image data is then supplied to the image signal processing unit ex455 via the synchronization bus ex470, and the encoded audio data is supplied to the audio signal processing unit ex454. The image signal processing unit ex455 decodes the image signal using a motion picture decoding method corresponding to the motion picture encoding method described in the above embodiments, and displays the image or still image contained in the linked motion picture file from the display unit ex458 via the display control unit ex459. The audio signal processing unit ex454 decodes the audio signal and outputs audio from the audio output unit ex457. Due to the increasing popularity of live streaming, depending on the user's situation, the audio reproduction may be unsuitable for certain social contexts. Therefore, as an initial value, it is preferable to have a structure that reproduces only the image data and not the sound signal, and reproduces the sound synchronously only when the user performs an operation such as clicking on the image data.
[0595] Furthermore, while the EX115 smartphone was used as an example here, as a terminal, three other installation methods can be considered besides transceiver terminals with both encoders and decoders: a transmitting terminal with only an encoder, and a receiving terminal with only a decoder. In digital broadcasting systems, the explanation assumes that multiplexed data, which multiplexes audio data into video data, is being received and transmitted. However, in addition to audio data, multiplexed data can also multiplex character data associated with the video. Alternatively, the video data itself can be received or transmitted instead of multiplexed data.
[0596] Furthermore, while the description assumes the CPU's main control unit (ex460) controls the encoding or decoding process, many terminals also possess GPUs. Therefore, a structure can be implemented that utilizes GPU performance to process larger regions simultaneously, using shared memory between the CPU and GPU (Graphics Processing Unit), or memory with shared address management. This reduces encoding time, ensures real-time performance, and achieves low latency. In particular, it is even more efficient to perform motion search, deblocking filtering, SAO (Sample Adaptive Offset), and transform / quantization processing on the GPU at the image level, rather than using the CPU.
[0597] Industrial applicability This disclosure can be used in encoding devices for encoding moving images, and can be applied to video conferencing systems, etc.
[0598] Explanation of reference numerals in the attached figures 100, 700 encoding devices 102 Divisions 104 Subtraction Department 106 Transformer 108 Quantitative Department 110 Entropy Coding Department 112, 204 Inverse Quantization Section Inverse Transformation Units 114 and 206 116, 208 Addition Department 118 and 210 memory blocks 120, 212 cyclic filter sections 122, 214 frame memory 124, 216 Intra-frame Prediction Unit Inter-frame prediction units 126 and 218 128, 220 Prediction and Control Department 130, 222 Prediction Parameter Generation Department Compressors 131, 133, 134, 731, and 733 132, 232, 732, 832 exporters 151, 251, 351 circuits 152, 252, 352 memory 200, 800 decoding devices 202 Entropy Decoding Department 224 Division Decision Department 231, 233, 235, 831, 833 decompressors 234 and 834 generators 236 synthesizer 300-bit stream generator
Claims
1. A decoding device, comprising: Memory; and The circuit is connected to the memory, wherein, The circuit is in operation. Decode from one or more streams (i) a reference image containing the face, (ii) geometric information corresponding to multiple frames of the captured motion image obtained by the camera and representing the geometric properties of the subject, and (iii) background information related to the background image. Based on the reference image, the geometric information, and the background information, a generative model is used to generate a synthetic facial motion image, which is a motion image containing the face and is obtained by synthesizing the background image.
2. The decoding device according to claim 1, wherein, The circuit inputs the reference image, the geometric information, and the background image into the generation model, and obtains the synthesized facial motion image from the generation model.
3. The decoding device according to claim 1, wherein, The circuit, The reference image and the geometric information are input into the generative model, and an intermediate facial motion image is obtained from the generative model. The intermediate facial motion image is a motion image containing the face and is a motion image without synthesized background image. The synthesized facial motion image is generated by embedding a corresponding region in the background image into the background region of the intermediate facial motion image.
4. The decoding device according to claim 3, wherein, The circuit, The intermediate facial motion image is segmented to obtain segmentation information representing the foreground and background regions within the intermediate facial motion image. The background region in the intermediate face motion image is determined using the segmentation information of the intermediate face motion image.
5. The decoding device according to claim 3, wherein, The circuit, Decoding from one or more streams represents the segmentation information of the foreground and background regions in the captured motion image. Using the segmented information of the captured motion image, the background region in the intermediate facial motion image is determined.
6. The decoding apparatus according to claim 3, wherein, A specified background color code is embedded in the background area of the reference image.
7. The decoding apparatus according to claim 6, wherein, The circuit determines the region in the intermediate face motion image that has the specified background color code as the background region in the intermediate face motion image.
8. The decoding apparatus according to claim 6 or 7, wherein, The circuit decodes background color code information representing the specified background color code from one or more streams.
9. The decoding apparatus according to claim 8, wherein, The background color code information will include a range of consecutive values, which will be represented by the specified background color code. The specified background color code is defined within the range represented by the background color code information.
10. The decoding apparatus according to claim 3, wherein, The circuit, Reference image segmentation information representing the foreground and background regions in the reference image is decoded from one or more streams. Using the segmentation information of the reference image, a specified background color code is embedded in the background region of the reference image.
11. The decoding apparatus according to any one of claims 1 to 7, wherein, The background image is an image prepared independently of the reference image and the captured motion image.
12. The decoding apparatus according to any one of claims 1 to 7, wherein, The circuit, The identifier of the background image is decoded into the background information. The identifier is used to select the background image from a plurality of background image candidates.
13. The decoding apparatus according to any one of claims 1 to 7, wherein, The background image is an image contained in the captured motion image, or an image obtained by synthesizing multiple images contained in the captured motion image.
14. The decoding apparatus according to any one of claims 1 to 7, wherein, The circuit decodes the reference image as the background information. The reference image is applied to the background image.
15. The decoding apparatus according to claim 13, wherein, When the background image contains a foreground region, the circuit interpolates the missing background region in the background image using the surrounding area of the foreground region in the background image or the background region in a past synthetic facial motion image.
16. The decoding apparatus according to claim 15, wherein, The circuit, The background image is segmented to obtain background image segmentation information representing the foreground region and the background region in the background image. Using the background image segmentation information, the foreground region and the background region in the background image are determined.
17. The decoding apparatus according to claim 15, wherein, The foreground region in the background image is embedded with a specified foreground color code.
18. The decoding apparatus according to claim 17, wherein, The circuit, Foreground color code information representing the specified foreground color code is decoded from one or more streams. The region in the background image that has the specified foreground color code is determined as the foreground region in the background image.
19. The decoding apparatus according to claim 18, wherein, The foreground color code information will be represented by a range of consecutive values as the specified foreground color code. The specified foreground color code is defined within the range represented by the foreground color code information.
20. The decoding apparatus according to any one of claims 1 to 7, wherein, The background image is decoded into an image from the access units in the one or more streams. The signal indicating that the background image exists in the access unit is decoded from the supplementary enhancement information (SEI) corresponding to the access unit containing the background image.
21. The decoding apparatus according to any one of claims 1 to 7, wherein, The background image is applied together to multiple frames of the synthesized facial motion image.
22. The decoding apparatus according to any one of claims 1 to 7, wherein, The circuit decodes at least one of background color code information representing a specified background color code and foreground color code information representing a specified foreground color code from the supplementary enhancement information SEI in the one or more streams.
23. The decoding apparatus according to any one of claims 1 to 7, wherein, The circuit decodes from the supplementary enhancement information SEI in the one or more streams at least one of the captured motion image segmentation information representing the foreground and background regions in the captured motion image, and the reference image segmentation information representing the foreground and background regions in the reference image.
24. A decoding device, comprising: Memory; and The circuit is connected to the memory, wherein, The circuit is in operation. Decode from one or more streams (i) a reference image containing the face and (ii) geometric information corresponding to multiple frames of the moving images obtained by the camera, representing the geometric properties of the subject. The reference image and the geometric information are input into the generation model, and an intermediate facial motion image containing the motion image of the face is obtained from the generation model. A synthetic facial motion image is generated by embedding a corresponding region from the reference image into the background region of the intermediate facial motion image.
25. An encoding device comprising: Memory; and The circuit is connected to the memory, wherein, The circuit is in operation. (i) a reference image, (ii) geometric information, and (iii) background information used to generate a synthetic facial motion image are encoded into one or more streams. The synthetic facial motion image is a motion image containing a face and is obtained by synthesizing a background image. The (i) reference image is an image containing the face. The (ii) geometric information corresponds to multiple frames of the captured motion image obtained by the camera and represents the geometric properties of the subject. The (iii) background information is related to the background image.
26. The encoding device according to claim 25, wherein, The circuit, The captured moving image is segmented to obtain segmentation information representing the foreground and background regions in the captured moving image. The segmented information of the captured motion image is encoded into one or more streams.
27. The encoding device according to claim 25, wherein, The circuit, The reference image is segmented to obtain reference image segmentation information representing the foreground and background regions in the reference image. Using the segmentation information of the reference image, a specified background color code is embedded in the background region of the reference image. The reference image, in which the specified background color code is embedded in the background area, is encoded.
28. The encoding device according to claim 27, wherein, The circuit encodes the background color code information representing the specified background color code into one or more streams.
29. The encoding device according to claim 28, wherein, The background color code information will include a range of consecutive values, which will be represented by the specified background color code. The specified background color code is defined within the range represented by the background color code information.
30. The encoding device according to any one of claims 27 to 29, wherein, The specified background color code is defined by a color code that occurs at a frequency below a threshold in the foreground region of the reference image.
31. The encoding device according to claim 25, wherein, The circuit, The reference image is segmented to obtain reference image segmentation information representing the foreground and background regions in the reference image. The reference image segmentation information is encoded into one or more streams.
32. The encoding device according to any one of claims 25 to 29, wherein, The background image is an image prepared independently of the reference image and the captured motion image.
33. The encoding device according to any one of claims 25 to 29, wherein, The circuit, Select the background image from a plurality of background image candidates. The identifier of the background image is encoded as the background information.
34. The encoding device according to any one of claims 25 to 29, wherein, The background image is an image contained in the captured motion image, or an image obtained by synthesizing multiple images contained in the captured motion image.
35. The encoding device according to any one of claims 25 to 29, wherein, The circuit encodes the reference image into the background information. The reference image is applied to the background image.
36. The encoding device according to claim 34, wherein, When the background image contains a foreground region, the circuit interpolates the missing background region in the background image using the surrounding area of the foreground region in the background image or the background region in other images contained in the captured motion image.
37. The encoding device according to claim 36, wherein, The circuit, The background image is segmented to obtain background image segmentation information representing the foreground region and the background region in the background image. Using the background image segmentation information, the foreground region and the background region in the background image are determined.
38. The encoding device according to claim 34, wherein, The circuit, The background image is segmented to obtain background image segmentation information representing the foreground and background regions in the background image. Using the background image segmentation information, a specified foreground color code is embedded into the foreground region of the background image.
39. The encoding device according to claim 38, wherein, The circuit encodes the foreground color code information representing the specified foreground color code into one or more streams.
40. The encoding device according to claim 39, wherein, The foreground color code information will be represented by a range of consecutive values as the specified foreground color code. The specified foreground color code is defined within the range represented by the foreground color code information.
41. The encoding device according to claim 38, wherein, The specified foreground color code is defined by a color code in the background region of the background image that occurs at a frequency below a threshold.
42. The encoding device according to any one of claims 25 to 29, wherein, The background image is encoded as an image into one or more access units of the stream. The signal indicating that the background image exists in the access unit is encoded into the supplementary enhancement information (SEI) corresponding to the access unit where the background image is encoded.
43. The encoding device according to any one of claims 25 to 29, wherein, The circuit encodes at least one of the background color code information representing the specified background color code and the foreground color code information representing the specified foreground color code into the supplementary enhancement information (SEI) in one or more streams.
44. The encoding device according to any one of claims 25 to 29, wherein, The circuit encodes at least one of the captured motion image segmentation information representing the foreground and background regions in the captured motion image, and the reference image segmentation information representing the foreground and background regions in the reference image, into supplementary enhancement information (SEI) in one or more streams.
45. An encoding device comprising: Memory; and The circuit is connected to the memory, wherein, The circuit is in operation. The (i) reference image and (ii) geometric information used to generate synthetic facial motion images are encoded into one or more streams, wherein the (i) reference image is an image containing a face, and the (ii) geometric information corresponds to multiple frames of the captured motion images obtained by the camera and represents the geometric properties of the subject. When generating the synthetic facial motion image, the reference image and the geometric information are used to input the reference image and the geometric information into the generation model, and to obtain an intermediate facial motion image containing the facial motion image from the generation model. The reference image is further used to generate the synthetic facial motion image by embedding the corresponding region in the reference image into the background region in the intermediate facial motion image.
46. A bitstream generation apparatus, comprising: Memory; and The circuit is connected to the memory, wherein, The circuit is in operation. Generate a bitstream containing (i) a reference image, (ii) geometric information, and (iii) background information for generating a synthetic facial motion image, wherein the synthetic facial motion image is a motion image containing a face and is obtained by synthesizing a background image, wherein (i) the reference image is an image containing the face, the (ii) geometric information corresponds to multiple frames of the captured motion image obtained by the camera and represents the geometric properties of the subject, and the (iii) background information is related to the background image.
47. A decoding method, wherein, Decode from one or more streams (i) a reference image containing the face, (ii) geometric information corresponding to multiple frames of the captured motion image obtained by the camera and representing the geometric properties of the subject, and (iii) background information related to the background image. Based on the reference image, the geometric information, and the background information, a generative model is used to generate a synthetic facial motion image, which is a motion image containing the face and is obtained by synthesizing the background image.
48. An encoding method, wherein, (i) a reference image, (ii) geometric information, and (iii) background information used to generate a synthetic facial motion image are encoded into one or more streams. The synthetic facial motion image is a motion image containing a face and is obtained by synthesizing a background image. The (i) reference image is an image containing the face. The (ii) geometric information corresponds to multiple frames of the captured motion image obtained by the camera and represents the geometric properties of the subject. The (iii) background information is related to the background image.