Encoding device, decoding device, encoding method, and decoding method

The encoding device and method enhances the efficiency of impersonation in video conferencing, enhancing the realism of impersonation in video conferencing by accurately identifying and preventing fraudulent impersonation through encoding and decoding processes.

WO2025204830A1PCT designated stage Publication Date: 2025-10-02PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/008955
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-26
Filing Date
2025-03-11
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently handling digital video data, particularly in identifying and preventing fraudulent impersonation in video conferencing scenarios where a second person's face is superimposed on a reference image without proper authentication.

Method used

An encoding device and method that encodes reference images and feature data from input videos, along with determination-related information to assess the likelihood of identity mismatch, allowing for accurate identification of impersonation.

Benefits of technology

Enables real-time detection of impersonation in video conferencing, enhancing the efficiency of impersonation detection of impersonation in video conferencing scenarios, enhancing the realism of the output video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025008955_02102025_PF_FP_ABST
    Figure JP2025008955_02102025_PF_FP_ABST
Patent Text Reader

Abstract

An encoding device (100) comprises a memory (152), and a circuit (151) connected to the memory (152). The circuit (151), in operation, encodes at least one reference image including the face of a first person into a bitstream (S501), encodes feature amount data into a bitstream that is extracted from an input video including the face of a second person that is the same as or different from the first person, for deforming at least one reference image to generate an output video during decoding (S502), and encodes determination-related information into a bitstream that pertains to determination of the possibility that the second person is different from the first person (S503).
Need to check novelty before this filing date? Find Prior Art

Description

Encoding device, decoding device, encoding method, and decoding method

[0001] The present disclosure relates to an encoding device and the like.

[0002] Video coding technology has progressed from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Codec). With this progress, there is a constant need to provide improvements and optimizations in video coding technology to handle the ever-increasing amount of digital video data in various applications. This disclosure relates to further advances, improvements, and optimizations in video coding.

[0003] Non-Patent Document 1 relates to an example of a conventional standard regarding the above-mentioned video coding technology.

[0004] H. 265 (ISO / IEC 23008-2 HEVC) / HEVC (High Efficiency Video Coding)

[0005] With regard to the above-mentioned encoding methods, it is desirable to propose new methods to improve encoding efficiency, improve image quality, reduce the amount of processing, reduce the circuit scale, or appropriately select elements or operations such as filters, block sizes, motion vectors, reference pictures or reference blocks.

[0006] The present disclosure provides a configuration or method that can contribute to one or more of, for example, improved coding efficiency, improved image quality, reduced processing amount, reduced circuit size, improved processing speed, and appropriate selection of elements or operations, etc. Note that the present disclosure may include a configuration or method that can contribute to benefits other than those described above.

[0007] For example, an encoding device according to one aspect of the present disclosure includes a memory and a circuit connected to the memory, and the circuit, in operation, encodes into a bitstream at least one reference image including a face of a first person, encodes into the bitstream feature data extracted from an input video including a face of a second person that is the same as or different from the first person, the feature data being used to transform the at least one reference image to generate an output video upon decoding, and encodes into the bitstream determination-related information that is information related to a determination of the possibility that the second person is different from the first person.

[0008] Each embodiment of the present disclosure, or a partial configuration or method thereof, enables at least one of, for example, improved coding efficiency, improved image quality, reduced encoding / decoding processing volume, reduced circuit size, or improved encoding / decoding processing speed. Alternatively, each embodiment of the present disclosure, or a partial configuration or method thereof, enables appropriate selection of components / operations such as filters, block sizes, motion vectors, reference pictures, and reference blocks in encoding and decoding. Note that the present disclosure also includes disclosure of configurations or methods that may provide benefits other than those described above. For example, a configuration or method that improves coding efficiency while suppressing an increase in processing volume.

[0009] Further advantages and benefits of certain aspects of the present disclosure will become apparent from the specification and drawings. While such advantages and / or benefits may be obtained by several embodiments and features described in the specification and drawings, not all of them necessarily need to be provided to obtain one or more advantages and / or benefits.

[0010] These general or specific aspects may be realized by a system, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be realized by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.

[0011] A configuration or method according to an aspect of the present disclosure may contribute to, for example, one or more of improved coding efficiency, improved image quality, reduced processing amount, reduced circuit size, improved processing speed, and appropriate selection of elements or operations, etc. Note that a configuration or method according to an aspect of the present disclosure may also contribute to benefits other than those described above.

[0012] FIG. 1 is a block diagram showing the configuration of an encoding / decoding system in a reference example. FIG. 2 is a block diagram showing the configuration of an encoding device in a reference example. FIG. 3 is a block diagram showing the configuration of a decoding device in a reference example. FIG. 4 is a block diagram showing an example configuration of an encoding / decoding system in an embodiment. FIG. 5 is a diagram showing an example of a hierarchical structure of data in a stream. FIG. 6 is a block diagram showing an example configuration of an encoding device in an embodiment. FIG. 7 is a flowchart showing an example of encoding processing in an embodiment. FIG. 8 is an example of facial landmarks derived from an image. FIG. 9 is another example of facial landmarks derived from an image. FIG. 10 is yet another example of facial landmarks derived from an image. FIG. 11 is a diagram showing an example of a bitstream layout candidate. FIG. 12 is a diagram showing another example of a bitstream layout candidate. FIG. 13 is a diagram showing yet another example of a bitstream layout candidate. FIG. 14 is a diagram showing yet another example of a bitstream layout candidate. FIG. 15 is a diagram showing an example of a bitstream layout candidate compliant with VVC. FIG. 16 is a diagram showing another example of a bitstream layout candidate compliant with VVC. FIG. 17 is a block diagram showing an example configuration of a decoding device in an embodiment. FIG. 18 is a flowchart showing an example of a decoding process in an embodiment. FIG. 19 is a conceptual diagram showing the decoding process at each time instance. FIG. 20 is a diagram showing examples of various neural networks that generate multiple images. FIG. 21 is a conceptual diagram showing an example of operation when a person in a driving video matches a person in a reference image. FIG. 22 is a conceptual diagram showing an example of operation when a person in a driving video is different from a person in the reference image. FIG. 23 is a block diagram showing an example of the configuration of a video calling system in a first aspect. FIG. 24 is a block diagram showing an example of the configuration of a video calling system in a second aspect. FIG. 25 is a flowchart showing an example of operation of a transmitting side in the second aspect. FIG. 26 is a flowchart showing an example of operation of a receiving side in the second aspect. FIG. 27 is a syntax diagram showing an example of syntax in the second aspect. FIG. 28 is a block diagram showing an example of the configuration of a video calling system in a third aspect.

[0043] Fig. 29 is a flowchart showing an example of the operation of a transmitting side in the third aspect. Fig. 30 is a flowchart showing an example of the operation of a receiving side in the third aspect. Fig. 31 is a syntax diagram showing an example of syntax in the third aspect. Fig. 32 is a block diagram showing an example of the configuration of a video calling system in the fourth aspect. Fig. 33 is a block diagram showing an example of the configuration of a video calling system in the fifth aspect. Fig. 34 is a block diagram showing an example of the configuration of a video calling system in the sixth aspect. Fig. 35 is a syntax diagram showing an example of the syntax in the sixth aspect. Fig. 36 is a conceptual diagram showing an example of a user interface in which a composite video and a reference image input mode are simultaneously presented. Fig. 37 is a conceptual diagram showing an example of a user interface in which a composite video and an authentication image are simultaneously presented. Fig. 38 is a conceptual diagram showing an example of a user interface in which a composite video based on a reference image and a composite video based on an authentication image are alternately presented. Fig. 39 is a flowchart showing an example of the operation of a video calling system in the seventh aspect. Fig. 40 is a conceptual diagram showing an example of the user interface of a video calling system in the seventh aspect. Fig. 41 is a block diagram showing an example of an implementation of an encoding device. Fig. 42 is a block diagram showing an implementation example of a decoding device. Fig. 43 is a flowchart showing an example of a basic operation of an encoding device. Fig. 44 is a flowchart showing an example of a basic operation of a decoding device. Fig. 45 is a diagram showing the overall configuration of a content supply system that realizes a content distribution service. Fig. 46 is a diagram showing an example of a display screen of a web page. Fig. 47 is a diagram showing an example of a display screen of a web page. Fig. 48 is a diagram showing an example of a smartphone. Fig. 49 is a block diagram showing an example of the configuration of a smartphone.

[0013] This disclosure relates to an encoding device, a decoding device, an encoding method, and a decoding method for facial reconstruction. In this disclosure, facial reconstruction refers to a process of mapping pose and facial expression onto a facial image of a target person while ensuring that the target person's identity is maintained. Facial reconstruction technology can be used in a variety of applications, from video conferencing to entertainment.

[0014] The present disclosure may be used in multimedia data coding related to facial reconstruction techniques that seek to enhance the realism of the output video produced, where the terms video, image, and motion picture may be used interchangeably.

[0015] For example, in a video conferencing application, a driving video containing one or more frames of a user is first captured by an encoding device and then transmitted to a decoding device, which is a receiving device for real-time communication, where the video is reconstructed and displayed. The user can select whether their face is represented by live video captured by a camera, by one or more pre-configured animated avatars, or by one or more images containing a preset face.

[0016] Facial reconstruction techniques have also been widely adopted in the entertainment industry, such as in creating advertisements, editing movie scenes, and extending to music videos. In these applications, facial expressions and poses from real-life videos of people are composited onto a target's face while ensuring the target's identity and appearance are preserved. These techniques are being refined to handle both cross-identity video reconstruction, where the driving video and the target's face belong to different people, and same-identity video reconstruction, where the driving video and the target's face belong to the same person.

[0017] With the growing popularity and increasing use of various social media applications, facial reproduction technology provides users with flexibility, convenience, and ease in creating uniquely customized user expressions and symbolizing their emotions and personalities. For example, driving videos can be real-time videos or pre-recorded videos. In such scenarios, various facial reproduction technologies have been proposed.

[0018] 1 is a block diagram showing the configuration of an encoding / decoding system according to a reference example. For example, the encoding / decoding system includes an encoding device 700 and a decoding device 800. First, the encoding device 700 receives a reference image of a target person and a driving video, and encodes and compresses the received image into one or more bitstreams. Next, the encoding device 700 transmits the compressed bitstreams to the decoding device 800 via a transmission channel. Finally, the decoding device 800 reconstructs an output video from the received bitstreams.

[0019] For example, the reference image represents the visual characteristics of the display of the output video, and the driving video serves to give motion to the visual characteristics of the display of the output video.

[0020] 2 is a block diagram showing the configuration of an encoding device 700 according to a reference example. In this example, the encoding device 700 includes a compressor 701, a deriver 702, and a compressor 703.

[0021] Compressor 701 first encodes at least one reference image using video compression techniques, which may be a frame from a driving video, a pre-captured image containing the face of a person of interest, or an avatar.

[0022] The deriver 702 then feeds multiple frames of the driving video to a neural network to derive latent information represented by a latent space. The compressor 703 compresses the latent information into one or more bitstreams using methods such as entropy coding. The latent information varies depending on the implementation and is not human-readable or easily understandable by humans. Finally, the bitstreams are transmitted to the decoder 800 via a transmission channel.

[0023] 3 is a block diagram showing the configuration of a decoding device 800 according to a reference example. In this example, the decoding device 800 includes a decompressor 801, a decompressor 802, a deriving device 803, and a generator 804.

[0024] First, the decompressor 801 decodes and reconstructs at least one reference image from the bitstream. The decompressor 801 then supplies the reference image to a derivator 803, which corresponds to the derivator 702 of the encoding device 700. The derivator 803 derives latent information in a latent space from the reference image. The decompressor 802 then decodes and reconstructs the latent information.

[0025] The latent information represents the distribution of features generated in the latent space using a neural network, and is not easily understood or interpreted by humans. Each type of neural network has its own unique representation of transmitted data.

[0026] The generator 804 then generates the output video from the multiple latent information using a neural network. For example, a generative adversarial network may be used to generate the output video. The generator 804 may also use the reference images to render the output video.

[0027] In the examples shown in Figures 1-3, the reference image may include the face of a first person. The driving video may include the face of a second person. The second person may be the same as the first person or may be different from the first person. Thus, for example, the appearance of another person may be used in a video conference.

[0028] However, there are cases where this use is not appropriate. For example, a second person different from the first person may be able to participate in a video conference using the appearance of the first person without being identified. Other participants in the video conference may then believe that they are videoconferenceing with the first person, not the second person. This may lead to fraudulent "spoofing." However, it may not be easy to identify potential spoofing based on the output video.

[0029] Therefore, the encoding device of Example 1 includes a memory and a circuit connected to the memory, which, in operation, encodes into a bitstream at least one reference image including a face of a first person, encodes into the bitstream feature data extracted from an input video including a face of a second person, the second person being the same as or different from the first person, the feature data being used to transform the at least one reference image to generate an output video upon decoding, and encodes into the bitstream determination-related information, which is information related to a determination of the possibility that the second person is different from the first person.

[0030] This may allow the encoding device to communicate to the decoding device the possibility that the second person corresponding to the information source of the feature data reflected in the reference image in generating the output video is different from the first person in the reference image. Therefore, it may be possible to appropriately communicate the possibility of impersonation. Therefore, it may be possible to appropriately identify the possibility of impersonation.

[0031] The encoding device of Example 2 may also be the encoding device of Example 1, wherein the decision-related information includes information based on a comparison of the at least one reference image and the input video.

[0032] This may allow for conveying information regarding the likelihood that the second person is different from the first person based on a comparison of the reference image and the input video, and thus may allow for conveying accurate information regarding the likelihood that the second person is different from the first person.

[0033] Furthermore, the encoding device of Example 3 may be the encoding device of Example 1 or 2, in which the determination-related information includes information based on whether the at least one reference image is an image included in the input video, and if the at least one reference image is an image included in the input video, it is determined that the possibility that the second person is different from the first person is lower than a predetermined possibility.

[0034] This may allow for conveying information regarding the likelihood that the second person is different from the first person based on whether the reference image is included in the input video, and thus may allow for efficient conveying of the likelihood that the second person is different from the first person.

[0035] Also, the encoding device of Example 4 may be the encoding device of Example 3, in which the determination-related information includes information based on a comparison of the at least one reference image with the input video if the at least one reference image is not an image included in the input video.

[0036] This may allow, when the reference image is not included in the input video, to convey information based on a comparison between the reference image and the input video regarding the likelihood that the second person is different from the first person, thereby allowing for efficient conveyance of accurate information regarding the likelihood that the second person is different from the first person.

[0037] Furthermore, the encoding device of Example 5 may be any of the encoding devices of Examples 1 to 4, in which the determination-related information includes information based on whether the at least one reference image is composed of multiple reference images, and if the at least one reference image is composed of multiple reference images, it is determined that the possibility that the second person is different from the first person is lower than a predetermined possibility.

[0038] This may allow for conveying information regarding the likelihood that the second person is different from the first person based on whether the reference image is comprised of multiple images, and thus may allow for efficient conveying of the likelihood that the second person is different from the first person.

[0039] Furthermore, the encoding device of Example 6 may be any of the encoding devices of Examples 1 to 5, in which the determination-related information includes information based on whether the at least one reference image is an image acquired in advance, and if the at least one reference image is an image acquired in advance, it is determined that the possibility that the second person is different from the first person is higher than a predetermined possibility.

[0040] This may allow for communicating information regarding the likelihood that the second person is different from the first person based on whether the reference image is a pre-acquired image, and thus may allow for efficiently communicating the likelihood that the second person is different from the first person.

[0041] Also, the encoding device of Example 7 may be any of the encoding devices of Examples 1 to 6, wherein the decision-related information includes information based on a comparison between the at least one reference image and an image in the input video.

[0042] This may allow information to be conveyed regarding the likelihood that the second person is different from the first person based on a comparison between the reference image and an image of the input video, and thus may allow accurate information regarding the likelihood that the second person is different from the first person to be conveyed efficiently.

[0043] Furthermore, the encoding device of Example 8 may be any of the encoding devices of Examples 1 to 6, in which the determination-related information includes information based on a comparison between the at least one reference image and an image at a predetermined timing in the input video.

[0044] This may allow information regarding the likelihood that the second person is different from the first person to be communicated at a predetermined timing based on a comparison between the reference image and the image of the input video, thereby allowing accurate information regarding the likelihood that the second person is different from the first person to be communicated efficiently.

[0045] The encoding device of Example 9 may be the encoding device of Example 8, in which the predetermined timing corresponds to the start of a randomly accessible unit.

[0046] This may allow information regarding the likelihood that the second person is different from the first person to be conveyed based on a comparison between the reference image and the image of the input video at a timing corresponding to the start of a randomly accessible unit, thereby allowing accurate information regarding the likelihood that the second person is different from the first person to be conveyed efficiently.

[0047] The encoding device of Example 10 may be the encoding device of Example 8, in which the predetermined timing corresponds to a designated interval.

[0048] This may make it possible to transmit information regarding the possibility that the second person is different from the first person based on a comparison between the reference image and the image of the input video at a timing corresponding to the specified interval, thereby making it possible to efficiently transmit accurate information regarding the possibility that the second person is different from the first person.

[0049] Furthermore, the encoding device of Example 11 may be any of the encoding devices of Examples 1 to 10, in which the judgment-related information includes an authentication image that is an image included in the input video.

[0050] This may make it possible to transmit an authentication image regarding the possibility that the second person is different from the first person. Therefore, it may be possible to transmit information that is effective in determining the possibility that the second person is different from the first person. Therefore, it may be possible to appropriately communicate the possibility that the second person is different from the first person.

[0051] Further, the decoding device of Example 12 includes a memory and a circuit connected to the memory, which, in operation, decodes from the bitstream at least one reference image including a face of a first person; decodes from the bitstream feature data extracted from an input video including a face of a second person, the second person being the same as or different from the first person, during encoding, the feature data being used to transform the at least one reference image to generate an output video; and decodes from the bitstream determination-related information, which is information related to determining whether the second person is likely to be different from the first person.

[0052] This may allow the encoding device to communicate to the decoding device the possibility that the second person corresponding to the information source of the feature data reflected in the reference image in generating the output video is different from the first person in the reference image. Therefore, it may be possible to appropriately communicate the possibility of impersonation. Therefore, it may be possible to appropriately identify the possibility of impersonation.

[0053] Also, the decoding device of Example 13 may be the decoding device of Example 12, in which the decision-related information includes information based on a comparison between the at least one reference image and the input video.

[0054] This may allow for conveying information regarding the likelihood that the second person is different from the first person based on a comparison of the reference image and the input video, and thus may allow for conveying accurate information regarding the likelihood that the second person is different from the first person.

[0055] Furthermore, the decoding device of Example 14 may be the decoding device of Example 12 or 13, in which the determination-related information includes information based on whether the at least one reference image is an image included in the input video, and if the at least one reference image is an image included in the input video, it is determined that the possibility that the second person is different from the first person is lower than a predetermined possibility.

[0056] This may allow for conveying information regarding the likelihood that the second person is different from the first person based on whether the reference image is included in the input video, and thus may allow for efficient conveying of the likelihood that the second person is different from the first person.

[0057] Also, the decoding device of Example 15 may be the decoding device of Example 14, wherein the determination-related information includes information based on a comparison between the at least one reference image and the input video when the at least one reference image is not an image included in the input video.

[0058] This may allow, when the reference image is not included in the input video, to convey information based on a comparison between the reference image and the input video regarding the likelihood that the second person is different from the first person, thereby allowing for efficient conveyance of accurate information regarding the likelihood that the second person is different from the first person.

[0059] Furthermore, the decoding device of Example 16 may be any of the decoding devices of Examples 12 to 15, in which the determination-related information includes information based on whether the at least one reference image is composed of multiple reference images, and if the at least one reference image is composed of multiple reference images, it is determined that the possibility that the second person is different from the first person is lower than a predetermined possibility.

[0060] This may allow for conveying information regarding the likelihood that the second person is different from the first person based on whether the reference image is comprised of multiple images, and thus may allow for efficient conveying of the likelihood that the second person is different from the first person.

[0061] Furthermore, the decoding device of Example 17 may be any of the decoding devices of Examples 12 to 16, in which the determination-related information includes information based on whether or not the at least one reference image is an image acquired in advance, and if the at least one reference image is an image acquired in advance, it is determined that the possibility that the second person is different from the first person is higher than a predetermined possibility.

[0062] This may allow for communicating information regarding the likelihood that the second person is different from the first person based on whether the reference image is a pre-acquired image, and thus may allow for efficiently communicating the likelihood that the second person is different from the first person.

[0063] Furthermore, the decoding device of Example 18 may be any of the decoding devices of Examples 12 to 17, wherein the determination-related information includes information based on a comparison between the at least one reference image and one image in the input video.

[0064] This may allow information to be conveyed regarding the likelihood that the second person is different from the first person based on a comparison between the reference image and an image of the input video, and thus may allow accurate information regarding the likelihood that the second person is different from the first person to be conveyed efficiently.

[0065] Furthermore, the decoding device of Example 19 may be any of the decoding devices of Examples 12 to 17, in which the judgment-related information includes information based on a comparison between the at least one reference image and an image at a predetermined timing in the input video.

[0066] This may allow information regarding the likelihood that the second person is different from the first person to be communicated at a predetermined timing based on a comparison between the reference image and the image of the input video, thereby allowing accurate information regarding the likelihood that the second person is different from the first person to be communicated efficiently.

[0067] Furthermore, the decoding device of Example 20 may be the decoding device of Example 19, in which the predetermined timing corresponds to the start of a randomly accessible unit.

[0068] This may allow information regarding the likelihood that the second person is different from the first person to be conveyed based on a comparison between the reference image and the image of the input video at a timing corresponding to the start of a randomly accessible unit, thereby allowing accurate information regarding the likelihood that the second person is different from the first person to be conveyed efficiently.

[0069] Furthermore, the decoding device of Example 21 may be the decoding device of Example 19, in which the predetermined timing corresponds to a designated interval.

[0070] This may make it possible to transmit information regarding the possibility that the second person is different from the first person based on a comparison between the reference image and the image of the input video at a timing corresponding to the specified interval, thereby making it possible to efficiently transmit accurate information regarding the possibility that the second person is different from the first person.

[0071] Furthermore, the decoding device of Example 22 may be the decoding device of any one of Examples 12 to 21, in which the judgment-related information includes an authentication image that is an image included in the input video.

[0072] This may make it possible to transmit an authentication image regarding the possibility that the second person is different from the first person. Therefore, it may be possible to transmit information that is effective in determining the possibility that the second person is different from the first person. Therefore, it may be possible to appropriately communicate the possibility that the second person is different from the first person.

[0073] Furthermore, the decoding device of Example 23 may be any of the decoding devices of Examples 12 to 22, wherein the circuit determines the likelihood that the second person is different from the first person based on the determination-related information.

[0074] This may allow an appropriate determination to be made based on determination-related information related to the determination of the possibility that the second person is different from the first person, and therefore may allow an appropriate determination of the possibility of impersonation.

[0075] Furthermore, the decoding device of Example 24 may be the decoding device of any one of Examples 12 to 23, in which the circuitry displays information based on the decision-related information together with the output video.

[0076] This may make it possible to visually notify information regarding the determination of the possibility that the second person is different from the first person, and therefore may make it possible to appropriately notify the possibility of impersonation.

[0077] Furthermore, the decoding device of Example 25 may be the decoding device of Example 22, wherein the circuit alternately displays the output video generated by transforming the at least one reference image based on the feature data and another output video generated by transforming the authentication image based on the feature data.

[0078] This may make it possible to visually notify information that is useful for determining whether the second person is different from the first person, and therefore may make it possible to appropriately notify the possibility of impersonation.

[0079] Furthermore, the encoding method of Example 26 is an encoding method that encodes at least one reference image including a face of a first person into a bit stream, encodes feature data extracted from an input video including a face of a second person who is the same as or different from the first person into the bit stream, the feature data being used to generate an output video by transforming the at least one reference image during decoding, and encodes determination-related information, which is information related to determining whether the second person is likely to be different from the first person, into the bit stream.

[0080] This may allow the encoding device to communicate to the decoding device the possibility that the second person corresponding to the information source of the feature data reflected in the reference image in generating the output video is different from the first person in the reference image. Therefore, it may be possible to appropriately communicate the possibility of impersonation. Therefore, it may be possible to appropriately identify the possibility of impersonation.

[0081] Further, the decoding method of Example 27 is a decoding method that decodes at least one reference image including a face of a first person from a bitstream, decodes from the bitstream feature data extracted from an input video including a face of a second person who is the same as or different from the first person during encoding, the feature data being used to transform the at least one reference image to generate an output video, and decodes from the bitstream determination-related information that is information related to determining whether the second person is likely to be different from the first person.

[0082] This may allow the encoding device to communicate to the decoding device the possibility that the second person corresponding to the information source of the feature data reflected in the reference image in generating the output video is different from the first person in the reference image. Therefore, it may be possible to appropriately communicate the possibility of impersonation. Therefore, it may be possible to appropriately identify the possibility of impersonation.

[0083] Furthermore, these comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0084] [Definition of Terms] As an example, each term may be defined as follows.

[0085] (1) Image: A unit of data made up of a set of pixels, consisting of pictures or blocks smaller than pictures, and includes both moving images and still images.

[0086] (2) Picture: A processing unit of an image composed of a set of pixels, and is sometimes called a frame or field.

[0087] (3) Block: A processing unit for a set containing a specific number of pixels, and can be named anything, as shown in the following examples. It can also be shaped anything, including, for example, a rectangle made up of M×N pixels, a square made up of M×M pixels, a triangle, a circle, or any other shape.

[0088] (Examples of blocks) Slice / tile / brick CTU / superblock / basic division unit VPDU / hardware processing division unit CU / processing block unit / prediction block unit (PU) / orthogonal transform block unit (TU) / unit Sub-block

[0089] (4) Pixel / Sample A pixel / sample is a minimum unit point that constitutes an image, and includes not only pixels at integer positions but also pixels at decimal positions generated based on pixels at integer positions.

[0090] (5) Pixel Value / Sample Value: A value inherent to a pixel, including not only brightness value, color difference value, and RGB gradation, but also depth value or binary values ​​of 0 and 1.

[0091] (6) Flags: In addition to one bit, flags may be multi-bit, for example, parameters or indexes of two or more bits. In addition, flags may be multi-valued using other bases as well as two values ​​using binary numbers.

[0092] (7) Signal: A signal that is symbolized or coded to transmit information, including discrete digital signals as well as analog signals that take continuous values.

[0093] (8) Stream / Bitstream: A digital data string or flow. A stream / bitstream may consist of a single stream or multiple streams divided into multiple layers. It also includes transmission by serial communication over a single transmission path as well as transmission by packet communication over multiple transmission paths.

[0094] (9) Difference / Difference In the case of scalar quantities, in addition to simple difference (x-y), it is sufficient to include difference calculations, including absolute value of difference (|x-y|), squared difference (x^2-y^2), square root of difference (√(x-y)), weighted difference (ax-by: a, b are constants), and offset difference (x-y+a: a is an offset).

[0095] (10) Sum In the case of a scalar quantity, in addition to simple sum (x + y), it is sufficient if a sum operation is included, including the absolute value of the sum (|x + y|), sum of squares (x^2 + y^2), square root of the sum (√(x + y)), weighted sum (ax + by: a, b are constants), and offset sum (x + y + a: a is an offset).

[0096] (11) Based on: This includes cases where factors other than the one being based on are taken into consideration. It also includes cases where a result is obtained directly or via an intermediate result.

[0097] (12) Using (used, using) This includes cases where elements other than the target of use are taken into account. It also includes cases where a result is obtained directly or via an intermediate result.

[0098] (13) Prohibit (forbid) This can be rephrased as not being allowed. Also, not prohibiting or being allowed does not necessarily mean obligation.

[0099] (14) Limit (restriction / restrict / restricted) This can be rephrased as not being permitted. Also, not prohibiting something or being permitted does not necessarily mean that it is an obligation. Furthermore, it is sufficient if something is partially prohibited in terms of quantity or quality, and it also includes cases where something is completely prohibited.

[0100] (15) Chroma: An adjective, denoted by the symbols Cb and Cr, that specifies that a sample array or a single sample represents one of two color difference signals associated with a primary color. Instead of the term chroma, the term chrominance can also be used.

[0101] (16) Luma: An adjective, denoted by the symbol or subscript Y or L, that specifies that a sample array or a single sample represents a monochrome signal associated with a primary color. Instead of the term luma, the term luminance may also be used.

[0102] [Description] In the drawings, the same reference numerals refer to the same or similar elements, and the sizes and relative positions of the elements in the drawings are not necessarily drawn to scale.

[0103] Hereinafter, embodiments will be described in detail with reference to the drawings. Note that the embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, the arrangement and connection of the components, steps, and the relationship and order of the steps shown in the following embodiments are merely examples and are not intended to limit the scope of the claims.

[0104] Below, embodiments of an encoding device and a decoding device will be described. The embodiments are examples of encoding devices and decoding devices to which the processes and / or configurations described in each aspect of the present disclosure can be applied. The processes and / or configurations can also be implemented in encoding devices and decoding devices different from the embodiments. For example, with regard to the processes and / or configurations applied to the embodiments, any of the following may be implemented.

[0105] (1) Any of the multiple components of the encoding device or decoding device of the embodiments described in each aspect of the present disclosure may be replaced or combined with other components described in any of the aspects of the present disclosure.

[0106] (2) In the encoding device or decoding device according to the embodiment, the functions or processes performed by some of the components of the encoding device or decoding device may be changed in any way, such as by adding, replacing, or deleting a function or process. For example, any function or process may be replaced with or combined with another function or process described in any of the aspects of the present disclosure.

[0107] (3) In the method implemented by the encoding device or decoding device according to the embodiment, some of the processes included in the method may be arbitrarily modified, such as by addition, replacement, deletion, etc. For example, any process in the method may be replaced with or combined with another process described in any of the aspects of the present disclosure.

[0108] (4) Some of the components constituting the encoding device or decoding device of the embodiment may be combined with components described in any of the aspects of the present disclosure, or may be combined with components having some of the functions described in any of the aspects of the present disclosure, or may be combined with components that perform some of the processing performed by the components described in each aspect of the present disclosure.

[0109] (5) A component having part of the functionality of the encoding device or decoding device of an embodiment, or a component that performs part of the processing of the encoding device or decoding device of an embodiment, may be combined or replaced with a component described in any of the aspects of the present disclosure, a component having part of the functionality described in any of the aspects of the present disclosure, or a component that performs part of the processing described in any of the aspects of the present disclosure.

[0110] (6) In the method implemented by the encoding device or decoding device of the embodiment, any of the multiple processes included in the method may be replaced or combined with the process described in any of the aspects of the present disclosure or any similar process.

[0111] (7) Some of the processes included in the method implemented by the encoding device or decoding device of the embodiment may be combined with the processes described in any of the aspects of the present disclosure.

[0112] (8) The implementation of the processes and / or configurations described in each aspect of the present disclosure is not limited to the encoding device or decoding device of the embodiments. For example, the processes and / or configurations may be implemented in a device used for a purpose other than video encoding or video decoding disclosed in the embodiments.

[0113] Fig. 4 is a block diagram showing the configuration of a coding / decoding system according to this embodiment. For example, the coding / decoding system includes a coding device 100 and a decoding device 200. The example in Fig. 4 is similar to the example in Fig. 1, but the specific configuration and processing of the coding device 100, the specific configuration and processing of the decoding device 200, and the bitstream are different from those in the example in Fig. 1.

[0114] 1, the encoding device 100 receives the reference image of the subject and the driving video, encodes and compresses them into one or more bitstreams, and then transmits the compressed bitstreams to the decoding device 200 using a transmission channel. Finally, the decoding device 200 reconstructs the output video from the received bitstreams.

[0115] The reference image is an example of first information of the present disclosure and represents visual features on the display of the output video, and the driving video plays a role in providing motion to the visual features on the display of the output video.

[0116] [Data Structure] Figure 5 is a diagram showing an example of a hierarchical structure of data in a stream. The stream includes, for example, a video sequence. This video sequence includes, for example, a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), supplemental enhancement information (SEI), and multiple pictures, as shown in (a) of Figure 5.

[0117] In a video composed of multiple layers, the VPS includes coding parameters common to multiple layers, and coding parameters related to multiple layers included in the video or to each individual layer.

[0118] The SPS includes parameters used for the sequence, i.e., encoding parameters that the decoding device 200 refers to in order to decode the sequence. For example, the encoding parameters may indicate the width or height of a picture. Note that there may be multiple SPSs.

[0119] The PPS includes parameters used for a picture, i.e., encoding parameters referenced by the decoding device 200 to decode each picture in a sequence. For example, the encoding parameters may include a reference value of the quantization width used in decoding the picture and a flag indicating the application of weighted prediction. Note that there may be multiple PPSs. Furthermore, the SPS and PPS may be simply referred to as parameter sets.

[0120] A picture may include a picture header and one or more slices, as shown in (b) of Fig. 5. The picture header includes coding parameters that the decoding device 200 references to decode the one or more slices.

[0121] As shown in (c) of Fig. 5, a slice includes a slice header and one or more bricks. The slice header includes coding parameters that are referenced by the decoding device 200 to decode the one or more bricks.

[0122] As shown in FIG. 5(d), a brick includes one or more coding tree units (CTUs).

[0123] Note that a picture may not contain slices, but may instead contain tile groups, where a tile group contains one or more tiles, and a brick may contain slices.

[0124] A CTU is also called a superblock or a basic division unit. As shown in (e) of Fig. 5, such a CTU includes a CTU header and one or more CUs (Coding Units). The CTU header includes coding parameters that the decoding device 200 references to decode the one or more CUs.

[0125] A CU may be divided into multiple small CUs. Furthermore, as shown in (f) of FIG. 5, a CU includes a CU header, prediction information, and residual coefficient information. The prediction information is information for predicting the CU, and the residual coefficient information is information indicating a prediction residual, which will be described later. A CU is basically the same as a PU (Prediction Unit) and a TU (Transform Unit), but may include multiple TUs smaller than the CU, for example, in an SBT, which will be described later. A CU may also be processed for each VPDU (Virtual Pipeline Decoding Unit) that constitutes the CU. A VPDU is a fixed unit that can be processed in one stage, for example, when performing pipeline processing in hardware.

[0126] Note that a stream may not have some of the layers shown in FIG. 5 . The order of these layers may be changed, or some layers may be replaced with other layers. A picture currently being processed by a device such as the encoding device 100 or the decoding device 200 is referred to as a current picture. If the processing is encoding, the current picture is synonymous with a picture to be encoded, and if the processing is decoding, the current picture is synonymous with a picture to be decoded. A block, such as a CU or CU, currently being processed by a device such as the encoding device 100 or the decoding device 200 is referred to as a current block. If the processing is encoding, the current block is synonymous with a block to be encoded, and if the processing is decoding, the current block is synonymous with a block to be decoded.

[0127] Here, a region in which parameters used for encoding are described may be referred to as a header region. For example, the header region is a region including an SEI. The header region may further include a VPS, an SPS, a PPS, an SEI, a picture header, a slice header, a CTU header, and a CU header.

[0128] Furthermore, for example, pictures can be classified into one of several types, including I-pictures, P-pictures, and B-pictures. I-pictures are intra-predicted pictures, which are coded and decoded without reference to other pictures. P-pictures are uni-predicted pictures, which can be coded and decoded with reference to one other picture. B-pictures are bi-predicted pictures, which can be coded and decoded with reference to two other pictures.

[0129] Furthermore, a moving image may be composed of a plurality of GOPs (Groups of Pictures). A GOP means a set of pictures. A GOP includes one or more I pictures. A GOP may include one or more P pictures, or one or more B pictures. A GOP may be a unit that allows video editing and random access. A GOP may have a fixed number of pictures, or may have a GOP structure in which I pictures, P pictures, and B pictures are arranged in a fixed order.

[0130] [Encoding Configuration and Processing] The encoding configuration and processing are shown below. Note that the encoding configuration and processing shown below correspond to an example of the encoding configuration and processing for face reproduction, and the encoding configuration and processing are not necessarily limited to the following example. In particular, although geometric attributes are used in the following example, the above-mentioned latent information or the like may also be used.

[0131] 6 is a block diagram showing the configuration of an encoding device 100 according to this embodiment. The encoding device 100 generates a bitstream from a reference image and a driving video. In this example, the encoding device 100 includes a compressor 101, a deriver 102, and a compressor 103. The compressor 101 and the compressor 103 may be integrated.

[0132] For example, the compressor 101 compresses the reference image by encoding the reference image, the deriving unit 102 derives geometric information indicating geometric attributes from the driving video, and the compressor 103 compresses the geometric information by encoding the geometric information.

[0133] A plurality of reference images may be used, a plurality of pieces of geometric information may be used, or a plurality of bitstreams may be used. The reference images are an example of first information in the present disclosure. The geometric information is an example of second information in the present disclosure. Note that the geometric information indicating the geometric attribute may also be simply expressed as the geometric attribute.

[0134] 7 is a flowchart showing an example of the encoding process according to this embodiment. For example, the encoding device 100 shown in FIG. 6 performs the encoding process shown in FIG.

[0135] In this example, the compressor 101 first encodes one or more pieces of first information into a bitstream (S101) that can be used to derive a facial identity of a first person. The facial identity of the first person may include information about at least one of hair, glasses, facial hair, eyebrows, eyes, mouth, nose, skin, facial contours, clothing, and accessories.

[0136] It should be noted that it is sufficient to transmit one or more pieces of first information only once, and the same first information may be used to generate each image in the decoding device 200. The first information may be encoded using a video codec method such as VVC. The first information may also be transmitted as a bitstream separate from the bitstreams of the multiple pieces of second information.

[0137] An example of the one or more first information is one or more reference images from which the facial identity of the first person can be derived by the neural network. The one or more reference images can be one or more frames of a driving video, one or more pre-set images containing the face of the first person, or one or more avatars containing a face. The first person does not have to be a real person, but can be an object that imitates a person, such as a living creature, structure, or creation.

[0138] For example, the reference image may be an image of an area including the face of the first person. The area including the face of the first person may include at least a portion of the face and may also include the periphery of the face. Specifically, the area including the face of the first person may include the head, neck, shoulders, etc. of the first person, or may include a portion of the upper body of the first person, or may include a hat, accessories, clothes, etc. worn by the first person. Furthermore, the area including the face of the first person may include an object in the periphery of the face of the first person, such as a bird sitting on the first person's shoulder.

[0139] One or more variations of the first information are identity vectors derived from one or more reference images using a neural network, which may be encoded and transmitted directly to the decoding device 200 or stored in a database from which the decoding device 200 can retrieve the identity vectors.

[0140] 7 , the deriver 102 then derives (S102) a plurality of pieces of second information from a plurality of images of the driving video. For example, the plurality of images of the driving video include a face of a second person. The second person may be the same person as the first person or a different person. Each piece of second information represents a decipherable geometric attribute of the face of the second person in a corresponding frame or image of the driving video at a time instance.

[0141] For example, the second information is geometric information indicating geometric attributes of a region including the face of the second person. The region including the face of the second person includes at least a portion of the face and may also include the periphery of the face. Specifically, the region including the face of the second person may include the head, neck, shoulders, etc. of the second person, or may include a portion of the upper body of the second person, or may include a hat, accessories, clothes, etc. worn by the second person. Furthermore, the region including the face of the second person may include an object in the periphery of the second person's face, such as a bird sitting on the second person's shoulder.

[0142] The second information may be derived using a neural network. The deriver 102 may include a neural network. An example of the second information is geometric information indicating geometric attributes of the face of the second person, and an example of the geometric attributes of the face is a facial landmark.

[0143] 8 shows an example of facial landmarks derived from an image, where the facial landmarks correspond to the eyes, eyebrows, nose, mouth, lips, chin, and points on the facial contour.

[0144] 9 shows another example of facial landmarks derived from an image, where the facial landmarks correspond to the eyes, eyebrows, nose, mouth, lips, chin, and more points on the facial contour.

[0145] 10 shows yet another example of facial landmarks derived from an image, where the facial landmarks correspond to the eyes, eyebrows, nose, mouth, lips, chin, cheeks, and more points on the facial contour.

[0146] Facial landmarks are derived from images containing faces. If more than one face is present in an image, landmarks may be derived from the largest face. These landmarks indicate the location of key facial features, such as edges or transitions in shape, in areas including the facial contours, eyes, eyebrows, nose, mouth, lips, and chin. These landmarks allow the geometric attributes of the face to be deciphered, which in turn allows for easy modification of the attributes to convey desired emotions and expressions.

[0147] In the example of FIG. 7, next, the compressor 103 compresses and encodes the plurality of pieces of second information into a bit stream using a method such as entropy coding (S103).

[0148] At each time instance, the encoding device 100 transmits new second information representing geometric attributes of the face of the second person derived from a frame of the driving video. The second information may be transmitted in the bitstream as a supplemental enhancement information (SEI) message or other information unit. The compressor 103 may or may not be the same as the compressor 101. The bitstream is transmitted to the decoding device 200 via a transmission channel.

[0149] By using decipherable geometric attributes of the face, the information used to generate the images can be more easily understood, and these attributes can be easily modified to controllably generate poses and expressions.

[0150] In another variation, the encoding device 100 may include an additional generator, i.e., the encoding method may include an additional generating step, in which the generator generates a plurality of images from one or more first information and a plurality of second information via a neural network, the generated plurality of images including a generated face of the first person.

[0151] Here, the neural network takes as input one or more first pieces of information and a plurality of second pieces of information, and maps decipherable facial geometric attributes of the second person to the facial identity of the first person, so that the encoding device 100 can also present output frames that are identical to those generated by the decoding device 200.

[0152] In this disclosure, the encoding device 100 encodes two types of information into a bitstream: first information used to derive a facial identity of a first person, and a plurality of second information representing geometric attributes of a second person.

[0153] Some bitstream layout candidates are shown below. The bitstream layout may be any one of the multiple bitstream layout candidates, or may be a combination of two or more of the multiple bitstream layout candidates.

[0154] 11 is a diagram showing an example of a bitstream layout candidate, in which the bitstream includes one piece of first information and multiple pieces of second information for one group of pictures (GOP).

[0155] 12 is a diagram showing another example of a bitstream layout candidate, in which the bitstream includes one piece of first information and multiple pieces of second information in the first GOP, and multiple pieces of second information in each of the remaining GOPs.

[0156] 13 shows another example of a bitstream layout candidate, in which a first bitstream includes one or more first information items and a second bitstream includes multiple second information items. For example, the first bitstream may be stored on a recording medium, and the second bitstream may be transmitted over a communication network.

[0157] 14 is a diagram showing yet another example of a bitstream layout candidate, in which the bitstream includes a plurality of first information items and a plurality of second information items in one GOP.

[0158] For example, when information about a plurality of people is signaled from encoding apparatus 100 to decoding apparatus 200, a plurality of pieces of first information corresponding to the plurality of people may be signaled as the information about the plurality of people.

[0159] Furthermore, for example, a plurality of pieces of first information corresponding to a plurality of patterns corresponding to a first person, such as a plurality of types of avatars corresponding to one person, may be signaled from the encoding device 100 to the decoding device 200. Then, the decoding device 200 may be able to switch between the plurality of patterns.

[0160] 15 is a diagram showing an example of a bitstream layout candidate that complies with the VVC codec. In this example, the first information corresponds to a reference image and is coded as a VVC intra-predicted picture, and the second information corresponds to facial landmark data and is coded into SEI. Each access unit includes image data (e.g., VVC intra) or facial landmark data.

[0161] 16 is a diagram showing another example of a candidate bitstream layout conforming to VVC. In this example, the first information corresponds to a reference image and is coded as a VVC intra-predicted picture, and the second information corresponds to facial landmark data and is coded into SEI. Each access unit may include both image data (e.g., VVC intra) and facial landmark data, or only facial landmark data.

[0162] The VVC codec may be replaced by other video, image, or 3D codecs, such as HEVC, AVC, AV1, SVC, EVC, JPEG, V3C, etc. The facial landmark data may be encoded in the VSEI or SEI present in any video, image, or 3D codec, or other information units in the bitstream. Each access unit may contain other NAL units from video coding layers or non-video coding layers.

[0163] In another example, the facial landmark data may be data of a new nal_unit_type in the video coding layer. In another example, a slice NAL containing a reference image may be included in the first access unit, and each subsequent access unit may have a slice NAL indicating that the respective image is a monochromatic image, such as a black background or a green background.

[0164] Furthermore, for the slice NAL of the subsequent access unit, all coding units may be coded in skip mode or by simply copying the image of the first access unit. Alternatively, for the slice NAL of the subsequent access unit, an image consisting of any texture that can be coded with a small amount of code may be coded. In this case, a parameter indicating that the slice NAL data is to be ignored may be coded.

[0165] In the above examples, the first information corresponding to the reference image is included at the beginning of the GOP. However, the first information may be included in the middle of the GOP instead of at the beginning. For example, the first information corresponding to the reference image may be included at the beginning of the GOP, and further, first information that is the same as or different from the first information included at the beginning of the GOP may be included in the middle of the GOP (for example, in the middle).

[0166] In the above example, the first information is coded as an intra-predicted picture. However, the first information may be coded as an inter-predicted picture instead of an intra-predicted picture. For example, first information corresponding to a reference image at the beginning of a GOP may be coded as an intra-predicted picture, and further, first information that is the same as or different from the first information coded at the beginning of the GOP may be coded as an inter-predicted picture in the middle (e.g., the middle) of the GOP.

[0167] Furthermore, in accordance with a conventional video coding scheme, each subsequent access unit may encode an image (picture) including the face of the first person or the second person at that time instance. More specifically, each subsequent access unit may encode a low-quality image (low-quality picture) including the face of the first person or the second person at that time instance that can be coded with a small amount of code, thereby optionally allowing for the reconstruction of a low-quality video.

[0168] Note that for each driving frame, a set of landmarks derived from that driving frame is transmitted.

[0169] Furthermore, facial landmarks are an example of geometric attributes within a region including a face, and the geometric attributes within a region including a face are not limited to facial landmarks. For example, geometric attributes may be represented by a polygon model instead of a point cloud. In a polygon model, the shape of an object is represented by a combination of multiple polygons. Furthermore, geometric attributes may be represented by other geometric models. Furthermore, geometric attributes may be represented by the positions of facial features.

[0170] Furthermore, geometric attributes may be expressed in three dimensions or two dimensions. For example, facial landmarks may be expressed in three dimensions or two dimensions. Although the above-described landmarks are basically expressed in three dimensions, they can also serve the same purpose even if they are expressed in two dimensions.

[0171] [Decoding Configuration and Processing] The decoding configuration and processing are shown below. Note that the decoding configuration and processing shown below correspond to an example of the decoding configuration and processing for face reproduction, and the decoding configuration and processing are not necessarily limited to the following example. In particular, although geometric attributes are used in the following example, the above-mentioned latent information or the like may also be used.

[0172] 17 is a block diagram showing the configuration of a decoding device 200 according to this embodiment. The decoding device 200 generates an output video consisting of a plurality of images from a bitstream. In this example, the decoding device 200 includes a decompressor 201, a decompressor 202, a deriver 203, and a generator 204. The decompressor 201 and the decompressor 202 may be integrated. Also, the deriver 203 and the generator 204 may be integrated.

[0173] For example, decompressor 201 decompresses the reference image by decoding the reference image, decompressor 202 decompresses the geometric information by decoding the geometric information indicating geometric attributes, deriver 203 derives the facial identity of the first person, and generator 204 generates a plurality of images that constitute the output video.

[0174] A plurality of reference images may be used, a plurality of pieces of geometric information may be used, or a plurality of bitstreams may be used. The reference images are an example of first information in the present disclosure. The geometric information is an example of second information in the present disclosure. Note that the geometric information indicating the geometric attribute may also be simply expressed as the geometric attribute.

[0175] 18 is a flowchart showing an example of the decoding process according to this embodiment. For example, the decoding device 200 shown in FIG. 17 performs the decoding process shown in FIG.

[0176] In this example, first, a decompressor 201 decodes one or more pieces of first information from a bitstream (S201). Then, a deriver 203 derives a facial identity of a first person using the one or more pieces of first information. The facial identity of the first person may include information about at least one of hair, glasses, facial hair, eyebrows, eyes, mouth, nose, skin, facial contours, clothing, and accessories.

[0177] It is sufficient to receive one or more pieces of first information only once, and the same first information may be used to generate each image in the decoding device 200. The first information may be decoded using a video codec method such as VVC. The first information may also be received as a bitstream separate from the bitstreams of the multiple pieces of second information.

[0178] An example of the one or more first information is one or more reference images from which the facial identity of the first person can be derived by the neural network. The one or more reference images can be one or more frames of a driving video, one or more pre-set images containing the face of the first person, or one or more avatars containing a face. The first person does not have to be a real person, but can be an object that imitates a person, such as a living creature, structure, or creation.

[0179] For example, the reference image may be an image of an area including the face of the first person. The area including the face of the first person may include at least a portion of the face and may also include the periphery of the face. Specifically, the area including the face of the first person may include the head, neck, shoulders, etc. of the first person, or may include a portion of the upper body of the first person, or may include a hat, accessories, clothes, etc. worn by the first person. Furthermore, the area including the face of the first person may include an object in the periphery of the face of the first person, such as a bird sitting on the first person's shoulder.

[0180] The one or more variations of the first information are identity vectors derived by the encoding device 100 using a neural network. The identity vectors may be encoded by the encoding device 100 and received by the decoding device 200 from the encoding device 100, or may be stored in a database by the encoding device 100 and retrieved from the database by the decoding device 200 during user authentication.

[0181] In the example of Figure 18, the decompressor 202 then decodes the plurality of pieces of second information from the bitstream into a decodable geometric space (S202). Each piece of decoded second information represents decodable geometric attributes of the face of a second person derived from a corresponding frame of the driving video at a corresponding time instance by the encoding device 100. Such attributes are easily understandable to humans. The second person may be the same person as the first person described above or a different person.

[0182] For example, the second information is geometric information indicating geometric attributes of a region including the face of the second person. The region including the face of the second person includes at least a portion of the face and may also include the periphery of the face. Specifically, the region including the face of the second person may include the head, neck, shoulders, etc. of the second person, or may include a portion of the upper body of the second person, or may include a hat, accessories, clothes, etc. worn by the second person. Furthermore, the region including the face of the second person may include an object in the periphery of the second person's face, such as a bird sitting on the second person's shoulder.

[0183] At each time instance, the decoding device 200 receives new second information representing geometric attributes derived from a frame of the driving video, which is used by the neural network to generate the image, and which may be transmitted in the bitstream as a supplemental enhancement information (SEI) message or other information unit.

[0184] An example of the plurality of pieces of second information is geometric information indicating geometric attributes, such as facial landmarks derived from one or more images including the face of the second person.

[0185] 8 is a diagram showing an example of facial landmarks decoded from the second information, in which the facial landmarks correspond to the eyes, eyebrows, nose, mouth, lips, chin, and a plurality of points on the facial contour.

[0186] 9 is a diagram showing another example of facial landmarks decoded from the second information, in which multiple facial landmarks correspond to the eyes, eyebrows, nose, mouth, lips, chin, and more points on the facial contour.

[0187] 10 is a diagram showing yet another example of facial landmarks decoded from the second information, in which the facial landmarks correspond to the eyes, eyebrows, nose, mouth, lips, chin, cheeks, and more points on the facial contour.

[0188] Facial landmarks are decoded from the second information. These landmarks indicate the locations of key facial features, such as edges or transitions in shapes in areas including the facial contours, eyes, eyebrows, nose, mouth, lips, and chin. These landmarks allow the geometric attributes of the face to be decoded, which can then be easily modified to convey desired emotions and expressions.

[0189] 18 , the generator 204 then generates a plurality of images based on the one or more pieces of decoded first information, the plurality of pieces of decoded second information, and a neural network (S203). Specifically, the generator 204 generates a plurality of images from face identities corresponding to the one or more pieces of decoded first information and the plurality of pieces of decoded second information via a neural network. The generated plurality of images includes the face of the first person.

[0190] Here, the neural network takes as input a facial identity derived from one or more decoded first pieces of information and a plurality of decoded second pieces of information, and maps decipherable geometric attributes of the second person's face to the first person's facial identity.

[0191] By using interpretable geometric attributes, the information used to generate the images can be more easily understood, and these attributes can be easily modified to controllably generate poses and expressions.

[0192] One or more first pieces of information may be used in place of or in addition to the facial identity in generating the plurality of images.

[0193] 19 is a conceptual diagram illustrating a decoding process at each time instance. In this example, the decoding device 200 receives compressed first information and compressed second information at the first time instance (t=0). The decoding device 200 then stores the compressed first information in its memory 252. The decoding device 200 also performs a decoding process on the compressed first information and the compressed second information of the first time instance (t=0), inputs the first information and the second information to a neural network, and generates an image of the first time instance (t=0).

[0194] At a subsequent time instance (t=T), the decoding device 200 receives the compressed second information and searches for and obtains the compressed first information from the memory 252 of the decoding device 200. The decoding device 200 then performs a decoding process on the compressed first information and the compressed second information for that time instance (t=T), inputs the first information and the second information into a neural network, and generates an image for that time instance (t=T).

[0195] The decoding device 200 may store the first information obtained by performing a decoding process on the compressed first information in the memory 252 of the decoding device 200. Then, at a subsequent time instance (t=T), the decoding device 200 may obtain the first information to which the decoding process has been applied from the memory 252, and input the first information obtained from the memory 252 to the neural network.

[0196] Figure 20 shows examples of various neural networks that generate multiple images. Specifically, Figure 20 shows a generative adversarial network, a variational autoencoder, a flow-based generative model, and a diffusion model.

[0197] A generative adversarial network (GAN) is a generative model that generates new data instances similar to input data by learning the features of the input data. Specifically, the unsupervised task in the generative model is converted into a supervised task by two submodels.

[0198] For example, the generator sub-model generates fake samples, the classifier sub-model discriminates between real inputs and fake samples generated by the generator sub-model, and an output image is generated through a minimax game that maximizes the classification probability of the classifier sub-model in assigning correct labels to the real inputs and fake samples while minimizing the distribution difference between the real inputs and fake samples.

[0199] In a variational autoencoder, the input data is first compressed into a multivariate latent distribution that reconstructs the data from the latent space as accurately as possible, resulting in efficient data compression and dimensionality reduction. In a flow-based generative model, the source distribution is transformed into the training data distribution through a sequence of one or more reversible transformations, which allows learning of the data distribution and accurate computation of the end-goal likelihood.

[0200] Diffusion models are also generative models. They generate new data instances that resemble the training data. First, they degrade the structure of the training data through repeated injection of perturbations and noise before initiating a denoising process in an attempt to recover the original data. As a result, the data is iteratively mapped to a latent distribution via a Markov chain, where the latent state at each step depends only on the latent state at the previous step. The data is then recovered through hierarchical denoising.

[0201] Alternatives to the neural networks described above may be any combination of the neural networks described above, other types of generative models, etc.

[0202] In this disclosure, the decoding device 200 decodes two types of information from the bitstream: first information used to derive a facial identity of a first person, and a plurality of second information representing geometric attributes of a second person.

[0203] Some bitstream layout candidates are shown below. The bitstream layout may be any one of the multiple bitstream layout candidates, or may be a combination of two or more of the multiple bitstream layout candidates.

[0204] 11 is a diagram showing an example of a bitstream layout candidate, in which the bitstream includes one piece of first information and multiple pieces of second information for one group of pictures (GOP).

[0205] 12 is a diagram showing another example of a bitstream layout candidate, in which the bitstream includes one piece of first information and multiple pieces of second information in the first GOP, and multiple pieces of second information in each of the remaining GOPs.

[0206] 13 shows another example of a bitstream layout candidate, in which a first bitstream includes one or more first information items and a second bitstream includes multiple second information items. For example, the first bitstream may be stored on a recording medium, and the second bitstream may be transmitted over a communication network.

[0207] 14 is a diagram showing yet another example of a bitstream layout candidate, in which the bitstream includes a plurality of first information items and a plurality of second information items in one GOP.

[0208] For example, when information about a plurality of people is signaled from encoding apparatus 100 to decoding apparatus 200, a plurality of pieces of first information corresponding to the plurality of people may be signaled as the information about the plurality of people.

[0209] Furthermore, for example, a plurality of pieces of first information corresponding to a plurality of patterns corresponding to a first person, such as a plurality of types of avatars corresponding to one person, may be signaled from the encoding device 100 to the decoding device 200. Then, the decoding device 200 may be able to switch between the plurality of patterns.

[0210] 15 shows an example of a bitstream layout candidate that complies with the VVC codec. In this example, the first information corresponds to a reference image and is decoded as a VVC intra-predicted picture, and the second information corresponds to facial landmark data and is decoded from SEI. Each access unit includes image data (e.g., VVC intra) or facial landmark data.

[0211] 16 is a diagram showing another example of a candidate bitstream layout conforming to VVC. In this example, the first information corresponds to a reference image and is decoded as a VVC intra-predicted picture, and the second information corresponds to facial landmark data and is decoded from SEI. Each access unit may include both image data (e.g., VVC intra) and facial landmark data, or only facial landmark data.

[0212] The VVC codec may be replaced by other video, image, or 3D codecs, such as HEVC, AVC, AV1, SVC, EVC, JPEG, V3C, etc. The facial landmark data may be decoded from the VSEI or SEI present in some video, image, or 3D codec, or other information units in the bitstream. Each access unit may contain other NAL units from video coding layers or non-video coding layers.

[0213] In another example, the facial landmark data may be data of a new nal_unit_type in the video coding layer. In another example, a slice NAL containing a reference image may be included in the first access unit, and each subsequent access unit may have a slice NAL indicating that the respective image is a monochromatic image, such as a black background or a green background.

[0214] Furthermore, for the slice NAL of the subsequent access unit, all coding units may be decoded in skip mode or by simply copying the image of the first access unit. Alternatively, for the slice NAL of the subsequent access unit, an image consisting of any texture that can be coded with a small amount of code may be decoded. In this case, a parameter indicating that the slice NAL data is to be ignored may be decoded.

[0215] In the above examples, the first information corresponding to the reference image is included at the beginning of the GOP. However, the first information may be included in the middle of the GOP instead of at the beginning. For example, the first information corresponding to the reference image may be included at the beginning of the GOP, and further, first information that is the same as or different from the first information included at the beginning of the GOP may be included in the middle of the GOP (for example, in the middle).

[0216] In the above example, the first information is decoded as an intra-predicted picture. However, the first information may be decoded as an inter-predicted picture instead of an intra-predicted picture. For example, first information corresponding to a reference image at the beginning of a GOP may be decoded as an intra-predicted picture, and further, first information that is the same as or different from the first information decoded at the beginning of the GOP may be decoded as an inter-predicted picture in the middle of the GOP (for example, in the middle).

[0217] Furthermore, in accordance with a conventional video decoding scheme, each subsequent access unit may decode an image (picture) including the face of the first person or the second person at that time instance. More specifically, each subsequent access unit may decode a low-quality image (low-quality picture) that includes the face of the first person or the second person at that time instance and can be coded with a small amount of code. This may optionally allow for reconstruction of low-quality video.

[0218] For each frame, the landmark set is decoded from the second information corresponding to that frame.

[0219] Furthermore, facial landmarks are an example of geometric attributes within a region including a face, and the geometric attributes within a region including a face are not limited to facial landmarks. For example, geometric attributes may be represented by a polygon model instead of a point cloud. In a polygon model, the shape of an object is represented by a combination of multiple polygons. Furthermore, geometric attributes may be represented by other geometric models. Furthermore, geometric attributes may be represented by the positions of facial features.

[0220] Furthermore, geometric attributes may be expressed in three dimensions or two dimensions. For example, facial landmarks may be expressed in three dimensions or two dimensions. Although the above-described landmarks are basically expressed in three dimensions, they can also serve the same purpose even if they are expressed in two dimensions.

[0221] Spoofing Next, methods and information for determining whether a person in a presented image (i.e., output image) matches the sender are described.

[0222] 21 is a conceptual diagram showing an example of operation when a person in a driving video matches a person in a reference image. For example, if the reference image and the driving video include person A, person A is presented in the decoded output video based on the reference image.

[0223] 22 is a conceptual diagram showing an example of operation when a person in a driving video is different from a person in a reference image. For example, if the reference image includes person A and the driving video includes person B, person A will still be presented in the output video.

[0224] In this way, person A is presented in both cases. Therefore, it is difficult for a viewer of the output video to determine whether the sender of the driving video conversation is person A or person B. In other words, person B can impersonate person A. In other words, for example, it is possible to tamper with person B's video into person A's video, thereby tampering with the sender.

[0225] As a countermeasure against the above-mentioned spoofing, in the following first, second and third aspects, the transmitting side generates information indicating whether the sender name has been tampered with and notifies the receiving side of the information. In addition, in the fourth and fifth aspects, the transmitting side generates, encodes and transmits authentication information, and the receiving side determines whether the sender name has been tampered with based on the authentication information.

[0226] Here, for example, the transmitting side includes the encoding device 100 and corresponds to one or more devices for generating a bitstream based on a reference image and a driving video and transmitting the bitstream, and the receiving side includes the decoding device 200 and corresponds to one or more devices for receiving the bitstream and generating an output video based on the bitstream.

[0227] The driving video may be referred to as an input video. The input video has one or more images corresponding to one or more frames. Feature data is extracted from the images in the input video. The feature data may correspond to latent information or geometric information.

[0228] For example, the deriving unit 102 identifies characteristic locations of a person's face from images in the input video and extracts feature data corresponding to the feature amounts and / or movement information of the characteristic locations. Then, the compressor 103 encodes (compresses) the feature data to generate encoded feature data.

[0229] A video (e.g., an output video) generated by transforming an image (e.g., a reference image) based on feature data may be referred to as a synthesized video.

[0230] 23 is a block diagram showing an example of the configuration of a video call system according to the first embodiment. The video call system includes an encoding device 100 and a decoding device 200. In this embodiment, the video call system also includes a reference image determiner 111. The reference image determiner 111 may be included in the encoding device 100.

[0231] First, on the transmitting side, a reference image and an input video are input to a reference image determiner 111. The reference image determiner 111 compares the reference image with the input video, determines whether or not a person included in the reference image matches a person included in the input video, and outputs the determination result.

[0232] The determination result is input to the encoding device 100. The encoding device 100 encodes the determination result as metadata in a header area such as an SEI, thereby multiplexing the determination result into a bitstream. The encoding device 100 also multiplexes the encoded data of the reference image and the encoded data of feature data extracted from the input video into the bitstream.

[0233] On the receiving side, the decoding device 200 decodes the determination result from the metadata and determines, based on the decoded determination result, whether the person included in the reference image matches the person included in the input video. This allows the decoding device 200 to determine whether the person presented in the output video matches the sender captured in the input video. Another application device may determine, based on the decoded determination result, whether the person in the reference image matches the person in the input video, i.e., whether the person in the output video matches the sender.

[0234] Furthermore, the decoding device 200 may perform processing based on the determination result, or may output the determination result together with the output video, and another application device may then perform processing based on the determination result.

[0235] Furthermore, for example, the decoding device 200 may determine to decode the video when the determination result is "match," and not to decode the video (specifically, the reference image and the feature data) when the determination result is "mismatch." Also, a flag, configuration, or process may be used to specify whether to enable a function for controlling whether to decode the video based on the determination result.

[0236] Furthermore, the decoding device 200 may notify the viewer whether the person displayed in the output video is the sender by presenting the determination result using text, an icon, or the like. Alternatively, the decoding device 200 may superimpose the determination result on the output video and present it simultaneously with the output video. Alternatively, another application device may perform presentation processing based on the determination result.

[0237] The reference image determiner 111 may determine whether or not the person shown in the output video is likely to be other than the sender based on a comparison between the reference image and the input video, and signal the determination result.

[0238] With the above configuration, even if the reference image contains a person other than the sender, the receiver or viewer can determine whether the person in the reference image matches the person in the input video, thereby making it possible to detect impersonation and the like.

[0239] Note that, before the processing of the reference image determiner 111, a value indicating whether the reference image is an image of the input video may be set in the parameter. If a value indicating that the reference image matches the image of the input video is set in the parameter, the reference image determiner 111 may skip the determination processing. On the other hand, if a value indicating that the reference image does not match the image of the input video is set in the parameter, the reference image determiner 111 may execute the determination processing and overwrite the parameter with the determination result.

[0240] This allows the determination process to be skipped if the reference image matches the image of the input video, thereby reducing the amount of processing required for the determination process.

[0241] 24 is a block diagram showing an example of the configuration of a video call system according to the second embodiment. The video call system includes an encoding device 100 and a decoding device 200. In this embodiment, the video call system also includes a reference image selector 112. The reference image selector 112 may be included in the encoding device 100.

[0242] First, at the transmitting side, at least one of an alternative image and an input video is input to a reference image selector 112. The reference image selector 112 selects and outputs a reference image from at least one of the alternative image and the input video. The reference image selector 112 may select a reference image based on a specification by a user (sender) or may select a reference image based on preset parameters.

[0243] The reference image selector 112 also outputs a reference image input mode indicating image input or video input as a selection result. For example, if a reference image corresponding to an alternative image is selected, specifically, if the alternative image is selected as the reference image, the reference image input mode indicates image input. For example, if a reference image corresponding to an input video is selected, specifically, if an image of the input video is selected as the reference image, the reference image input mode indicates video input.

[0244] The selection result is input to the encoding device 100. The encoding device 100 encodes the selection result as metadata in a header area such as an SEI, thereby multiplexing the selection result into a bitstream. The encoding device 100 also multiplexes the encoded data of the reference image and the encoded data of feature data extracted from the input video into the bitstream.

[0245] The input video corresponds to the driving video described above. In other words, the input video is used to extract feature data. That is, for example, if an image from the input video is selected as a reference image, the person included in the reference image will be identical to the person who is the source of the feature data, making it less likely to be impersonated.

[0246] An alternative image refers to an image that is different from the image of the input video. The alternative image may be an image that is prepared in advance before the input video is acquired. Alternatively, one or more images, i.e., videos, may be used as the alternative image. For example, if an alternative image is selected as the reference image, the person included in the reference image will be different from the person in the source of the feature data, which may lead to the possibility of impersonation. In other words, in this case, the possibility of impersonation is not necessarily low.

[0247] On the receiving side, the decoding device 200 obtains the selection result by decoding the selection result from the metadata. The decoding device 200 can determine whether the input is an image or a video based on the selection result. The decoding device 200 may also output the selection result to an application device. Alternatively, the decoding device 200 may output the determination result to the application device.

[0248] This allows the receiving side to determine whether the reference image is an alternative image or an image from the input video. If the reference image is an alternative image, it can be determined that the person in the reference image is likely to be different from the sender corresponding to the person in the input video, i.e., there is a possibility of impersonation. If the reference image is an image from the input video, it can be determined that the person in the reference image is likely to match the sender corresponding to the person in the input video, i.e., there is a low possibility of impersonation.

[0249] The decoding device 200 may switch between executing and skipping the decoding process based on the selection result or the determination result, or may present the selection result or the determination result to the viewer using text, an icon, or the like. Alternatively, the decoding device 200 may superimpose the selection result or the determination result on the output video and present it simultaneously with the output video. Another application device may perform a presentation process based on the selection result or the determination result.

[0250] 25 is a flowchart showing an example of the operation of the transmitting side in the second aspect. For example, the reference image selector 112 selects an image of the input video or an alternative image as a reference image, thereby selecting a video input corresponding to the image of the input video or an image input corresponding to the alternative image.

[0251] If video input is selected (Yes in S121), the encoding device 100 sets the reference image input mode to video input (S122). If image input is selected (No in S121), the encoding device 100 sets the reference image input mode to image input (S123). The encoding device 100 then multiplexes the reference image input mode as metadata into the bitstream (S124).

[0252] 26 is a flowchart showing an example of the operation of the receiving side in the second aspect. For example, the decoding device 200 acquires metadata from a bitstream (S221). Then, the decoding device 200 analyzes the metadata to acquire a reference image input mode.

[0253] If the reference image input mode indicates video input (Yes in S222), the decoding device 200 determines that there is a high possibility that the person in the reference image matches the person in the input video, that is, that there is a low possibility of tampering (S223). Here, if the decoding device 200 can determine that there is no possibility of tampering, it may determine that there is no possibility of tampering.

[0254] On the other hand, if the reference image input mode indicates image input (No in S222), the decoding device 200 determines that the person in the reference image may not match the person in the video input, i.e., that there is a possibility of tampering (S224).

[0255] Then, the decoding device 200 presents the determination result (S225). For example, the decoding device 200 may output the determination result together with the output video.

[0256] 27 is a syntax diagram showing an example of syntax in the second aspect. Specifically, the example shows a syntax for transmitting a reference image input mode or the like as metadata in a header area of ​​an SEI or the like.

[0257] For example, the number of reference images is signaled. Then, a reference image input mode is transmitted for each reference image. Correspondence identification information is also signaled. The correspondence identification information indicates the range in which this metadata is valid. For example, the correspondence identification information indicates a sequence number, frame number, or divided data number corresponding to this metadata. The correspondence identification information may also indicate whether the content of the reference image input mode, etc. in this metadata has been updated. And, if the content has been updated, the updated content may be indicated in this metadata.

[0258] The reference image input mode indicates image input or video input. For example, if a reference image corresponding to an alternative image is selected, the reference image input mode indicates image input. Also, for example, if a reference image corresponding to an input video is selected, the reference image input mode indicates video input.

[0259] 28 is a block diagram showing an example configuration of a video call system according to a third embodiment. The video call system includes an encoding device 100 and a decoding device 200. In this embodiment, the video call system also includes both a reference image determiner 111 and a reference image selector 112. The reference image determiner 111 and the reference image selector 112 may be included in the encoding device 100.

[0260] First, at the transmitting side, at least one of an alternative image and an input video is input to a reference image selector 112. The reference image selector 112 selects and outputs a reference image from at least one of the alternative image and the input video. The reference image selector 112 may select a reference image based on a specification by a user (sender) or may select a reference image based on preset parameters.

[0261] The reference image selector 112 also outputs a reference image input mode indicating image input or video input as a selection result. For example, if a reference image corresponding to an alternative image is selected, specifically, if the alternative image is selected as the reference image, the reference image input mode indicates image input. For example, if a reference image corresponding to an input video is selected, specifically, if an image of the input video is selected as the reference image, the reference image input mode indicates video input.

[0262] The selection result is input to the reference image determiner 111. Based on the selection result, the reference image determiner 111 determines whether or not to perform reference image determination. For example, in the case of video input, the reference image determiner 111 determines not to perform reference image determination. On the other hand, in the case of image input, the reference image determiner 111 determines to perform reference image determination. In this case, the reference image determiner 111 performs reference determination processing and outputs the determination result.

[0263] The selection results and the determination results are input to the encoding device 100. The encoding device 100 encodes the selection results and the determination results as metadata in a header area such as an SEI, thereby multiplexing the selection results and the determination results into a bitstream. The encoding device 100 also multiplexes the coded data of the reference image and the coded data of feature data extracted from the input video into the bitstream.

[0264] In the case of video input, the determination result does not need to be coded or multiplexed into the bitstream.

[0265] On the receiving side, the decoding device 200 decodes the selection result from the metadata and determines whether the input is a video or an image based on the selection result. In the case of a video input, the decoding device 200 determines that the person in the reference image is likely to match the person in the input video. In the case of an image input, the decoding device 200 decodes the determination result from the metadata and determines, based on the determination result, whether the person in the reference image matches the sender corresponding to the person in the input video.

[0266] With the above configuration, the reference image determiner 111 can determine whether or not to perform reference image determination based on the selection result. This makes it possible to omit the reference image determination in the case of video input, thereby reducing the amount of processing.

[0267] The reference image input mode does not have to be included in the bitstream as a selection result. In the case of video input, the reference image determination result indicating "match" may be included in the bitstream as a determination result.

[0268] 29 is a flowchart showing an example of the operation of the transmitting side in the third aspect. For example, the reference image selector 112 selects a video input (i.e., an image of the input video) or an image input (a substitute image) as the reference image.

[0269] In the case of image input (No in S131), the reference image determiner 111 performs reference image determination (S132). That is, in this case, the reference image determiner 111 determines whether or not the person in the reference image matches the person in the input video.

[0270] If the person in the reference image matches the person in the input video (Yes in S133), the encoding device 100 sets the reference image input mode to image input and sets the reference image determination result to match (S134). On the other hand, if the person in the reference image does not match the person in the input video (No in S133), the encoding device 100 sets the reference image input mode to image input and sets the reference image determination result to mismatch (S135).

[0271] In the case of video input (Yes in S131), the reference image determiner 111 does not perform reference image determination (S136), and the encoding device 100 sets the reference image input mode to video input (S137).

[0272] Then, the encoding device 100 multiplexes the reference image input mode and the reference image determination result as metadata into the bitstream (S138). Note that in the case of video input, the reference image determination result does not need to be multiplexed into the bitstream.

[0273] 30 is a flowchart showing an example of the operation of the receiving side in the third aspect. For example, the decoding device 200 acquires metadata from the bitstream (S231). Then, the decoding device 200 analyzes the metadata to acquire a reference image input mode and, in the case of image input, further acquires a reference image determination result.

[0274] Furthermore, decoding device 200 determines the likelihood that a person in the reference image matches a person in the input video based on the reference image input mode, or based on the reference image input mode and the reference image determination result (S232). Specifically, if the reference image input mode indicates video input, decoding device 200 determines that there is a high likelihood that a person in the reference image matches a person in the input video. If the reference image input mode indicates image input, decoding device 200 determines the likelihood that a person in the reference image matches a person in the video input based on the reference image determination result.

[0275] Then, the decoding device 200 presents the determination result (S233). For example, the decoding device 200 may output the determination result together with the output video.

[0276] Fig. 31 is a syntax diagram showing an example of syntax in the third aspect. Specifically, an example of syntax is shown for transmitting a reference image input mode, a reference image determination result, and the like as metadata in a header area such as SEI. In the example of Fig. 31, compared to the example of Fig. 27, when the reference image input mode indicates image input as the selection result, the reference image determination result is signaled. In this case, the person type of the reference image and authentication information may also be signaled.

[0277] The reference image determination result indicates whether the person in the reference image matches the person in the video input. The person type in the reference image indicates the attributes of the person projected onto the reference image. Specifically, the person type in the reference image may indicate whether the person in the reference image is an actual human, an animal, or an avatar. The person type in the reference image may also indicate gender, age, nationality, or the like. The authentication information may indicate, for example, authenticated identification information, an email address, or a contact ID.

[0278] The person type or authentication information of the reference image may also be used in determining the likelihood that the person in the reference image matches the person (i.e., the sender) in the input video.

[0279] [Supplementary Notes on the First, Second, and Third Aspects] The processing of the reference image determiner 111 and the reference image selector 112 is not limited to the examples shown in the first, second, and third aspects.

[0280] The determination or selection may be made by the user, without relying on the reference image determiner 111 and the reference image selector 112. The result of the determination or selection made by the user may be indicated in the metadata. In this case, the metadata may indicate whether the determination was made by the user or the reference image determiner 111, or whether the selection was made by the user or the reference image selector 112.

[0281] By not using the reference image determiner 111 or the reference image selector 112, it may be possible to reduce the amount of processing in the video call system.

[0282] The reference image determiner 111 may use AI or the like to perform image analysis to determine the possibility of a person matching.

[0283] Instead of a person, a possibility that an object in the reference image matches an object in the input video may be determined for an object other than a person, such as an animal, other static or dynamic object, CG-generated data, or image data generated by generative AI.

[0284] The reference image determiner 111 may also extract any one image from the input video and compare the reference image with the arbitrary one image extracted from the input video. Alternatively, the reference image determiner 111 may sequentially extract multiple images at predetermined intervals and sequentially compare the multiple images with the reference image.

[0285] The predetermined timing may be, for example, the timing of the first frame of a randomly accessible unit in image coding (a randomly accessible I-frame). In other words, the predetermined timing may correspond to the start of the randomly accessible unit.

[0286] Alternatively, the predetermined timing may correspond to a specified interval. For example, it may be specified that an image is extracted from the input video once every 0.5 seconds. The interval may also be variable. Randomly varying the extraction timing may provide more secure operation.

[0287] As described above, an image may be extracted from the input video at a predetermined timing, and a reference image determination may be performed at the timing when the image is extracted.

[0288] In addition, the above example shows a video conference, so it is assumed that the input video is real-time data, and that if an image of the input video is selected as the reference image, there is a low possibility of tampering.

[0289] However, the substitute image and input video may be pre-acquired data or real-time data, and pre-acquired data may be more susceptible to tampering and therefore may be determined to be susceptible to tampering, while real-time data may be less susceptible to tampering and therefore may be determined to be less susceptible to tampering.

[0290] Here, real-time data is data obtained by directly capturing a person using a camera during encoding. For example, the encoding device 100 may detect whether the camera is directly capturing a person. If the camera is directly capturing a person, the encoding device 100 may determine that the substitute image or input video obtained from the camera is real-time data. On the other hand, if the camera is not directly capturing a person, the encoding device 100 may determine that the substitute image or input video obtained from the camera is not real-time data.

[0291] Additionally, it is assumed that the input video may be real-time data. If the reference image is real-time data, it may be determined that there is a low possibility of tampering. That is, in this case, it may be determined that the person in the reference image matches the person in the input video. On the other hand, if the reference image is data acquired in advance, it may be determined that there is a possibility of tampering. That is, in this case, it may be determined that there is a possibility that the person in the reference image is different from the person in the input video.

[0292] Here, the pre-acquired data may be data acquired before the acquisition of the input video.

[0293] For example, the reference image determiner 111 determines whether the person presented in the output video matches the sender captured in the input video by determining whether the input is an image or a video, or by other methods.

[0294] The determination of image input or video input refers to a determination of whether the reference image is an image included in the input video or whether the reference image is an image different from the image included in the input video. Multiple images may be used as reference images. This may increase the accuracy of determining whether a person in the reference image matches a person in the input video, and may increase the reliability of determining whether there is a possibility of tampering.

[0295] Furthermore, if the number of reference images is greater than 1, it is assumed that the plurality of reference images constitute a video. Therefore, the reference image determiner 111 may determine whether the number of reference images is greater than 1.

[0296] Then, when the number of reference images is greater than 1, the reference image determiner 111 may determine that there is a high possibility that the person in the reference image matches the person in the input video, that is, there is a low possibility of tampering. On the other hand, when the number of reference images is 1, the reference image determiner 111 may determine that there is a high possibility that the person in the reference image is different from the person in the input video, that is, there is a possibility of tampering.

[0297] In the above determination, even when multiple images extracted from a video different from the input video are used as multiple reference images, it may be determined that the person in the reference image is likely to match the person in the input video, i.e., that the possibility of tampering is low. However, the determination process is simplified.

[0298] [Fourth Aspect] Fig. 32 is a block diagram showing an example configuration of a video call system according to the fourth aspect. The video call system includes an encoding device 100 and a decoding device 200. In this aspect, the video call system also includes an authentication image generator 113. The authentication image generator 113 may be included in the encoding device 100. The decoding device 200 also includes a decoding processor 221, a reference image determiner 222, and a video generator 223. In particular, in this aspect, an authentication image is transmitted, and reference image determination is performed on the receiving side.

[0299] First, on the transmitting side, an input video is input to authentication image generator 113. Authentication image generator 113 generates an authentication image by extracting an image from the input video as the authentication image. Encoding device 100 encodes the authentication image and transmits the bitstream containing the encoded image. Encoding device 100 may use an image encoding method or a moving image encoding method to encode the authentication image. Encoding device 100 may encode the authentication image as metadata in a header area such as SEI.

[0300] The authentication image generator 113 may extract any one image from the input video as the authentication image, or may extract each of a plurality of images as the authentication image at a predetermined timing.

[0301] The predetermined timing may be, for example, the timing of the first frame of a randomly accessible unit in image coding (a randomly accessible I-frame). In other words, the predetermined timing may correspond to the start of the randomly accessible unit.

[0302] Alternatively, the predetermined timing may correspond to a specified interval, e.g., extracting an image from the input video once every 0.5 seconds. The interval may also be variable, e.g., randomly varying the extraction timing may provide more secure operation.

[0303] The encoding device 100 may also divide and transmit the authentication image. For example, the encoding device 100 may divide the authentication image corresponding to one frame into multiple tiles or multiple slices and transmit the authentication image data multiple times. Then, the reference image determination may be performed once the authentication image corresponding to one frame is completed on the receiving side. This enables encoding, transmission, and decoding in units of tiles or slices. Therefore, it may be possible to start decoding earlier and reduce processing delays.

[0304] It should be noted that in addition to images extracted from the input video, images extracted from one or more reference images may also be transmitted as authentication images.

[0305] On the receiving side, a decoding processor 221 decodes the reference image and feature data as well as the authentication image. The authentication image and the reference image are input to a reference image determiner 222. The reference image determiner 222 determines whether the person included in the reference image matches the person included in the authentication image.

[0306] The decoding device 200 may switch between performing and skipping the decoding process for the feature data, etc., based on the determination result. The decoding device 200 may present the determination result to the viewer using text, an icon, etc. The decoding device 200 may superimpose the determination result on the output video and present it simultaneously with the output video. The decoding device 200 may output a reference image or an authentication image and present it to the viewer. Another application device may perform the presentation process based on the determination result.

[0307] As described above, an authentication image is generated on the transmitting side and transmitted. This allows the receiving side to perform reference image judgment. Also, the transmitting side can reduce the judgment process.

[0308] Furthermore, as described above, an authentication image may be presented. This allows a viewer of the output video to visually compare the output video with the authentication image and determine whether the people match. In other words, it is possible to provide the viewer with information for determining whether the person in the reference image matches the person in the input video (i.e., the sender). In this case, the determination process does not need to be performed in the decoding device 200, and the decoding device 200 does not need to include the reference image determiner 222.

[0309] Alternatively, for example, the receiving side may determine whether or not reference image determination is necessary, and may then perform the determination process if necessary, thereby reducing the amount of processing on the receiving side.

[0310] [Fifth Aspect] Fig. 33 is a block diagram showing an example configuration of a video calling system according to the fifth aspect. The video calling system includes an encoding device 100 and a decoding device 200. In this aspect, the video calling system also includes an authentication image generator 113. The authentication image generator 113 may be included in the encoding device 100. The decoding device 200 also includes a decoding processor 221 and a video generator 223.

[0311] The configuration of the video call system in this embodiment is generally the same as that of the video call system in the fourth embodiment, except that the reference image determiner 222 is not included in the video call system. In this embodiment, the receiving side is presented with both information (specifically, an image or video) generated based on the reference image and information (specifically, an image or video) generated based on the authentication image.

[0312] Specifically, on the receiving side, the authentication image, the reference image, and the feature data decoded by the decoding processor 221 are input to the video generator 223 .

[0313] The video generator 223 switches the output by generating and outputting a video from the reference image and feature data at certain times and generating and outputting a video from the authentication image and feature data at other times. This allows a viewer of the output video to compare the video based on the reference image with the video based on the authentication image and determine whether the people match. In other words, it is possible to determine whether the person in the reference image matches the sender corresponding to the person in the input video.

[0314] In the above example, the reference image information and the authentication image information are displayed alternately. However, the reference image information and the authentication image information may be displayed simultaneously. For example, the reference image information may be displayed at a certain time, and the reference image information and the authentication image information may be displayed simultaneously at another time. Furthermore, the reference image information and the authentication image information may be displayed superimposed on each other.

[0315] [Sixth Aspect] Fig. 34 is a block diagram showing an example configuration of a video call system according to the sixth aspect. The video call system includes an encoding device 100 and a decoding device 200. In this aspect, the video call system further includes another encoding device 300. This aspect may include the video call system described in at least one of the first to fifth aspects. The encoding device 100 and the decoding device 200 correspond to the encoding device 100 and the decoding device 200 described in at least one of the first to fifth aspects.

[0316] The encoding device 300 encodes and transmits normal video (real video). The encoding device 100 generates data for a composite video, encodes the data, and transmits the data. Here, the composite video corresponds to an output video generated by, for example, changing a reference image based on feature data of an input video. That is, the encoding device 100 encodes and transmits the reference image and feature data. The decoding device 200 receives the encoded data from the encoding device 300 and the encoded data from the encoding device 100 and decodes the data.

[0317] The multiplexing and demultiplexing of the bitstream may be performed according to a predetermined format or multiplexing scheme.

[0318] The encoding device 300 and the encoding device 100 may be the same device. That is, the same device may operate as both the encoding device 300 and the encoding device 100. The encoding device 300 and the encoding device 100 may be different devices. Furthermore, each of one or more devices may operate as at least one of the encoding device 300 and the encoding device 100.

[0319] The encoding device 300 transmits a bitstream containing a parameter indicating whether the encoded data is real video data or not. Here, the bitstream may conform to a multiplexing format.

[0320] The encoding device 100 transmits a bitstream containing a parameter indicating whether the encoded data is composite video data or not. Here, the bitstream may conform to a multiplexing format.

[0321] For example, the encoding device 300 includes a video encoder 314 and a multiplexer 315. The video encoder 314 encodes the input video. The multiplexer 315 multiplexes the encoded data, parameters, etc. into a bitstream.

[0322] Furthermore, for example, the encoding device 100 includes a composite video encoder 114 and a multiplexer 115. The composite video encoder 114 encodes the reference image, feature data of the input video, and the person type. The multiplexer 115 multiplexes the encoded data, parameters, etc. into a bitstream.

[0323] Also, for example, the decoding device 200 includes a demultiplexer 224, a decoding processor 221, and a presenter 225. The demultiplexer 224 performs a demultiplexing process to obtain coded data, parameters, etc. from the bitstream. The decoding processor 221 decodes the data to reconstruct video. The presenter 225 presents the video.

[0324] The presenter 225 may be a display device that presents the video by displaying the video. The presenter 225 may be a separate device that is not included in the decoding device 200 and may be referred to as a presentation device or a display device.

[0325] FIG. 35 is a syntax diagram showing an example of syntax in the sixth aspect. Specifically, an example of syntax for a multiplexing format is shown. In this example, the multiplexing format includes a loop performed for each encoding device. For example, each of the multiple encoding devices may transmit information to a server device. Then, the server device may compile the information transmitted from the multiple encoding devices to generate the information in the multiplexing format shown in FIG. 35 .

[0326] 35 includes a video type indicating whether the encoded data is real video data or synthetic video data, which makes it possible to notify the decoding device 200 of the type of data the encoded data is.

[0327] Alternatively, the multiplexing format may include the encoding method used to encode the video. For example, the encoding device 300 may indicate a video codec method such as HEVC or VVC used to encode the video as the encoding method. Alternatively, the encoding device 100 may indicate an encoding method for the composite video as the encoding method. This makes it possible to notify the decoding device 200 of the type of encoded data.

[0328] Fig. 36 is a conceptual diagram showing an example of a user interface in which a composite video and a reference image input mode are simultaneously presented. The example of Fig. 36 corresponds to the example of a user interface presented in the video call system shown in Fig. 34. In particular, a presentation screen for a many-to-many video call is shown here. In a one-to-one video call, the video of one of the other parties is presented in the same manner as in the example of Fig. 36.

[0329] For example, the decoding device 200 (or the application device) displays the determination result as to whether an avatar, real video, or synthetic video is used in the encoding device 100. This allows the user on the receiving side to recognize the operating mode of the transmitting side.

[0330] Regarding the determination of whether an avatar is used, the decoding device 200 may receive the character type of the reference image and make the determination based on whether the character type indicates an avatar. Alternatively, the decoding device 200 may make the determination using other parameters indicating whether an avatar is used, rather than the character type parameters. The determination may also be made using information on whether an avatar is used that is set in advance and shared by the encoding device 100 and the decoding device 200.

[0331] The determination of whether the video is real or not and the determination of whether the video is synthetic or not may be made based on the notified video type, or may be made based on whether the encoding format is a synthetic video encoding format or a normal video encoding format.

[0332] In the case of a synthetic video, a reference image input mode may be presented to indicate a video input or an image input. Alternatively, a reference image determination result may be presented to indicate whether the person displayed on the screen matches the sender. This allows the user to recognize whether the person displayed on the screen is the same as the sender.

[0333] In this example, the information is presented in text, but the presentation method is not limited to this. The information may be presented using background, text color, sound, icons, image movement, animation, etc.

[0334] Furthermore, instead of information on whether the person presented on the screen is the same as the sender, information may be presented such as whether the person presented on the screen may or may not be the same as the sender.

[0335] 37 is a conceptual diagram showing an example of a user interface in which a composite video and an authentication image are presented simultaneously. The example of FIG. 37 corresponds to an example in the fourth mode in which an authentication image is presented simultaneously with a composite video. By presenting the authentication image together with the composite video, the user can visually determine whether the reference image matches the authentication image. The authentication image may be displayed constantly or for a fixed period of time. Alternatively, the authentication image may be displayed in response to a user's designation.

[0336] 38 is a conceptual diagram showing an example of a user interface in which a composite video based on a reference image and a composite video based on an authentication image are presented alternately. The example of FIG. 38 corresponds to an example of the fifth aspect. By periodically presenting a composite video based on an authentication image, it becomes possible for a user to visually determine whether the person in the presented video is the same person as the sender.

[0337] The composite video based on the reference image corresponds to the output video generated by transforming the reference image based on the feature data of the input video, and the composite video based on the authentication image corresponds to the output video generated by transforming the authentication image based on the feature data of the input video.

[0338] The above example allows the user to visually determine whether the person in the video being presented is the same person as the sender, making it possible to prevent spoofing of synthesized videos, etc.

[0339] The decoding device 200 may also have a mode for issuing an alert if the person in the video being presented is not the same as the sender, or a mode for rejecting an incoming call or telephone call, and the mode may be set by the user.

[0340] The encoded data may be multiplexed using a predetermined multiplexing method and transmitted. The encoded data may be transmitted to a storage device and stored in the storage device.

[0341] The multiplexed data may include, in addition to the encoded data of video data or the encoded data of synthetic video, encoded data of media such as audio, 3D data, subtitles, application data, and files, SEI, reference time information, etc. Then, the demultiplexer 224 may extract the encoded data, SEI, time information, etc. from the multiplexed data.

[0342] As a multiplexing method or a file format, for example, ISOBMFF, MPEG-DASH, MMT, MPEG-2 TS Systems, RTP, glTF, etc., which are transmission methods based on ISOBMFF, may be used. Metadata included in the SEI may be stored in a separate syntax in the file format.

[0343] For example, the metadata may be stored in a "moov" box in ISOBMFF. Alternatively, a new box may be defined in the "moov" box, and the metadata may be stored in the new box. Furthermore, the metadata may be stored in a file such as MPD (Media Presentation Description) in MPEG-DASH.

[0344] [Seventh Aspect] Whether or not to perform the determination process, selection process, and presentation process shown in the first to sixth aspects may be switched depending on the modes of the transmitting and receiving sides.

[0345] For example, the modes in which the processes shown in the first to sixth aspects are performed may be expressed as "secure modes" for detecting spoofing. Whether or not to make a call in secure mode may be set at the start of a call or at the time of initial setup.

[0346] 39 is a flowchart showing an example of the operation of the video call system in the seventh aspect. This example corresponds to an example in which whether or not to make a call in secure mode is set at the start of the call.

[0347] In this example, at the start of a call, the receiving terminal corresponding to the decoding device 200 receives data from the transmitting terminal corresponding to the encoding device 100. The receiving terminal then determines whether the received data is composite video data (S301). If the received data is not composite video data (No in S301), the processes shown in the first to sixth aspects regarding the determination of the possibility that the person in the reference image is different from the person in the input video are not performed (S306).

[0348] If the received data is composite video data (Yes in S301), the receiving terminal acquires from the user a designation as to whether or not to conduct the call in secure mode. Then, the receiving terminal determines whether or not to conduct the call in secure mode in accordance with the designation acquired from the user (S302). If the call is not conducted in secure mode (No in S302), the processes shown in the first to sixth aspects regarding the determination of the possibility that the person in the reference image is different from the person in the input video are not performed (S306).

[0349] If the call is made in secure mode (Yes in S302), the privacy setting of the transmitting terminal is confirmed (S303). Specifically, the receiving terminal requests the transmitting terminal to disclose information about the input video. In response to the request for information disclosure, the transmitting terminal acquires a designation from the user as to whether or not to permit information disclosure. The transmitting terminal then responds to the receiving terminal as to whether or not to permit information disclosure.

[0350] If the disclosure of information is denied (No in S304), the processes shown in the first to sixth aspects regarding the determination of the possibility that the person in the reference image is different from the person in the input video are not performed (S306).If the disclosure of information is permitted (Yes in S304), the processes shown in the first to sixth aspects regarding the determination of the possibility that the person in the reference image is different from the person in the input video are performed (S305).

[0351] Fig. 40 is a conceptual diagram showing an example of a user interface of the video call system in the seventh aspect. For example, to determine whether or not to make a call in secure mode, a confirmation message as shown in the upper left of Fig. 40 is displayed on the screen of the receiving terminal (S302). If the call is to be made in secure mode, a confirmation message as shown in the right of Fig. 40 is displayed on the screen of the transmitting terminal (S303). Furthermore, if information disclosure is rejected at the transmitting terminal, a confirmation message as shown in the lower left of Fig. 40 is displayed at the receiving terminal.

[0352] When the secure mode is set to off, that is, when high security is not required, the processes of the first to sixth aspects may be omitted, thereby reducing the amount of processing.

[0353] Furthermore, in all of the first to sixth aspects, input video information is used. However, from the viewpoint of privacy, it may be inappropriate to obtain input video information without the user's permission. Therefore, when one user wishes to conduct a call in secure mode, the video call system may allow the other user to select whether or not to permit disclosure of input video information.

[0354] The example screen on the right side of Fig. 40 corresponds to an example of a screen for this selection. This makes it possible to prevent unintentional invasion of privacy. Whether or not to permit disclosure of input video information may be selected for each call. Alternatively, a "privacy mode" may be set as a default so that information disclosure is not permitted.

[0355] The security level may be different between the mode in which the sender performs reference image judgment and the mode in which the sender presents an authentication image. In the mode in which reference image judgment is performed, the authentication image is not sent, which may achieve both prevention of spoofing and protection of the sender's privacy.

[0356] Each of the reference image determiner 111, the reference image selector 112, the authentication image generator 113, and the multiplexer 115 may or may not be included in the encoding device 100. Furthermore, the encoding device 100 may include an encoding processor that performs encoding processing for each piece of information. Furthermore, each of the reference image determiner 222, the video generator 223, the demultiplexer 224, and the presenter 225 may or may not be included in the decoding device 200.

[0357] Furthermore, the encoding device 100 may be, be included in, or include a transmitting device. The transmitting device may include a transmitter for transmitting the bitstream. Transmitting the bitstream may correspond to transmitting the bitstream to a recording medium, i.e., writing the bitstream to the recording medium.

[0358] Furthermore, the decoding device 200 may be a receiving device, may be included in a receiving device, or may include a receiving device. The receiving device may include a receiver that receives a bitstream. Receiving the bitstream may correspond to receiving the bitstream from a recording medium, i.e., reading the bitstream from the recording medium.

[0359] [Implementation Example] As described above, video generation based on facial feature data of a person is performed. The video generation is realized, for example, by a device including a memory and a circuit connected to the memory. One example of this device stores an input video in the memory, retrieves the input video stored in the memory using a circuit, and generates an output video based on the input video.

[0360] 41 is a block diagram showing an example implementation of the encoding device 100. In this example, the encoding device 100 includes a circuit 151 and a memory 152.

[0361] The circuit 151 is a circuit that processes information and has access to the memory 152. For example, the circuit 151 is a dedicated or general-purpose electronic circuit that encodes input video. The circuit 151 may be a processor such as a CPU. The circuit 151 may also be a collection of multiple electronic circuits. For example, the circuit 151 may fulfill the roles of multiple components of the encoding device 100 in the above-mentioned examples, excluding the component for storing information.

[0362] Memory 152 is a dedicated or general-purpose memory that stores information for encoding input video by circuitry 151. Memory 152 may be an electronic circuit and may be connected to circuitry 151. Alternatively, memory 152 may be included in circuitry 151.

[0363] The memory 152 may be a collection of multiple electronic circuits. The memory 152 may be a magnetic disk, an optical disk, or the like, and may be expressed as a storage or a recording medium. The memory 152 may be a non-volatile memory or a volatile memory.

[0364] For example, the memory 152 may store an input video to be encoded, an encoded bitstream, a reference image used in encoding the input video, metadata such as a determination result and a selection result, etc. The memory 152 may also store a program that causes the circuit 151 to encode the reference image, feature data, and determination-related information.

[0365] It should be noted that not all of the components of the encoding device 100 in the above examples may be implemented, and not all of the above-described processes may be performed, in the encoding device 100. Some of the components may be included in another device, and some of the above-described processes may be performed by another device.

[0366] 42 is a block diagram showing an example implementation of the decoding device 200. In this example, the decoding device 200 includes a circuit 251 and a memory 252.

[0367] The circuit 251 is a circuit that performs information processing and is a circuit that can access the memory 252. For example, the circuit 251 is a dedicated or general-purpose electronic circuit that decodes a stream. The circuit 251 may be a processor such as a CPU. The circuit 251 may also be a collection of multiple electronic circuits. For example, the circuit 251 may fulfill the roles of multiple components of the decoding device 200 in the above-described examples, excluding the component for storing information.

[0368] The memory 252 is a dedicated or general-purpose memory that stores information for decoding the stream by the circuit 251. The memory 252 may be an electronic circuit and may be connected to the circuit 251. Alternatively, the memory 252 may be included in the circuit 251.

[0369] The memory 252 may be a collection of multiple electronic circuits. The memory 252 may be a magnetic disk, an optical disk, or the like, and may be expressed as a storage or a recording medium. The memory 252 may be a non-volatile memory or a volatile memory.

[0370] For example, the memory 252 may store an output video or a bitstream. The memory 252 may also store a program for the circuit 251 to decode the stream. The memory 252 may also store metadata such as a reference image used to generate the output video, a determination result, and a selection result.

[0371] Note that not all of the components of the decoding device 200 in the above examples may be implemented, and not all of the above-described processes may be performed, in the decoding device 200. Some of the components may be included in another device, and some of the above-described processes may be executed by another device.

[0372] 43 is a flowchart showing an example of the basic operation of the encoding device 100. In the operation of this example, the circuit 151 of the encoding device 100 uses the memory 152 to perform the following.

[0373] Specifically, the circuit 151 encodes, into a bitstream, at least one reference image including a face of a first person (S501). The circuit 151 also encodes, into a bitstream, feature data extracted from an input video including a face of a second person, the same as or different from the first person, and the feature data is used to generate an output video by transforming the at least one reference image during decoding (S502). The circuit 151 also encodes, into a bitstream, determination-related information, which is information related to determining whether the second person is likely to be different from the first person (S503).

[0374] The above-described operation of the encoding device 100 may enable the encoding device 100 to communicate to the decoding device 200 the possibility that the second person corresponding to the information source of the feature data reflected in the reference image in generating the output video is different from the first person in the reference image. Therefore, it may be possible to appropriately communicate the possibility of impersonation. Therefore, it may be possible to appropriately identify the possibility of impersonation.

[0375] 44 is a flowchart showing an example of the basic operation of the decoding device 200. In the operation of this example, the circuit 251 of the decoding device 200 uses the memory 252 to perform the following.

[0376] Specifically, the circuit 251 decodes from the bitstream at least one reference image including a face of a first person (S601). The circuit 251 also decodes from the bitstream feature data extracted from an input video including a face of a second person, the second person being the same as or different from the first person, during encoding, and the feature data is used to generate an output video by transforming the at least one reference image (S602). The circuit 251 also decodes from the bitstream determination-related information, which is information related to determining whether the second person is likely to be different from the first person (S603).

[0377] The above-described operation of the decoding device 200 may enable the encoding device 100 to communicate to the decoding device 200 the possibility that the second person corresponding to the information source of the feature data reflected in the reference image in generating the output video is different from the first person in the reference image. Therefore, it may be possible to appropriately communicate the possibility of impersonation. Therefore, it may be possible to appropriately identify the possibility of impersonation.

[0378] The determination of the possibility that the second person is different from the first person may be performed by the encoding device 100, the decoding device 200, or another device. The determination-related information may be information indicating a determination result of the possibility that the second person is different from the first person, or may be information for determining the possibility that the second person is different from the first person. Furthermore, the possibility that the second person is different from the first person may be replaced with the possibility that the second person is the same as the first person.

[0379] For example, the determination-related information may include information based on a comparison of at least one reference image with the input video, which may allow for conveying information based on the comparison of the reference image with the input video regarding the likelihood that the second person is different from the first person, thereby allowing for conveying accurate information regarding the likelihood that the second person is different from the first person.

[0380] In addition, the information based on the comparison between the reference image and the input video may indicate a result obtained by determining the likelihood that the second person is different from the first person based on the comparison between the reference image and the input video.

[0381] For example, the determination-related information may include information based on whether at least one reference image is an image included in the input video, and if at least one reference image is an image included in the input video, it may be determined that the likelihood that the second person is different from the first person is lower than a predetermined probability.

[0382] This may allow for conveying information regarding the likelihood that the second person is different from the first person based on whether the reference image is included in the input video, and thus may allow for efficient conveying of the likelihood that the second person is different from the first person.

[0383] The information based on whether the reference image is an image included in the input video may be information indicating whether the reference image is an image included in the input video, or may indicate a result obtained by determining the likelihood that the second person is different from the first person based on whether the reference image is an image included in the input video.

[0384] The predetermined probability may be an average probability that the second person is different from the first person, or an overall average probability regardless of conditions. Alternatively, the predetermined probability may be a pre-specified probability such as 50%. Alternatively, the predetermined probability may be the probability that the second person is different from the first person when the reference image is not an image included in the input video.

[0385] Furthermore, for example, the determination-related information may include information based on a comparison between at least one reference image and the input video when the at least one reference image is not included in the input video. This may make it possible to convey information based on a comparison between the reference image and the input video regarding the likelihood that the second person is different from the first person when the reference image is not included in the input video. Therefore, it may be possible to efficiently convey accurate information regarding the likelihood that the second person is different from the first person.

[0386] Here, when the reference image is an image included in the input video, the determination related information does not need to include information based on a comparison between the reference image and the input video.

[0387] Furthermore, for example, the determination-related information may include information based on whether at least one reference image is composed of multiple reference images, and if at least one reference image is composed of multiple reference images, it may be determined that the likelihood that the second person is different from the first person is lower than a predetermined likelihood.

[0388] This may allow for conveying information regarding the likelihood that the second person is different from the first person based on whether the reference image is comprised of multiple images, and thus may allow for efficient conveying of the likelihood that the second person is different from the first person.

[0389] The information based on whether the reference image is composed of multiple reference images may be information indicating whether the reference image is composed of multiple reference images, or may indicate a result obtained by determining the likelihood that the second person is different from the first person based on whether the reference image is composed of multiple reference images.

[0390] Furthermore, the predetermined possibility may be an average possibility that the second person is different from the first person, or an overall average possibility regardless of conditions. Alternatively, the predetermined possibility may be a pre-specified possibility such as 50%. Alternatively, the predetermined possibility may be the possibility that the second person is different from the first person when the reference image is not composed of multiple reference images.

[0391] Furthermore, for example, the determination-related information may include information based on whether at least one reference image is a previously acquired image, and if at least one reference image is a previously acquired image, it may be determined that the likelihood that the second person is different from the first person is higher than a predetermined probability.

[0392] This may allow for communicating information regarding the likelihood that the second person is different from the first person based on whether the reference image is a pre-acquired image, and thus may allow for efficiently communicating the likelihood that the second person is different from the first person.

[0393] The information based on whether the reference image is an image acquired in advance may be information indicating whether the reference image is an image acquired in advance, or may indicate a result obtained by determining the likelihood that the second person is different from the first person based on whether the reference image is an image acquired in advance.

[0394] The predetermined probability may be an average probability that the second person is different from the first person, or an overall average probability regardless of conditions. Alternatively, the predetermined probability may be a predetermined probability such as 50%. Alternatively, the predetermined probability may be the probability that the second person is different from the first person when the reference image is not an image acquired in advance.

[0395] Also, the reference image being acquired in advance may correspond to the reference image being acquired prior to the input video during encoding.

[0396] For example, the determination-related information may include information based on a comparison between at least one reference image and an image in the input video. This may allow for conveying information based on the comparison between the reference image and an image in the input video regarding the likelihood that the second person is different from the first person. Therefore, it may allow for efficient conveyance of accurate information regarding the likelihood that the second person is different from the first person.

[0397] Here, the decision-related information may include information based on a comparison of at least one reference image with only one image in the input video.

[0398] Furthermore, for example, the determination-related information may include information based on a comparison between at least one reference image and an image at a predetermined timing in the input video. This may make it possible to convey information regarding the possibility that the second person is different from the first person based on a comparison between the reference image and an image in the input video at a predetermined timing. Therefore, it may be possible to efficiently convey accurate information regarding the possibility that the second person is different from the first person.

[0399] Here, the determination-related information may include information based on a comparison between at least one reference image and each image at a predetermined timing in the input video.

[0400] Furthermore, for example, the predetermined timing may correspond to the start of a randomly accessible unit. This may make it possible to convey information regarding the possibility that the second person is different from the first person based on a comparison between the reference image and the image of the input video at a timing corresponding to the start of the randomly accessible unit. This may make it possible to efficiently convey accurate information regarding the possibility that the second person is different from the first person.

[0401] Furthermore, for example, the predetermined timing may correspond to a specified interval. This may make it possible to transmit information regarding the possibility that the second person is different from the first person based on a comparison between the reference image and the image of the input video at the timing corresponding to the specified interval. Therefore, it may be possible to efficiently transmit accurate information regarding the possibility that the second person is different from the first person.

[0402] Furthermore, for example, the determination-related information may include an authentication image, which is an image included in the input video. This may make it possible to convey the authentication image regarding the possibility that the second person is different from the first person. Therefore, it may be possible to convey information that is effective in determining the possibility that the second person is different from the first person. Therefore, it may be possible to appropriately convey the possibility that the second person is different from the first person.

[0403] Furthermore, for example, the circuit 251 may determine the possibility that the second person is different from the first person based on the determination-related information. This may allow an appropriate determination to be made based on the determination-related information related to the determination of the possibility that the second person is different from the first person. Therefore, it may allow an appropriate determination of the possibility of impersonation.

[0404] Also, for example, the circuit 251 may display information based on the determination-related information together with the output video. This may allow visual notification of information related to the determination that the second person may be different from the first person. Therefore, it may allow appropriate notification of the possibility of impersonation.

[0405] For example, the circuit 251 may alternately display an output video generated by modifying at least one reference image based on the feature data and another output video generated by modifying the authentication image based on the feature data. This may visually notify information useful for determining whether the second person is different from the first person. Therefore, it may be possible to appropriately notify the possibility of impersonation.

[0406] Note that the circuit 251 may display the output video, etc. via a display device inside or outside the decoding device 200. That is, the circuit 251 may display the output video, etc. by causing a display device to display the output video, etc.

[0407] Alternatively, the encoding device 100 may include an input terminal, an entropy encoder, and an output terminal. The operations performed by the circuit 151 may be performed by the entropy encoder. Data used in the operation of the entropy encoder may be input to the input terminal. Data obtained by the operation of the entropy encoder may be output from the output terminal.

[0408] Alternatively, the decoding device 200 may include an input terminal, an entropy decoder, and an output terminal. The operations performed by the circuit 251 may be performed by the entropy decoder. Data used in the operation of the entropy decoder may be input to the input terminal. Data obtained by the operation of the entropy decoder may be output from the output terminal.

[0409] [Other Examples] The encoding device 100 and the decoding device 200 in each of the above-described examples may be used as an image encoding device and an image decoding device, or as a video encoding device and a video decoding device, respectively. Furthermore, multiple components included in the encoding device 100 and multiple components included in the decoding device 200 may perform corresponding operations.

[0410] Additionally, the term "encode" may be substituted with terms such as "store," "include," "write," "write," "signal," "send," "notify," or "preserve," and these terms may be interchangeable. For example, encoding information may mean including the information in a bitstream. Also, encoding information into a bitstream may mean encoding the information to generate a bitstream that includes the encoded information.

[0411] Furthermore, the term "decode" may be replaced with terms such as "read," "decode," "read," "load," "derive," "obtain," "receive," "extract," or "restore," and these terms may be interchangeable. For example, decoding information may mean obtaining information from a bitstream. Decoding information from a bitstream may mean decoding the bitstream to obtain information contained in the bitstream.

[0412] Furthermore, at least some of the examples described above may be used as an encoding method, a decoding method, an entropy encoding method, an entropy decoding method, or some other method.

[0413] Each component may be configured with dedicated hardware, or may be realized by executing a software program suitable for that component. Each component may be realized by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.

[0414] Specifically, each of the encoding device 100 and the decoding device 200 may include a processing circuit and a storage device electrically connected to and accessible from the processing circuit. For example, the processing circuit corresponds to the circuit 151 or 251, and the storage device corresponds to the memory 152 or 252.

[0415] The processing circuit includes at least one of dedicated hardware and a program execution unit, and executes processing using a storage device. If the processing circuit includes a program execution unit, the storage device stores the software program executed by the program execution unit.

[0416] An example of the above-mentioned software program is a bitstream. The bitstream includes an encoded image and a syntax for performing a decoding process to decode the image. The bitstream causes the decoding device 200 to decode the image by executing a process based on the syntax. Furthermore, for example, software for realizing the above-mentioned encoding device 100 or decoding device 200 is a program such as the following.

[0417] For example, the program may cause a computer to execute an encoding method that encodes at least one reference image including a face of a first person into a bitstream, encodes feature data extracted from an input video including a face of a second person who is the same as or different from the first person into the bitstream, the feature data being used to transform the at least one reference image to generate an output video upon decoding, and encodes determination-related information, which is information related to determining whether the second person is likely to be different from the first person, into the bitstream.

[0418] Furthermore, for example, the program may cause a computer to execute a decoding method that decodes at least one reference image including a face of a first person from a bitstream, decodes from the bitstream feature data extracted from an input video including a face of a second person that is the same as or different from the first person during encoding, the feature data being used to transform the at least one reference image to generate an output video, and decodes from the bitstream determination-related information that is information related to determining whether the second person is likely to be different from the first person.

[0419] Furthermore, each of the above-described components may be a circuit. These circuits may form a single circuit as a whole, or each may be a separate circuit. Furthermore, each component may be realized by a general-purpose processor or a dedicated processor.

[0420] Furthermore, a process performed by a specific component may be performed by another component. The order in which the processes are performed may be changed, or multiple processes may be performed in parallel. Any two or more of the multiple examples of the present disclosure may be appropriately combined and implemented. The encoding / decoding device may include the encoding device 100 and the decoding device 200.

[0421] Furthermore, ordinal numbers such as "first" and "second" used in the description may be changed as appropriate. Furthermore, new ordinal numbers may be assigned to components or removed. Furthermore, these ordinal numbers may be assigned to elements in order to identify them, and may not correspond to a meaningful order.

[0422] Also, for example, a phrase "at least one of" a first element, a second element, and a third element corresponds to the first element, the second element, the third element, or any combination thereof.

[0423] Although the aspects of the encoding device 100 and the decoding device 200 have been described above based on a number of examples, the aspects of the encoding device 100 and the decoding device 200 are not limited to these examples. As long as they do not deviate from the spirit of the present disclosure, various modifications conceivable by those skilled in the art to each example, or configurations constructed by combining components of different examples, may also be included within the scope of the aspects of the encoding device 100 and the decoding device 200.

[0424] One or more aspects disclosed herein may be implemented in combination with at least a part of other aspects of the present disclosure. Also, some processes shown in the flowcharts of one or more aspects disclosed herein, some configurations of devices, some syntax, etc. may be implemented in combination with other aspects.

[0425] [Implementation and Application] In each of the above embodiments, each of the functional or operational blocks can typically be realized by an MPU (micro processing unit), memory, etc. Furthermore, the processing by each of the functional blocks may be realized as a program execution unit such as a processor that reads and executes software (programs) recorded on a recording medium such as a ROM. The software may be distributed. The software may be recorded on various recording media such as semiconductor memory. It is also possible to realize each functional block by hardware (dedicated circuitry).

[0426] The processing described in each embodiment may be realized by centralized processing using a single device (system), or may be realized by distributed processing using multiple devices. Furthermore, the processor that executes the program may be a single processor or multiple processors. That is, centralized processing or distributed processing may be performed.

[0427] The aspects of the present disclosure are not limited to the above examples, and various modifications are possible, and these modifications are also included within the scope of the aspects of the present disclosure.

[0428] Furthermore, application examples of the video coding method (image coding method) or video decoding method (image decoding method) shown in each of the above embodiments and various systems for implementing the application examples will be described below. Such systems may be characterized by having an image coding device using the image coding method, an image decoding device using the image decoding method, or an image coding / decoding device that includes both. Other configurations of such systems can be appropriately changed depending on the situation.

[0429] [Example of Use] Fig. 45 shows the overall configuration of an appropriate content supply system ex100 that realizes a content distribution service. The area where communication services are provided is divided into cells of a desired size, and base stations ex106, ex107, ex108, ex109, and ex110, which are fixed wireless stations in the illustrated example, are installed in each cell.

[0430] In this content supply system ex100, devices such as a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 are connected to the Internet ex101 via an Internet service provider ex102 or a communication network ex104 and base stations ex106 to ex110. The content supply system ex100 may connect a combination of any of the above devices. In various implementations, the devices may be connected to each other directly or indirectly via a telephone network, short-range wireless communication, or the like, without going through the base stations ex106 to ex110. Furthermore, the streaming server ex103 may be connected to devices such as the computer ex111, the game console ex112, the camera ex113, the home appliance ex114, and the smartphone ex115 via the Internet ex101, etc. The streaming server ex103 may also be connected to a terminal in a hotspot on an airplane ex117 via a satellite ex116.

[0431] Note that wireless access points, hot spots, etc. may be used instead of the base stations ex106 to ex110. Furthermore, the streaming server ex103 may be directly connected to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or may be directly connected to the airplane ex117 without going through the satellite ex116.

[0432] The camera ex113 is a device such as a digital camera that can take still images and videos. The smartphone ex115 is a smartphone, mobile phone, or PHS (Personal Handyphone System) that supports mobile communication systems such as 2G, 3G, 3.9G, 4G, and the upcoming 5G.

[0433] The home appliance ex114 is a refrigerator, a device included in a home fuel cell cogeneration system, or the like.

[0434] In the content supply system ex100, a terminal having a photographing function is connected to a streaming server ex103 via a base station ex106 or the like, thereby enabling live streaming and the like. In live streaming, a terminal (such as a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, a smartphone ex115, or a terminal on an airplane ex117) may perform the encoding process described in each of the above embodiments on still image or video content captured by a user using the terminal, may multiplex the video data obtained by encoding with audio data obtained by encoding audio corresponding to the video, and may transmit the obtained data to the streaming server ex103. In other words, each terminal functions as an image encoding device according to one aspect of the present disclosure.

[0435] Meanwhile, the streaming server ex103 streams the transmitted content data to the requesting client. The client is a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, a smartphone ex115, a terminal on an airplane ex117, or the like, which is capable of decoding the encoded data. Each device that receives the distributed data decodes and plays back the received data. That is, each device may function as an image decoding device according to one aspect of the present disclosure.

[0436] [Distributed Processing] The streaming server ex103 may also be multiple servers or multiple computers that process, record, and distribute data in a distributed manner. For example, the streaming server ex103 may be implemented using a CDN (Content Delivery Network), where content distribution is achieved through a network connecting numerous edge servers distributed around the world. In a CDN, a physically nearby edge server is dynamically assigned depending on the client. Content is then cached and distributed to that edge server, thereby reducing delays. Furthermore, when certain types of errors occur or communication conditions change due to increased traffic, processing can be distributed among multiple edge servers, the distribution entity can be switched to another edge server, or distribution can be continued by bypassing the failed portion of the network, thereby achieving high-speed and stable distribution.

[0437] In addition to the distributed processing of the distribution itself, the encoding of captured data may be performed by each terminal, by the server, or by multiple terminals. For example, encoding generally involves two processing loops. The first loop detects the complexity of the image or the amount of code for each frame or scene. The second loop maintains image quality while improving encoding efficiency. For example, a terminal may perform the first encoding process, and the server that receives the content may perform the second encoding process, thereby improving content quality and efficiency while reducing the processing load on each terminal. In this case, if there is a request to receive and decode the data in near real time, the data encoded by a terminal can be received and played back by another terminal, enabling more flexible real-time distribution.

[0438] As another example, the camera ex113 or the like extracts features from an image, compresses the data related to the features as metadata, and transmits the compressed data to the server. The server performs compression according to the meaning (or importance of the content) of the image, for example, by determining the importance of an object from the features and switching the quantization precision accordingly. The feature data is particularly effective in improving the accuracy and efficiency of motion vector prediction when the server re-compresses the image. Alternatively, the terminal may perform simple encoding such as VLC (variable length coding), and the server may perform encoding with a high processing load such as CABAC (context-adaptive binary arithmetic coding).

[0439] As another example, in a stadium, shopping mall, factory, or the like, there may be multiple pieces of video data that have been shot by multiple terminals of almost the same scene. In this case, using the multiple terminals that shot the video and, as necessary, other terminals and servers that did not shoot the video, encoding processes are assigned to each of them, for example, in units of GOPs (Group of Pictures), pictures, or tiles obtained by dividing a picture, for distributed processing. This reduces delays and achieves better real-time performance.

[0440] Since multiple video data are of almost the same scene, the server may manage and / or instruct the video data shot by each terminal to be mutually referential. The server may also receive encoded data from each terminal and change the reference relationships between multiple data, or correct or replace the pictures themselves and re-encode them. This allows for the generation of streams with improved quality and efficiency for each piece of data.

[0441] Furthermore, the server may perform transcoding to change the encoding method of the video data before distributing it. For example, the server may convert an MPEG-based encoding method into a VP-based encoding method (e.g., VP9), or convert H.264 to H.265.

[0442] In this way, the encoding process can be performed by a terminal or one or more servers. Therefore, although the following uses terms such as "server" or "terminal" to refer to the entity performing the process, some or all of the processing performed by the server may be performed by the terminal, and some or all of the processing performed by the terminal may be performed by the server. The same applies to the decoding process.

[0443] [3D, Multi-Angle] Images or videos of different scenes or the same scene taken from different angles using multiple devices such as cameras ex113 and / or smartphones ex115 that are approximately synchronized with each other are increasingly being integrated and used. The videos taken by each device are integrated based on the relative positional relationship between the devices obtained separately, or on areas where feature points in the videos match.

[0444] The server may not only encode two-dimensional video, but also encode still images automatically or at a time specified by the user based on scene analysis of the video and transmit them to the receiving terminal. Furthermore, if the server can acquire the relative positional relationship between the capturing terminals, it can generate a three-dimensional shape of the scene based on not only two-dimensional video but also video of the same scene captured from different angles. The server may separately encode three-dimensional data generated by a point cloud or the like, or may select or reconstruct the video to be transmitted to the receiving terminal from video captured by multiple terminals based on the results of recognizing or tracking people or objects using the three-dimensional data.

[0445] In this way, the user can enjoy a scene by arbitrarily selecting each video corresponding to each shooting terminal, or can enjoy content in which a video from a selected viewpoint is cut out from 3D data reconstructed using multiple images or videos. Furthermore, together with the video, sound may also be collected from multiple different angles, and the server may multiplex the sound from a specific angle or space with the corresponding video and transmit the multiplexed video and sound.

[0446] In recent years, content that associates the real world with a virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has also become popular. In the case of VR images, the server creates viewpoint images for the right eye and left eye, respectively, and may perform encoding that allows reference between each viewpoint video using Multi-View Coding (MVC) or the like, or may encode them as separate streams without referencing each other. When decoding the separate streams, it is preferable to play them in synchronization with each other so that a virtual three-dimensional space is reproduced according to the user's viewpoint.

[0447] In the case of AR images, the server superimposes virtual object information in the virtual space onto camera information in the real space based on the three-dimensional position or the movement of the user's viewpoint. The decoding device may acquire or store virtual object information and three-dimensional data, generate a two-dimensional image according to the movement of the user's viewpoint, and smoothly connect the two-dimensional image to create superimposed data. Alternatively, the decoding device may send the movement of the user's viewpoint to the server in addition to a request for virtual object information. The server may create superimposed data according to the movement of the viewpoint received from the three-dimensional data stored on the server, encode the superimposed data, and distribute it to the decoding device. Note that the superimposed data may have an α value indicating transparency in addition to RGB, and the server may set the α value of parts other than the object created from the three-dimensional data to 0, etc., to encode the parts in a transparent state. Alternatively, the server may generate data by setting a predetermined RGB value as the background, like a chromakey, and using the background color for parts other than the object.

[0448] Similarly, the decoding of distributed data may be performed by each client terminal, by the server, or by multiple terminals. For example, one terminal may first send a reception request to the server, and then other terminals may receive and decode content according to the request, after which the decoded signal is transmitted to a device having a display. By distributing the processing and selecting appropriate content regardless of the capabilities of the communication terminals themselves, high-quality data can be reproduced. As another example, large-sized image data may be received on a TV or other device, and only a portion of the picture, such as a tile into which the picture is divided, may be decoded and displayed on the viewer's personal device. This allows the viewer to share the overall picture while checking their own area of ​​responsibility or an area of ​​interest in more detail.

[0449] In situations where multiple short-, medium-, or long-range wireless communications are available indoors and outdoors, seamless content reception may be possible using distribution system standards such as MPEG-DASH (Dynamic Adaptive Streaming over HTTP). Users may freely select and switch in real time between decoding devices or display devices, such as their own terminals and indoor / outdoor displays. Furthermore, decoding can be performed while switching between decoding and display devices using their own location information, etc. This allows information to be mapped and displayed on a part of the wall or ground of a neighboring building with an embedded display device while the user is traveling to their destination. It is also possible to switch the bit rate of received data based on the accessibility of the encoded data on the network, such as when the encoded data is cached on a server that can be quickly accessed from the receiving terminal or copied to an edge server in a content delivery service.

[0450] [Web Page Optimization] FIG. 46 is a diagram showing an example of a display screen of a web page on a computer ex111 or the like. FIG. 47 is a diagram showing an example of a display screen of a web page on a smartphone ex115 or the like. As shown in FIGS. 46 and 47 , a web page may include multiple link images that are links to image content, and the appearance of the link images may differ depending on the device used to view the page. When multiple link images are visible on the screen, the display device (decoding device) may display a still image or I-picture contained in each content as a link image until the user explicitly selects the link image, or until the link image approaches the center of the screen or the entire link image is within the screen, or may display a video such as a GIF animation using multiple still images or I-pictures, or may receive only the base layer and decode and display the video.

[0451] When a link image is selected by a user, the display device performs decoding while giving top priority to the base layer. Note that if the HTML (HyperText Markup Language) constituting the web page contains information indicating that the content is scalable, the display device may decode up to the enhancement layer. Furthermore, to ensure real-time performance, before selection or when the communication bandwidth is very limited, the display device decodes and displays only forward-reference pictures (I pictures, P pictures, and forward-reference-only B pictures), thereby reducing the delay between the decoding time of the first picture and the display time (the delay from the start of content decoding to the start of display). Furthermore, the display device may intentionally ignore the picture reference relationships and roughly decode all B pictures and P pictures using forward reference, and then perform normal decoding as the number of received pictures increases over time.

[0452] [Autonomous Driving] When transmitting and receiving still image or video data such as two-dimensional or three-dimensional map information for automatic driving or driving assistance of a vehicle, the receiving terminal may receive weather or construction information as meta information in addition to image data belonging to one or more layers, and may associate and decode these. Note that the meta information may belong to a layer, or may simply be multiplexed with the image data.

[0453] In this case, since a vehicle, drone, airplane, or the like including the receiving terminal is moving, the receiving terminal can transmit location information of the receiving terminal, thereby realizing seamless reception and decoding while switching between base stations ex106 to ex110. Furthermore, the receiving terminal can dynamically switch how much meta information to receive or how much to update map information depending on the user's selection, the user's situation, and / or the state of the communication bandwidth.

[0454] In the content supply system ex100, the client can receive, decode, and play back encoded information sent by a user in real time.

[0455] [Distribution of Personal Content] The content supply system ex100 also allows for unicast or multicast distribution of not only high-quality, long-duration content from video distribution companies, but also low-quality, short-duration content from individuals. It is expected that such personal content will continue to increase in the future. To improve the quality of personal content, the server may perform editing before encoding. This can be achieved, for example, using the following configuration.

[0456] During shooting, either in real time or after accumulating and shooting, the server performs recognition processing such as detecting shooting errors, scene search, semantic analysis, and object detection from the original image data or encoded data. Based on the recognition results, the server manually or automatically corrects out-of-focus or camera shake, deletes less important scenes such as scenes with lower brightness or out-of-focus compared to other pictures, emphasizes object edges, changes color, and performs other editing. The server then encodes the edited data based on the editing results. It is also known that viewing rates decrease if the shooting time is too long. Therefore, the server may automatically clip not only less important scenes as described above but also scenes with little movement, based on the image processing results, so that the content falls within a specific time range depending on the shooting time. Alternatively, the server may generate and encode a digest based on the results of the semantic analysis of the scenes.

[0457] Personal content may contain content that, if left as is, violates copyright, moral rights, or portrait rights, and may cause the scope of sharing to exceed the intended scope, resulting in inconvenience to individuals. Therefore, for example, the server may intentionally defocus images of people's faces on the periphery of the screen or the interior of a house before encoding. Furthermore, the server may recognize whether the image to be encoded contains the face of a person other than a pre-registered person, and if so, perform processing such as blurring the face. Alternatively, as pre- or post-processing before encoding, the user may specify a person or background area they wish to modify in the image for copyright or other reasons. The server may replace the specified area with another image or blur the focus. If the image contains a person, the server may track the person in the video and replace the image of the person's face.

[0458] Because viewing personal content with small data volumes requires high real-time performance, the decoding device first receives the base layer as a top priority, and then decodes and plays it back, depending on the bandwidth. The decoding device may also receive an enhancement layer during this time, and if the content is played back more than once, such as when playback is looped, it may play back high-quality video including the enhancement layer. A stream that has undergone scalable encoding in this way can provide an experience in which the video appears rough when not selected or when viewing begins, but gradually becomes smarter and the image quality improves. In addition to scalable encoding, a similar experience can also be provided when a rough stream played the first time and a second stream that is encoded with reference to the first video are configured as a single stream.

[0459] [Other Application Examples] Furthermore, these encoding or decoding processes are generally performed by the LSIex500 possessed by each terminal. The LSI (large scale integration circuitry) ex500 (see FIG. 45) may be a single-chip or multi-chip configuration. Furthermore, video encoding or decoding software may be embedded in some kind of recording medium (such as a CD-ROM, flexible disk, or hard disk) readable by the computer ex111, and the encoding or decoding process may be performed using that software. Furthermore, if the smartphone ex115 is equipped with a camera, video data captured by the camera may be transmitted. This video data is data encoded and processed by the LSIex500 possessed by the smartphone ex115.

[0460] The LSIex500 may be configured to download and activate application software. In this case, the terminal first determines whether it supports the content encoding method or has the capability to execute a specific service. If the terminal does not support the content encoding method or does not have the capability to execute a specific service, the terminal downloads a codec or application software, and then acquires and plays the content.

[0461] Furthermore, at least one of the moving image encoding device (image encoding device) or moving image decoding device (image decoding device) of each of the above embodiments can be incorporated into a digital broadcasting system, not limited to the content supply system ex100 via the Internet ex101. Since multiplexed data in which video and audio are multiplexed is transmitted and received over broadcast radio waves using a satellite or the like, there is a difference in that it is more suited to multicast than the content supply system ex100, which has a configuration that is easy to use for unicast, but similar applications are possible with regard to encoding and decoding processes.

[0462] [Hardware Configuration] Figure 48 is a diagram showing further details of the smartphone ex115 shown in Figure 45. Also, Figure 49 is a diagram showing an example configuration of the smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves to and from the base station ex110, a camera unit ex465 capable of capturing video and still images, and a display unit ex458 for displaying video captured by the camera unit ex465 and decoded data of the video and the like received by the antenna ex450. The smartphone ex115 further includes an operation unit ex466 such as a touch panel, an audio output unit ex457 such as a speaker for outputting voice or sound, an audio input unit ex456 such as a microphone for inputting voice, a memory unit ex467 capable of storing captured video or still images, recorded voice, received video or still images, encoded data such as email, or decoded data, and a slot unit ex464 that is an interface with a SIM (Subscriber Identity Module) ex468 for identifying a user and authenticating access to various data including the network. Note that an external memory may be used instead of the memory unit ex467.

[0463] A main control unit ex460 that comprehensively controls the display unit ex458 and operation unit ex466, etc., is connected to a power supply circuit unit ex461, an operation input control unit ex462, a video signal processing unit ex455, a camera interface unit ex463, a display control unit ex459, a modulation / demodulation unit ex452, a multiplexing / separation unit ex453, an audio signal processing unit ex454, a slot unit ex464, and a memory unit ex467 via a synchronization bus ex470.

[0464] When the power key is turned on by a user's operation, the power supply circuit unit ex461 starts up the smartphone ex115 to an operable state and supplies power to each unit from the battery pack.

[0465] The smartphone ex115 performs processes such as telephone calls and data communications under the control of a main control unit ex460 having a CPU, ROM, RAM, etc. During a call, an audio signal collected by an audio input unit ex456 is converted into a digital audio signal by an audio signal processing unit ex454, subjected to spectrum spread processing by a modulation / demodulation unit ex452, subjected to digital-to-analog conversion and frequency conversion processing by a transmission / reception unit ex451, and the resulting signal is transmitted via an antenna ex450. The received data is also amplified and subjected to frequency conversion and analog-to-digital conversion processing, subjected to spectrum despreading processing by a modulation / demodulation unit ex452, and converted into an analog audio signal by an audio signal processing unit ex454, which is then output from an audio output unit ex457. During data communication mode, text, still images, or video data is sent to the main control unit ex460 via an operation input control unit ex462 based on operations on the main unit's operation unit ex466, etc. Similar transmission and reception processing is performed. When transmitting video, still images, or video and audio in the data communication mode, the video signal processing unit ex455 compresses and encodes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 using the moving image encoding method described in each of the above embodiments, and sends the encoded video data to the multiplexing / demultiplexing unit ex453. The audio signal processing unit ex454 encodes the audio signal collected by the audio input unit ex456 while the camera unit ex465 is capturing video or still images, and sends the encoded audio data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded video data and encoded audio data using a predetermined method, and modulates and converts the multiplexed video data and audio data in the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, before transmitting the multiplexed video data and audio data via the antenna ex450.

[0466] In order to decode the multiplexed data received via the antenna ex450, such as when receiving video attached to an email or chat, or video linked to a web page, the multiplexing / separation unit ex453 separates the multiplexed data into a video data bit stream and an audio data bit stream, and supplies the encoded video data to the video signal processing unit ex455 and the encoded audio data to the audio signal processing unit ex454 via the synchronization bus ex470. The video signal processing unit ex455 decodes the video signal using a video decoding method corresponding to the video encoding method described in each of the above embodiments, and the video or still image contained in the linked video file is displayed on the display unit ex458 via the display control unit ex459. The audio signal processing unit ex454 decodes the audio signal, and audio is output from the audio output unit ex457. As real-time streaming becomes increasingly common, audio playback may be socially inappropriate depending on the user's situation. Therefore, it is preferable that the initial setting be a configuration in which only the video data is played without playing the audio signal, and audio may be played in sync only when the user performs an operation such as clicking on the video data.

[0467] Although the smartphone ex115 has been used as an example, three other implementation formats are possible: a transmitting / receiving terminal having both an encoder and a decoder, a transmitting terminal having only an encoder, and a receiving terminal having only a decoder. In the digital broadcasting system, multiplexed data in which audio data is multiplexed with video data is received or transmitted. However, in addition to audio data, text data related to the video may also be multiplexed into the multiplexed data. Furthermore, the video data itself may be received or transmitted instead of the multiplexed data.

[0468] Although the main control unit ex460 including a CPU has been described as controlling the encoding or decoding process, various terminals often include a GPU (Graphics Processing Unit). Therefore, a configuration may be adopted in which a memory shared by the CPU and GPU, or a memory whose addresses are managed for common use, is used to take advantage of the GPU's performance to process a large area in a batch. This shortens the encoding time, ensures real-time performance, and achieves low latency. It is particularly efficient to perform motion estimation, deblocking filtering, SAO (Sample Adaptive Offset), and transformation / quantization processes in a batch, such as by picture, by the GPU rather than by the CPU.

[0469] The present disclosure is applicable to encoding devices that encode information related to video, and is applicable to video conferencing systems, etc.

[0470] 100, 300, 700 Encoding device 101, 103, 701, 703 Compressor 102, 203, 702, 803 Deriver 111, 222 Reference image determiner 112 Reference image selector 113 Authentication image generator 114 Synthesis video encoder 115, 315 Multiplexer 151, 251 Circuit 152, 252 Memory 200, 800 Decoding device 201, 202, 801, 802 Decompressor 204, 804 Generator 221 Decoding processor 223 Video generator 224 Demultiplexer 225 Presenter 314 Video encoder

Claims

1. An encoding device comprising: a memory; and a circuit connected to the memory, the circuit operating to: encode at least one reference image including a face of a first person into a bitstream; encode into the bitstream feature data extracted from an input video including a face of a second person, the second person being the same as or different from the first person, the feature data being used to transform the at least one reference image to generate an output video upon decoding; and encode into the bitstream determination-related information, the determination being information related to a determination of the likelihood that the second person is different from the first person.

2. The encoding device of claim 1, wherein the decision-related information includes information based on a comparison of the at least one reference image with the input video.

3. The encoding device described in claim 1, wherein the determination-related information includes information based on whether or not the at least one reference image is an image included in the input video, and if the at least one reference image is an image included in the input video, it is determined that the possibility that the second person is different from the first person is lower than a predetermined possibility.

4. The encoding device according to claim 3, wherein the decision-related information includes information based on a comparison between the at least one reference image and the input video if the at least one reference image is not an image included in the input video.

5. The encoding device of claim 1, wherein the determination-related information includes information based on whether the at least one reference image is composed of multiple reference images, and if the at least one reference image is composed of multiple reference images, it is determined that the possibility that the second person is different from the first person is lower than a predetermined possibility.

6. The encoding device of claim 1, wherein the determination-related information includes information based on whether or not the at least one reference image is an image acquired in advance, and if the at least one reference image is an image acquired in advance, it is determined that the possibility that the second person is different from the first person is higher than a predetermined possibility.

7. The encoding device of claim 1, wherein the decision-related information comprises information based on a comparison of the at least one reference image with an image in the input video.

8. The encoding device according to claim 1, wherein the decision-related information includes information based on a comparison between the at least one reference image and an image at a predetermined timing in the input video.

9. The encoding device according to claim 8, wherein the predetermined timing corresponds to the start of a randomly accessible unit.

10. The encoding device according to claim 8, wherein the predetermined timing corresponds to a designated interval.

11. The encoding device according to claim 1, wherein the judgment-related information includes an authentication image that is an image included in the input video.

12. A decoding device comprising: a memory; and a circuit connected to the memory, which, in operation, decodes from the bitstream at least one reference image including a face of a first person; decodes from the bitstream feature data extracted from an input video including a face of a second person, the second person being the same as or different from the first person, during encoding, the feature data being used to transform the at least one reference image to generate an output video; and decodes from the bitstream determination-related information, which is information related to determining whether the second person is likely to be different from the first person.

13. The decoding device of claim 12, wherein the decision-related information includes information based on a comparison of the at least one reference image with the input video.

14. The decoding device described in claim 12, wherein the determination-related information includes information based on whether or not the at least one reference image is an image included in the input video, and if the at least one reference image is an image included in the input video, it is determined that the possibility that the second person is different from the first person is lower than a predetermined possibility.

15. The decoding device of claim 14, wherein the decision-related information includes information based on a comparison of the at least one reference image with the input video if the at least one reference image is not an image included in the input video.

16. The decoding device described in claim 12, wherein the determination-related information includes information based on whether the at least one reference image is composed of multiple reference images, and if the at least one reference image is composed of multiple reference images, it is determined that the possibility that the second person is different from the first person is lower than a predetermined possibility.

17. The decoding device described in claim 12, wherein the determination-related information includes information based on whether or not the at least one reference image is an image acquired in advance, and if the at least one reference image is an image acquired in advance, it is determined that the possibility that the second person is different from the first person is higher than a predetermined possibility.

18. The decoding device of claim 12, wherein the decision-related information includes information based on a comparison of the at least one reference image with an image in the input video.

19. The decoding device according to claim 12, wherein the decision-related information includes information based on a comparison between the at least one reference image and an image at a predetermined timing in the input video.

20. The decoding device according to claim 19, wherein the predetermined timing corresponds to the start of a randomly accessible unit.

21. The decoding device according to claim 19, wherein the predetermined timing corresponds to a designated interval.

22. The decoding device according to claim 12, wherein the determination-related information includes an authentication image that is an image included in the input video.

23. The decoding device according to any one of claims 12 to 22, wherein the circuit determines the likelihood that the second person is different from the first person based on the determination-related information.

24. A decoding device according to any one of claims 12 to 22, wherein the circuitry displays information based on the decision-related information together with the output video.

25. The decoding device of claim 22, wherein the circuit alternately displays the output video generated by modifying the at least one reference image based on the feature data and another output video generated by modifying the authentication image based on the feature data.

26. A method for encoding at least one reference image including a face of a first person into a bitstream; encoding feature data extracted from an input video including a face of a second person, the second person being the same as or different from the first person, into the bitstream, the feature data being used to generate an output video by transforming the at least one reference image during decoding; and encoding determination-related information into the bitstream, the determination being information related to a determination of the possibility that the second person is different from the first person.

27. A decoding method comprising: decoding at least one reference image including a face of a first person from a bitstream; decoding feature data from the bitstream, the feature data being extracted during encoding from an input video including a face of a second person, the second person being the same as or different from the first person, and used to generate an output video by transforming the at least one reference image; and decoding determination-related information from the bitstream, the determination being information related to a determination of the likelihood that the second person is different from the first person.

Citation Information

Patent Citations

  • Image processing apparatus, computer program, video call system, and image processing method

    JP2019201360A

  • Video conferencing method

    US20200358983A1