Video decoding device, video decoding method, and program
The video decoding apparatus addresses the inefficiencies in encoding emotional stamps by accurately mapping and displaying them, enhancing the encoding efficiency and transmission quality of face videos.
Patent Information
- Application Number
- JP2023216567
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-07-03
AI Technical Summary
Conventional video decoding techniques fail to consider the display location of various emotional stamps applied to face videos, leading to inefficiencies in encoding and transmission.
A video decoding apparatus and method that includes units for decoding image texture, expression and stamp control information, generating face mesh and map information, and predicting interpolation frames, thereby improving encoding efficiency by accurately mapping and displaying emotional stamps.
Enhances the encoding efficiency of face videos with emotional stamps by accurately determining their display locations and movements, resulting in improved transmission quality.
Smart Images

Figure 2025099694000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a video decoding apparatus, a video decoding method, and a program.
Background Art
[0002] Non-Patent Document 1 discloses GFVC (Generative Face Video Codec).
[0003] GFVC is specialized for the transmission of face videos. It decodes information that controls the movement of the mouth, blinking of the eyes, rotation of the head, and movement of the head that make up a person's expression, projects it onto a three-dimensional semantic face space, and synthesizes it with separately decoded face key frames to generate face interpolation frames, enabling efficient transmission of face videos.
Prior Art Documents
Non-Patent Documents
[0004]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] In Non-Patent Document 1, interpolation frames can be generated by synthesizing face images with key frames and control information for each part such as the tilt of the eyes, mouth, and head, realizing efficient transmission. However, there is a problem in that other emotional stamps applied to the face are not considered.
[0006] In addition, various types of emotion stamps have been proposed, and their display locations vary depending on the type. However, conventional decoding techniques have a problem in that such considerations for the display location are not made.
[0007] Therefore, the present invention has been made in view of the above-described problems, and an object thereof is to provide a video decoding apparatus, a video decoding method, and a program that can improve the encoding efficiency of a face video with an emotion stamp added thereto.
Means for Solving the Problems
[0008] A first feature of the present invention is a video decoding apparatus including: a first decoding unit that decodes an image texture from coded information of a key frame of a face video; a second decoding unit that decodes stamp control information including expression control information for controlling an expression in the face video, stamp identification information for identifying a stamp, and display area information indicating a display area of the stamp, from coded information of an interpolation frame; a face mesh generation unit that creates face mesh information based on the expression control information; an accumulation unit that accumulates map information defining a method of mapping the stamp to the interpolation frame; a stamp map generation unit that generates the map information of a next frame based on the stamp control information and the past map information accumulated in the accumulation unit; a face space generation unit that generates face space information, which is information abstractly expressed for expressing the expression and the stamp in the interpolation frame, based on the face mesh information and the map information generated by the stamp map generation unit; and an inter-frame prediction unit that predicts the interpolation frame based on the image texture and the face space information.
[0009] A second feature of the present invention is a video decoding method, which includes steps of decoding an image texture from the coded information of a key frame of a face video, decoding stamp control information including expression control information for controlling an expression in the face video, stamp identification information for identifying a stamp, and display area information indicating a display area of the stamp from the coded information of an interpolation frame, creating face mesh information based on the expression control information, accumulating map information defining a method of mapping the stamp to the interpolation frame, generating the map information of a next frame based on the stamp control information and the past map information stored, generating face space information which is information abstractly expressed for expressing the expression and the stamp in the interpolation frame based on the face mesh information and the map information generated by the stamp map generation unit, and predicting the interpolation frame based on the image texture and the face space information.
[0010] A third feature of the present invention is a program for causing a computer to function as a video decoding apparatus, the video decoding apparatus including a first decoding unit for decoding an image texture from the coded information of a key frame of a face video, a second decoding unit for decoding stamp control information including expression control information for controlling an expression in the face video, stamp identification information for identifying a stamp, and display area information indicating a display area of the stamp from the coded information of an interpolation frame, a face mesh generation unit for creating face mesh information based on the expression control information, an accumulation unit for accumulating map information defining a method of mapping the stamp to the interpolation frame, a stamp map generation unit for generating the map information of a next frame based on the stamp control information and the past map information stored in the accumulation unit, a face space generation unit for generating face space information which is information abstractly expressed for expressing the expression and the stamp in the interpolation frame based on the face mesh information and the map information generated by the stamp map generation unit, and an inter-frame prediction unit for predicting the interpolation frame based on the image texture and the face space information.
Advantages of the Invention
[0011] According to the present invention, it is possible to provide a video decoding apparatus, a video decoding method, and a program that can improve the encoding efficiency of face videos with emotion stamps added thereto.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Modes for Carrying Out the Invention
[0013] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that the components in the following embodiments can be appropriately replaced with existing components, etc., and various variations including combinations with other existing components are possible. Therefore, the description of the following embodiments does not limit the content of the invention described in the claims.
[0014] <First Embodiment> Hereinafter, with reference to FIGS. 1 to 5, a video decoding apparatus 100 according to the first embodiment of the present invention will be described.
[0015] As shown in FIG. 1, the video decoding apparatus 100 according to the present embodiment includes a first decoding unit 101, a second decoding unit 102, a face mesh generation unit 103, a stamp map generation unit 104, a storage unit 105, a face space generation unit 106, and an inter-frame prediction unit 107.
[0016] The first decoding unit 101 is configured to decode an image texture from the coded information of the key frame of the face video and output it to the inter-frame prediction unit 107.
[0017] The second decoding unit 102 is configured to decode expression control information and stamp control information from the coded information of the interpolation frame, output the expression control information to the face mesh generation unit 103, and output the stamp control information to the stamp map generation unit 104.
[0018] Here, the stamp control information is information for controlling a stamp (emotion stamp), and includes stamp identification information for identifying the type of the stamp and display area information indicating the display area of the stamp.
[0019] Here, the stamp identification information is a character string such as a Unicode or HTML code that corresponds one-to-one with the stamp.
[0020] FIGS. 2 to 4 show lists of the display areas of the stamps for each coordinate system. FIG. 2 shows a list of the display areas of the stamps on the face surface (face coordinate system), FIG. 3 shows a list of the display areas of the stamps on the face surface (face coordinate system), and FIG. 4 shows a list of the display areas of the stamps in the background (background coordinate system). As shown in FIGS. 2 to 4, the display areas of the stamps in the image are limited.
[0021] As shown in FIG. 2, in the face coordinate system, the display area information corresponding to the forehead is H1, HR1, HR2, HL1, HL2; the display area information corresponding to the eyes is PR1, PL1; the display area information corresponding to the area around the eyes is ER1, ER2, ER3, EL1, EL2, EL3; the display area information corresponding to the nose is N1, N2; the display area information corresponding to the cheeks is CR1, CR2, CL1, CL2; the display area information corresponding to the mouth is M1; and the display area information corresponding to the periphery of the face is O1, OR1, OL1.
[0022] Also, in the background coordinate system, the display area information is only B1 indicating the entire surface, and in the foreground coordinate system, the display area information is F1, F2, F3, F4 indicating the coordinates that divide the screen into four equal parts.
[0023] That is, the stamp control information is information that integrates the stamp identification information and the display area information for each coordinate system.
[0024] Also, the expression control information is information for controlling the expression in the face video. As shown in FIG. 5, for example, the expression control information includes "Eye Blinking" for controlling the blinking of the eyes, "Mouth Motion" for controlling the movement of the mouth, "Head Rotation" for controlling the rotation of the head, "Head Translation" for controlling the size of the head, and the like.
[0025] Here, the expression control information may be three-dimensional information or two-dimensional information.
[0026] As shown in FIG. 5, the face mesh generation unit 103 is configured to generate face mesh information based on the expression control information output from the second decoding unit 102.
[0027] The storage unit 105 is configured to store the map information generated by the stamp map generation unit 104 in time series.
[0028] Here, the map information is map information that defines a method for mapping stamps to interpolation frames. For example, the map information indicates at which position in the interpolation frame each stamp is to be displayed, and in the case of a moving stamp, what state the stamp corresponds to in its time-series movement.
[0029] The stamp map generation unit 104 is configured to generate map information for the next frame based on the stamp control information and the past map information stored in the storage unit 105.
[0030] Specifically, the stamp map generation unit 104 calculates the type and display area of the stamp from the stamp control information output from the second decoding unit 102.
[0031] When the type of the stamp is a moving stamp, the stamp map generation unit 104 calculates the position of the stamp in the interpolation frame from the past map information output from the storage unit 105.
[0032] Note that the stamp map generation unit 104 is configured to generate map information for all the stamps to be displayed based on the stamp identification information, the display area information, and the past map information.
[0033] The face space generation unit 106 is configured to generate face space information based on the face mesh information output from the face mesh generation unit 103 and the map information generated by the stamp map generation unit 104.
[0034] Here, the face space information is information abstractly expressed for expressing expressions and stamps in the interpolation frame.
[0035] Specifically, the face space information is not visual information such as a face image, but is composed of basic structures and features such as map information indicating the position and movement of the stamp, and face mesh information indicating the expressions of the face such as eyes, mouth, and head.
[0036] The inter-frame prediction unit 107 is configured to predict an interpolated frame based on the image texture output from the first decoding unit 101 and the face space information output from the face space generation unit 106.
[0037] Also, the inter-frame prediction unit 107 is configured to output decoded pixels of the key frame and the interpolated frame.
[0038] Specifically, as shown in FIG. 5, the inter-frame prediction unit 107 is configured to decode the pixels of the interpolated frame based on the face space information with reference to the image texture. In the example of FIG. 5, the description of the stamp is omitted for the sake of explanation.
[0039] According to such a video decoding apparatus 100, the encoding efficiency of the face video with a stamp (emotion stamp) can be improved.
[0040] The above-described video decoding apparatus 100 may be realized by a program that causes a computer to execute each function (each process).
Industrial Applicability
[0041] According to the present embodiment, for example, since an overall improvement in service quality can be realized in moving image communication, it is possible to contribute to Goal 9 of the Sustainable Development Goals (SDGs) led by the United Nations, "Build resilient infrastructure, promote sustainable industrialization and foster innovation."
Explanation of Signs
[0042] 100... Video decoding apparatus 101... First decoding unit 102... Second decoding unit 103... Face mesh generation unit 104... Stamp map generation unit 105... Storage unit 106... Face space generation unit 107... Inter-frame prediction unit
Claims
1. A first decoding unit that decodes an image texture from the code information of the key frame of the face video; A second decoding unit that decodes, from the code information of the interpolation frame, expression control information for controlling the expression in the face video, stamp identification information for identifying a stamp, and stamp control information including display area information indicating the display area of the stamp; A face mesh generation unit that creates face mesh information based on the expression control information; A storage unit that stores map information defining a method of mapping the stamp to the interpolation frame; A stamp map generation unit that generates the map information of the next frame based on the stamp control information and the past map information stored in the storage unit; A face space generation unit that generates face space information, which is information abstractly expressed to represent the expression and the stamp in the interpolation frame, based on the face mesh information and the map information generated by the stamp map generation unit; A video decoding apparatus comprising an inter-frame prediction unit that predicts the interpolation frame based on the image texture and the face space information.
2. The video decoding apparatus according to claim 1, wherein the expression control information is three-dimensional or two-dimensional information.
3. The video decoding apparatus according to claim 1, wherein the stamp map generation unit generates the map information of all the stamps to be displayed based on the stamp identification information, the display area information, and the past map information.
4. The video decoding apparatus according to claim 1, wherein the face space information is composed of the display area of the stamp and the face mesh information.
5. A step of decoding an image texture from the code information of the key frame of the face video; A step of decoding, from the code information of the interpolation frame, expression control information for controlling the expression in the face video, stamp identification information for identifying a stamp, and stamp control information including display area information indicating the display area of the stamp; A step of creating face mesh information based on the expression control information; A step of storing map information defining a method of mapping the stamp to the interpolation frame; A step of generating the map information of the next frame based on the stamp control information and the past map information stored; A step of generating face space information, which is information abstractly expressed to represent the expression and the stamp in the interpolation frame, based on the face mesh information and the map information generated by the stamp map generation unit; A video decoding method characterized by including a step of predicting the interpolation frame based on the image texture and the face space information. **Claim 6** A program that causes a computer to function as a video decoding device, wherein the video decoding device includes a first decoding unit that decodes an image texture from coded information of a key frame of a face video; a second decoding unit that decodes stamp control information including expression control information for controlling an expression in the face video, stamp identification information for identifying a stamp, and display area information indicating a display area of the stamp, from coded information of an interpolation frame; a face mesh generation unit that creates face mesh information based on the expression control information; a storage unit that stores map information defining a method of mapping the stamp to the interpolation frame; a stamp map generation unit that generates the map information of the next frame based on the stamp control information and the past map information stored in the storage unit; a face space generation unit that generates face space information, which is information abstractly expressed to represent the expression and the stamp in the interpolation frame, based on the face mesh information and the map information generated by the stamp map generation unit; and a frame interpolation unit that predicts the interpolation frame based on the image texture and the face space information.