Video decoding device, video decoding method, and program
The video decoding device and method address the inefficiencies in encoding face videos with emotional stamps by incorporating advanced units for decoding and predicting emotional content, resulting in improved encoding efficiency.
Patent Information
- Application Number
- PCT/JP2024/039124
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-11-01
- Publication Date
- 2025-06-26
AI Technical Summary
Conventional video decoding techniques fail to consider emotional stamps attached to faces during the decoding process, leading to inefficiencies in encoding face videos with emotional content.
A video decoding device and method that includes units for decoding image textures and stamp control information, generating face mesh and stamp map information, and predicting interpolation frames based on abstract face space information, thereby improving encoding efficiency for face videos with emotional stamps.
The proposed solution enhances the encoding efficiency of face videos with emotional stamps by accurately handling and predicting the display locations and movements of emotional stamps within the decoding process.
Smart Images

Figure JP2024039124_26062025_PF_FP_ABST
Abstract
Description
Video decoding device, video decoding method and program
[0001] The present invention relates to a video decoding device, a video decoding method, and a program.
[0002] Non-Patent Document 1 discloses GFVC (Generative Face Video Codec).
[0003] GFVC is specialized for transmitting facial images, and decodes the information that controls mouth movements, eye blinks, head rotations, and head movements that make up human facial expressions, projects it into a three-dimensional semantic facial space, and synthesizes it with separately decoded facial keyframes to generate interpolated facial frames, enabling efficient transmission of facial images.
[0004] Bolin Chen et al. "Interactive Face Video Coding: A Generative Compression Framework," in arXiv - CS - Computer Vision and Pattern Recognition, 2023.
[0005] In Non-Patent Document 1, an interpolated frame can be generated by combining a facial image with key frames and control information for each body part, such as the eyes, mouth, and head tilt, thereby achieving efficient transmission. However, there is another problem in that the emotion stamps attached to the face are not taken into consideration.
[0006] Furthermore, various emotion stamps have been proposed, and the display locations differ depending on the type. However, conventional decoding techniques have had the problem of not taking such display locations into consideration.
[0007] Therefore, the present invention has been made in consideration of the above-mentioned problems, and aims to provide a video decoding device, a video decoding method, and a program that can improve the coding efficiency of facial videos to which emotion stamps have been added.
[0008] a storage unit that stores map information that defines a method for mapping the stamps to the interpolated frame; a stamp map generation unit that generates the map information for a next frame based on the stamp control information and the past map information stored in the storage unit; a face space generation unit that generates face space information, which is information abstractly expressed to represent the face and the stamp in the interpolated frame, based on the face mesh information and the map information generated by the stamp map generation unit; and an inter-frame prediction unit that predicts the interpolated frame based on the image texture and the face space information.
[0009] a step of generating face mesh information based on the face mesh information; a step of storing map information that defines a method for mapping the stamps to the interpolated frame; a step of generating face space information, which is information abstractly expressed to represent the faces and the stamps in the interpolated frame, based on the face mesh information and the map information generated by the stamp map generation unit; and a step of predicting the interpolated frame based on the image texture and the face space information.
[0010] a storage unit that stores map information that defines a method for mapping the stamps to the interpolated frame; a stamp map generation unit that generates the map information for a next frame based on the stamp control information and the past map information stored in the storage unit; a face space generation unit that generates face space information, which is information abstractly expressed to represent the face and the stamp in the interpolated frame, based on the face mesh information and the map information generated by the stamp map generation unit; and an inter-frame prediction unit that predicts the interpolated frame based on the image texture and the face space information.
[0011] According to the present invention, it is possible to provide a video decoding device, a video decoding method, and a program that can improve the coding efficiency of a face video to which an emotion stamp is added.
[0012] Fig. 1 is a diagram showing an example of functional blocks of a video decoding device 100 according to an embodiment. Fig. 2 is a diagram showing an example of a list of stamp display areas. Fig. 3 is a diagram showing an example of a list of stamp display areas. Fig. 4 is a diagram showing an example of a list of stamp display areas. Fig. 5 is a diagram for explaining processing in the video decoding device 100 according to an embodiment.
[0013] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that components in the following embodiments can be replaced with existing components as appropriate, and various variations, including combinations with other existing components, are possible. Therefore, the description of the following embodiments does not limit the content of the invention described in the claims. <First Embodiment> Hereinafter, a video decoding device 100 according to a first embodiment of the present invention will be described with reference to FIGS. 1 to 5.
[0014] As shown in FIG. 1, the video decoding device 100 according to this embodiment includes a first decoding unit 101, a second decoding unit 102, a face mesh generation unit 103, a stamp map generation unit 104, a storage unit 105, a face space generation unit 106, and an inter-frame prediction unit 107.
[0015] The first decoding unit 101 is configured to decode image texture from the coded information of the face video key frame and output the decoded image texture to the inter-frame prediction unit 107 .
[0016] The second decoding unit 102 is configured to decode facial expression control information and stamp control information from the code information of the interpolated frame, output the facial expression control information to the face mesh generation unit 103, and output the stamp control information to the stamp map generation unit 104.
[0017] Here, the stamp control information is information for controlling a stamp (emotion stamp), and includes stamp identification information for identifying the type of stamp, and display area information indicating the display area of the stamp.
[0018] Here, the stamp identification information is a character string such as Unicode or HTML code that corresponds one-to-one with the stamp.
[0019] Figures 2 to 4 show a list of stamp display areas for each coordinate system. Figure 2 shows a list of stamp display areas on the face surface (face coordinate system), Figure 3 shows a list of stamp display areas on the face surface (face coordinate system), and Figure 4 shows a list of stamp display areas in the background (background coordinate system). As shown in Figures 2 to 4, the stamp display area within an image is limited.
[0020] As shown in FIG. 2, in the face coordinate system, display area information corresponding to the forehead is H1, HR1, HR2, HL1, HL2, display area information corresponding to the eyes is PR1, PL1, display area information corresponding to the area around the eyes is ER1, ER2, ER3, EL1, EL2, EL3, display area information corresponding to the nose is N1, N2, display area information corresponding to the cheeks is CR1, CR2, CL1, CL2, display area information corresponding to the mouth is M1, and display area information corresponding to the area around the face is O1, OR1, OL1.
[0021] In the background coordinate system, the display area information is only B1, which indicates the entire screen, and in the foreground coordinate system, the display area information is F1, F2, F3, and F4, which indicate coordinates that divide the screen into four equal parts.
[0022] That is, the stamp control information is information that integrates stamp identification information and display area information for each coordinate system.
[0023] Furthermore, the facial expression control information is information for controlling facial expressions in a facial image. As shown in Fig. 5, for example, the facial expression control information includes "Eye Blinking" for controlling eye blinking, "Mouth Motion" for controlling mouth movement, "Head Rotation" for controlling head rotation, and "Head Translation" for controlling head size.
[0024] Here, the facial expression control information may be three-dimensional information or two-dimensional information.
[0025] As shown in FIG. 5, the face mesh generating unit 103 is configured to generate face mesh information based on the facial expression control information output from the second decoding unit 102.
[0026] The storage unit 105 is configured to store the map information generated by the stamp map generation unit 104 in chronological order.
[0027] Here, the map information is information that defines a method for mapping stamps onto an interpolated frame, such as information indicating which stamp is displayed at which position in the interpolated frame, and, in the case of a moving stamp, which state the stamp is currently in as it moves over time.
[0028] The stamp map generating unit 104 is configured to generate map information for the next frame based on the stamp control information and past map information stored in the storage unit 105 .
[0029] Specifically, the stamp map generating unit 104 calculates the type and display area of the stamp from the stamp control information output from the second decoding unit 102 .
[0030] If the type of stamp is a moving stamp, the stamp map generating unit 104 calculates the position of the stamp in the interpolated frame from the past map information output from the storage unit 105 .
[0031] The stamp map generating unit 104 is configured to generate map information for all stamps to be displayed based on the stamp identification information, display area information, and past map information.
[0032] The face space generation unit 106 is configured to generate face space information based on the face mesh information output from the face mesh generation unit 103 and the map information generated by the stamp map generation unit 104 .
[0033] Here, the face spatial information is information that is abstractly expressed to represent facial expressions and stamps in the interpolated frames.
[0034] Specifically, facial space information is not visual information such as facial images, but is composed of basic structures and features such as map information showing the position and movement of stamps, and facial mesh information showing facial expressions such as the eyes, mouth, and head.
[0035] The inter-frame prediction unit 107 is configured to predict an interpolated frame based on the image texture output from the first decoding unit 101 and the face space information output from the face space generation unit 106 .
[0036] The inter-frame prediction unit 107 is configured to output decoded pixels of the key frames and interpolated frames.
[0037] Specifically, as shown in Fig. 5, the inter-frame prediction unit 107 is configured to decode pixels of an interpolated frame based on face space information with reference to image texture. Note that stamps are omitted from the example of Fig. 5 for ease of explanation.
[0038] According to the video decoding device 100, it is possible to improve the coding efficiency of face video to which a stamp (emotion stamp) is added.
[0039] The above-described video decoding device 100 may be realized as a program that causes a computer to execute each function (each step).
[0040] According to this embodiment, for example, it is possible to improve the overall service quality in video communication, which makes it possible to contribute to Goal 9 of the Sustainable Development Goals (SDGs) led by the United Nations, which is to "Develop resilient infrastructure, promote sustainable industrialization and foster innovation."
[0041] REFERENCE SIGNS LIST 100... Video decoding device 101... First decoding unit 102... Second decoding unit 103... Face mesh generating unit 104... Stamp map generating unit 105... Storage unit 106... Face space generating unit 107... Inter-frame prediction unit
Claims
1. A video decoding device comprising: a first decoding unit that decodes image texture from code information of a key frame of a facial image; a second decoding unit that decodes, from code information of an interpolated frame, facial expression control information that controls facial expressions in the facial image, and stamp control information including stamp identification information that identifies a stamp and display area information that indicates the display area of the stamp; a facial mesh generation unit that creates facial mesh information based on the facial expression control information; a storage unit that stores map information that defines a method of mapping the stamp to the interpolated frame; a stamp map generation unit that generates the map information for the next frame based on the stamp control information and the past map information stored in the storage unit; a facial space generation unit that generates facial space information, which is information abstractly expressed to express the facial expression and the stamp in the interpolated frame, based on the facial mesh information and the map information generated by the stamp map generation unit; and an inter-frame prediction unit that predicts the interpolated frame based on the image texture and the facial space information.
2. A video decoding device according to claim 1, wherein the facial expression control information is three-dimensional or two-dimensional information.
3. The video decoding device according to claim 1, characterized in that the stamp map generation unit generates the map information for all stamps to be displayed based on the stamp identification information, the display area information and the past map information.
4. A video decoding device according to claim 1, characterized in that the face space information is composed of a display area of the stamp and the face mesh information.
5. A video decoding method comprising the steps of: decoding an image texture from code information of a key frame of a facial image; decoding, from code information of an interpolated frame, facial expression control information for controlling facial expressions in the facial image, and stamp control information including stamp identification information for identifying a stamp and display area information indicating a display area of the stamp; creating facial mesh information based on the facial expression control information; accumulating map information defining a method for mapping the stamp to the interpolated frame; generating the map information for a next frame based on the stamp control information and the accumulated past map information; generating facial space information, which is information abstractly expressed to represent the facial expression and the stamp in the interpolated frame, based on the facial mesh information and the map information generated by the stamp map generation unit; and predicting the interpolated frame based on the image texture and the facial space information.
6. A program for causing a computer to function as a video decoding device, the video decoding device comprising: a first decoding unit that decodes image texture from coding information of a key frame of a facial image; a second decoding unit that decodes, from coding information of an interpolated frame, facial expression control information that controls facial expressions in the facial image, and stamp control information including stamp identification information that identifies a stamp and display area information that indicates a display area of the stamp; a facial mesh generation unit that creates facial mesh information based on the facial expression control information; a storage unit that stores map information that defines a method of mapping the stamp to the interpolated frame; a stamp map generation unit that generates the map information for the next frame based on the stamp control information and the past map information stored in the storage unit; a facial space generation unit that generates facial space information, which is information abstractly expressed to express the facial expressions and the stamp in the interpolated frame, based on the facial mesh information and the map information generated by the stamp map generation unit; and an inter-frame prediction unit that predicts the interpolated frame based on the image texture and the facial space information.
Citation Information
Patent Citations
Face video coding method and device, and face video decoding method and device
CN114422795A
Decoding device, encoding device, decoding method, and encoding method
WO2023195426A1