Facial video encoding method, decoding method and device
By calculating the information difference value in facial video encoding and decoding and dynamically updating the reference frame, the problem of poor facial video reconstruction quality when facial movement and expression changes are large, and the encoding and decoding quality and reconstruction effect are improved.
Patent Information
- Application Number
- CN202210085252.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-25
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-01-25
AI Technical Summary
In the existing video encoding and decoding algorithm, in facial videos, when facial movements and expressions change greatly, the difference between the reconstruction frame and the original video frame is large, resulting in poor quality of facial video reconstruction.
By calculating the information difference value between the current face video frame and the initial reference face video frame, if the difference value is greater than the preset threshold, the current face video frame is used as a new reference face video frame for encoding other video frames.
The quality of facial video encoding and decoding is improved, and the problem of poor facial video reconstruction quality caused by large information difference between the initial reference facial video frame and the facial video frame to be encoded is reduced, so that a higher quality reconstruction facial video frame is obtained.
Smart Images

Figure CN114205585B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of computer technologies, and in particular, to a facial video encoding method, a decoding method, and an apparatus therefor. Background Art
[0002] With the continuous development of video encoding and decoding technologies, in order to improve video encoding and decoding performance, a variety of video encoding and decoding algorithms have emerged. For example: traditional video encoding and decoding algorithms using methods such as block-based motion estimation and discrete cosine transform; end-to-end video encoding and decoding algorithms based on deep learning, and so on.
[0003] For facial videos, when the changes in information such as facial movements and expressions are small, using the above existing video encoding and decoding algorithms for video frame reconstruction results in better quality of the reconstructed frames. However, when the changes in information such as facial movements and expressions are large, such as when there is a large movement or rotation of the head, the difference between the reconstructed frame and the original video frame is large, and the quality of facial video reconstruction is poor. Summary of the Invention
[0004] In view of this, embodiments of the present application provide a facial video encoding method, a decoding method, and an apparatus therefor to at least partially solve the above problems.
[0005] According to a first aspect of embodiments of the present application, a facial video encoding method is provided, including:
[0006] Calculating an information difference value between a current facial video frame to be encoded and an initial reference facial video frame in a reference frame list, where the information difference value represents the degree of difference between the information included in the current facial video frame and the information included in the initial reference facial video frame;
[0007] If the information difference value is greater than a preset threshold, using the current facial video frame as a newly added reference facial video frame, where the newly added reference facial video frame is used to encode other video frames to obtain a first facial video bitstream.
[0008] According to a second aspect of embodiments of the present application, a facial video decoding method is provided, including:
[0009] Obtaining a facial video bitstream;
[0010] Decoding the facial video bitstream to obtain target driving information of a target facial video frame to be encoded and target identification information indicating whether the target facial video is a newly added reference facial video frame;
[0011] If the target facial video frame is not a newly added reference facial video frame, respectively obtaining reference driving information of multiple reference facial video frames in a reference frame list;
[0012] Calculate the information difference value between the current facial video frame and each reference facial video frame based on the target driving information and each reference driving information; the information difference value characterizes the difference degree between the information contained in the current facial video frame and the information contained in the reference facial video frame.
[0013] Use the reference facial video frame corresponding to the minimum information difference value as the target reference facial video frame, and obtain a reconstructed facial video frame based on the target reference facial video frame and the target driving information.
[0014] According to the third aspect of the embodiments of the present application, there is provided a facial video encoding device, including:
[0015] An information difference value calculation module, configured to calculate the information difference value between the current facial video frame to be encoded and the initial reference facial video frame in the reference frame list, and the information difference value characterizes the difference degree between the information contained in the current facial video frame and the information contained in the initial reference facial video frame.
[0016] A first encoding module, configured to, if the information difference value is greater than a preset threshold, use the current facial video frame as a new reference facial video frame, and the new reference facial video frame is used to encode other video frames to obtain a first facial video bitstream.
[0017] According to the fourth aspect of the embodiments of the present application, there is provided an electronic device, including: a processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete communication with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the facial video encoding method as described in the first aspect, or the operations corresponding to the facial video decoding method as described in the second aspect.
[0018] According to the fifth aspect of the embodiments of the present application, there is provided a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the facial video encoding method as described in the first aspect, or the facial video decoding method as described in the second aspect.
[0019] According to the sixth aspect of the embodiments of the present application, there is provided a computer program product, including computer instructions, and the computer instructions instruct a computing device to perform the operations corresponding to the facial video encoding method as described in the first aspect, or the operations corresponding to the facial video decoding method as described in the second aspect.
[0020] According to the facial video encoding method, decoding method and device provided by the embodiments of the present application, during the encoding process of the current facial video frame, the current facial video frame is not directly encoded or decoded based on the initial reference facial video frames in the reference frame list. Instead, the information difference value between the current facial video frame and the initial reference facial video frames is first calculated. If the information difference value is large (greater than the preset threshold), the current facial video frame is no longer encoded or decoded based on the initial reference facial video frames. Instead, the current facial video frame is used as a new (added) reference facial video frame to encode other video frames, obtaining the first facial video bitstream. In the embodiments of the present application, the actual reference facial video frames used in the encoding and decoding process are determined based on the information difference degree between the current facial video frame and the initial reference facial video frames. When the difference degree is large, the current facial video frame is used as a new reference facial video frame, rather than relying on the initial reference facial video frames for encoding and decoding. Therefore, the quality of facial video encoding and decoding can be improved, and a higher-quality reconstructed facial video frame can be obtained. Description of the Drawings
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0022] Figure 1 It is a schematic framework diagram of an encoding and decoding method based on depth video;
[0023] Figure 2 It is a flowchart of the steps of a facial video encoding method according to Embodiment 1 of the present application;
[0024] Figure 3 For Figure 2 It is a schematic diagram of a scenario example in the shown embodiment;
[0025] Figure 4 It is a flowchart of the steps of a facial video encoding method according to Embodiment 3 of the present application;
[0026] Figure 5 It is a flowchart of the steps of a facial video decoding method according to Embodiment 4 of the present application;
[0027] Figure 6 For Figure 5 It is a schematic diagram of a scenario example in the shown embodiment;
[0028] Figure 7 It is a structural block diagram of a facial video encoding device according to Embodiment 5 of the present application;
[0029] Figure 8 It is a structural block diagram of a facial video decoding device according to Embodiment 6 of the present application;
[0030] Figure 9 It is a schematic structural diagram of an electronic device according to Embodiment 7 of the present application. Specific embodiments
[0031] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art shall fall within the protection scope of the embodiments of the present application.
[0032] See Figure 1 , Figure 1 It is a schematic framework diagram of an encoding and decoding method based on depth video generation. The main principle of this method is to deform the reference frame based on the motion of the frame to be encoded to obtain the reconstructed frame corresponding to the frame to be encoded. The following will describe the basic framework of the encoding and decoding method based on depth video generation in conjunction with Figure 1 to illustrate:
[0033] In the first step, the encoding stage, the encoder uses a key point extractor to extract the target key point information of the target facial video frame to be encoded and encodes the target key point information; at the same time, a traditional image encoding method (such as VVC, HEVC, etc.) is used to encode the reference facial video frame.
[0034] In the second step, the decoding stage, the motion estimation module in the decoder extracts the reference key point information of the reference facial video frame through the key point extractor; and performs dense motion estimation based on the reference key point information and the target key point information to obtain a dense motion estimation map and an occlusion map, where the dense motion estimation map represents the relative motion relationship between the target facial video frame and the reference facial video frame in the feature domain represented by the key point information; the occlusion map represents the degree of occlusion of each pixel point in the target facial video frame.
[0035] In the third step, the decoding stage, the generation module in the decoder performs deformation processing on the reference facial video frame based on the dense motion estimation map to obtain a deformation processing result, and then multiplies the deformation processing result by the occlusion map to output the reconstructed facial video frame.
[0036] Figure 1In the method shown above, the reference frame is usually a fixed facial video frame, such as the first frame of the facial video, etc. When the change amplitude of the facial motion, expression and other information between the facial video frames is small, the above method is used for video frame reconstruction, and the quality of the reconstructed frame is good. However, when the change amplitude of the facial motion, expression and other information is large, such as when there is a large movement or rotation of the head, if the current facial video frame is still encoded and decoded based on the above fixed reference frame, the difference between the reconstructed frame and the original video frame is large, and the quality of the facial video reconstruction is poor.
[0037] In the embodiments of the present application, based on the information difference degree between the current facial video frame and the initial reference facial video frame, the reference facial video frame actually used in the encoding and decoding process is determined. When the difference degree is large, the current facial video frame is used as the new reference facial video frame, and the encoding and decoding no longer rely on the initial reference facial video frame, which can improve the quality of facial video encoding and decoding, and thus obtain a higher-quality reconstructed facial video frame.
[0038] The following further describes the specific implementation of the embodiments of the present application with reference to the accompanying drawings of the embodiments of the present application.
[0039] Embodiment 1
[0040] Refer to Figure 2 , Figure 2 which is a step flowchart of a facial video encoding method according to Embodiment 1 of the present application. Specifically, the facial video encoding method provided in this embodiment includes the following steps:
[0041] Step 202, calculate the information difference value between the current facial video frame to be encoded and the initial reference facial video frame in the reference frame list.
[0042] Among them, the information difference value characterizes the difference degree between the information contained in the current facial video frame and the information contained in the initial reference facial video frame.
[0043] In the embodiments of the present application, no specific calculation method is limited for calculating the above information difference value. For example: the mean square error can be calculated based on the pixel values of each pixel point in the current facial video frame and the pixel values of each pixel point in the initial reference facial video frame, so as to obtain the pixel-level information difference value; or the feature extraction can be performed on the current facial video frame and the initial reference facial video frame respectively to obtain: the feature that can characterize the key feature information of the current facial video frame, and the feature that can characterize the key feature information of the initial reference facial video frame, and then based on the above two features, the mean square error (MSE) calculation is performed to obtain the information difference value in the feature domain, and so on.
[0044] Step 204, if the information difference value is greater than a preset threshold, use the current facial video frame as a newly added reference facial video frame, and the newly added reference facial video frame is used to encode other video frames to obtain a first facial video bitstream.
[0045] If the information difference value is greater than the preset threshold, it indicates that the information difference between the current facial video frame and the initial reference facial video frame is relatively large. At this time, if the initial reference facial video frame is still used as the reference frame to perform encoding and decoding operations on the current facial video frame, the information difference between the reconstructed facial video frame obtained and the current facial video frame is relatively large, and it is difficult to guarantee the reconstruction quality of the facial video frame.
[0046] Therefore, in the embodiments of the present application, when it is determined that the information difference between the current facial video frame and the initial reference facial video frame is relatively large, the current facial video frame is used as a newly added reference facial video frame to encode other video frames to obtain a first facial video bitstream. This can reduce the problem of poor facial video reconstruction quality caused by encoding and decoding the current facial video frame based on the initial reference facial video frame when the information difference between the initial reference facial video frame and the facial video frame to be encoded is relatively large, improve the quality of facial video reconstruction, and thus obtain a reconstructed facial video frame of higher quality.
[0047] Specifically, in the following manner, for the current facial video frame used as the newly added reference facial video frame, it can be encoded in the following manner:
[0048] Use the newly added reference facial video frame as a key frame, and independently encode the newly added reference facial video frame, encoding the newly added reference facial video frame with relatively small quantization distortion, and the complete data of the newly added reference facial video frame is retained during the encoding process. Correspondingly, during decoding, only the newly added reference facial video frame itself is required to complete the decoding process.
[0049] Subsequently, when encoding other facial video frames after the current facial video, it can be performed based on the above-mentioned newly added reference facial video frame (current facial video) to improve the encoding quality, and thus obtain a reconstructed facial video frame of higher quality.
[0050] See Figure 3 , Figure 3 which is a schematic diagram of the scenario corresponding to Embodiment 1 of the present application. Hereinafter, with reference to Figure 3 the schematic diagram shown, a specific scenario example will be used to illustrate the embodiments of the present application:
[0051] Obtain the current facial video frame a to be encoded and the initial reference facial video frame a0 in the reference frame list (the number of initial reference facial video frames included in the reference frame list can be 1 or multiple, Figure 3Only one initial reference facial video frame is taken as an example for illustration, Figure 3 The illustrated scenario example does not limit the facial video encoding provided by the embodiments of the present application). Calculate the information difference value between a and a0. When the calculated information difference value is large (greater than the preset threshold), then use a as a new reference facial video frame to encode other facial video frames to obtain the first facial video bitstream.
[0052] In the embodiments of the present application, during the encoding process of the current facial video frame, the current facial video frame is not directly encoded and decoded based on the initial reference facial video frames in the reference frame list. Instead, first calculate the information difference value between the current facial video frame and the initial reference facial video frame. If the information difference value is large (greater than the preset threshold), then no longer encode and decode the current facial video frame based on the initial reference facial video frame. Instead, use the current facial video frame as a new (added) reference facial video frame to encode other video frames to obtain the first facial video bitstream. In the embodiments of the present application, the actual reference facial video frame used in the encoding and decoding process is determined based on the information difference degree between the current facial video frame and the initial reference facial video frame. When the difference degree is large, then use the current facial video frame as a new reference facial video frame and no longer rely on the initial reference facial video frame for encoding and decoding. Therefore, the quality of facial video encoding and decoding can be improved, and then a higher-quality reconstructed facial video frame can be obtained.
[0053] The facial video encoding method provided in Embodiment 1 of the present application can be executed by a video encoding end (encoder) for encoding a facial video file to compress the digital bandwidth of the facial video file. It can be applied to a variety of different scenarios, such as: storage and streaming of conventional facial-related video games. Specifically: the game video frames can be encoded by the facial video encoding method provided in the embodiments of the present application to form a corresponding video bitstream for storage and transmission in a video stream service or other similar applications; another example: low-latency scenarios such as video conferencing and video live streaming. Specifically: the facial video data collected by a video capture device can be encoded by the facial video encoding method provided in the embodiments of the present application to form a corresponding video bitstream and sent to a conference terminal, and the conference terminal decodes the video bitstream to obtain the corresponding facial video picture; yet another example: virtual reality scenarios. The facial video data collected by a video capture device can be encoded by the facial video encoding method provided in the embodiments of the present application to form a corresponding video bitstream and sent to virtual reality-related devices (such as VR virtual glasses, etc.). The VR device decodes the video bitstream to obtain the corresponding facial video picture and realizes the corresponding VR function based on the facial video picture, and so on.
[0054] Embodiment 2
[0055] Based on the solution of the first embodiment above, optionally, in an embodiment of the present application, after the above step 204, the following may further be included:
[0056] Add the newly added reference facial video frame to the reference frame list to update the reference frame list.
[0057] In the embodiment of the present application, the number of reference facial video frames included in the reference frame list is constant (which can be set according to the actual situation). Therefore, when the current facial video frame is determined as the newly added reference facial video frame, the reference frame list needs to be updated to obtain the updated reference frame list, so as to perform subsequent encoding and decoding operations on the facial video frames.
[0058] Specifically, if the number of initial reference facial video frames is 1, the method for updating the reference frame list may be:
[0059] Add the newly added reference facial video frame to the reference frame list and delete the initial reference facial video frame to obtain the updated reference frame list.
[0060] If the number of initial reference facial video frames is multiple, the method for updating the reference frame list may be: add the newly added reference facial video frame to the reference frame list and delete one initial reference facial video frame from all the initial reference facial video frames to obtain the updated reference frame list. Specifically, the initial reference facial video frame may be deleted according to the information difference value between each initial reference facial video and the newly added reference facial video frame. For example, the initial reference facial video with the largest information difference value may be deleted, or the initial reference facial video frame with the earliest timestamp may be deleted according to the timestamps of each initial reference facial video, and so on.
[0061] In addition, the facial video encoding may further include:
[0062] If the information difference value is less than or equal to the preset threshold, encode the current facial video frame based on the initial reference facial video frame to obtain the second facial video bitstream.
[0063] In the embodiment of the present application, the specific manner of encoding the current facial video frame based on the initial reference facial video frame to obtain the second facial video bitstream and decoding the second facial video bitstream based on the initial reference facial video frame is not limited, and any existing method for encoding and decoding the current frame with the help of the reference frame may be adopted, which will not be elaborated here.
[0064] In the second embodiment above, during the encoding of the current facial video frame, first calculate the information difference value between the current facial video frame and the initial reference facial video frame. If the information difference value is large (greater than the preset threshold), then use the current facial video frame as a new (added) reference facial video frame for encoding and decoding operations, so as to obtain the reconstructed facial video frame corresponding to the current facial video frame; if the information difference value is small, then based on the initial reference facial video frame, perform encoding and decoding operations on the current facial video frame, so as to obtain the corresponding reconstructed facial video frame. Therefore, it is possible to avoid the problem of poor quality of facial video reconstruction caused by still performing encoding and decoding operations on the current facial video frame based on the initial reference facial video frame when the information difference between the initial reference facial video frame and the current facial video frame is large, improve the quality of facial video reconstruction, and thus obtain a higher-quality reconstructed facial video frame.
[0065] The facial video encoding method of this embodiment can be executed by any suitable electronic device with data capabilities, including but not limited to: servers, PCs, etc.
[0066] Refer to Figure 4 , Figure 4 FIG. is a flowchart of the steps of a facial video encoding method according to Embodiment 3 of the present application. Specifically, in this embodiment, the number of initial reference facial video frames in the reference frame list is multiple, where the number of initial reference facial video frames can be set in advance according to actual needs, and no specific limit is imposed on the specific numerical value here. The provided facial video encoding method includes the following steps:
[0067] Step 402, calculate the information difference value between the current facial video frame to be encoded and each initial reference facial video frame in the reference frame list respectively. If there is a candidate information difference value less than or equal to the preset threshold among all the information difference values, then execute Step 404; if all the information difference values are greater than the preset threshold, then execute Step 408.
[0068] Further, the following two different methods can be used to calculate the information difference value between the current facial video frame and the initial reference facial video frame:
[0069] The first method: Based on the pixel values of each pixel point in the current facial video frame to be encoded and the pixel values of each pixel point in the initial reference facial video frame, perform mean-square error (MSE) calculation to obtain the information difference value between the current facial video frame and the initial reference facial video frame.
[0070] The second method is: extracting features from the current facial video frame to be encoded to obtain compact features of the current facial video frame; extracting features from the initial reference facial video frame in the reference frame list to obtain compact features of the initial reference facial video frame; performing mean square error calculation based on the compact features of the current facial video frame and the compact features of the initial reference facial video frame to obtain the information difference value between the current facial video frame and the initial reference facial video frame in the reference frame list. The compact features of the current facial video frame represent key feature information in the facial video frame; the compact features of the initial reference facial video frame represent key feature information in the initial reference facial video frame.
[0071] Furthermore, when feature extraction is performed on facial video frames, it can be performed based on a deep learning model. Specifically: the current facial video frame to be encoded can be input into a pre-trained feature extraction model, so that the feature extraction model outputs compact features of the current facial video frame; and the initial reference facial video frame in the reference frame list can be input into the feature extraction model, so that the feature extraction model outputs compact features of the initial reference facial video frame.
[0072] In the embodiment of the present application, the specific structure and parameters of the feature extraction model are not limited and can be set according to actual conditions.
[0073] Comparing the above two different ways of calculating the information difference value, the first way is to calculate the information difference value from the pixel level, so the calculation result is more accurate. The second way is to calculate the information difference value from the perspective of the extracted features, that is, to calculate the information difference value from the perspective of the feature domain. Since the features are extracted from the facial video frames, the facial video frames are downsampled, so the calculation amount is small and the calculation efficiency is high.
[0074] Step 404: determine the information difference value with the minimum value from the candidate information difference values as the target information difference value.
[0075] Step 406: Based on the initial reference facial video frame corresponding to the target information difference value, the current facial video frame is encoded to obtain a second facial video bitstream. At this point, the encoding process ends.
[0076] In an embodiment of the present application, when there are multiple candidate information difference values that are less than or equal to a preset threshold, in order to improve the quality of facial video encoding and decoding, an initial reference facial video frame corresponding to the minimum information difference value (target information difference value) can be selected from multiple initial reference facial video frames as the actual reference frame, and then based on the actual reference frame, the current facial video frame is encoded to obtain a second facial video bitstream.
[0077] In the embodiments of the present application, there is no limitation on the specific manner of encoding the current facial video frame to obtain the second facial video bitstream based on the initial reference facial video frame corresponding to the target information difference value, and decoding the second facial video bitstream based on the initial reference facial video frame corresponding to the target information difference value. Any existing method for encoding and decoding the current frame with the help of a reference frame can be adopted, which will not be elaborated here.
[0078] Step 408: Use the current facial video frame as a newly added reference facial video frame, and the newly added reference facial video frame is used to encode other video frames to obtain the first facial video bitstream.
[0079] If all information difference values are greater than the preset threshold, it indicates that the information differences between the current facial video frame and each initial reference facial video frame are relatively large. At this time, using any one of the initial reference facial video frames as the reference frame for encoding and decoding the current facial video frame, the information difference between the reconstructed facial video frame and the current facial video frame is relatively large, and it is difficult to guarantee the reconstruction quality of the facial video frame.
[0080] Therefore, in the embodiments of the present application, when it is determined that the information differences between the current facial video frame and all initial reference facial video frames are relatively large, the current facial video frame is used as a newly added reference facial video frame to encode other video frames to obtain the first facial video bitstream. This can improve the quality of facial video reconstruction, and thus obtain a reconstructed facial video frame with higher quality.
[0081] Step 410: Add the newly added reference facial video frame to the reference frame list, and delete the initial reference facial video frame with the earliest timestamp to obtain the updated reference frame list.
[0082] After updating the initial reference facial video frames in the reference frame list, when encoding other facial video frames after the current facial video frame, it can be performed based on the reference facial video frames in the above-mentioned updated reference frame list to improve the encoding quality, and thus obtain a reconstructed facial video frame with higher quality.
[0083] In the embodiments of the present application, there is no limitation on the execution order of step 408 and step 410. Step 408 can be executed first, and then step 410; step 410 can also be executed first, and then step 408; step 408 and step 410 can also be executed in parallel. In the embodiments of the present application, only the example of executing step 408 first and then step 410 is used for illustration, which does not constitute a limitation on the embodiments of the present application.
[0084] In the embodiment of the present application, during the encoding of the current facial video frame, first calculate the information difference value between the current facial video frame and each initial reference facial video frame. If all the information difference values are large (greater than the preset threshold), then use the current facial video frame as a new (added) reference facial video frame for encoding and decoding operations; if there are smaller information difference values, then based on the initial reference facial video frame corresponding to the minimum difference value, perform encoding and decoding operations on the current facial video frame to obtain the corresponding reconstructed facial video frame. Therefore, it is possible to avoid the problem of poor quality of facial video reconstruction caused by still performing encoding and decoding operations on the current facial video frame based on the initial reference facial video frame when the information difference between the initial reference facial video frame and the current facial video frame is large; in addition, when there are smaller information difference values, the encoding and decoding operations on the current facial video frame are based on the initial reference facial video frame corresponding to the minimum difference value. Since the difference between the initial reference facial video frame corresponding to the minimum difference value and the previous facial video frame is the smallest, therefore, based on the initial reference facial video frame with the smallest difference for subsequent encoding and decoding operations, the quality of facial video reconstruction is relatively high, and thus the quality of the obtained reconstructed facial video frame is also relatively high.
[0085] Embodiment 4
[0086] Refer to Figure 5 , Figure 5 which is a flowchart of the steps of a facial video decoding method according to Embodiment 4 of the present application. Specifically, the facial video decoding method provided in this embodiment includes the following steps:
[0087] Step 502, obtain the facial video bitstream.
[0088] Step 504, decode the facial video bitstream to obtain the target driving information of the target facial video frame to be encoded and the target identification information indicating whether the target facial video is a newly added reference facial video frame.
[0089] The target driving information may be information representing the key feature information of the target facial video frame. Among them, the key feature information may include: facial feature position information, pose information, expression information, etc. In the embodiment of the present application, the specific form of the target driving information is not limited. For example: it may be the explicitly represented key point information extracted by a facial key point extractor, or it may be a compact feature matrix (vector) with a smaller data volume and richer representation information extracted by a feature extraction model, etc.
[0090] The target identification information in this step is generated by the encoding end during the encoding of the target facial video frame and sent to the decoding end. The encoding end can determine whether the target facial video frame is a newly added reference facial video frame based on the information difference value between the target facial video frame and the initial reference facial video frame in the reference frame list. Specifically, if all the information difference values are greater than the preset threshold, it is determined that the target facial video frame is a newly added reference facial video frame, and this newly added reference facial video frame can be used for encoding other video frames; if there is an information difference value less than or equal to the preset threshold, it is determined that the target facial video frame is a non-newly added reference facial video frame, and the target facial video frame can be encoded based on the initial reference facial video frame in the reference frame list to obtain a facial video bitstream and sent to the decoding end.
[0091] Step 506, if the target facial video frame is a non-newly added reference facial video frame, obtain the reference driving information of multiple reference facial video frames in the reference frame list respectively.
[0092] The reference driving information can be information representing the key feature information of the reference facial video frame. Among them, the key feature information can include: facial feature position information, pose information, expression information, and so on. In the embodiments of the present application, the specific form of the reference driving information is not limited. For example, it can be the explicitly represented key point information extracted by a facial key point extractor, or it can be a compact feature matrix (vector) with a smaller amount of implicitly represented data and richer representation information extracted by a feature extraction model, and so on.
[0093] When the target facial video frame is a non-newly added reference facial video frame, it indicates that the target facial video frame is encoded based on the initial reference facial video frame in the reference frame list. Since there are multiple initial reference facial video frames in the reference frame list, it is necessary to determine the actual reference facial video frame from them: the target reference facial video frame.
[0094] In addition, when the target facial video frame is a newly added reference facial video frame, it indicates that the target facial video frame is independently encoded, that is, the target facial video frame is encoded with relatively small quantization distortion, and the complete data of the target facial video frame is retained during the encoding process. Correspondingly, during decoding, only the decoding method corresponding to the encoding method needs to be used for decoding.
[0095] Step 508, based on the target driving information and each reference driving information, calculate the information difference value between the target facial video frame and each reference facial video frame.
[0096] Among them, the information difference value represents the degree of difference between the information contained in the target facial video frame and the information contained in the reference facial video frame.
[0097] Further, in the embodiments of the present application, the information difference value can be calculated in the following manner:
[0098] Based on the target driving information and each reference driving information, mean square error calculation is performed to obtain the information difference value between the target facial video frame and each reference facial video frame.
[0099] Step 510: Use the reference facial video frame corresponding to the minimum information difference value as the target reference facial video frame, and based on the target reference facial video frame and the target driving information, obtain the reconstructed facial video frame.
[0100] Specifically, in the embodiments of the present application, the specific decoding form for obtaining the reconstructed facial video frame is not limited, and it can be based on Figure 4 the encoding method used to obtain the facial video bitstream in step 406 in the illustrated embodiment, and the corresponding decoding method is used for decoding to obtain the reconstructed facial video frame.
[0101] Refer to Figure 6 , Figure 6 which is a schematic diagram of the scenario corresponding to Embodiment 1 of the present application. Hereinafter, with reference to the Figure 6 schematic diagram shown, a specific scenario example will be used to illustrate the embodiments of the present application:
[0102] Obtain the facial video bitstream; decode the facial video bitstream to obtain the target driving information D of the target facial video frame and the target identification information, where the target identification information indicates that the target facial video frame is a non-new reference facial video frame; obtain the reference driving information of each reference facial video frame in the reference frame list: D1, D2, and D3 (the number of reference facial video frames included in the reference frame list can be any integer greater than 1, Figure 6 only three reference facial video frames are taken as an example for illustration here); based on D, D1, D2, and D3, calculate the information difference values d1 (the difference between D and D1), d2 (the difference between D and D2), and d3 (the difference between D and D3) between the target facial video frame and each reference facial video frame, and use the reference facial video frame D1 corresponding to the minimum information difference value d1 as the target reference facial video frame, and based on the target reference facial video frame and the target driving information, obtain the reconstructed facial video frame.
[0103] In the embodiments of the present application, during the decoding process, if it is determined that the target facial video frame is not a newly added reference facial video frame, based on the target driving information of the target facial video frame and the reference driving information of each reference facial video frame, the information difference value between the current facial video frame and each reference facial video frame is calculated. Then, from the multiple reference facial video frames in the reference frame list, the reference facial video frame corresponding to the minimum information difference value is used as the target reference facial video frame to obtain the reconstructed facial video frame. Since the finally determined target reference facial video frame is the reference facial video frame with the smallest information difference from the current facial video frame, therefore, based on the above target reference facial video frame for facial video frame reconstruction, the quality of the reconstructed facial video frame is also the highest, improving the reconstruction quality of the facial video frame.
[0104] Embodiment 5
[0105] See Figure 7 , Figure 7 is a structural block diagram of a facial video encoding device according to Embodiment 5 of the present application. The facial video encoding device provided by the embodiments of the present application includes:
[0106] The first information difference value calculation module 702 is configured to calculate the information difference value between the current facial video frame to be encoded and the initial reference facial video frame in the reference frame list, and the information difference value represents the difference degree between the information included in the current facial video frame and the information included in the initial reference facial video frame;
[0107] The first encoding module 704 is configured to, if the information difference value is greater than a preset threshold, use the current facial video frame as a newly added reference facial video frame, and the newly added reference facial video frame is used to encode other video frames to obtain the first facial video bitstream.
[0108] Optionally, in some embodiments, the first information difference value calculation module 702 is specifically configured to:
[0109] Based on the pixel values of each pixel point in the current facial video frame to be encoded and the pixel values of each pixel point in the initial reference facial video frame, mean square error calculation is performed to obtain the information difference value between the current facial video frame and the initial reference facial video frame in the reference frame list.
[0110] Optionally, in some embodiments, the first information difference value calculation module 702 is specifically configured to:
[0111] Extract features from the current facial video frame to be encoded to obtain the compact features of the current facial video frame, and the compact features of the current facial video frame represent the key feature information in the facial video frame;
[0112] Feature extraction is performed on the initial reference facial video frame in the reference frame list to obtain the compact feature of the initial reference facial video frame, and the compact feature of the initial reference facial video frame characterizes the key feature information in the initial reference facial video frame;
[0113] Based on the compact feature of the current facial video frame and the compact feature of the initial reference facial video frame, mean square error calculation is performed to obtain the information difference value between the current facial video frame and the initial reference facial video frame in the reference frame list.
[0114] Optionally, in some embodiments, when the first information difference value calculation module 702 performs the step of feature extraction on the current facial video frame to be encoded to obtain the compact feature of the current facial video frame, it specifically is used for:
[0115] Input the current facial video frame to be encoded into a pre-trained feature extraction model, so that the feature extraction model outputs the compact feature of the current facial video frame;
[0116] When the first information difference value calculation module 702 performs the step of feature extraction on the initial reference facial video frame in the reference frame list to obtain the compact feature of the initial reference facial video frame, it specifically is used for:
[0117] Input the initial reference facial video frame in the reference frame list into the feature extraction model, so that the feature extraction model outputs the compact feature of the initial reference facial video frame.
[0118] Optionally, in some embodiments, the facial video encoding device further includes:
[0119] A second encoding module, configured to, if the information difference value is less than or equal to a preset threshold, encode the current facial video frame based on the initial reference facial video frame to obtain a second facial video bitstream.
[0120] Optionally, in some embodiments, if the number of initial reference facial video frames is multiple; the first information difference value calculation module 702 specifically is used for:
[0121] Calculate the information difference values between the current facial video frame to be encoded and each initial reference facial video frame in the reference frame list respectively;
[0122] The first encoding module 704, specifically configured to, if the information difference values between the current facial video frame and each initial reference facial video frame in the reference frame list are all greater than the preset threshold, use the current facial video frame as a new reference facial video frame, and the new reference facial video frame is used to encode other video frames to obtain a first facial video bitstream;
[0123] A second encoding module, specifically configured to, if there is a candidate information difference value less than or equal to a preset threshold among all information difference values, encode the current facial video frame based on the initial reference facial video frame corresponding to the candidate information difference value to obtain a second facial video bitstream.
[0124] Optionally, in some embodiments, if the number of candidate information difference values is multiple, the second encoding module is specifically configured to:
[0125] If there is a candidate information difference value less than or equal to a preset threshold among all information difference values, determine the information difference value with the smallest numerical value among the candidate information difference values as the target information difference value;
[0126] Encode the current facial video frame based on the initial reference facial video frame corresponding to the target information difference value to obtain a second facial video bitstream.
[0127] Optionally, in some embodiments, if the information difference value is greater than the preset threshold, the facial video encoding device further includes:
[0128] A reference frame list update module, configured to add the newly added reference facial video frame to the reference frame list to update the reference frame list.
[0129] Optionally, in some embodiments, if the number of initial reference facial video frames is 1, the reference frame list update module is specifically configured to add the newly added reference facial video frame to the reference frame list and delete the initial reference facial video frame to obtain an updated reference frame list.
[0130] Optionally, in some embodiments, if the number of initial reference facial video frames is multiple, the reference frame list update module is specifically configured to add the newly added reference facial video frame to the reference frame list and delete the initial reference facial video frame with the earliest timestamp to obtain an updated reference frame list.
[0131] The facial video encoding device in this embodiment is used to implement the corresponding facial video encoding method in the foregoing multiple method embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here. In addition, the function implementation of each module in the facial video encoding device in this embodiment can be referred to the description of the corresponding part in the foregoing method embodiments, which will not be elaborated here either.
[0132] Embodiment Six
[0133] See Figure 8 , Figure 8 is a structural block diagram of a facial video decoding device according to Embodiment Six of the present application. The facial video decoding device provided in the embodiment of the present application includes:
[0134] A video bitstream acquisition module 802 for acquiring a facial video bitstream;
[0135] A decoding module 804 for decoding the facial video bitstream to obtain target driving information of a target facial video frame to be encoded and target identification information indicating whether the target facial video is a newly added reference facial video frame;
[0136] A reference driving information acquisition module 806 for, if the target facial video frame is not a newly added reference facial video frame, acquiring reference driving information of multiple reference facial video frames in a reference frame list respectively;
[0137] A second information difference value calculation module 808 for calculating an information difference value between the target facial video frame and each reference facial video frame based on the target driving information and each reference driving information; the information difference value represents the degree of difference between the information contained in the target facial video frame and the information contained in the reference facial video frame;
[0138] A reconstructed facial video frame obtaining module 810 for using the reference facial video frame corresponding to the minimum information difference value as the target reference facial video frame and obtaining a reconstructed facial video frame based on the target reference facial video frame and the target driving information.
[0139] Optionally, in some embodiments, the second information difference value calculation module 808 is specifically configured to: perform mean square error calculation based on the target driving information and each reference driving information to obtain an information difference value between the target facial video frame and each reference facial video frame.
[0140] The facial video decoding device in this embodiment is used to implement the corresponding facial video decoding method in the foregoing multiple method embodiments and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here. In addition, the function implementation of each module in the facial video decoding device in this embodiment can refer to the description of the corresponding part in the foregoing method embodiments, which will not be elaborated here either.
[0141] Embodiment Seven
[0142] Refer to Figure 9 , which shows a schematic structural diagram of an electronic device according to Embodiment Five of the present application. The specific implementation of the electronic device is not limited in the specific embodiments of the present application.
[0143] As Figure 9 shown, the conference terminal may include: a processor 902, a communication interface 904, a memory 906, and a communication bus 908.
[0144] Wherein:
[0145] The processor 902, the communication interface 904, and the memory 906 communicate with each other via the communication bus 908.
[0146] The communication interface 904 is used to communicate with other electronic devices or servers.
[0147] The processor 902 is used to execute the program 910, and specifically can execute the above-mentioned facial video encoding method or the relevant steps in the embodiments of the facial video decoding method.
[0148] Specifically, the program 910 may include program codes, and the program codes include computer operation instructions.
[0149] The processor 902 may be a CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application. One or more processors included in the intelligent device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.
[0150] The memory 906 is used to store the program 910. The memory 906 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.
[0151] Specifically, the program 910 can be used to cause the processor 902 to perform the following operations: calculate the information difference value between the current facial video frame to be encoded and the initial reference facial video frame in the reference frame list, where the information difference value characterizes the difference degree between the information included in the current facial video frame and the information included in the initial reference facial video frame; if the information difference value is greater than a preset threshold, use the current facial video frame as a new reference facial video frame, and the new reference facial video frame is used to encode other video frames to obtain the first facial video bitstream.
[0152] Or,
[0153] Program 910 can be specifically used to cause the processor 902 to perform the following operations: obtain a facial video bitstream; decode the facial video bitstream to obtain target driving information of a target facial video frame to be encoded and target identification information indicating whether the target facial video is a newly added reference facial video frame; if the target facial video frame is not a newly added reference facial video frame, respectively obtain reference driving information of multiple reference facial video frames in a reference frame list; based on the target driving information and each reference driving information, calculate an information difference value between the target facial video frame and each reference facial video frame; the information difference value represents the degree of difference between the information contained in the target facial video frame and the information contained in the reference facial video frame; use the reference facial video frame corresponding to the minimum information difference value as the target reference facial video frame, and based on the target reference facial video frame and the target driving information, obtain a reconstructed facial video frame.
[0154] For the specific implementation of each step in Program 910, reference can be made to the corresponding steps and descriptions in the corresponding steps and units in the above-mentioned facial video encoding method or facial video decoding method embodiments, which will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding process descriptions in the foregoing method embodiments, which will not be repeated here.
[0155] Through the electronic device of this embodiment, during the encoding process of the current facial video frame, the current facial video frame is not directly encoded and decoded based on the initial reference facial video frames in the reference frame list. Instead, the information difference value between the current facial video frame and the initial reference facial video frames is first calculated. If the information difference value is large (greater than a preset threshold), then the current facial video frame is not encoded and decoded based on the initial reference facial video frames. Instead, the current facial video frame is used as a new (newly added) reference facial video frame to encode other video frames to obtain a first facial video bitstream. In the embodiments of the present application, the reference facial video frame actually used in the encoding and decoding process is determined based on the degree of information difference between the current facial video frame and the initial reference facial video frames. When the degree of difference is large, the current facial video frame is used as a new reference facial video frame instead of relying on the initial reference facial video frames for encoding and decoding. Therefore, the quality of facial video encoding and decoding can be improved, and a higher-quality reconstructed facial video frame can be obtained.
[0156] The embodiments of the present application also provide a computer program product, including computer instructions, and the computer instructions instruct a computing device to perform the operations corresponding to any one of the above-mentioned multiple method embodiments.
[0157] It should be noted that according to the needs of implementation, each component / step described in the embodiments of the present application can be split into more components / steps, or two or more components / steps or partial operations of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present application.
[0158] The methods according to the embodiments of the present application described above can be implemented in hardware, firmware, or be implemented as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or be implemented as computer code originally stored in a remote recording medium or a non-transitory machine-readable medium and downloaded through a network and to be stored in a local recording medium, so that the methods described herein can be stored on such a software process on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as RAM, ROM, flash memory, etc.) that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, the processor, or the hardware, the facial video encoding method, or the facial video decoding method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the facial video encoding method, or the facial video decoding method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the facial video encoding method, or the facial video decoding method shown herein.
[0159] Those of ordinary skill in the art can realize that the units and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.
[0160] The above embodiments are only used to illustrate the embodiments of the present application, rather than to limit the embodiments of the present application. Those of ordinary skill in the relevant technical field can also make various changes and modifications without departing from the spirit and scope of the embodiments of the present application. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present application. The patent protection scope of the embodiments of the present application shall be defined by the claims.
Claims
1. A facial video encoding method, comprising: Calculating an information difference value between a current facial video frame to be encoded and one or more initial reference facial video frames in a reference frame list, where the information difference value characterizes the degree of difference between the information contained in the current facial video frame and the information contained in the initial reference facial video frame; When obtaining one such information difference value and the information difference value is greater than a preset threshold, or when obtaining multiple such information difference values and all the multiple information difference values are greater than the preset threshold, using the current facial video frame as a new reference facial video frame, where the new reference facial video frame is used to encode other video frames to obtain a first facial video bitstream.
2. The method according to claim 1, wherein, The calculating of the information difference value between the current facial video frame to be encoded and the initial reference facial video frames in the reference frame list includes: Based on the pixel values of each pixel point in the current facial video frame to be encoded and the pixel values of each pixel point in the initial reference facial video frame, performing mean square error calculation to obtain the information difference value between the current facial video frame and the initial reference facial video frames in the reference frame list.
3. The method according to claim 1, wherein The calculating of the information difference value between the current facial video frame to be encoded and the initial reference facial video frames in the reference frame list includes: Performing feature extraction on the current facial video frame to be encoded to obtain a compact feature of the current facial video frame, where the compact feature of the current facial video frame characterizes the key feature information in the facial video frame; Performing feature extraction on the initial reference facial video frames in the reference frame list to obtain a compact feature of the initial reference facial video frames, where the compact feature of the initial reference facial video frames characterizes the key feature information in the initial reference facial video frames; Based on the compact feature of the current facial video frame and the compact feature of the initial reference facial video frames, performing mean square error calculation to obtain the information difference value between the current facial video frame and the initial reference facial video frames in the reference frame list.
4. The method according to claim 3, wherein, The performing of feature extraction on the current facial video frame to be encoded to obtain a compact feature of the current facial video frame includes: Inputting the current facial video frame to be encoded into a pre-trained feature extraction model so that the feature extraction model outputs the compact feature of the current facial video frame; The performing of feature extraction on the initial reference facial video frames in the reference frame list to obtain a compact feature of the initial reference facial video frames includes: Inputting the initial reference facial video frames in the reference frame list into the feature extraction model so that the feature extraction model outputs the compact feature of the initial reference facial video frames.
5. The method according to claim 1, wherein The method further includes: If the information difference value is less than or equal to the preset threshold, encoding the current facial video frame based on the initial reference facial video frame to obtain a second facial video bitstream.
6. The method according to claim 5, wherein If the number of the initial reference facial video frames is multiple; The calculating of the information difference value between the current facial video frame to be encoded and the initial reference facial video frames in the reference frame list includes: Respectively calculating the information difference values between the current facial video frame to be encoded and each of the initial reference facial video frames in the reference frame list; If the information difference value is less than or equal to the preset threshold, encoding the current facial video frame based on the initial reference facial video frame to obtain a second facial video bitstream, including: If there is a candidate information difference value less than or equal to the preset threshold among all the information difference values, encoding the current facial video frame based on the initial reference facial video frame corresponding to the candidate information difference value to obtain a second facial video bitstream.
7. The method according to claim 6, wherein If the number of the candidate information difference values is multiple, the encoding the current facial video frame based on the initial reference facial video frame corresponding to the candidate information difference value to obtain a second facial video bitstream includes: Determining the information difference value with the smallest numerical value among the candidate information difference values as the target information difference value; Encoding the current facial video frame based on the initial reference facial video frame corresponding to the target information difference value to obtain a second facial video bitstream.
8. The method according to claim 6, wherein If the information difference value is greater than the preset threshold, the method further includes: Adding the newly added reference facial video frame to the reference frame list to update the reference frame list.
9. The method according to claim 8, wherein If the number of the initial reference facial video frames is 1, The adding the newly added reference facial video frame to the reference frame list to update the reference frame list includes: Adding the newly added reference facial video frame to the reference frame list and deleting the initial reference facial video frame to obtain an updated reference frame list.
10. The method according to claim 8, wherein, If the number of the initial reference facial video frames is multiple, The adding the newly added reference facial video frame to the reference frame list to update the reference frame list includes: Adding the newly added reference facial video frame to the reference frame list and deleting the initial reference facial video frame with the earliest timestamp to obtain an updated reference frame list.
11. A facial video decoding method, including: Obtaining a facial video bitstream; Decoding the facial video bitstream to obtain target driving information of a target facial video frame to be encoded and target identification information indicating whether the target facial video is a newly added reference facial video frame; If the target facial video frame is not a newly added reference facial video frame, respectively obtaining reference driving information of multiple reference facial video frames in the reference frame list; Calculating an information difference value between the target facial video frame and each reference facial video frame based on the target driving information and each reference driving information; the information difference value represents the difference degree between the information included in the target facial video frame and the information included in the reference facial video frame; Taking the reference facial video frame corresponding to the minimum information difference value as the target reference facial video frame, and obtaining a reconstructed facial video frame based on the target reference facial video frame and the target driving information.
12. The method according to claim 11, wherein, The calculating an information difference value between the target facial video frame and each reference facial video frame based on the target driving information and each reference driving information includes: Performing mean square error calculation based on the target driving information and each reference driving information to obtain an information difference value between the target facial video frame and each reference facial video frame.
13. A facial video encoding device, including: The first information difference value calculation module is configured to calculate an information difference value between a current facial video frame to be encoded and one or more initial reference facial video frames in a reference frame list, where the information difference value characterizes the degree of difference between the information included in the current facial video frame and the information included in the initial reference facial video frame; The first encoding module is configured to, when obtaining one such information difference value and the information difference value is greater than a preset threshold, or when obtaining multiple such information difference values and all the multiple information difference values are greater than the preset threshold, use the current facial video frame as a new reference facial video frame, and the new reference facial video frame is used to encode other video frames to obtain a first facial video bitstream.
14. An electronic device, comprising: A processor, a memory, a communication interface, and a communication bus, where the processor, the memory, and the communication interface complete communication with each other through the communication bus; The memory is configured to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the facial video encoding method according to any one of claims 1-10, or operations corresponding to the facial video decoding method according to any one of claims 11-12.
Citation Information
Patent Citations
Video encoding and decoding method and device
CN113573063A
Digital video signal, a method for encoding of a digital video signal and a digital video signal encoder
US20120300030A1