A method, apparatus, electronic device, and storage medium for optimizing key points.

By calculating the prior loss and scale loss of key points in the expression transfer model, the expression transfer algorithm is optimized, which solves the problem of the influence of scale and depth on the standard key points and improves the accuracy and richness of expression transfer.

CN120746873BActive Publication Date: 2025-12-02HANGZHOU QIUGUOJIHUA TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511202832.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-12-02
Estimated Expiration
2045-08-26

AI Technical Summary

Technical Problem

Existing facial expression transfer algorithms, after introducing scale parameters, cause errors in the scale and depth of the standardized key points, affecting the richness of the expressions the network learns, and the effect is not ideal, especially when the dataset is small.

Method used

By acquiring image data, we extract algorithm parameters using an expression transfer model, calculate implicit keypoints, and calculate the prior loss of keypoints based on offset parameters with a mean of zero in the depth dimension. We then combine scale loss to optimize the expression transfer model and limit the influence of depth and scale on the standardized keypoints.

Benefits of technology

It effectively avoids the negative impact of scale and depth on key points of the standard, optimizes the expression transfer model, and improves the accuracy and richness of expression transfer results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120746873B_ABST
    Figure CN120746873B_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, electronic device, and storage medium for optimizing keypoints. The method includes: acquiring image data, wherein the image data includes a source image and a driving image; extracting algorithm parameters from the source image and driving image using an expression transfer model, and using these parameters to calculate implicit keypoints in the source image and driving image, so as to generate expression transfer results using the implicit keypoints in the source image and driving image; calculating the keypoint prior loss for the implicit keypoints of the driving image based on the offset parameters of the driving image after the depth dimension is constrained to zero in the algorithm parameters, with the mean of the depth dimension being zero; calculating the scale loss using the scale of the driving image in the algorithm parameters; and including the keypoint prior loss and scale loss in the total loss of the model, and using them to optimize the expression transfer model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of facial expression transfer technology, and in particular to a method and apparatus for optimizing key points, an electronic device, and a storage medium. Background Technology

[0002] Expression transfer algorithms aim to synthesize realistic videos or images using a single source image and a driving video or image, transferring the facial expressions from the driving video or image to the source image. Current traditional methods primarily employ the Face vid2vid algorithm for expression transfer. Specifically, this involves calculating implicit keypoints and using these keypoints to generate the expression-transfer image.

[0003] However, the Face vid2vid algorithm ignores scale, which leads to scaling being incorporated into facial expression deformation, increasing training difficulty. To address this, LivePortrait introduces scale into its motion transformation calculations, adding a scale parameter to the algorithm's parameters. However, this also introduces new problems. Since the offset parameters in the algorithm include offsets in the x, y, and z directions, changes in scale affect these offsets, resulting in very large scale and depth of the canonical keypoints in the algorithm.

[0004] While analyzing large amounts of data at different scales can improve the performance of canonical keypoints, the results are less than ideal when the dataset is too small. Specifically, inputting the implicit keypoints, canonical keypoints, and facial features from the source image into a distortion field estimator to obtain motion feature deformation results in a jumbled output when fed into the generator. This is highly detrimental to the network's ability to learn rich facial expressions. Summary of the Invention

[0005] In view of the shortcomings of the prior art, this application provides a method, apparatus, electronic device and storage medium for optimizing key points, so as to solve the problem that the introduction of scale influence on the key points in the prior art leads to errors in the generated results.

[0006] To achieve the above objectives, this application provides the following technical solution:

[0007] The first aspect of this application provides an optimization method for standardizing key points, including:

[0008] Acquire image data; wherein the image data includes a source image and a driving image;

[0009] The expression transfer model extracts algorithm parameters from the source image and the driving image, and uses the algorithm parameters to calculate the implicit key points of the source image and the driving image, so as to generate expression transfer results using the implicit key points of the source image and the driving image.

[0010] Based on the offset parameters of the driving graph after the depth dimension is restricted to zero in the algorithm parameters, the implicit keypoints of the driving graph are subjected to keypoint prior loss calculation with the mean of the depth dimension being zero.

[0011] The scale loss is obtained by calculating the scale of the driving graph in the algorithm parameters;

[0012] The prior loss of the key points and the scale loss are included in the total loss of the model, and the expression transfer model is optimized based on the total loss of the model.

[0013] Optionally, in the above-described optimization method for key specification points, acquiring image data includes:

[0014] Generate multi-angle photos of different people using the StyleGan model;

[0015] The multi-angle photos are interpolated to synthesize a continuous video extended image, and the continuous video extended image is used to generate initial image data;

[0016] Human face masking is performed on each group of images in the image data according to random probability; wherein each group of images includes a source image and a driving image;

[0017] The human figures in each set of images are extracted, enlarged, or cropped to obtain human portrait images;

[0018] The portrait image is synthesized into an image of target resolution using a background of preset color; wherein, the target resolution is the resolution of the output image of the expression transfer model.

[0019] Optionally, in the above-described optimization method for key specification points, before inputting the image data into the expression transfer model, the method further includes:

[0020] Gaussian noise is added to the source image in the image data according to a set probability.

[0021] Optionally, in the above-described method for optimizing key points, the step of extracting algorithm parameters from the source image and the driving image using an expression transfer model, and calculating the implicit key points of the source image and the driving image using the algorithm parameters, so as to generate expression transfer results using the implicit key points of the source image and the driving image, includes:

[0022] The image data is input into the expression transfer model to extract target parameters from the source image and the driving image, and to extract facial features and standardized key points from the source image; wherein, the target parameters include rotation parameters, offset parameters, scale, and expression deformation parameters;

[0023] Using the target parameters of the source graph and the standardized key points, the implicit key points of the source graph are calculated, and using the target parameters of the driving graph and the standardized key points, the implicit key points of the driving graph are calculated.

[0024] Based on the implicit keypoints of the source image, the implicit keypoints of the driving image, and the facial features, an expression transfer result is generated.

[0025] Optionally, in the above-described optimization method for keypoints, the step of calculating the keypoint prior loss for the implicit keypoints of the driving graph based on the offset parameters of the driving graph after the depth dimension in the algorithm parameters is restricted to zero, with the mean of the depth dimension being zero, includes:

[0026] The depth dimension of the offset parameter of the driving graph in the algorithm parameters is restricted to zero;

[0027] Using the offset parameters of the driving graph after the depth dimension is restricted to zero and the target keypoints in the implicit keypoints of the driving graph, the keypoint prior loss is calculated with the depth dimension mean being zero; wherein, the target keypoints refer to all keypoints except the guiding keypoints.

[0028] Optionally, in the above-described optimization method for canonical keypoints, the step of calculating the keypoint prior loss using the offset parameters of the driving graph after limiting the depth dimension to zero and the target keypoints in the implicit keypoints of the driving graph, with the depth dimension mean being zero, includes:

[0029] Calculate the squared distance between every two target key points in the implicit key points of the driving graph to obtain the squared distances corresponding to each target key point;

[0030] For each of the target key points, the maximum distance corresponding to each target key point is obtained by subtracting the difference between the squares of the distances corresponding to the target key point and the maximum value among zero from the distance threshold.

[0031] The maximum distance corresponding to the target key point, the difference between the mean depth dimension and the depth dimension value of the target key point are summed, and the summation result is combined with the offset parameter of the driving graph with the depth dimension restricted to zero to obtain the key point prior loss.

[0032] Optionally, in the above-described optimization method for key points, the step of calculating the scale loss of the driving graph to obtain the scale loss includes:

[0033] The difference between 1 and the scale of the driving graph is calculated by adding an activation function, and the scale loss is output based on the calculated difference; wherein, when the difference is greater than 1, the difference is output as the scale loss; when the difference is not greater than 1, zero is output as the scale loss.

[0034] A second aspect of this application provides an optimization apparatus for standardizing key points, comprising:

[0035] An image acquisition unit is used to acquire image data; wherein the image data includes a source image and a driving image;

[0036] The expression transfer unit is used to extract algorithm parameters from the source image and the driving image through the expression transfer model, and use the algorithm parameters to calculate the implicit key points of the source image and the driving image, so as to generate expression transfer results using the implicit key points of the source image and the driving image;

[0037] The prior loss calculation unit is used to calculate the key point prior loss of the implicit key points of the driving graph based on the offset parameters of the driving graph after the depth dimension is restricted to zero in the algorithm parameters, and with the mean of the depth dimension being zero.

[0038] The scale loss calculation unit is used to calculate the scale of the driving graph in the algorithm parameters to obtain the scale loss;

[0039] An optimization unit is used to include the key point prior loss and the scale loss in the total model loss, and to optimize the expression transfer model based on the total model loss.

[0040] Optionally, in the above-described optimization apparatus for key specification points, the image acquisition unit includes:

[0041] The image generation unit is used to generate multi-angle photos of different people using the StyleGAN model;

[0042] An extension unit is used to interpolate the multi-angle photos to synthesize a continuous video extended image, and to generate initial image data using the continuous video extended image;

[0043] The image matting unit is used to perform portrait matting on each group of images in the image data according to random probability; wherein, each group of images includes a source image and a driving image;

[0044] The scaling unit is used to enlarge or crop the human portrait cutouts from each group of images to obtain human portrait images;

[0045] A synthesis unit is used to synthesize the portrait image into an image of a target resolution using a background of a preset color; wherein, the target resolution is the resolution of the output image of the expression transfer model.

[0046] Optionally, the optimization device for the aforementioned key specifications further includes:

[0047] Gaussian noise is added to the source image in the image data according to a set probability.

[0048] Optionally, in the above-mentioned optimization device for key specification points, the expression transfer unit includes:

[0049] The parameter extraction unit is used to input the image data into the expression transfer model, extract the target parameters of the source image and the driving image, and extract facial features and standardized key points from the source image; wherein, the target parameters include rotation parameters, offset parameters, scale, and expression deformation parameters;

[0050] The key point calculation unit is used to calculate the implicit key points of the source graph using the target parameters of the source graph and the standard key points, and to calculate the implicit key points of the driving graph using the target parameters of the driving graph and the standard key points.

[0051] The result generation unit is used to generate expression transfer results based on the implicit keypoints of the source image, the implicit keypoints of the driving image, and the facial features.

[0052] Optionally, in the above-mentioned optimization device for key specifications, the prior loss calculation unit includes:

[0053] A limiting unit is used to limit the depth dimension of the offset parameter of the driving graph in the algorithm parameters to zero;

[0054] The prior loss calculation subunit is used to calculate the prior loss of key points by using the offset parameters of the driving graph after the depth dimension is restricted to zero and the target key points in the implicit key points of the driving graph, with the mean of the depth dimension being zero; wherein, the target key points refer to each key point other than the guiding key points.

[0055] Optionally, in the above-mentioned optimization device for key specifications, the prior loss calculation subunit includes:

[0056] The distance calculation unit is used to calculate the squared distance between every two target key points in the implicit key points of the driving graph, so as to obtain the squared distances corresponding to each target key point;

[0057] The distance determination unit is used to select a distance threshold and subtract the difference between the squares of the distances corresponding to the target key point and the maximum value among zero for each target key point to obtain the maximum distance corresponding to each target key point.

[0058] The result calculation unit is used to sum the maximum distance corresponding to the target key point, the difference between the mean depth dimension and the depth dimension value of the target key point, and combine the summation result with the offset parameter of the driving graph whose depth dimension is restricted to zero to obtain the key point prior loss.

[0059] Optionally, in the above-mentioned optimization device for key specifications, the scale loss calculation unit includes:

[0060] The scale loss calculation subunit is used to calculate the difference between 1 and the scale of the driving graph by adding an activation function, and output the scale loss according to the calculated difference; wherein, when the difference is greater than 1, the difference is output as the scale loss; when the difference is not greater than 1, zero is output as the scale loss.

[0061] A third aspect of this application provides an electronic device, comprising:

[0062] Memory and processor;

[0063] The memory is used to store programs;

[0064] The processor is used to execute the program, which, when executed, is specifically used to implement the optimization method for the specification key points as described in any of the above.

[0065] A fourth aspect of this application provides a computer storage medium for storing a computer program, which, when executed by a processor, is used to implement the optimization method for specification key points as described in any of the preceding claims.

[0066] This application provides an optimization method for standardizing keypoints, acquiring image data including a source image and a driving image. Algorithm parameters are extracted from the source and driving images using an expression transfer model, and implicit keypoints in the source and driving images are calculated using these parameters to generate expression transfer results. Next, based on the offset parameters of the driving image with its depth dimension constrained to zero in the algorithm parameters, a keypoint prior loss is calculated for the implicit keypoints of the driving image, with a mean depth dimension of zero. By setting the mean depth dimension to zero and constraining the depth dimension of the offset parameters of the driving image to zero, the influence of scale parameters on standardizing keypoints is effectively avoided. Then, a loss is calculated for the scale of the driving image in the algorithm parameters to obtain the scale loss. This scale loss constrains the scale, avoiding its influence. The keypoint prior loss and scale loss are included in the total model loss, and the expression transfer model is optimized based on the total model loss. Since the keypoint prior loss and scale loss take into account the influence of depth and scale on canonical keypoints, the influence of scale on canonical keypoints in the expression transfer model can be resolved by continuously optimizing the model, thereby achieving optimization of canonical keypoints. Attached Figure Description

[0067] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0068] Figure 1 A flowchart of a method for optimizing key points is provided in this application embodiment;

[0069] Figure 2 A flowchart illustrating a method for scaling image data, provided in an embodiment of this application;

[0070] Figure 3 A flowchart for generating expression transfer results using an expression transfer model is provided as an embodiment of this application;

[0071] Figure 4 A flowchart illustrating a method for calculating prior loss of key points, provided as an embodiment of this application;

[0072] Figure 5 A schematic diagram of the architecture of an optimization device for standardizing key points provided in this application embodiment;

[0073] Figure 6 This is a schematic diagram of the architecture of an electronic device provided in an embodiment of this application. Detailed Implementation

[0074] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0075] In this application, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0076] This application provides an optimization method for standardizing key points, such as... Figure 1 As shown, it includes the following steps:

[0077] S101. Acquire image data.

[0078] The image data includes a source image and a driving image, which can drive the expression in the source image through the driving image, that is, transfer the expression in the driving image to the source image.

[0079] Specifically, in order to effectively optimize the key points of the specification, the acquired image data may include multiple sets of data, and each set of data includes a source image and a driving image.

[0080] Optionally, in another embodiment of this application, one specific implementation of step S101 is as follows: Figure 2 As shown, it includes the following steps:

[0081] S201. Use the StyleGan model to generate multi-angle photos of different people, interpolate the multi-angle photos to synthesize a continuous video extended image, and use the continuous video extended image to generate the initial image data.

[0082] It should be noted that the StyleGAN model is a variant of Generative Adversarial Networks (GANs). It is an unsupervised learning model capable of generating high-resolution, realistic images while allowing control over the style of the generated images. Therefore, in this embodiment, using the StyleGAN model can more conveniently generate high-resolution, multi-angle photographs of multiple different people, thus addressing the issue of uniform portrait scale to some extent.

[0083] To enable effective training on the same person ID, multi-angle photos are further interpolated to synthesize continuous video extended images, resulting in multiple photos of each person for training under the same task ID. Then, each image from the synthesized continuous video extended images is combined with a driving graph to obtain the initial image data.

[0084] It should be noted that, to address the issue of limited image data resulting in a uniform scale for human figures, the scale of the acquired image data is manipulated to enrich the representation of figures. Furthermore, this process simultaneously adjusts the image data to a resolution that meets the model's requirements. Therefore, further scale processing of the initially acquired images is necessary to obtain images with more scales.

[0085] S202. Perform portrait cutout on each group of images in the image data according to random probability.

[0086] Each set of images includes a source image and a driving image.

[0087] It should be noted that when performing expression transfer using the expression transfer model, a source image and a driving image are used for processing. Since the model's input dimensions are static, the source image and driving image must undergo consistent operations to ensure that the variables are the same. Therefore, each set of images is processed as a single unit.

[0088] In order to increase the scale of the figures and solve the problem of the uniformity of the figures, the figures in each group of images in the image data are first cut out according to random probability, so as to randomly cut out the figures.

[0089] S203. Enlarge or crop the human portrait cutouts from each group of images to obtain human portrait images.

[0090] Specifically, by enlarging or cropping the cutout of a human figure, the scale of the figure can be altered, enriching the overall composition. Optionally, the enlargement and cropping ratios of the cutout can be randomized.

[0091] S204. Use a preset background color to synthesize the portrait image into an image of the target resolution.

[0092] The target resolution is the resolution of the output image of the expression transfer model.

[0093] It's important to note that since the portrait image only includes the person, a background needs to be added to restore it to a complete image. To avoid the influence of the background, a preset background color is used to composite the portrait image into a complete image; that is, a fixed-color background is blended with the portrait image to form a complete image. Furthermore, to ensure the image conforms to the output of the expression transfer model, it is composited to a target resolution. For example, a 512x512 resolution image is composited, meaning the image resolution is fixed at 512x512.

[0094] Optionally, in order to enhance the upsampling generation capability of the generative network without altering the image, in another embodiment of this application, after executing step S103, the following can be further performed:

[0095] Gaussian noise is added to the source image in the image data according to a set probability.

[0096] S102. Extract algorithm parameters from the source image and the driving image using the expression transfer model, and use the algorithm parameters to calculate the implicit key points of the source image and the driving image, so as to generate expression transfer results using the implicit key points of the source image and the driving image.

[0097] In this embodiment, the expression transfer model primarily uses the LivePortrait algorithm, which introduces a scale. However, to eliminate the impact of the introduced scale on the canonical key points, the canonical key points need to be optimized accordingly. Specifically, the canonical key points need to be continuously optimized based on the expression transfer model's performance when transferring expressions from the image. Therefore, image data is input into the expression transfer model multiple times for processing and optimization, i.e., steps S103 to S106 are executed multiple times.

[0098] Specifically, image data is input into the expression transfer model. The model extracts the parameters required for the algorithm's calculation from the source and driving images, i.e., it extracts the algorithm parameters. Then, the expression transfer model calculates these parameters to obtain the implicit keypoints of the source and driving images. Finally, using the implicit keypoints of the source and driving images, combined with the facial features of the source image from the algorithm parameters, the expression transfer result is generated, i.e., the expression from the driving image is transferred to the source image.

[0099] Optionally, in another embodiment of this application, one specific implementation of step S103 is as follows: Figure 3 As shown, it includes:

[0100] S301. Input the image data into the expression transfer model, extract the target parameters of the source image and the driving image, and extract facial features and standardized key points from the source image.

[0101] The target parameters include rotation parameters, offset parameters, scale, and facial expression deformation parameters.

[0102] Specifically, facial features and standardized key points are extracted from the source image, and target parameters for both the source and driving images are extracted separately. More specifically, information can be extracted from relevant parts of the expression transfer model architecture. For example, facial features of the portrait in the source image are extracted using an appearance feature extractor; rotation and offset parameters (i.e., head rotation and movement parameters) are extracted using a head pose estimation network; and expression deformation parameters are extracted using an expression deformation estimation network, etc.

[0103] S302. Using the target parameters and standard key points of the source graph, calculate the implicit key points of the source graph, and using the target parameters and standard key points of the driving graph, calculate the implicit key points of the driving graph.

[0104] Specifically, the rotation parameters of the source image are multiplied by the canonical keypoints, and the resulting product is added to the facial distortion parameters. The sum is then multiplied by the scale, and finally, the product is added to the offset parameters to obtain the implicit keypoints of the source image. The calculation method for the implicit keypoints of the driving image is the same as that of the source image, except that it requires the use of the target parameters of the driving image for calculation.

[0105] Therefore, the implicit keypoints of the source graph and the implicit keypoints of the driving graph are calculated as follows:

[0106]

[0107] Where, x s and x d Represent the implicit keypoints of the source graph and driving graph, respectively; x c,s To standardize keypoints, their dimension is k×3, where k is the number of keypoints. R s and R d Rotation parameters for the two graphs. , This represents the facial expression deformation in the source and driving images. s and t d This represents the offset parameter, which has a dimension of 1×3, meaning it includes offsets in the x, y, and z directions.

[0108] S303. Based on the implicit keypoints of the source image, the implicit keypoints of the driving image, and facial features, generate expression transfer results.

[0109] Specifically, the implicit keypoints of the source image, the implicit keypoints of the driving image, and the face features are input into the distortion field estimator in the model to obtain motion feature deformation, which is then fed into the generator to obtain the result.

[0110] S103. Based on the offset parameters of the driving graph after the depth dimension in the algorithm parameters is restricted to zero, calculate the key point prior loss for the implicit key points of the driving graph with the mean of the depth dimension being zero.

[0111] The depth dimension refers to the dimension of the z-axis in the coordinate system.

[0112] It's worth noting that the loss function in face vid2vid includes the keypoint prior loss. LivePortrait is an improvement on face vid2vid, so it also uses this loss function. The keypoint prior loss limits x by calculating the distance matrix of the keypoints. d In a dimension N×2¹×3, a penalty is applied if the squared distance between any two keypoints is greater than 0.1, otherwise a loss is incurred. The mean value along the z-axis is 0.33. This is because when obtaining x... s and x d Then, it is fed into the torsion field estimator along with f_s to obtain motion features, and finally the generator obtains the result. This process is similar to orthogonal projection, removing the influence of the z-depth, so the result is obviously normal. Explain the key specification point x. c,s The errors in the generated results are closely related to depth and scale. Based on this, we studied and optimized the Keypointprior loss.

[0113] Therefore, in this embodiment, the mean value in the z-dimensional dimension is modified from 0.33 to 0.0 when calculating the prior loss for keypoints. Furthermore, since the offset parameter identifies the offset dimension as 1×3 (xyz), which also involves depth, it is further combined with the offset parameter of the driving graph. However, its depth dimension needs to be restricted to zero first, that is, the z-dimensional value of the offset parameter of the driving graph is restricted to 0.0. This eliminates the influence of depth on the canonical keypoints.

[0114] Optionally, in another embodiment of this application, one specific implementation of step S103 includes:

[0115] The depth dimension of the offset parameter of the driving graph in the algorithm parameters is restricted to zero. Then, using the offset parameter of the driving graph with the depth dimension restricted to zero and the target key points in the implicit key points of the driving graph, the key point prior loss is calculated with the depth dimension mean being zero.

[0116] Among them, the target key points refer to all key points other than the guiding key points.

[0117] It should be noted that LivePortrait adds prior guidance for keypoint training. Specifically, 10 keypoints are used for guidance on the face. The keypoint prior loss is designed so that the squared distance between any two points must be greater than 0.1, otherwise a penalty is applied. However, the distances between some facial keypoints, such as those in the eye region, are inherently very small. Furthermore, these 10 guidance keypoints are compared to ground truth using precise coordinates from the 2D image. If these 10 points are not removed, requiring both compliant spacing and precise coordinates would prevent the loss from decreasing. Therefore, to avoid the influence of these keypoints, the guidance keypoint loss is removed in this embodiment. However, it should be noted that this only excludes the loss caused by the distance between these keypoints in the initial loss calculation; the loss caused by the distance between other points and the keypoints needs to be calculated separately and included in the total loss.

[0118] Therefore, in this embodiment, the keypoint prior loss is calculated only for the distances between target keypoints in the implicit keypoints of the driving graph. Furthermore, the mean of the depth dimension is set to zero during calculation. Finally, an offset parameter of the driving graph with a depth dimension constrained to zero is added to obtain the keypoint prior loss.

[0119] Optionally, in another embodiment of this application, a method for calculating the prior loss of target key points is as follows: Figure 4 As shown, it includes:

[0120] S401. Calculate the squared distance between every two target key points in the implicit key points of the driving graph to obtain the squared distances corresponding to each target key point.

[0121] Specifically, in this embodiment, only the square of the distance between every two target key points is calculated, thereby obtaining the square of each distance corresponding to each target key point.

[0122] S402. For each target key point, select the maximum value among the differences of the squares of the distances corresponding to the target key point and zero by subtracting the distance threshold from the difference of the squares of the distances corresponding to the target key point.

[0123] Since the distance between two points needs to be limited to a distance threshold, the larger the difference between the distance threshold and the squared distance corresponding to the target keypoint, the greater the penalty. All distances need to be limited to the distance threshold, so penalizing the maximum difference in the squared distances corresponding to the target keypoints can more effectively regulate all distances to be limited to the distance threshold. Therefore, the maximum value among the differences in the squared distances corresponding to the target keypoints is determined. If this maximum value is positive, it is determined as the maximum distance corresponding to that target keypoint. If the maximum value is negative (less than zero), then 0 is determined as the maximum distance corresponding to that target keypoint.

[0124] S403. Sum the difference between the maximum distance and the mean depth dimension of the target key point and the depth dimension value of the target key point, and combine the summation result with the offset parameter of the driving graph with the depth dimension restricted to zero to obtain the key point prior loss.

[0125] S104. Calculate the loss of the scale of the driving graph in the algorithm parameters to obtain the scale loss.

[0126] It should be noted that, to avoid impacting key points of the specification on scale, the scale in this embodiment is limited to be greater than a certain threshold. Therefore, a scale loss function for the driving graph is added to calculate the scale loss. When the scale is less than the threshold, a corresponding loss is generated, which can be used to adjust for the existing impact. The smaller the scale, the greater the loss. If the scale is not less than the threshold, the loss is zero.

[0127] Optionally, in another embodiment of this application, one specific implementation of step S105 includes:

[0128] The difference between 1 and the scale of the driving graph is calculated by adding an activation function, and the scale loss is output based on the calculated difference.

[0129] Specifically, when the difference is less than 1, the difference is output as the scaling loss. When the difference is not less than 1, zero is output as the scaling loss. That is, in this embodiment, the scaling threshold is set to 1 to effectively eliminate the influence of scaling.

[0130] It should be noted that the activation function (ReLU function) only retains values ​​greater than 0; for values ​​not greater than 0, it directly outputs 0. Therefore, in this embodiment, the added activation function calculates the difference between 1 and the scale of the driving graph, and outputs the scale loss based on the calculated difference. Thus, when the difference is less than 1, the output scale loss is this difference.

[0131] S105. Incorporate the key point prior loss and scale loss into the total model loss, and optimize the expression transfer model based on the total model loss.

[0132] It should be noted that since the keypoint prior loss and scale loss take into account the influence of depth and scale on the canonical keypoints, they are included in the total model loss. This means that the keypoint prior loss and scale loss are summed with other original losses to obtain the total model loss. Then, by optimizing the expression transfer model based on the total model loss, the obtained canonical keypoints can be optimized.

[0133] This application provides an optimization method for standardizing keypoints, acquiring image data including a source image and a driving image. Algorithm parameters are extracted from the source and driving images using an expression transfer model, and implicit keypoints in the source and driving images are calculated using these parameters to generate expression transfer results. Next, based on the offset parameters of the driving image with its depth dimension constrained to zero in the algorithm parameters, a keypoint prior loss is calculated for the implicit keypoints of the driving image, with a mean depth dimension of zero. By setting the mean depth dimension to zero and constraining the depth dimension of the offset parameters of the driving image to zero, the influence of scale parameters on standardizing keypoints is effectively avoided. Then, a loss is calculated for the scale of the driving image in the algorithm parameters to obtain the scale loss. This scale loss constrains the scale, avoiding its influence. The keypoint prior loss and scale loss are included in the total model loss, and the expression transfer model is optimized based on the total model loss. Since the keypoint prior loss and scale loss take into account the influence of depth and scale on canonical keypoints, the influence of scale on canonical keypoints in the expression transfer model can be resolved by continuously optimizing the model, thereby achieving optimization of canonical keypoints.

[0134] Another embodiment of this application provides an optimization apparatus for standardizing key points, such as... Figure 5 As shown, it includes:

[0135] Image acquisition unit 501 is used to acquire image data.

[0136] The image data includes source images and driving images.

[0137] The expression transfer unit 502 is used to extract algorithm parameters from the source image and the driving image through the expression transfer model, and to use the algorithm parameters to calculate the implicit key points of the source image and the driving image, so as to generate expression transfer results using the implicit key points of the source image and the driving image.

[0138] The prior loss calculation unit 503 is used to calculate the key point prior loss of the implicit key points of the driving graph based on the offset parameters of the driving graph after the depth dimension is restricted to zero in the algorithm parameters, and with the mean of the depth dimension being zero.

[0139] The scale loss calculation unit 504 is used to calculate the scale loss of the driving graph in the algorithm parameters to obtain the scale loss.

[0140] The optimization unit 505 is used to incorporate the key point prior loss and scale loss into the total model loss, and optimize the expression transfer model based on the total model loss.

[0141] Optionally, in the optimization apparatus for specification key points in another embodiment of this application, the image acquisition unit includes:

[0142] The image generation unit is used to generate multi-angle photos of different people using the StyleGAN model.

[0143] The extension unit is used to interpolate multi-angle photos to synthesize a continuous video extended image, and to generate initial image data using the continuous video extended image.

[0144] The matting unit is used to perform portrait matting on each group of images in the image data according to random probabilities. Each group of images includes a source image and a driving image.

[0145] The scale processing unit is used to enlarge or crop the human portrait cutouts from each group of images to obtain human portrait images.

[0146] The synthesis unit is used to synthesize a portrait image into an image of a target resolution using a background of a preset color. The target resolution is the resolution of the output image of the expression transfer model.

[0147] Optionally, in another embodiment of the specification key point optimization apparatus of this application, the apparatus further includes:

[0148] Gaussian noise is added to the source image in the image data according to a set probability.

[0149] Optionally, in the optimization apparatus for specification key points in another embodiment of this application, the expression transfer unit includes:

[0150] The parameter extraction unit is used to input image data into the expression transfer model, extract target parameters from the source and driving images, and extract facial features and normalized key points from the source image. The target parameters include rotation parameters, offset parameters, scale, and expression deformation parameters.

[0151] The key point calculation unit is used to calculate the implicit key points of the source graph using the target parameters and canonical key points of the source graph, and to calculate the implicit key points of the driving graph using the target parameters and canonical key points of the driving graph.

[0152] The result generation unit is used to generate expression transfer results based on the implicit keypoints of the source image, the implicit keypoints of the driving image, and facial features.

[0153] Optionally, in the above-mentioned optimization device for key specifications, the prior loss calculation unit includes:

[0154] A constraint unit is used to constrain the depth dimension of the offset parameter of the driving graph in the algorithm parameters to zero.

[0155] The prior loss calculation subunit is used to calculate the prior loss of key points by utilizing the offset parameters of the driving graph after the depth dimension is restricted to zero and the target key points in the implicit key points of the driving graph, with the mean of the depth dimension being zero.

[0156] Among them, the target key points refer to all key points other than the guiding key points.

[0157] Optionally, in the optimization apparatus for specification key points in another embodiment of this application, the prior loss calculation unit includes:

[0158] The distance calculation unit is used to calculate the squared distance between every two target key points in the implicit key points of the driving graph, and obtain the squared distances corresponding to each target key point.

[0159] The distance determination unit is used to select the maximum distance for each target key point by subtracting the difference between the squares of the distances corresponding to the target key point and the maximum value among zero from the distance threshold.

[0160] The result calculation unit is used to sum the difference between the maximum distance and the mean depth dimension of the target key point and the depth dimension value of the target key point, and combine the summation result with the offset parameter of the driving graph with the depth dimension restricted to zero to obtain the key point prior loss.

[0161] Optionally, in the optimization apparatus for specification key points in another embodiment of this application, the scale loss calculation unit includes:

[0162] The scale loss calculation subunit calculates the difference between 1 and the scale of the driving graph using an added activation function, and outputs the scale loss based on the calculated difference. Specifically, if the difference is greater than 1, the difference is output as the scale loss; if the difference is not greater than 1, zero is output as the scale loss.

[0163] It should be noted that the specific working process of each unit provided in the above embodiments of this application can be referred to the implementation process of the corresponding steps in the above method embodiments, and will not be repeated here.

[0164] Another embodiment of this application provides an electronic device, such as... Figure 6 As shown, it includes:

[0165] Memory 601 and processor 602.

[0166] The memory 601 is used to store the program.

[0167] The processor 602 is used to execute the program stored in the memory 601, which, when executed, is specifically used to implement the optimization method for the specification key points provided in any of the above embodiments.

[0168] Another embodiment of this application provides a computer storage medium for storing a computer program, which, when executed by a processor, is used to implement the optimization method for specification key points as provided in any of the above embodiments.

[0169] Computer storage media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0170] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0171] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for optimizing key points, characterized in that, include: Acquire image data; wherein the image data includes a source image and a driving image; The expression transfer model extracts algorithm parameters from the source image and the driving image, and uses the algorithm parameters to calculate the implicit key points of the source image and the driving image, so as to generate expression transfer results using the implicit key points of the source image and the driving image. Based on the offset parameters of the driving graph after the depth dimension is restricted to zero in the algorithm parameters, the implicit keypoints of the driving graph are subjected to keypoint prior loss calculation with the mean of the depth dimension being zero. The difference between 1 and the scale of the driving graph is calculated by adding an activation function, and the scale loss is output based on the calculated difference; wherein, when the difference is greater than 0, the difference is output as the scale loss; when the difference is not greater than 0, zero is output as the scale loss. The prior loss of the key points and the scale loss are included in the total model loss, and the expression transfer model is optimized based on the total model loss.

2. The method according to claim 1, characterized in that, The acquisition of image data includes: Generate multi-angle photos of different people using the StyleGan model; The multi-angle photos are interpolated to synthesize a continuous video extended image, and the continuous video extended image is used to generate initial image data; Human face masking is performed on each group of images in the image data according to random probability; wherein each group of images includes a source image and a driving image; The human figures in each set of images are extracted, enlarged, or cropped to obtain human portrait images; The portrait image is synthesized into an image of target resolution using a background of preset color; wherein, the target resolution is the resolution of the output image of the expression transfer model.

3. The method according to claim 1, characterized in that, Before inputting the image data into the facial expression transfer model, the process further includes: Gaussian noise is added to the source image in the image data according to a set probability.

4. The method according to claim 1, characterized in that, The step of extracting algorithm parameters from the source image and the driving image using an expression transfer model, and calculating implicit keypoints in the source image and the driving image using the algorithm parameters, to generate expression transfer results using the implicit keypoints in the source image and the driving image, includes: The image data is input into the expression transfer model to extract target parameters from the source image and the driving image, and to extract facial features and standardized key points from the source image; wherein, the target parameters include rotation parameters, offset parameters, scale, and expression deformation parameters; Using the target parameters of the source graph and the standardized key points, the implicit key points of the source graph are calculated, and using the target parameters of the driving graph and the standardized key points, the implicit key points of the driving graph are calculated. Based on the implicit keypoints of the source image, the implicit keypoints of the driving image, and the facial features, an expression transfer result is generated.

5. The method according to claim 1, characterized in that, The offset parameters of the driving graph, after the depth dimension is restricted to zero in the algorithm parameters, are used to calculate the keypoint prior loss for the implicit keypoints of the driving graph, with the mean depth dimension being zero. This includes: The depth dimension of the offset parameter of the driving graph in the algorithm parameters is restricted to zero; Using the offset parameters of the driving graph after the depth dimension is restricted to zero and the target keypoints in the implicit keypoints of the driving graph, the keypoint prior loss is calculated with the depth dimension mean being zero; wherein, the target keypoints refer to all keypoints except the guiding keypoints.

6. The method according to claim 5, characterized in that, The offset parameters of the driving graph after limiting the depth dimension to zero, and the target keypoints in the implicit keypoints of the driving graph, are used to calculate the keypoint prior loss with a depth dimension mean of zero, including: Calculate the squared distance between every two target key points in the implicit key points of the driving graph to obtain the squared distances corresponding to each target key point; For each of the target key points, the maximum distance corresponding to each target key point is obtained by subtracting the difference between the squares of the distances corresponding to the target key point and the maximum value among zero from the distance threshold. The maximum distance corresponding to the target key point, the difference between the mean depth dimension and the depth dimension value of the target key point are summed, and the summation result is combined with the offset parameter of the driving graph with the depth dimension restricted to zero to obtain the key point prior loss.

7. An optimization device for standardizing key points, characterized in that, include: An image acquisition unit is used to acquire image data; wherein the image data includes a source image and a driving image; The expression transfer unit is used to extract algorithm parameters from the source image and the driving image through the expression transfer model, and use the algorithm parameters to calculate the implicit key points of the source image and the driving image, so as to generate expression transfer results using the implicit key points of the source image and the driving image; The prior loss calculation unit is used to calculate the key point prior loss of the implicit key points of the driving graph based on the offset parameters of the driving graph after the depth dimension is restricted to zero in the algorithm parameters, and with the mean of the depth dimension being zero. The scale loss calculation unit is used to calculate the difference between 1 and the scale of the driving graph through an added activation function, and output the scale loss according to the calculated difference; wherein, when the difference is greater than 0, the difference is output as the scale loss; when the difference is not greater than 0, zero is output as the scale loss. An optimization unit is used to include the key point prior loss and the scale loss in the total model loss, and to optimize the expression transfer model based on the total model loss.

8. An electronic device, characterized in that, include: Memory and processor; The memory is used to store programs; The processor is used to execute the program, which, when executed, is specifically used to implement the optimization method for specification key points as described in any one of claims 1 to 6.

9. A computer storage medium, characterized in that, Used to store a computer program, which, when executed by a processor, is used to implement the optimization method for specification key points as described in any one of claims 1 to 6.