Face pinching result generation method and device, equipment and storage medium
By using a pre-trained parameter translator, the features of the target image are converted into face-pinching parameters, which solves the problem that users need to manually input a large number of parameters in the prior art, and improves the efficiency of face-pinching results.
Patent Information
- Application Number
- CN202510196540.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-27
AI Technical Summary
The existing face-pinching system requires users to manually enter or select a large number of face-pinching parameters, resulting in low generation efficiency.
By inputting the target image into a pre-trained parameter translator, the features of the target image are extracted and the target face pinching parameters matching these features are output, and the face pinching results matching the target image are directly generated.
There is no need for users to set massive face pinching parameters one by one, which significantly improves the efficiency of face pinching results.
Smart Images

Figure CN120047585A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and more particularly, to a method, apparatus, device, and storage medium for generating a face pinching result. Background Art
[0002] With the rise of role-playing games and short video platforms, users' demand for quickly generating realistic and personalized virtual avatars has been increasing. More and more platforms have started to provide intelligent face pinching services for users, supporting users to use the face pinching system provided in the platform to generate customized face pinching results (such as three-dimensional virtual user characters or two-dimensional user avatars, etc.), so that users can use the face pinching result as their virtual user image in the platform.
[0003] Currently, existing face pinching systems usually rely on a variety of face pinching parameters manually input or selected by users (such as eye size, pupil color, eye distance, nose bridge height, etc.) to automatically generate a face pinching result that matches the above-mentioned various face pinching parameters. However, due to the large variety of face pinching parameters and the need for users to have certain experience in adjusting face pinching parameters and aesthetic ability to generate a delicate face pinching result of a person's face, the generation efficiency of the face pinching result is relatively low. Summary of the Invention
[0004] In view of this, the present application provides a method, apparatus, device, and storage medium for generating a face pinching result. As long as the user provides a target image representing the face pinching requirement, it can assist the face pinching system to generate a target face pinching result that matches the above-mentioned target image, so that the user does not need to set a large number of face pinching parameters one by one, effectively improving the generation efficiency of the face pinching result.
[0005] To make the above objects, features, and advantages of the present application more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows.
[0006] In a first aspect, an embodiment of the present application provides a method for generating a face pinching result, and the generating method includes:
[0007] Input the target image into a pre-trained parameter translator, extract the features of the target image through the parameter translator to obtain the target image features of the target image, and output, through the parameter translator, target face pinching parameters that match the target image features;
[0008] Input the target face pinching parameters into the face pinching system, and generate, through the face pinching system, a face pinching result that matches the target face pinching parameters as the target face pinching result corresponding to the target image.
[0009] In a second aspect, an embodiment of the present application provides a device for generating a face pinching result, and the generating device includes:
[0010] A parameter generation unit, configured to input a target image into a pre-trained parameter translator, extract features of the target image through the parameter translator to obtain target image features of the target image, and output target face pinching parameters matching the target image features through the parameter translator;
[0011] An image generation unit, configured to input the target face pinching parameters into a face pinching system, and generate a face pinching result matching the target face pinching parameters through the face pinching system as the target face pinching result corresponding to the target image.
[0012] In a third aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the above method for generating a face pinching result are implemented.
[0013] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps of the above method for generating a face pinching result are executed.
[0014] The technical solutions provided by the embodiments of the present application may include the following beneficial effects:
[0015] A method, device, equipment, and storage medium for generating a face pinching result provided by an embodiment of the present application input a target image into a pre-trained parameter translator, extract features of the target image through the parameter translator to obtain target image features of the target image, and output target face pinching parameters matching the target image features through the parameter translator; input the target face pinching parameters into a face pinching system, and generate a face pinching result matching the target face pinching parameters through the face pinching system as the target face pinching result corresponding to the target image. Through the above generation method, the present application only requires the user to provide a target image representing the face pinching requirement, and can assist the face pinching system to generate a target face pinching result matching the above target image, so that the user does not need to set a large number of face pinching parameters one by one, effectively improving the generation efficiency of the face pinching result. Description of the Drawings
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0017] Figure 1Shows a schematic flowchart of a method for generating a face pinching result provided by an embodiment of the present application;
[0018] Figure 2 Shows a schematic flowchart of a method for generating a target face pinching result provided by an embodiment of the present application;
[0019] Figure 3 Shows a schematic flowchart of a self-supervised training method for an encoding and decoding model provided by an embodiment of the present application;
[0020] Figure 4 Shows a schematic flowchart of a training method for a parameter generation module provided by an embodiment of the present application;
[0021] Figure 5 Shows a schematic structural diagram of a device for generating a face pinching result provided by an embodiment of the present application;
[0022] Figure 6 Is a schematic structural diagram of an electronic device 600 provided by an embodiment of the present application. Detailed implementation manners
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. It should be understood that the accompanying drawings in the present application are only for the purposes of illustration and description, and are not used to limit the protection scope of the present application. In addition, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in the present application show operations implemented according to some embodiments of the present application. It should be understood that the operations in the flowchart may not be implemented in sequence, and steps without a logical context relationship may be reversed or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of the present application.
[0024] In addition, the described embodiments are only some embodiments of the present application, rather than all embodiments. The components of the embodiments of the present application usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application claimed, but only represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.
[0025] It should be noted that the term "including" will be used in the embodiments of the present application to indicate the existence of the features stated hereinafter, but does not exclude adding other features.
[0026] Currently, existing face sculpting systems usually rely on various face sculpting parameters manually input or selected by users (such as eye size, pupil color, eye distance, nose bridge height, etc.) to automatically generate face sculpting results that match the above-mentioned various face sculpting parameters. However, due to the large variety of face sculpting parameters and the need for users to have certain experience in adjusting face sculpting parameters and aesthetic ability to generate delicate face sculpting results, the generation efficiency of face sculpting results is relatively low.
[0027] Based on this, the embodiments of the present application provide a method, device, equipment and storage medium for generating face sculpting results. As long as the user provides a target image representing the face sculpting requirement, it can assist the face sculpting system to generate a target face sculpting result that matches the above-mentioned target image, enabling the user to avoid setting a large number of face sculpting parameters one by one and effectively improving the generation efficiency of face sculpting results.
[0028] In one of the embodiments of the present application, a method for generating face sculpting results can run on a terminal device or a server. Among them, the terminal device can be a local terminal device. When the method for generating face sculpting results runs on the server, the generation method can be implemented and executed based on a cloud interaction system, where the cloud interaction system includes a server and a client device (i.e., the terminal device).
[0029] For the convenience of understanding the embodiments of the present application, the following will introduce in detail a method, device, equipment and storage medium for generating face sculpting results provided by the embodiments of the present application.
[0030] Refer to Figure 1 as shown Figure 1 shows a schematic flowchart of a method for generating face sculpting results provided by the embodiments of the present application. Among them, the generation method includes steps S101 - S102; specifically:
[0031] S101, input the target image into a pre-trained parameter translator, extract the features of the target image through the parameter translator to obtain the target image features of the target image, and output target face sculpting parameters that match the target image features through the parameter translator.
[0032] S102, input the target face sculpting parameters into the face sculpting system, and generate a face sculpting result that matches the target face sculpting parameters as the target face sculpting result corresponding to the target image through the face sculpting system.
[0033] In the method for generating the above-mentioned face pinching result provided by the embodiment of the present application, the target image is input into a pre-trained parameter translator. The parameter translator extracts features from the target image to obtain the target image features of the target image, and outputs target face pinching parameters that match the target image features through the parameter translator. The target face pinching parameters are input into a face pinching system, and a face pinching image that matches the target face pinching parameters is output through the face pinching system as the target face pinching image corresponding to the target image. Through the above-mentioned generation method, the present application only requires the user to provide a target image representing the face pinching requirement, and can assist the face pinching system to generate a target face pinching image that matches the above-mentioned target image, so that the user does not need to set a large number of face pinching parameters one by one, effectively improving the generation efficiency of the face pinching image.
[0034] The following is an exemplary description of each step in the method for generating the above-mentioned face pinching result provided by the embodiment of the present application:
[0035] S101, input the target image into a pre-trained parameter translator, extract features from the target image through the parameter translator to obtain the target image features of the target image, and output target face pinching parameters that match the target image features through the parameter translator.
[0036] In the embodiment of the present application, different from the prior art that requires the user to manually input or select a large number of face pinching parameters in the face pinching system to represent their own face pinching requirements, the embodiment of the present application supports the user to use a target image that can represent their own face pinching requirements to replace the above traditional face pinching parameter input method. Among them, in order to reduce the impact on the existing face pinching system and reduce the software and hardware configuration costs, the face pinching result generation method provided by the embodiment of the present application does not involve improving the existing face pinching system itself (that is, the face pinching system still generates a face pinching result that matches the input face pinching parameters), but uses a pre-trained set of image feature extraction modules and parameter generation modules as the parameter translator, translates the target image input by the user into the target face pinching parameters required by the face pinching system through the above parameter translator, and outputs the translated target face pinching parameters to the face pinching system, so that the face pinching system can automatically generate a target face pinching result that matches the above target image according to the input target face pinching parameters, thereby enabling the user to not need to set a large number of face pinching parameters one by one, effectively improving the generation efficiency of the face pinching result.
[0037] It should be noted that the above target image is only used to represent the user's face pinching requirement through the image content. That is, the user only needs to select an image whose image content matches the face pinching requirement as the target image according to their own face pinching requirement. Moreover, the image style of the above target image will not affect the face pinching result generated by the face pinching system (equivalent to the generation style of the face pinching result being determined by the face pinching system itself, and the embodiments of the present application do not involve improving or training the face pinching system itself, etc.). Therefore, the image style of the target image can be a real face image or an anime character image in the second dimension style. The specific image style and specific image content of the target image are not limited in the embodiments of the present application.
[0038] For example, taking the case where the user wants to obtain a face pinching image of a male virtual image as an example, the user can choose to use a real male face image as the above target image, or can choose to use a second dimension character image of a male virtual character (such as a male game character or a male anime character, etc.) as the above target image. The image style of the target image does not need to be limited in any way and can be freely selected by the user.
[0039] Specifically, when using the above parameter translator (i.e., a pre-trained set of image feature extraction modules and parameter generation modules) to translate the target image input by the user into the target face pinching parameters required by the face pinching system, inside the above parameter translator, in the embodiments of the present application, the above target image is first input into the pre-trained image feature extraction module, and the image feature extraction module extracts features from the input above target image, and outputs the image features of the above target image as the target image features (that is, the target image features output by the image feature extraction module for the target image), so that the subsequent parameter generation module can generate face pinching parameters matching the above target image (that is, face pinching parameters matching the above target image features) according to the target image features.
[0040] It should be noted that in the embodiments of the present application, it is only necessary to ensure that the above image feature extraction module can implement the image feature extraction function (that is, it can extract image features from the input target image and output the target image features corresponding to the target image). Among them, the above image feature extraction module can be a pre-trained image feature extraction network (such as a convolutional neural network, etc.), or can be a pre-trained image encoder (equivalent to the above image feature extraction module can include the image encoder in the pre-trained encoding and decoding model). The specific structural type of the above image feature extraction module is not limited in the embodiments of the present application.
[0041] Here, inside the above-mentioned parameter translator, the above-mentioned image feature extraction module inputs the target image features of the above-mentioned target image into a pre-trained parameter generation module, and the parameter generation module generates face pinching parameters that match the input above-mentioned target image features as the target face pinching parameters that match the above-mentioned target image (that is, obtains the target face pinching parameters output by the parameter generation module for the target image features).
[0042] Specifically, the above-mentioned target face pinching parameters may include, but are not limited to: face feature parameters such as eye size, eye distance, nose bridge height, mouth opening amplitude, mandible height, etc., and makeup feature parameters such as eyebrow color, thickness, skin gloss, pupil type, blush type, eyeshadow type, face decoration type, etc. The specific parameter types and specific parameter quantities of the above-mentioned target face pinching parameters (which is also equivalent to the parameter types and parameter quantities of the face pinching parameters that the parameter generation module can generate) are not limited in any way in the embodiments of the present application.
[0043] S102, input the above-mentioned target face pinching parameters into the face pinching system, and generate a face pinching result that matches the above-mentioned target face pinching parameters through the face pinching system as the target face pinching result corresponding to the above-mentioned target image.
[0044] Here, the embodiments of the present application do not make any changes to the face pinching system. That is, the face pinching system still relies on the input target face pinching parameters to generate a face pinching result that matches the above-mentioned target face pinching parameters (that is, the above-mentioned target face pinching result). It's just that the above-mentioned target face pinching parameters are no longer manually input into the face pinching system one by one by the user, but are directly obtained by translating the target image input by the user through the above-mentioned parameter translator (that is, a pre-trained set of image feature extraction modules and parameter generation modules). Therefore, in the embodiments of the present application, the user only needs to provide a target image representing the face pinching requirement to assist the face pinching system in generating a target face pinching image that matches the above-mentioned target image, enabling the user to avoid setting a large number of face pinching parameters one by one, effectively improving the generation efficiency of the face pinching result. And since the implementation of the embodiments of the present application does not require changing the face pinching system itself, it is also conducive to reducing the impact on the existing face pinching system, reducing the software and hardware configuration costs, and facilitating the popularization and implementation of the technical solutions of the present application.
[0045] Specifically, the face pinching system is a system that supports users to customize and create character images. Among them, the face pinching system can adjust the control parameters corresponding to each face pinching component in the system according to the input above-mentioned target face pinching parameters, so as to generate a face pinching result that matches the input above-mentioned target face pinching parameters (that is, the target face pinching result).
[0046] In the embodiments of the present application, since no modifications are made to the face pinching system itself, the above-mentioned target face pinching result specifically belongs to a three-dimensional virtual character model or a two-dimensional virtual character image depending on the face pinching system itself. That is, if the face pinching system is originally used to generate a three-dimensional virtual character model that matches the input face pinching parameters, then when performing the above-mentioned step S102, the face pinching system can generate a three-dimensional virtual character model that matches the input target face pinching parameters as the above-mentioned target face pinching result; if the face pinching system is originally used to generate a two-dimensional virtual character image that matches the input face pinching parameters, then when performing the above-mentioned step S102, the face pinching system can generate a two-dimensional virtual character image that matches the input target face pinching parameters as the above-mentioned target face pinching result; for the specific result type of the above-mentioned target face pinching result, the embodiments of the present application do not make any limitations.
[0047] It should be noted that the above-mentioned face pinching system can be an Avatar face pinching system or other types of face pinching systems. For the specific system type of the above-mentioned face pinching system, the embodiments of the present application do not make any limitations.
[0048] Specifically, for the specific implementation processes of the above-mentioned steps S101 - S102, taking the above-mentioned target image as an anime character image of a girl as an example, Figure 2 FIG. shows a schematic flowchart of a method for generating a target face pinching result provided by an embodiment of the present application. As Figure 2 shown, on the side of software and hardware configuration, the method for generating a face pinching result provided by an embodiment of the present application does not require changing the face pinching system itself. Only by connecting a parameter translator composed of a pre-trained group of image feature extraction modules and parameter generation modules to the front end of the face pinching system can it support users to use the image input method to replace the cumbersome input method of manually inputting face pinching parameters.
[0049] In addition, as Figure 2 shown, on the side of the specific generation process of the face pinching result, the method for generating a face pinching result provided by an embodiment of the present application only needs to obtain a target image input by the user as the source data input by the user, translate the target image input by the user into the target face pinching parameters required by the face pinching system through the above-mentioned parameter translator, and output the translated target face pinching parameters to the face pinching system, so that the face pinching system can automatically generate a target face pinching result that matches the above-mentioned target image according to the input target face pinching parameters, thereby enabling the user to avoid setting a large number of face pinching parameters one by one and effectively improving the generation efficiency of the face pinching result.
[0050] The following will separately describe in detail the specific implementation processes of the above-mentioned steps in the embodiments of the present application:
[0051] For the image feature extraction module in the above parameter translator, when using the image encoder in the pre-trained codec model as the actual image feature extraction module in the parameter translator in step S101 above, as an optional embodiment, a conventional codec model training method can be adopted to train the above codec model to obtain the image encoder in the trained codec model as the above image feature extraction module; for example, a large number of face images are collected as the model training data set, and the face images in the model training data set are input into the image encoder of the codec model one by one to obtain the image encoding features (i.e., image encoding results) output by the image encoder for the above face images. The above image encoding features output by the image encoder are input into the decoder of the codec model to obtain the decoding results output by the decoder. According to the image loss between the decoding results and the above face images, the above codec model is trained until the above codec model converges.
[0052] In addition, as another optional embodiment, a self-supervised training method can also be adopted to train the above codec model. Under this optional embodiment, Figure 3 shows a schematic flowchart of a self-supervised training method for a codec model provided by an embodiment of the present application, as Figure 3 shown, the self-supervised training method includes steps S301-S305; specifically:
[0053] S301, collect face images of multiple different styles as the training data set of the codec model.
[0054] Here, to ensure that the trained image encoder (i.e., the above image feature extraction module) can be applied to extract image features from input images of any style (that is, for input images of any style, the image encoder can output image feature representations with a unified feature dimension), when performing self-supervised training on the above codec model, face images of multiple different styles such as real face images and anime face images in the second-dimensional style can be collected as the training data set of the above codec model.
[0055] It should be noted that the training dataset of the above encoding and decoding model may also include face images sampled from the above face shaping system (i.e., the specific face shaping system actually applied in the above step S102). Among them, the face images are determined according to the historical face shaping results generated by the above face shaping system. For example, if the above face shaping system is originally used to generate a three-dimensional virtual character model matching the input face shaping parameters, the historical three-dimensional virtual character model generated by the face shaping system (i.e., the historical face shaping result belongs to a three-dimensional virtual model) can be collected from the face shaping system, and the face image of the above historical three-dimensional virtual character model can be automatically intercepted as the above face image; if the above face shaping system is originally used to generate a two-dimensional virtual character image matching the input face shaping parameters, the historical two-dimensional virtual character image generated by the face shaping system (i.e., the historical face shaping result belongs to a two-dimensional image) can be collected from the face shaping system as the above face image; regarding the style type and the number of face images included in the above training dataset, the embodiments of the present application do not make any limitations.
[0056] S302. For the face images in the training dataset, divide the face images into multiple image units, and randomly mask the multiple image units according to a preset mask ratio to obtain a masked image of the face images.
[0057] Here, in the self-supervised training stage, it is not necessary to annotate the face images in the above training dataset. Among them, for each face image in the above training dataset, it can be divided into multiple image units with an image size equal to the above unit image size according to a preset unit image size; or it can be divided into multiple image units with an image number equal to the above division number according to a preset division number; regarding whether it is necessary to evenly divide the face images according to the same unit image size during the image division process, the embodiments of the present application do not make a mandatory limitation.
[0058] Here, the above mask ratio can be flexibly adjusted according to actual self-supervised training requirements. For example, the above mask ratio can be preset to 80% (i.e., randomly mask 80% of the image units among all the divided image units), or the above mask ratio can be preset to 70% (i.e., randomly mask 70% of the image units among all the divided image units). Regarding the specific value of the above mask ratio, the embodiments of the present application do not make any limitations.
[0059] Exemplarily, taking the example of segmenting the face image into multiple image units with the same size as the preset unit image size, if the preset unit image size is 16*16, for each face image in the training dataset, the face image can be segmented into multiple 16*16 image units, and according to the preset masking ratio of 80%, randomly mask 80% of the image units among all the segmented image units to obtain the masked image of the face image.
[0060] S303. Input the masked image into the image encoder of the codec model to obtain the image features output by the image encoder for the masked image.
[0061] Here, the image encoder can be a VIT (Vision Transformer) encoder or other types of image encoders. For the specific encoder type of the above image encoder, the embodiments of the present application do not make any limitations.
[0062] Exemplarily, taking the above image encoder as a VIT encoder as an example, input the above masked image into the VIT encoder. The VIT encoder can transform the input masked image into f i ∈R 257×768 features. Among them, as an optional embodiment, it can be preset to select the first global feature among 257 features as the image features finally output by the VIT encoder for the above masked image (i.e., f ∈ R 1×768 ), so that the trained VIT encoder can output image feature representations with consistent feature dimension distributions for face images of any style input.
[0063] S304. Input the image features into the decoder of the codec model to obtain the reconstructed image output by the decoder for the image features.
[0064] Here, in the codec model, the specific structure of the decoder only needs to be ensured to match the specific structure of the above image encoder. For example, if a VIT encoder is selected as the above image encoder, a VIT decoder can be correspondingly selected as the decoder in the above codec model. For the specific structure of the above decoder, the embodiments of the present application also do not make any limitations.
[0065] Specifically, the image encoder can input the image features output for the above masked image into the above decoder. The decoder can obtain the decoding result as the reconstructed image corresponding to the above masked image by performing decoding processing on the input image features.
[0066] S305. Adjust the model parameters of the encoding and decoding model according to the image loss between the reconstructed image and the face image, to obtain the encoding and decoding model including the adjusted model parameters.
[0067] Here, since the model training objective of the encoding and decoding model is mainly to train the decoder to output the reconstructed image that can restore the face image before the mask as much as possible according to the above image features output by the image encoder, therefore, during the model training process of the encoding and decoding model, the training loss of the model belongs to the reconstruction loss. The L2 loss function can be preferentially used to calculate the image loss between the above reconstructed image and the above face image (that is, the face image corresponding to the above mask image before mask processing), that is, the mean square error between the above reconstructed image and the above face image can be calculated as the image loss between the above reconstructed image and the above face image.
[0068] It should be noted that, in addition to the L2 loss function, other types of loss functions can also be used to calculate the image loss between the above reconstructed image and the above face image. The specific type of loss function actually used is not limited in the embodiments of the present application.
[0069] Specifically, since the training data set contains multiple face images, multiple groups of face images and reconstructed images can be obtained. Among them, after calculating the image loss between each group of face images and reconstructed images, the sum of the image losses between each group of face images and reconstructed images can be calculated as the overall model loss of the encoding and decoding model, and the encoding and decoding model is trained until the encoding and decoding model converges (equivalent to continuously adjusting the model parameters of the encoding and decoding model until the above overall model loss reaches the minimum), and the converged encoding and decoding model is obtained as the trained encoding and decoding model (that is, the encoding and decoding model including the adjusted model parameters), so that the image encoder in the trained encoding and decoding model can be obtained as the pre-trained image feature extraction module in the above step S101.
[0070] For the parameter generation module in the above parameter translator, in an alternative embodiment, Figure 4 shows a schematic flowchart of a training method for a parameter generation module provided by an embodiment of the present application, as Figure 4 shown, the training method includes steps S401 - S404; specifically:
[0071] S401. Randomly sample multiple pairs of data of original face - shaping parameters and original face - shaping images from the face - shaping system.
[0072] Here, when training the parameter generation module in the parameter translator, considering that the face sculpting styles of the face sculpting results generated by different face sculpting systems may vary, in order to improve the matching degree between the face sculpting parameters generated by the parameter generation module and the face sculpting system in actual application, multiple sets of paired data of original face sculpting parameters and original face sculpting images can be randomly sampled from the face sculpting system in step S103 above (i.e., the face sculpting system in actual application); among them, the original face sculpting image in each set of paired data is determined according to the face sculpting result generated by the above face sculpting system according to the original face sculpting parameters in this set of paired data.
[0073] Exemplarily, for the original face sculpting parameters in each set of paired data, if the above face sculpting system is originally used to generate a three-dimensional virtual character model matching the input face sculpting parameters, the historical three-dimensional virtual character model generated by the face sculpting system according to the input original face sculpting parameters can be collected from the face sculpting system, and the face image of the historical three-dimensional virtual character model can be automatically intercepted as the above original face sculpting image in the same set of paired data; if the above face sculpting system is originally used to generate a two-dimensional virtual character image matching the input face sculpting parameters, the historical two-dimensional virtual character image generated by the face sculpting system according to the input original face sculpting parameters can be collected from the face sculpting system as the above original face sculpting image in the same set of paired data.
[0074] Specifically, after randomly sampling multiple sets of paired data from the above face sculpting system, as an optional embodiment, the deformed paired data (such as the paired data where the original face sculpting image has a distorted face) can also be removed, and only the paired data where the original face sculpting image is reasonable and natural is retained to improve the generation accuracy of the face sculpting parameters by the parameter generation module.
[0075] It should be noted that the present application embodiment does not make any limitations on the specific number of sampling groups of the above paired data and the specific parameter types and parameter quantities of the above original face sculpting parameters included in each set of paired data.
[0076] S402, input the original face sculpting image in the paired data into the pre-trained image feature extraction module, and obtain the face sculpting image features output by the image feature extraction module for the original face sculpting image.
[0077] Here, when training the parameter generation module in the parameter translator, it is necessary to fix the model parameters of the image feature extraction module in the parameter translator to remain unchanged. Therefore, when training the parameter generation module, the pre-trained image feature extraction module can be used to perform image feature extraction on the original face sculpting image in each set of paired data (equivalent to not adjusting the model parameters of the image feature extraction module during the model training stage of the parameter generation module).
[0078] Specifically, the specific implementation manner of step S402 is the same as that of the foregoing step S101, and the repeated parts will not be elaborated here.
[0079] S403. Input the face pinching image features into the parameter generation module, and obtain the predicted face pinching parameters output by the parameter generation module for the face pinching image features.
[0080] Here, the parameter generation module can be a deep learning model with a multi-layer fully connected layer structure (such as a deep learning model with a two-layer fully connected layer structure), or a neural network learning model with other structures. The present application embodiment does not make any limitation on the specific model structure corresponding to the above parameter generation module.
[0081] Specifically, the specific implementation manner of step S403 is the same as that of the foregoing step S102, and the repeated parts will not be elaborated here.
[0082] S404. Adjust the model parameters of the parameter generation module according to the error loss between the predicted face pinching parameters and the original face pinching parameters in the paired data, and obtain the parameter generation module including the adjusted model parameters.
[0083] Here, for the predicted face pinching parameters output by the parameter generation module, the above predicted face pinching parameters can be classified into the following two categories according to whether the parameter values are continuous values:
[0084] Type 1: Continuous face pinching parameters with continuous parameter values. For example, face pinching parameters such as eye size, eye distance, nose bridge height, mouth opening amplitude, jaw height, eyebrow color, thickness, and skin glossiness have continuous values that can be continuously adjusted within a parameter value range interval (such as a value range interval of 0-1); among them, the parameter value range intervals corresponding to continuous face pinching parameters of different parameter types can be different.
[0085] Type 2: Discrete face pinching parameters with discrete parameter values. For example, face pinching parameters such as pupil type, blush type, eyeshadow type, and face flower type have specific parameter values selected from a finite number of one or more discrete numerical values.
[0086] In the embodiments of the present application, considering that different types of predicted face pinching parameters are applicable to different types of loss functions, therefore, in the case of classifying the predicted face pinching parameters into the above two categories of continuous face pinching parameters and discrete face pinching parameters, as an optional embodiment, the error loss between the above predicted face pinching parameters and the original face pinching parameters in the paired data can be calculated according to the method shown in the following steps a1-a3. Specifically:
[0087] Step a1: For the continuous face pinching parameters in the predicted face pinching parameters, determine, from the original face pinching parameters, a first original face pinching parameter of the same parameter type as the continuous face pinching parameter, and calculate the absolute error loss between the continuous face pinching parameter and the first original face pinching parameter.
[0088] Here, the absolute error loss is also known as the L1 loss, which is mainly used to calculate the sum of the absolute errors between the predicted value (i.e., the continuous face pinching parameter in the above predicted face pinching parameters) and the true value (i.e., the first original face pinching parameter above). That is, the L1 loss function can be used to calculate the absolute error loss between the continuous face pinching parameter and the first original face pinching parameter as the error loss between the continuous face pinching parameter and the first original face pinching parameter.
[0089] Step a2: For the discrete face pinching parameters in the predicted face pinching parameters, determine, from the original face pinching parameters, a second original face pinching parameter of the same parameter type as the discrete face pinching parameter, and calculate the classification loss between the discrete face pinching parameter and the second original face pinching parameter.
[0090] Here, since the specific parameter values of the discrete face pinching parameters are equivalent to the parameter value results predicted by the parameter generation module from a finite number of discrete values, the classification loss function can be used to calculate the classification loss between the discrete face pinching parameter and the second original face pinching parameter as the error loss between the discrete face pinching parameter and the second original face pinching parameter.
[0091] Step a3: Determine the overall training loss corresponding to the parameter generation module according to the absolute error loss and the classification loss, and adjust the model parameters of the parameter generation module according to the overall training loss to obtain the parameter generation module including the adjusted module parameters.
[0092] Here, the sum of the absolute error loss and the classification loss can be calculated as the overall training loss corresponding to the parameter generation module, and the parameter generation module is trained according to the calculated overall training loss until the parameter generation module converges (equivalent to continuously adjusting the model parameters of the parameter generation module until the above overall model loss reaches the minimum), and the converged parameter generation module is obtained as the trained parameter generation module (i.e., the parameter generation module including the adjusted model parameters), so that the pre-trained parameter generation module in step S102 above can be obtained.
[0093] Regarding the training method shown in the above steps S401 - S404, it should be noted that: in the existing technical solution, in the first training stage, a neural renderer that can generate a face - sculpting image matching the specified face - sculpting parameters according to the input specified face - sculpting parameters needs to be trained using the face - sculpting images and face - sculpting parameters sampled in the face - sculpting system; then, in the second training stage, a pre - trained face recognition model is additionally introduced. At this time, first obtain the initial face - sculpting image generated by the neural renderer trained in the first training stage according to the input initial face - sculpting parameters (which can be randomly generated face - sculpting parameters) and the original image that the user specifies for face - sculpting. Then, input the above - mentioned initial face - sculpting image and the above - mentioned original image into the above - mentioned face recognition model respectively. According to the recognition error loss between the two face feature recognition results output by the face recognition model for the above - mentioned initial face - sculpting image and the above - mentioned original image respectively, the above - mentioned initial face - sculpting parameters are adjusted and optimized in the reverse direction until specific face - sculpting parameters matching the above - mentioned original image specified by the user are obtained; that is to say, the existing technical solution also needs to use the neural renderer trained in the first training stage to execute the optimization step regarding the initial face - sculpting parameters (i.e., the face - sculpting parameters input into the neural renderer).
[0094] In the embodiment of the present application, however, the parameter generation module is trained according to the error loss between the predicted face - sculpting parameters output by the parameter generation module and the real face - sculpting parameters actually generated by the face - sculpting system (i.e., the original face - sculpting parameters in the above - mentioned paired data), so as to obtain a parameter translator composed of a pre - trained image feature extraction module and a parameter generation module. Among them, this parameter translator belongs to an end - to - end feed - forward deep - learning model. Therefore, the trained above - mentioned parameter translator can directly generate target face - sculpting parameters matching the target image according to the target image input by the user, and the generated above - mentioned target face - sculpting parameters can also be directly input into the existing face - sculpting system (equivalent to that the trained parameter translator can be directly docked with the user side and the face - sculpting system). Thus, compared with the above - mentioned existing technical solution, on the one hand, the embodiment of the present application does not need to additionally introduce a pre - trained face recognition model in the training stage, and on the other hand, it does not need to use the trained parameter translator to perform any subsequent optimization steps regarding the face - sculpting parameters; furthermore, it enables the embodiment of the present application to effectively improve the model training efficiency of the parameter translator.
[0095] Based on the above-mentioned method for generating a face pinching result provided by the embodiments of the present application, the target image is input into a pre-trained parameter translator. The parameter translator extracts features from the target image to obtain the target image features of the target image, and outputs target face pinching parameters that match the target image features through the parameter translator. The target face pinching parameters are input into a face pinching system, and the face pinching system generates a face pinching result that matches the target face pinching parameters as the target face pinching result corresponding to the target image. Through the above generation method, the present application only requires the user to provide a target image representing the face pinching requirement, and can assist the face pinching system to generate a target face pinching result that matches the above target image, so that the user does not need to set a large number of face pinching parameters one by one, effectively improving the generation efficiency of the face pinching result.
[0096] Based on the same inventive concept, the present application also provides a device for generating a face pinching result corresponding to the above-mentioned method for generating a face pinching result. Since the principle of solving problems by the device for generating a face pinching result in the embodiments of the present application is similar to that of the above-mentioned method for generating a face pinching result in the embodiments of the present application, the implementation of the device for generating a face pinching result can refer to the implementation of the above-mentioned method for generating a face pinching result, and the repeated parts will not be elaborated.
[0097] Referring to Figure 5 as shown in Figure 5 FIG. shows a schematic structural diagram of a device for generating a face pinching result provided by an embodiment of the present application. Among them, the generating device includes:
[0098] A parameter generation unit 501, configured to input a target image into a pre-trained parameter translator, extract features from the target image through the parameter translator to obtain the target image features of the target image, and output target face pinching parameters that match the target image features through the parameter translator;
[0099] An image generation unit 502, configured to input the target face pinching parameters into a face pinching system, and generate a face pinching result that matches the target face pinching parameters as the target face pinching result corresponding to the target image through the face pinching system.
[0100] In an optional implementation manner, the parameter translator includes: a pre-trained image feature extraction module and a parameter generation module. Among them, the parameter generation unit 501 is specifically configured to:
[0101] Input the target image into the pre-trained image feature extraction module to obtain the target image features output by the image feature extraction module for the target image;
[0102] Input the target image features into the pre-trained parameter generation module to obtain the target face pinching parameters output by the parameter generation module for the target image features.
[0103] In an alternative embodiment, the image feature extraction module includes: an image encoder in a pre-trained codec model.
[0104] In an alternative embodiment, the generating device further includes a first training unit; wherein, the first training unit is configured to:
[0105] Collect face images of multiple different styles as the training data set of the codec model;
[0106] For the face images in the training data set, segment the face images into multiple image units, and randomly mask the multiple image units according to a preset mask ratio to obtain a masked image of the face images;
[0107] Input the masked image into the image encoder of the codec model to obtain image features output by the image encoder for the masked image;
[0108] Input the image features into the decoder of the codec model to obtain a reconstructed image output by the decoder for the image features;
[0109] Adjust the model parameters of the codec model according to the image loss between the reconstructed image and the face image to obtain the codec model including the adjusted model parameters.
[0110] In an alternative embodiment, the training data set of the codec model includes face pinching images sampled from the face pinching system; wherein, the face pinching images are determined according to historical face pinching results generated by the face pinching system.
[0111] In an alternative embodiment, the generating device further includes a second training unit; wherein, the second training unit is configured to:
[0112] Randomly sample multiple sets of paired data of original face pinching parameters and original face pinching images from the face pinching system; wherein, the original face pinching images are determined according to face pinching results generated by the face pinching system according to the original face pinching parameters;
[0113] Input the original face pinching images in the paired data into the pre-trained image feature extraction module to obtain face pinching image features output by the image feature extraction module for the original face pinching images;
[0114] Input the face pinching image features into the parameter generation module to obtain predicted face pinching parameters output by the parameter generation module for the face pinching image features;
[0115] Adjust the model parameters of the parameter generation module according to the error loss between the predicted face pinching parameters and the original face pinching parameters in the paired data, to obtain the parameter generation module including the adjusted model parameters.
[0116] In an alternative embodiment, the predicted face pinching parameters include: continuous face pinching parameters with continuous parameter values and discrete face pinching parameters with discrete parameter values.
[0117] In an alternative embodiment, when adjusting the model parameters of the parameter generation module according to the error loss between the predicted face pinching parameters and the original face pinching parameters in the paired data, the second training unit is configured to:
[0118] For the continuous face pinching parameters in the predicted face pinching parameters, determine, from the original face pinching parameters, a first original face pinching parameter of the same parameter type as the continuous face pinching parameters, and calculate the absolute error loss between the continuous face pinching parameters and the first original face pinching parameters;
[0119] For the discrete face pinching parameters in the predicted face pinching parameters, determine, from the original face pinching parameters, a second original face pinching parameter of the same parameter type as the discrete face pinching parameters, and calculate the classification loss between the discrete face pinching parameters and the second original face pinching parameters;
[0120] Determine the overall training loss corresponding to the parameter generation module according to the absolute error loss and the classification loss, and adjust the model parameters of the parameter generation module according to the overall training loss, to obtain the parameter generation module including the adjusted model parameters.
[0121] Based on the above-mentioned face pinching result generation device provided by the embodiments of the present application, input a target image into a pre-trained parameter translator, extract the target image features of the target image through the parameter translator, obtain the target face pinching parameters that match the target image features through the parameter translator; input the target face pinching parameters into the face pinching system, and generate a face pinching result that matches the target face pinching parameters through the face pinching system as the target face pinching result corresponding to the target image. Through the above generation method, the present application only requires the user to provide a target image representing the face pinching requirement, and can assist the face pinching system to generate a target face pinching result that matches the above target image, so that the user does not need to set a large number of face pinching parameters one by one, effectively improving the generation efficiency of the face pinching result.
[0122] Based on the same inventive concept, this application also provides an electronic device corresponding to the above-mentioned method for generating a face pinching result. Since the principle of solving problems by the electronic device in the embodiments of this application is similar to that of the above-mentioned method for generating a face pinching result in the embodiments of this application, the implementation of the electronic device can refer to the implementation of the above-mentioned method for generating a face pinching result, and the repeated parts will not be elaborated.
[0123] Figure 6 FIG. 4 is a schematic structural diagram of an electronic device 600 provided in an embodiment of this application, including: a processor 601, a memory 602, and a bus 603. The memory 602 stores machine-readable instructions executable by the processor 601. When the electronic device runs a method for generating a face pinching result as in the embodiment, the processor 601 communicates with the memory 602 through the bus 603, and the processor 601 executes the machine-readable instructions. Among them, when the processor 601 executes the machine-readable instructions, the following steps are implemented, specifically:
[0124] Input the target image into a pre-trained parameter translator, extract features of the target image through the parameter translator to obtain the target image features of the target image, and output, through the parameter translator, target face pinching parameters matching the target image features;
[0125] Input the target face pinching parameters into a face pinching system, and generate, through the face pinching system, a face pinching result matching the target face pinching parameters as the target face pinching result corresponding to the target image.
[0126] In an optional implementation manner, the parameter translator includes: a pre-trained image feature extraction module and a parameter generation module. Among them, when inputting the target image into the pre-trained parameter translator, extracting features of the target image through the parameter translator to obtain the target image features of the target image, and outputting, through the parameter translator, target face pinching parameters matching the target image features, the processor 601 is configured to:
[0127] Input the target image into the pre-trained image feature extraction module to obtain the target image features output by the image feature extraction module for the target image;
[0128] Input the target image features into the pre-trained parameter generation module to obtain the target face pinching parameters output by the parameter generation module for the target image features.
[0129] In an optional implementation manner, the image feature extraction module includes: an image encoder in a pre-trained codec model.
[0130] In an optional implementation manner, the processor 601 is further configured to:
[0131] Collect face images of multiple different styles as the training data set of the encoding and decoding model;
[0132] For the face images in the training data set, segment the face images into multiple image units, and randomly mask the multiple image units according to a preset mask ratio to obtain a masked image of the face images;
[0133] Input the masked image into the image encoder of the encoding and decoding model to obtain the image features output by the image encoder for the masked image;
[0134] Input the image features into the decoder of the encoding and decoding model to obtain the reconstructed image output by the decoder for the image features;
[0135] Adjust the model parameters of the encoding and decoding model according to the image loss between the reconstructed image and the face images to obtain the encoding and decoding model including the adjusted model parameters.
[0136] In an alternative embodiment, the training data set of the encoding and decoding model includes face-morphing images sampled from the face-morphing system; wherein, the face-morphing images are determined according to the historical face-morphing results generated by the face-morphing system.
[0137] In an alternative embodiment, the processor 601 is further configured to:
[0138] Randomly sample multiple sets of paired data of original face-morphing parameters and original face-morphing images from the face-morphing system; wherein, the original face-morphing images are determined according to the face-morphing results generated by the face-morphing system according to the original face-morphing parameters;
[0139] Input the original face-morphing images in the paired data into the pre-trained image feature extraction module to obtain the face-morphing image features output by the image feature extraction module for the original face-morphing images;
[0140] Input the face-morphing image features into the parameter generation module to obtain the predicted face-morphing parameters output by the parameter generation module for the face-morphing image features;
[0141] Adjust the model parameters of the parameter generation module according to the error loss between the predicted face-morphing parameters and the original face-morphing parameters in the paired data to obtain the parameter generation module including the adjusted model parameters.
[0142] In an alternative embodiment, the predicted face-morphing parameters include: continuous face-morphing parameters with continuous parameter values and discrete face-morphing parameters with discrete parameter values.
[0143] In an alternative embodiment, when adjusting the model parameters of the parameter generation module according to the error loss between the predicted face pinching parameters and the original face pinching parameters in the paired data, the processor 601 is configured to:
[0144] For the continuous face pinching parameters in the predicted face pinching parameters, determine, from the original face pinching parameters, a first original face pinching parameter of the same parameter type as the continuous face pinching parameters, and calculate the absolute error loss between the continuous face pinching parameters and the first original face pinching parameter;
[0145] For the discrete face pinching parameters in the predicted face pinching parameters, determine, from the original face pinching parameters, a second original face pinching parameter of the same parameter type as the discrete face pinching parameters, and calculate the classification loss between the discrete face pinching parameters and the second original face pinching parameter;
[0146] Determine the overall training loss corresponding to the parameter generation module according to the absolute error loss and the classification loss, and adjust the model parameters of the parameter generation module according to the overall training loss to obtain the parameter generation module including the adjusted model parameters.
[0147] By the above electronic device provided in the embodiments of the present application, input a target image into a pre-trained parameter translator, extract features of the target image through the parameter translator to obtain target image features of the target image, and output target face pinching parameters matching the target image features through the parameter translator; input the target face pinching parameters into a face pinching system, and generate a face pinching result matching the target face pinching parameters through the face pinching system as the target face pinching result corresponding to the target image. Through the above generation method, the present application only requires the user to provide a target image representing the face pinching requirement, and can assist the face pinching system to generate a target face pinching result matching the above target image, so that the user does not need to set a large number of face pinching parameters one by one, effectively improving the generation efficiency of the face pinching result.
[0148] Based on the same inventive concept, the embodiments of the present application further provide a computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, the processor executes the following steps:
[0149] Input a target image into a pre-trained parameter translator, extract features of the target image through the parameter translator to obtain target image features of the target image, and output target face pinching parameters matching the target image features through the parameter translator;
[0150] Input the target face sculpting parameters into the face sculpting system, and generate a face sculpting result that matches the target face sculpting parameters through the face sculpting system as the target face sculpting result corresponding to the target image.
[0151] In an alternative embodiment, the parameter translator includes: a pre-trained image feature extraction module and a parameter generation module. Among them, when inputting the target image into the pre-trained parameter translator, extracting the features of the target image through the parameter translator to obtain the target image features of the target image, and outputting the target face sculpting parameters that match the target image features through the parameter translator, the processor is used for:
[0152] Input the target image into the pre-trained image feature extraction module to obtain the target image features output by the image feature extraction module for the target image;
[0153] Input the target image features into the pre-trained parameter generation module to obtain the target face sculpting parameters output by the parameter generation module for the target image features.
[0154] In an alternative embodiment, the image feature extraction module includes: an image encoder in a pre-trained codec model.
[0155] In an alternative embodiment, the processor is further used for:
[0156] Collect face images of multiple different styles as the training dataset of the codec model;
[0157] For the face images in the training dataset, segment the face images into multiple image units, and randomly mask the multiple image units according to a preset mask ratio to obtain the masked images of the face images;
[0158] Input the masked images into the image encoder of the codec model to obtain the image features output by the image encoder for the masked images;
[0159] Input the image features into the decoder of the codec model to obtain the reconstructed images output by the decoder for the image features;
[0160] Adjust the model parameters of the codec model according to the image loss between the reconstructed images and the face images to obtain the codec model including the adjusted model parameters.
[0161] In an alternative embodiment, the training dataset of the encoding and decoding model includes face pinching images sampled from the face pinching system; wherein, the face pinching images are determined according to historical face pinching results generated by the face pinching system.
[0162] In an alternative embodiment, the processor is further configured to:
[0163] Randomly sample multiple sets of paired data of original face pinching parameters and original face pinching images from the face pinching system; wherein, the original face pinching images are determined according to face pinching results generated by the face pinching system according to the original face pinching parameters;
[0164] Input the original face pinching images in the paired data into the pre-trained image feature extraction module to obtain face pinching image features output by the image feature extraction module for the original face pinching images;
[0165] Input the face pinching image features into the parameter generation module to obtain predicted face pinching parameters output by the parameter generation module for the face pinching image features;
[0166] Adjust the model parameters of the parameter generation module according to the error loss between the predicted face pinching parameters and the original face pinching parameters in the paired data to obtain the parameter generation module including the adjusted model parameters.
[0167] In an alternative embodiment, the predicted face pinching parameters include: continuous face pinching parameters with continuous parameter values and discrete face pinching parameters with discrete parameter values.
[0168] In an alternative embodiment, when adjusting the model parameters of the parameter generation module according to the error loss between the predicted face pinching parameters and the original face pinching parameters in the paired data, the processor is configured to:
[0169] For the continuous face pinching parameters in the predicted face pinching parameters, determine first original face pinching parameters of the same parameter type as the continuous face pinching parameters from the original face pinching parameters, and calculate the absolute error loss between the continuous face pinching parameters and the first original face pinching parameters;
[0170] For the discrete face pinching parameters in the predicted face pinching parameters, determine second original face pinching parameters of the same parameter type as the discrete face pinching parameters from the original face pinching parameters, and calculate the classification loss between the discrete face pinching parameters and the second original face pinching parameters;
[0171] Determine the overall training loss corresponding to the parameter generation module according to the absolute error loss and the classification loss, and adjust the model parameters of the parameter generation module according to the overall training loss to obtain the parameter generation module including the adjusted model parameters.
[0172] By using the computer-readable storage medium provided in the embodiments of the present application, input the target image into a pre-trained parameter translator, extract the target image features of the target image through the parameter translator, and output the target face morphing parameters matching the target image features through the parameter translator; input the target face morphing parameters into the face morphing system, and generate a face morphing result matching the target face morphing parameters through the face morphing system as the target face morphing result corresponding to the target image. Through the above generation method, the present application only requires the user to provide a target image representing the face morphing requirement, and can assist the face morphing system to generate a target face morphing result matching the above target image, so that the user does not need to set a large number of face morphing parameters one by one, effectively improving the generation efficiency of the face morphing result.
[0173] In the embodiments of the present application, when the computer-readable storage medium is run by a processor, it can also execute other machine-readable instructions to execute the face morphing result generation method described in other parts of the embodiments. For the specific steps and principles of the executed generation method, refer to the description of the method-side embodiments, which will not be repeated here.
[0174] In the embodiments provided in the present application, it should be understood that the disclosed system and method can be implemented in other ways. The system embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of the system or unit can be in electrical, mechanical or other forms.
[0175] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0176] In addition, each functional unit in the embodiments provided in the present application can be integrated into one processing unit, or each unit exists physically alone, or two or more units can be integrated into one unit.
[0177] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0178] It should be noted that: similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In addition, the terms "first", "second", "third", etc. are only used for descriptive distinction and cannot be understood as indicating or implying relative importance.
[0179] Finally, it should be noted that: the above-mentioned embodiments are only specific implementation manners of this application, used to illustrate the technical solution of this application, rather than limiting it. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed in this application can still modify the technical solutions recorded in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A method for generating a face pinching result, characterized in that: The generation method comprises: Inputting a target image into a pre-trained parameter translator, extracting features of the target image through the parameter translator to obtain target image features of the target image, and outputting target face pinching parameters matching the target image features through the parameter translator; The target face pinching parameters are input into a face pinching system, and a face pinching result matching the target face pinching parameters is generated by the face pinching system as the target face pinching result corresponding to the target image.
2. The generation method according to claim 1, characterized in that: The parameter translator includes: a pre-trained image feature extraction module and a parameter generation module, wherein the target image is input into the pre-trained parameter translator, the target image is extracted by the parameter translator to obtain the target image features of the target image, and the target face pinching parameters matching the target image features are output by the parameter translator, including: Inputting the target image into the pre-trained image feature extraction module to obtain the target image features output by the image feature extraction module for the target image; The target image feature is input into the pre-trained parameter generation module to obtain the target face pinching parameter output by the parameter generation module for the target image feature.
3. The generation method according to claim 2, characterized in that: The image feature extraction module includes: an image encoder in a pre-trained encoding and decoding model.
4. The generation method according to claim 3, characterized in that: The training method of the encoding and decoding model includes: Collecting a variety of facial images of different styles as training data sets for the encoding and decoding model; For a face image in the training data set, the face image is divided into a plurality of image units, and the plurality of image units are randomly masked according to a preset mask ratio to obtain a mask image of the face image; Inputting the mask image into the image encoder of the encoding and decoding model to obtain image features output by the image encoder for the mask image; Inputting the image features into a decoder of the encoding and decoding model to obtain a reconstructed image output by the decoder for the image features; According to the image loss between the reconstructed image and the face image, the model parameters of the encoding and decoding model are adjusted to obtain the encoding and decoding model including the adjusted model parameters.
5. The generation method according to claim 3, characterized in that: The training data set of the encoding and decoding model includes face pinching images sampled from the face pinching system; wherein the face pinching images are determined according to historical face pinching results generated by the face pinching system.
6. The generation method according to claim 2, characterized in that: The training method of the parameter generation module includes: Randomly sampling from the face pinching system to obtain multiple sets of paired data of original face pinching parameters and original face pinching images; wherein the original face pinching images are determined according to the face pinching results generated by the face pinching system according to the original face pinching parameters; Inputting the original face-pinching image in the pairing data into the pre-trained image feature extraction module to obtain the face-pinching image features output by the image feature extraction module for the original face-pinching image; Inputting the face pinching image feature into the parameter generation module to obtain the predicted face pinching parameter output by the parameter generation module for the face pinching image feature; According to the error loss between the predicted face pinching parameters and the original face pinching parameters in the pairing data, the model parameters of the parameter generation module are adjusted to obtain the parameter generation module including the adjusted model parameters.
7. The generation method according to claim 6, characterized in that: The predicted face pinching parameters include: continuous face pinching parameters whose parameter values are continuous values and discrete face pinching parameters whose parameter values are discrete values.
8. The generation method according to claim 7, characterized in that: The adjusting the model parameters of the parameter generation module according to the error loss between the predicted face pinching parameters and the original face pinching parameters in the paired data includes: For the continuous face pinching parameter in the predicted face pinching parameter, determining, from the original face pinching parameters, a first original face pinching parameter of the same parameter type as that of the continuous face pinching parameter, and calculating an absolute error loss between the continuous face pinching parameter and the first original face pinching parameter; For the discrete face pinching parameters in the predicted face pinching parameters, determining, from the original face pinching parameters, second original face pinching parameters of the same parameter type as that of the discrete face pinching parameters, and calculating a classification loss between the discrete face pinching parameters and the second original face pinching parameters; According to the absolute error loss and the classification loss, the overall training loss corresponding to the parameter generation module is determined, and according to the overall training loss, the model parameters of the parameter generation module are adjusted to obtain the parameter generation module including the adjusted model parameters.
9. A device for generating face pinching results, characterized in that: The generating device comprises: a parameter generating unit, configured to input a target image into a pre-trained parameter translator, extract features of the target image through the parameter translator to obtain target image features of the target image, and output target face pinching parameters matching the target image features through the parameter translator; An image generation unit is used to input the target face pinching parameters into a face pinching system, and generate a face pinching result matching the target face pinching parameters as the target face pinching result corresponding to the target image through the face pinching system.
10. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate via the bus, and when the machine-readable instructions are executed by the processor, the steps of the method for generating a face pinching result as described in any one of claims 1 to 8 are performed.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for generating a face pinching result as described in any one of claims 1 to 8 are executed.