Face pinching data processing method and device, electronic equipment and medium
By using pre-trained literary and graphic models and parameter translators, the problems of poor results and time-consuming in existing intelligent face-pinching techniques are solved, and the fast and diverse face-pinching effects are achieved, and the rendering styles of different face-pinching systems are adapted.
Patent Information
- Application Number
- CN202510196591.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-30
AI Technical Summary
The existing text-based intelligent face-pinching technology has a general effect and is time-consuming, and it is difficult to adapt to the rendering styles of different face-pinching systems, resulting in poor style adaptability.
The pre-trained target text image model processes the text description input by the user, generates images that match the style of the face-pinching system, and uses the target parameter translator to convert the image into face-pinching parameters and sends it to the face-pinching system to generate face-pinching results that meet the description.
It realizes the rapid generation of face pinching effects that conform to text descriptions, reduces the style adaptability of face pinching, improves processing efficiency, and provides diverse face pinching results.
Smart Images

Figure CN120070671A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of game technologies, and in particular, to a method, apparatus, electronic device, and medium for processing face pinching data. Background Art
[0002] In the modern digital entertainment industry, character personalization has become one of the important factors in enhancing the user experience. With the rise of role-playing games (RPGs) and short video platforms, the user's demand for quickly generating realistic and personalized virtual avatars is increasing. Traditional face pinching systems usually rely on cumbersome manual adjustments and a large number of parameter regulations, which not only consume time but also require users to have certain technical and aesthetic abilities.
[0003] With the development of artificial intelligence technology, text-based intelligent face pinching technology has also begun to be proposed. This technology can quickly customize and create the character image desired by the user according to the text description input by the user, providing a highly personalized creation experience for the user. The existing text-based intelligent face pinching technology solutions generally have average effects and take a long time. Summary of the Invention
[0004] In view of this, the purpose of the present application is to provide a method for processing face pinching data, which can obtain a face pinching effect that matches the face pinching system through one inference based on the description text, solve the style adaptability problem during face pinching, and reduce the time consumption of face pinching.
[0005] In a first aspect, a method for processing face pinching data provided by an embodiment of the present application includes:
[0006] Obtain the target face pinching text description input by the user;
[0007] Process the target face pinching text description through a pre-trained target text-to-image model to obtain a target face pinching image that matches the target face pinching text description and the face pinching style of the target face pinching system;
[0008] Input the target face pinching image into a target parameter translator that matches the target text-to-image model, and process the target face pinching image through the target parameter translator to determine the target face pinching parameters for the target face pinching system;
[0009] Send the target face pinching parameters to the target face pinching system so that the target face pinching system generates a target face pinching result that conforms to the face pinching text description based on the target face pinching parameters.
[0010] In a second aspect, the present application also provides a device for processing face pinching data, and the device includes:
[0011] An acquisition module, configured to acquire a target face sculpting text description of a user for a target face sculpting system; the target face sculpting system corresponds to a pre-trained target text-to-image model and a target parameter translator;
[0012] A first processing module, configured to process the target face sculpting text description through the pre-trained target text-to-image model to obtain a target face sculpting image that matches the target face sculpting text description and the face sculpting style of the target face sculpting system;
[0013] A second processing module, configured to input the target face sculpting image into the target parameter translator, and process the target face sculpting image through the target parameter translator to determine target face sculpting parameters for the target face sculpting system;
[0014] A sending module, configured to send the target face sculpting parameters to the target face sculpting system, so that the target face sculpting system generates a target face sculpting result that conforms to the face sculpting text description based on the target face sculpting parameters.
[0015] In a third aspect, an electronic device is further provided. The electronic device includes: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the face sculpting data processing method are executed.
[0016] In a fourth aspect, a computer-readable storage medium stores a computer program. When the computer program is run by a processor, the steps of the face sculpting data processing method are executed.
[0017] In the embodiments of the present application, a method, apparatus, electronic device, and medium for processing face pinching data are provided. Obtain the target face pinching text description of the user for the target face pinching system; process the target face pinching text description through a pre-trained target text-to-image model to obtain a target face pinching image that matches the target face pinching text description and the face pinching style of the target face pinching system; input the target face pinching image into the target parameter translator, and process the target face pinching image through the target parameter translator to determine the target face pinching parameters for the target face pinching system; send the target face pinching parameters to the target face pinching system, so that the target face pinching system generates a target face pinching result that conforms to the face pinching text description based on the target face pinching parameters; solve the problem of style adaptability during face pinching through a text-to-image model that matches the style of the target face pinching system. The text-to-image model that matches the style ensures that images of a specific style are always generated and smoothly used as the input of the subsequent parameter translator, ensuring that the parameter translator will not introduce translation errors due to changes in the image style distribution; decompose and decouple the time-consuming iterative optimization into text-to-image inference generation and parameter translation, and adopt a feed-forward end-to-end method to directly generate face pinching parameters from the input text, and then combine image parameter translation to improve the processing efficiency of intelligent face pinching; the text-to-image model can generate multiple groups of images with consistent semantics but different details by adjusting the latent variables based on the same text input. By virtue of the characteristics of the text-to-image model itself, diverse images that conform to the text description can be generated through the given text, so that for the same text, multiple different face pinching results that conform to the description can be generated; BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 Shows the method flowchart of the face pinching data processing method described in the embodiments of the present application;
[0020] Figure 2 Shows the framework diagram of the face pinching data processing method described in the embodiments of the present application;
[0021] Figure 3 Shows the flowchart of the training method of the target text-to-image model described in the embodiments of the present application;
[0022] Figure 4 Shows the method flowchart of another face pinching data processing method described in the embodiments of the present application;
[0023] Figure 5The structural schematic diagram of the face pinching data processing device according to the embodiments of the present application is shown;
[0024] Figure 6 The structural schematic diagram of the electronic device according to the embodiments of the present application is shown. Detailed implementation manners
[0025] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. It should be understood that the accompanying drawings in the present application are only for the purposes of illustration and description, and are not used to limit the protection scope of the present application. In addition, it should be understood that the schematic drawings are not drawn according to the actual scale. The flowcharts used in the present application show the operations implemented according to some embodiments of the present application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and the steps without logical context relationships may be reversed or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of the present application.
[0026] In addition, the described embodiments are only some embodiments of the present application, rather than all embodiments. The components of the embodiments of the present application usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application to be protected, but only represents the selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.
[0027] It should be noted that the term "including" will be used in the embodiments of the present application to indicate the existence of the features stated hereinafter, but does not exclude adding other features.
[0028] In the modern digital entertainment industry, character personalization has become one of the important factors to enhance the user experience. With the rise of role-playing games (RPGs) and short video platforms, users' demand for quickly generating realistic and personalized virtual avatars is increasing. Traditional face pinching systems usually rely on cumbersome manual adjustments and a large number of parameter settings, which not only consume time but also require users to have certain technical and aesthetic abilities.
[0029] With the development of artificial intelligence technology, text-based intelligent face pinching technology has also begun to be proposed. This technology can quickly create customized character images according to the text descriptions input by users, providing users with a highly personalized creation experience. The existing text-based intelligent face pinching technology solutions generally have average effects and take a long time.
[0030] The defects of some existing intelligent face sculpting technology solutions based on text are as follows. First, since the rendering styles of different face sculpting systems are diverse, such as realistic, cartoonish, etc. For the human faces rendered by systems with different styles, if the CLIP pre-training stage has not seen enough samples of similar styles, the CLIP image encoder may not be able to extract accurate semantic information. In this case, the face sculpting obtained through this solution is often not accurate enough. Second, through the iterative optimization method, it often requires many iterations to find a set of relatively close results, and this process usually takes a relatively long time. Third, since CLIP itself is a relatively weak supervision signal, using only CLIP often results in deformed and abstract results. To make the results stable, existing solutions often need to impose some regularization constraints (such as using PCA). However, although this ensures that normal face shapes can be sculpted, these face shapes usually tend to a fixed pattern, greatly reducing the diversity of face sculpting and affecting the user experience.
[0031] Based on this, in the embodiments of the present application, a face sculpting data processing method, device, electronic device, and medium are provided. Obtain the target face sculpting text description of the user for the target face sculpting system; process the target face sculpting text description through the pre-trained target text-to-image model to obtain a target face sculpting image that matches the target face sculpting text description and the face sculpting style of the target face sculpting system; input the target face sculpting image into the target parameter translator matched with the target text-to-image model, and process the target face sculpting image through the target parameter translator to determine the target face sculpting parameters for the target face sculpting system; send the target face sculpting parameters to the target face sculpting system so that the target face sculpting system generates a target face sculpting result that conforms to the face sculpting text description based on the target face sculpting parameters; solve the style adaptability problem during face sculpting through a text-to-image model that matches the style of the target face sculpting system. The text-to-image model that matches the style ensures that images of a specific style are always generated and smoothly used as the input of the subsequent parameter translator, ensuring that the parameter translator will not introduce translation errors due to changes in the image style distribution; decouple the time-consuming iterative optimization into text-to-image inference generation and parameter translation, and adopt a feed-forward end-to-end method to directly generate face sculpting parameters from the input text, and then combine image parameter translation to improve the processing efficiency of intelligent face sculpting; the text-to-image model can generate multiple groups of images with consistent semantics but different details by adjusting the latent variables based on the same text input. By virtue of the characteristics of the text-to-image model itself, diverse images that conform to the text description can be generated through the given text, so that for the same text, multiple different face sculpting results that conform to the description can be generated.
[0032] The face pinching data processing method in the embodiment of the present application can be run on a terminal device or a server. Among them, the terminal device can be a local terminal device (such as a local touch terminal). When the face pinching data processing method is run on a server, it can be a cloud game.
[0033] In an optional implementation, cloud gaming refers to a gaming method based on cloud computing. In the operation mode of cloud gaming, the operating entity of the game program and the entity presenting the game screen are separated. The storage and operation of the deployment method of the game unit are completed on the cloud gaming server. The cloud gaming client is used for receiving and sending data and presenting the game screen. For example, the cloud gaming client can be a display device with data transmission function close to the user side, such as a mobile terminal, a television, a computer, a handheld computer, etc.; but the terminal device for deploying the game unit is a cloud gaming server in the cloud. When playing the game, the user operates the cloud gaming client to send an operation instruction to the cloud gaming server. The cloud gaming server runs the game according to the operation instruction, encodes and compresses the game screen and other data, and returns it to the cloud gaming client through the network. Finally, the cloud gaming client decodes and outputs the game screen.
[0034] In another optional embodiment, the terminal device may be a local terminal device. The local terminal device stores a game program and is used to present a game screen. The local terminal device is used to interact with the user through a graphical user interface, that is, the game program is downloaded and installed and run by an electronic device in a conventional manner. The local terminal device may provide the graphical user interface to the user in a variety of ways, for example, it may be rendered and displayed on a display screen of the local terminal device, or provided to the user through a holographic projection. For example, the local terminal device may include a display screen and a processor, the display screen is used to present a graphical user interface, the graphical user interface includes a game screen, and the processor is used to run the game, generate a graphical user interface, and control the display of the graphical user interface on the display screen.
[0035] Please refer to Figure 1 , Figure 1 A method flow chart of the face pinching data processing method according to an embodiment of the present application is shown; Figure 1 As shown, the method includes the following steps S101-S104:
[0036] S101, obtaining a text description of a target face pinching input by a user;
[0037] S102, processing the target face pinching text description through a pre-trained target text-based graph model to obtain a target face pinching image that matches the target face pinching text description and matches the face pinching style of the target face pinching system;
[0038] S103. Input the target face - pinching image into the target parameter translator that matches the target text - to - image model, and process the target face - pinching image through the target parameter translator to determine the target face - pinching parameters for the target face - pinching system;
[0039] S104. Send the target face - pinching parameters to the target face - pinching system, so that the target face - pinching system generates a target face - pinching result that conforms to the face - pinching text description based on the target face - pinching parameters.
[0040] Please refer to Figure 2 , Figure 2 which shows the framework diagram of the face - pinching data processing method described in the embodiments of the present application; as Figure 2 shown, first, in the way of text - to - image, use text to generate a 2D photo of the face - pinching character image, and then convert the 2D photo into face - pinching parameters; based on this, the target face - pinching system includes a text - to - image module that can generate a character photo in the rendering style of the target face - pinching system according to the input text and a parameter translator module that can convert the system character photo into face - pinching parameters; by connecting these two modules in series, according to the input text description, directly generate face - pinching parameters that conform to the description end - to - end.
[0041] In the step S101, obtain the target face - pinching text description input by the user.
[0042] The target face - pinching system corresponds to a pre - trained target text - to - image model and a pre - trained target parameter translator.
[0043] The target face - pinching system is specifically an interaction module in the game that allows players to customize the facial features of the character. The target face - pinching system generates a personalized virtual image by adjusting parameters (such as eye size, nose bridge height).
[0044] Specifically, the target face - pinching system internally includes multiple parameter controllers, and the parameter controllers are used to adjust parameters.
[0045] The text - to - image model, that is, a text - to - image generation model (such as StableDiffusion) fine - tuned based on the game art style, inputs a text description (such as "big eyes, long eyelashes, black hair") and outputs a character concept map that conforms to the game painting style.
[0046] The target parameter translator is used to map the image to the parameters of the target face - pinching system. Its specific structure is a neural network. The input is the image generated by the text - to - image model, and the output is the parameter value that can be parsed by the game engine; the target parameter translator converts the target face - pinching image that matches the face - pinching style of the target face - pinching system into the face - pinching parameters required by the target face - pinching system.
[0047] In the embodiments of the present application, the target text-to-image model corresponding to the target face sculpting system, that is, the style of the image generated by the target text-to-image model matches the style of the virtual character of the target face sculpting system.
[0048] In some embodiments, the obtaining of the target face sculpting text description for the target face sculpting system includes:
[0049] Obtaining the target face sculpting text description of the text type directly input based on the terminal device; or,
[0050] Obtaining the target face sculpting text description obtained by converting other types of face sculpting descriptions.
[0051] In an alternative embodiment, the other types of face sculpting descriptions include at least one of the following:
[0052] Face sculpting descriptions of the voice type, face sculpting descriptions of the picture type, face sculpting descriptions of the cross-language type.
[0053] In practical applications, a graphical user interface is displayed on the terminal device running the game. The graphical user interface includes a first interaction control. After triggering the first interaction control through the terminal device, a face sculpting description of the game character for the target face sculpting system can be input.
[0054] Here, the face sculpting description of the game character for the target face sculpting system input by the user through the terminal device can be multi-modal and cross-language. For example, directly inputting a text description, inputting a voice description, inputting an image including a text description, inputting a text description in another language, and so on.
[0055] When the user inputs the other types of face sculpting descriptions, they are converted into text descriptions by the conversion module, so that the target text-to-image model generates target face sculpting images based on the text descriptions.
[0056] For the directly input text description, specifically, directly input the target face sculpting text description in the text box displayed in the graphical user interface.
[0057] For the input voice description, specifically, the player holds down the voice button and speaks out the requirements, and the conversion module converts them into text descriptions in real time.
[0058] For the input picture type face sculpting description, exemplarily, the player uploads a picture including a text description, and the conversion module extracts the text description in real time.
[0059] For the cross-language type face sculpting description, exemplarily, the player inputs an English description, and the conversion module converts it into a Chinese description in real time.
[0060] In some embodiments, before the user inputs a face pinching description for a game character of a target face pinching system through a terminal device, an input mode is determined in response to an input module selection operation.
[0061] Specifically, a second interactive control is displayed in the graphical user interface of the terminal device, and the second interactive control includes a text input box (placeholder prompt: "Please enter a description such as 'sweet girl'"); a voice button (microphone icon, long press to record); a picture upload button (camera / album icon), etc. The input mode is determined based on the operation of the text input box, language button or picture upload button.
[0062] After the user inputs the selected type of face pinching description through the terminal device, the face pinching text description is uploaded to the server in response to the upload operation, so that the server obtains the user's target face pinching text description for the target face pinching system.
[0063] In some embodiments, the target face-pinching text description represents the user's requirements for the appearance of the virtual character.
[0064] In some embodiments, the target face pinching text description includes descriptions in multiple dimensions.
[0065] The descriptions of the multiple different dimensions include: face shape, facial features, makeup, role, type, etc.
[0066] The type refers to the type of character effect, such as cute, mature, strong, etc.
[0067] The face shapes include oval face, round face, square face and the like.
[0068] The following is a text description of the target face pinching system: "Facial features": "oval face, a pair of long and slender eyes, a high nose bridge, a delicate nose, round ears, and a plump cherry mouth"; "Makeup": "The grayish-white skin is smooth and delicate, the brown-black eyebrows are particularly slender and beautiful, the thick double eyelids, the slightly curled black eyelashes are long and dense, the blue pupils seem to contain two crystal clear snowflakes, and the brown-red lips are sexy and charming", "Type": "Beautiful"; "Role": "Waiter"; "Brief description": "Black hair, grayish-white skin, blue and transparent eyes, a beautiful waitress".
[0069] In some embodiments, when the target face-pinching system has multiple styles, such as cartoon style and realistic style, each style corresponds to a set of target text model and target parameter translator.
[0070] Based on this, before or after obtaining the target face pinching text description of the user for the target face pinching system, the method further includes:
[0071] In response to a selection instruction for multiple face - shaping styles of a target face - shaping system, determine the target face - shaping style of the target face - shaping system; different face - shaping styles of the target face - shaping system correspond to different text - to - image models and parameter translators;
[0072] Determine the target text - to - image model and the target parameter translator corresponding to the target face - shaping style of the target face - shaping system.
[0073] In some embodiments, one or more options of face - shaping styles are displayed in the graphical user interface of the terminal device to the user; the options can be a list, an icon, or a preview image;
[0074] In response to the user's selection operation on the target option, determine the selection instruction for multiple face - shaping styles of the target face - shaping system.
[0075] In step S102, process the target face - shaping text description through the pre - trained target text - to - image model to obtain a target face - shaping image that matches the target face - shaping text description and the face - shaping style of the target face - shaping system.
[0076] Please refer to Figure 3 , Figure 3 which shows the flowchart of the training method of the target text - to - image model described in the embodiments of the present application; as Figure 2 shown, the target text - to - image model is trained based on the following steps S201 - S202:
[0077] S301. Obtain a target training set; the target training set includes first - sample face - shaping images that conform to the face - shaping style of the target face - shaping system, and sample text descriptions corresponding to the first - sample face - shaping images;
[0078] S302. Train an initial text - to - image model based on the target training set to obtain a target text - to - image model corresponding to the target face - shaping system.
[0079] That is to say, the target text - to - image model is obtained by fine - tuning based on the initial text - to - image model and in combination with a training set that conforms to the face - shaping style of the target face - shaping system. Thus, the text - to - image model can be conveniently and quickly extended to various face - shaping systems with different styles, and the fine - tuned target text - to - image model can be optimized for a specific face - shaping style, so as to generate images highly consistent with the style of the system.
[0080] The first - sample face - shaping images that conform to the face - shaping style of the target face - shaping system can be face - shaping images created by professional designers that conform to the system style, images of virtual characters already existing in the database of the target face - shaping system, face - shaping images shared by players in an interactive community, and face - shaping images obtained after being enhanced using AI technology.
[0081] The face - shaping image enhanced by using AI technology, specifically, the face - shaping image obtained by transforming the existing face - shaping image based on AI technology.
[0082] Similarly, the sample text description corresponding to the first sample face - shaping image includes descriptions in multiple different dimensions.
[0083] The multiple different - dimension descriptions include: face shape, facial features, makeup, character, type, and so on.
[0084] By including descriptions in multiple dimensions, such as facial features, makeup, and character of a human face, the text - to - image model can provide a richer and more detailed character image.
[0085] It should be noted that in the current gameplay, makeup is an important part of the facial features of a character, but the existing face - shaping system has insufficient perception of makeup.
[0086] Based on this, in the embodiments of the present application, at least some of the human faces in the constructed target training set among the first sample face - shaping images include makeup, and the sample text description corresponding to the first sample face - shaping image includes a description corresponding to the makeup, so as to increase the perception of the target text - to - image model obtained by fine - tuning for makeup.
[0087] In some embodiments, the initial text - to - image model can be the StableDiffusion model.
[0088] The StableDiffusion model is simply referred to as the SD model.
[0089] Based on the training set of the target face - shaping system, an SD model that generates images in the style of the target face - shaping system can be obtained. This SD model can utilize the knowledge of the original SD itself and can also ensure that it always generates images in a specific style, smoothly serving as the input of the subsequent parameter translator, ensuring that the parameter translator will not introduce translation errors due to changes in the image style distribution, thereby ensuring the face - shaping effect.
[0090] In the step S201, obtain a target training set; the target training set includes first sample face - shaping images that conform to the face - shaping style of the target face - shaping system and the sample text descriptions corresponding to the first sample face - shaping images.
[0091] Before obtaining the target training set, construct a small training set (about 6000 samples) with precise annotations for the target face - shaping system. This training set includes rendered face - shaping images and text descriptions of facial features, makeup, style, etc. of the human faces in the images.
[0092] In some embodiments, the obtaining of the target training set includes:
[0093] Obtain the first sample face - sculpting image rendered by the target face - sculpting system;
[0094] Determine the sample text description corresponding to each first sample face - sculpting image;
[0095] Construct the target training set based on the first sample face - sculpting image and the sample text description.
[0096] In some embodiments, the determining the sample text description corresponding to each first sample face - sculpting image includes:
[0097] Input the first sample face - sculpting image into the trained image - to - text model for processing to obtain the content text description of the first sample face - sculpting image;
[0098] Adjust the content text description based on the adjustment instruction for the content text description of the first sample face - sculpting image to obtain the sample text description of the first sample face - sculpting image.
[0099] That is to say, the annotation corresponding to the first sample face - sculpting image (the annotation is the sample text description) is first automatically annotated by the trained image - to - text model and then manually verified.
[0100] Exemplarily, the image - to - text model can be the gpt4o model.
[0101] S202. Train the initial text - to - image model based on the target training set to obtain the target text - to - image model corresponding to the target face - sculpting system.
[0102] Fine - tune the StableDiffusion model on the constructed dataset to train an SD that can stably generate images in the rendering style of the target face - sculpting system through text description; thus enabling the obtained SD to always generate face - sculpting images in the style of the target face - sculpting system, and being able to utilize the rich semantic information of the original SD to achieve a good generalization effect; this training process also improves the SD model's perception of makeup details, and for text descriptions involving makeup details, it can generate more accurate face - sculpting images.
[0103] In some embodiments, when fine - tuning the text - to - image model based on the constructed dataset, specifically, the training the initial text - to - image model based on the target training set to obtain the target text - to - image model corresponding to the target face - sculpting system includes:
[0104] Process the sample text description corresponding to the first sample face - sculpting image in the target training set to obtain multiple sample enhanced text descriptions corresponding to this first sample face - sculpting image; wherein, each first sample face - sculpting image corresponds to one sample text description and multiple sample enhanced text descriptions;
[0105] Train an initial text-to-image model based on the first sample face sculpting image, the sample text description corresponding to the first sample face sculpting image, and multiple sample enhanced text descriptions to obtain the target text-to-image model corresponding to the target face sculpting system.
[0106] First, when training the SD model, the model needs to understand the same image under different expressions, so as to improve the diversity and accuracy of the generated images, increase the generalization ability of the text-to-image model, and ensure that the trained model will not perform poorly due to slight differences in text descriptions. Second, when the user inputs a description of the character image, they will not input it sequentially according to the preset dimensions, nor will they input all the dimensions. In other words, the user's description of the character image will not be as accurate and comprehensive as the sample text description. Therefore, enhance the sample text description corresponding to the first sample face sculpting image.
[0107] The same first sample face sculpting image corresponds to multiple sample enhanced text descriptions, where the multiple sample enhanced text descriptions include sample enhanced text descriptions after deleting some descriptions and single-dimension text descriptions.
[0108] The sample enhanced text description after deleting some descriptions is the sample enhanced text description obtained by randomly extracting and combining the sample text description corresponding to the first sample face sculpting image.
[0109] The single-dimension text description is the text description of a single dimension. For example, a text description only for makeup, a text description only for face shape, or a text description only for facial features.
[0110] In some embodiments, in the face sculpting data processing method, processing the sample text description corresponding to the first sample face sculpting image in the target training set to obtain multiple sample enhanced text descriptions corresponding to the first sample face sculpting image includes:
[0111] Randomly extract and combine the sample text description corresponding to the first sample face sculpting image to obtain multiple sample enhanced text descriptions corresponding to the first sample face sculpting image;
[0112] and / or
[0113] Extract the single-dimension text description corresponding to each dimension in the sample text description corresponding to the first sample face sculpting image to obtain multiple sample enhanced text descriptions corresponding to the first sample face sculpting image.
[0114] Randomly extract the sample text descriptions corresponding to the first sample face pinching images and combine them. In this way, the annotation corresponding to the first sample face pinching image has stronger randomness. Among the combined sample enhanced text descriptions, there are various descriptions such as incomplete face pinching descriptions, sample enhanced text descriptions with different orders of face pinching descriptions in different dimensions, and sample enhanced text descriptions lacking some dimensions, thereby enhancing the generalization ability of the adjusted text-to-image model.
[0115] Exemplarily, for the target face pinching text description of the target face pinching system: "Facial features": "oval face, a pair of long and slender big eyes, a high nose bridge, a delicate and beautiful nose, round small ears, a plump cherry-like small mouth"; "Makeup": "The grayish-white skin is smooth and delicate, the brownish-black single-line eyebrows are extremely long and beautiful, thick double eyelids, slightly upturned black eyelashes that are long and dense, and the blue pupils seem to contain two crystal-clear snowflakes, and the brownish-red lip smear is sexy and charming", "Type": "beautiful and charming"; "Role": "waitress"; "Brief description": "black hair, grayish-white skin, blue and translucent eyes, a beautiful and charming waitress", after random extraction and combination, it becomes the following description.
[0116] Specifically, the processed description one: Facial shape: oval face; Facial features: round small ears, a plump cherry-like small mouth; Makeup: grayish-white skin, blue pupils, brownish-red lip smear; Type: beautiful and charming; Role: waitress; Brief description: black hair, grayish-white skin, blue and translucent eyes, a beautiful and charming waitress.
[0117] Some descriptions (such as thick double eyelids) and adjectives (such as smooth and delicate, etc.) are deleted in description one.
[0118] The processed description two: Makeup: The brownish-black single-line eyebrows are extremely long and beautiful; Type: beautiful and charming; Role: waitress; Brief description: grayish-white skin, blue and translucent eyes, a beautiful and charming waitress.
[0119] The descriptions of the two dimensions of facial shape and facial features are missing in the processed description two.
[0120] Single-dimension description one: "Facial features": "oval face, a pair of long and slender big eyes, a high nose bridge, a delicate and beautiful nose, round small ears, a plump cherry-like small mouth".
[0121] Single-dimension description two: "Makeup": "The grayish-white skin is smooth and delicate, the brownish-black single-line eyebrows are extremely long and beautiful, thick double eyelids, slightly upturned black eyelashes that are long and dense, and the blue pupils seem to contain two crystal-clear snowflakes, and the brownish-red lip smear is sexy and charming".
[0122] In some embodiments, training an initial text-to-image model based on a first sample face sculpting image, a sample text description corresponding to the first sample face sculpting image, and multiple sample enhanced text descriptions to obtain a target text-to-image model corresponding to the target face sculpting system, includes:
[0123] Obtain a first sample face sculpting image in the target training set and a single-dimensional text description corresponding to the first sample face sculpting image;
[0124] In some training rounds of training the initial text-to-image model, train based on the first sample face sculpting image in the target training set and the single-dimensional text corresponding to the first sample face sculpting image.
[0125] That is to say, in the fine-tuning process of the text-to-image model, in some training rounds, train using the first sample face sculpting image and the single-dimensional text description corresponding to the first sample face sculpting image for single-dimensional perception training; thereby increasing the text-to-image model's perception ability of single-dimensional features; for example, use the first sample face sculpting image and the makeup text description corresponding to the first sample face sculpting image for model training, thereby increasing the perception ability of makeup. Single-dimensional training enables the model to generate more refined and realistic details on specific features while ensuring that the overall style of the image matches the style of the target face sculpting system. For example, after training, the model can better restore the length, color, and density of eyelashes, or the shape and gloss of lips, or the color, gloss, and area of blusher, thereby improving the overall realism and detail expressiveness of the image.
[0126] In step S103, input the target face sculpting image into a target parameter translator matched with the target text-to-image model, and process the target face sculpting image through the target parameter translator to determine target face sculpting parameters for the target face sculpting system.
[0127] The target face sculpting parameters are the parameters required for the target face sculpting system to generate a game character's face image and are also parameter values that the target face sculpting system can parse.
[0128] The target parameter translator is trained based on paired face sculpting parameters and second sample face sculpting images of the target face sculpting system.
[0129] The second sample face sculpting image is an image rendered by the target face sculpting system.
[0130] Or rather, the target parameter translator is trained based on a paired dataset of face sculpting parameters and corresponding rendered images within the target face sculpting system.
[0131] Through the face sculpting system, a large number of random samplings are carried out on the face sculpting parameters, and paired data of the face sculpting parameters and the rendered images of the face sculpting system can be obtained. Using these data, a parameter translator is trained, which can input the rendered photos of the face sculpting system and output the corresponding face sculpting parameters.
[0132] Exemplarily, the parameter translator uses ResNet-50 as the backbone network. After ResNet, a fully connected layer is used to map the feature output by ResNet into face sculpting parameters.
[0133] In this way, only by using the face sculpting system, without other additional annotations, a parameter translator that can accurately convert an image of a specific face sculpting style into face sculpting parameters can be obtained, realizing the accurate conversion from image to parameters.
[0134] In the step S104, the target face sculpting parameters are sent to the target face sculpting system, so that the target face sculpting system generates a target face sculpting result that conforms to the face sculpting text description based on the target face sculpting parameters.
[0135] The target face sculpting result is a rendered image generated by the target face sculpting system based on the target face sculpting parameters.
[0136] In some embodiments, the generated rendered image may not meet the user's requirements. For example, the skin is not white enough, the face shape is not round enough, etc. Based on this, please refer to Figure 4 After sending the target face sculpting parameters to the target face sculpting system, so that the target face sculpting system generates a target face sculpting result that conforms to the face sculpting text description based on the target face sculpting parameters, the method further includes the following steps S401-S402:
[0137] S401. Receive a parameter adjustment instruction for the parameter controller of the target face sculpting system;
[0138] S402. Send the face adjustment instruction to the target face sculpting system, so that the parameter controller of the target face sculpting system adjusts the generated target face sculpting result based on the parameter adjustment instruction to obtain an adjusted target face sculpting result.
[0139] That is to say, when the virtual character effect generated by the user based on the text description does not meet their requirements, there is no need to perform feedback adjustment through the text-to-image model, which has a low adjustment efficiency; moreover, the model cannot accurately perceive the aesthetics of specific users and is difficult to fine-tune. The user can manually fine-tune each parameter controller based on the face sculpting result for modification, so as to combine the two means of intelligent generation and manual adjustment to make the adjusted target face sculpting result meet the user's aesthetics.
[0140] The technical solution for intelligently generating a face - pinching result based on text description in the embodiments of the present application is a feed - forward end - to - end technical solution. The feed - forward solution only requires one inference during application. Compared with the iterative optimization solution that needs to execute dozens of iterations, it can greatly improve efficiency. Moreover, by utilizing the diversity and accuracy generated by the existing text - to - image model itself, for the same text, the method can generate completely different face - pinching effects that conform to the same text description, greatly enriching the diversity of face - pinching.
[0141] Based on the same inventive concept, in the embodiments of the present application, there is also provided a face - pinching data processing device corresponding to the face - pinching data processing method. Since the principle of problem - solving of the device in the embodiments of the present application is similar to that of the above - mentioned face - pinching data processing method in the embodiments of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be elaborated.
[0142] Please refer to Figure 5 , Figure 5 FIG. shows a schematic structural diagram of a face - pinching data processing device described in the embodiments of the present application; the face - pinching data processing device includes:
[0143] An acquisition module 501, configured to acquire a target face - pinching text description of a user for a target face - pinching system; the target face - pinching system corresponds to a pre - trained target text - to - image model and a target parameter translator;
[0144] A first processing module 502, configured to process the target face - pinching text description through the pre - trained target text - to - image model to obtain a target face - pinching image that matches the target face - pinching text description and the face - pinching style of the target face - pinching system;
[0145] A second processing module 503, configured to input the target face - pinching image into the target parameter translator matched with the target text - to - image model, and process the target face - pinching image through the target parameter translator to determine target face - pinching parameters for the target face - pinching system;
[0146] A sending module 504, configured to send the target face - pinching parameters to the target face - pinching system, so that the target face - pinching system generates a target face - pinching result that conforms to the face - pinching text description based on the target face - pinching parameters.
[0147] In some embodiments, the face - pinching data processing device further includes a training module, and the training module is configured to train the target text - to - image model based on the following steps:
[0148] Acquire a target training set; the target training set includes first sample face - pinching images that conform to the face - pinching style of the target face - pinching system, and sample text descriptions corresponding to the first sample face - pinching images;
[0149] Training the initial text-to-image model based on the target training set to obtain the target text-to-image model corresponding to the target face morphing system.
[0150] In some embodiments, in the face morphing data processing device, the sample text description corresponding to the first sample face morphing image includes descriptions in multiple different dimensions.
[0151] In some embodiments, in the face morphing data processing device, the multiple different dimensions of descriptions include: face shape, facial features, makeup, character, type.
[0152] In some embodiments, in the face morphing data processing device, when the training module trains the initial text-to-image model based on the target training set to obtain the target text-to-image model corresponding to the target face morphing system, it is specifically configured to:
[0153] Process the sample text description corresponding to the first sample face morphing image in the target training set to obtain multiple sample enhanced text descriptions corresponding to the first sample face morphing image; wherein, each first sample face morphing image corresponds to a sample text description and multiple sample enhanced text descriptions;
[0154] Train the initial text-to-image model based on the first sample face morphing image, the sample text description corresponding to the first sample face morphing image, and the multiple sample enhanced text descriptions to obtain the target text-to-image model corresponding to the target face morphing system.
[0155] In some embodiments, in the face morphing data processing device, when the training module processes the sample text description corresponding to the first sample face morphing image in the target training set to obtain multiple sample enhanced text descriptions corresponding to the first sample face morphing image, it is specifically configured to:
[0156] Randomly extract and combine the sample text description corresponding to the first sample face morphing image to obtain multiple sample enhanced text descriptions corresponding to the first sample face morphing image;
[0157] And / or,
[0158] Extract the single-dimension text descriptions corresponding to each dimension in the sample text description corresponding to the first sample face morphing image to obtain multiple sample enhanced text descriptions corresponding to the first sample face morphing image.
[0159] In some embodiments, in the face morphing data processing device, when the training module trains the initial text-to-image model based on the first sample face morphing image, the sample text description corresponding to the first sample face morphing image, and the multiple sample enhanced text descriptions to obtain the target text-to-image model corresponding to the target face morphing system, it is specifically configured to:
[0160] Obtain the first sample face pinching image in the target training set and the single-dimensional text description corresponding to the first sample face pinching image;
[0161] In some training rounds of training the initial text-to-image model, train based on the first sample face pinching image in the target training set and the single dimension text corresponding to the first sample face pinching image.
[0162] In some embodiments, in the face pinching data processing device, in the target training set, at least some of the faces in the first sample face pinching images include makeup, and the corresponding sample text description of the first sample face pinching image includes a description corresponding to the makeup.
[0163] In some embodiments, the training module in the face pinching data processing device, when obtaining the target training set, is specifically configured to:
[0164] Obtain the first sample face pinching image rendered by the target face pinching system;
[0165] Determine the sample text description corresponding to each first sample face pinching image;
[0166] Based on the first sample face pinching image and the sample text description, construct the target training set.
[0167] In some embodiments, the training module in the face pinching data processing device, when determining the sample text description corresponding to each first sample face pinching image, is specifically configured to:
[0168] Input the first sample face pinching image into the trained image-to-text model for processing to obtain the content text description of the first sample face pinching image;
[0169] Based on the adjustment instruction for the content text description of the first sample face pinching image, adjust the content text description to obtain the sample text description of the first sample face pinching image.
[0170] In some embodiments, in the face pinching data processing device, the target parameter translator is trained based on the paired face pinching parameters and the second sample face pinching image of the target face pinching system.
[0171] In some embodiments, the obtaining module in the face pinching data processing device, when obtaining the target face pinching text description of the user for the target face pinching system, is specifically configured to:
[0172] Obtain the target face pinching text description of the text type directly input based on the terminal device; or,
[0173] Obtain the target face pinching text description converted from other types of face pinching descriptions.
[0174] In some embodiments, in the described face pinching data processing device, the other types of face pinching descriptions include at least one of the following:
[0175] Face pinching descriptions of the voice type, face pinching descriptions of the picture type, and face pinching descriptions of the cross - language type.
[0176] In some embodiments, the described face pinching data processing device further includes:
[0177] An adjustment module, configured to, after sending the target face pinching parameters to the target face pinching system so that the target face pinching system generates a target face pinching result that conforms to the face pinching text description based on the target face pinching parameters, receive a parameter adjustment instruction for the parameter controller of the target face pinching system; and send the face adjustment instruction to the target face pinching system so that the parameter controller of the target face pinching system adjusts the generated target face pinching result based on the parameter adjustment instruction to obtain an adjusted target face pinching result.
[0178] In some embodiments, the described face pinching data processing device further includes:
[0179] A determination module, configured to determine the target face pinching style of the target face pinching system in response to a selection instruction for multiple face pinching styles of the target face pinching system; different face pinching styles of the target face pinching system correspond to different text - to - image generation models and parameter translators;
[0180] Determine the target text - to - image generation model and the target parameter translator corresponding to the target face pinching style of the target face pinching system.
[0181] Based on the same inventive concept, embodiments of the present application also provide an electronic device corresponding to the face pinching data processing method. Since the principle of solving problems by the electronic device in the embodiments of the present application is similar to the above - mentioned face pinching data processing method in the embodiments of the present application, the implementation of the electronic device can refer to the implementation of the method, and repeated parts will not be described again.
[0182] Please refer to Figure 6 , in some embodiments, an electronic device is further provided. The electronic device 600 includes: a processor 602, a memory 601, and a bus. The memory 601 stores machine - readable instructions executable by the processor 602. When the electronic device 600 runs, the processor 602 communicates with the memory 601 through the bus. When the machine - readable instructions are executed by the processor 602, the steps of the face pinching data processing method are executed, specifically as follows:
[0183] Obtain the target face pinching text description of the user for the target face pinching system; the target face pinching system corresponds to a pre - trained target text - to - image generation model and a target parameter translator;
[0184] Process the target face - shaping text description through a pre - trained text - to - image model for the target, to obtain a target face - shaping image that matches the target face - shaping text description and the face - shaping style of the target face - shaping system;
[0185] Input the target face - shaping image into a target parameter translator matched with the target text - to - image model, and process the target face - shaping image through the target parameter translator to determine target face - shaping parameters for the target face - shaping system;
[0186] Send the target face - shaping parameters to the target face - shaping system, so that the target face - shaping system generates a target face - shaping result that conforms to the face - shaping text description based on the target face - shaping parameters.
[0187] In some embodiments, the processor further performs the following steps to train the target text - to - image model:
[0188] Obtain a target training set; the target training set includes first - sample face - shaping images that conform to the face - shaping style of the target face - shaping system, and sample text descriptions corresponding to the first - sample face - shaping images;
[0189] Train an initial text - to - image model based on the target training set to obtain a target text - to - image model corresponding to the target face - shaping system.
[0190] In some embodiments, the sample text description corresponding to the first - sample face - shaping image includes descriptions in multiple different dimensions.
[0191] In some embodiments, the multiple different - dimension descriptions include: face shape, facial features, makeup, character, type.
[0192] In some embodiments, when the processor performs the step of training the initial text - to - image model based on the target training set to obtain a target text - to - image model corresponding to the target face - shaping system, it specifically performs the following steps:
[0193] Process the sample text description corresponding to the first - sample face - shaping image in the target training set to obtain multiple sample enhanced text descriptions corresponding to the first - sample face - shaping image; wherein, each first - sample face - shaping image corresponds to a sample text description and multiple sample enhanced text descriptions;
[0194] Train the initial text - to - image model based on the first - sample face - shaping image, the sample text description corresponding to the first - sample face - shaping image, and the multiple sample enhanced text descriptions to obtain a target text - to - image model corresponding to the target face - shaping system.
[0195] In some embodiments, in the face pinching data processing device, when the training module processes the sample text description corresponding to the first sample face pinching image in the target training set to obtain multiple sample enhanced text descriptions corresponding to the first sample face pinching image, it specifically is used for:
[0196] Randomly extract and combine the sample text descriptions corresponding to the first sample face pinching image to obtain multiple sample enhanced text descriptions corresponding to the first sample face pinching image;
[0197] And / or,
[0198] Extract the single-dimensional text descriptions corresponding to each dimension in the sample text description corresponding to the first sample face pinching image to obtain multiple sample enhanced text descriptions corresponding to the first sample face pinching image.
[0199] In some embodiments, when the processor executes the step of training the initial text-to-image model based on the first sample face pinching image, the sample text description corresponding to the first sample face pinching image, and multiple sample enhanced text descriptions to obtain the target text-to-image model corresponding to the target face pinching system, it specifically executes the following steps:
[0200] Obtain the first sample face pinching image in the target training set and the single-dimensional text description corresponding to the first sample face pinching image;
[0201] In some partial training rounds of training the initial text-to-image model, train based on the first sample face pinching image in the target training set and the single-dimensional text corresponding to the first sample face pinching image.
[0202] In some embodiments, the human faces in at least some of the first sample face pinching images in the target training set include makeup, and the sample text description corresponding to the first sample face pinching image includes a description corresponding to the makeup.
[0203] In some embodiments, when the processor executes the step of obtaining the target training set, it specifically executes the following steps:
[0204] Obtain the first sample face pinching image rendered by the target face pinching system;
[0205] Determine the sample text description corresponding to each first sample face pinching image;
[0206] Based on the first sample face pinching image and the sample text description, construct the target training set.
[0207] In some embodiments, when the processor executes the step of determining the sample text description corresponding to each first sample face pinching image, it specifically executes the following steps:
[0208] Input the first sample face morphing image into the trained image-to-text model for processing to obtain the content text description of the first sample face morphing image;
[0209] Based on the adjustment instruction for the content text description of the first sample face morphing image, adjust the content text description to obtain the sample text description of the first sample face morphing image.
[0210] In some embodiments, the target parameter translator is trained based on the paired face morphing parameters of the target face morphing system and the second sample face morphing image.
[0211] In some embodiments, when the processor executes the step of obtaining the target face morphing text description of the user for the target face morphing system, it specifically executes the following steps:
[0212] Obtain the target face morphing text description of the text type directly input based on the terminal device; or,
[0213] Obtain the target face morphing text description converted from other types of face morphing descriptions.
[0214] In some embodiments, the other types of face morphing descriptions include at least one of the following:
[0215] Face morphing descriptions of the voice type, face morphing descriptions of the picture type, face morphing descriptions of the cross-language type.
[0216] In some embodiments, after the processor executes the step of sending the target face morphing parameters to the target face morphing system so that the target face morphing system generates a target face morphing result that conforms to the face morphing text description based on the target face morphing parameters, it further executes the following steps:
[0217] Receive a parameter adjustment instruction for the parameter controller of the target face morphing system;
[0218] Send the face adjustment instruction to the target face morphing system so that the parameter controller of the target face morphing system adjusts the generated target face morphing result based on the parameter adjustment instruction to obtain an adjusted target face morphing result.
[0219] In some embodiments, the processor further executes the following steps:
[0220] In response to a selection instruction for multiple face morphing styles of the target face morphing system, determine the target face morphing style of the target face morphing system; different face morphing styles of the target face morphing system correspond to different text-to-image models and parameter translators;
[0221] Determine the target text-to-image model and the target parameter translator corresponding to the target face morphing style of the target face morphing system.
[0222] Based on the same inventive concept, embodiments of the present application also provide a computer-readable storage medium corresponding to the face pinching data processing method. Since the principle of solving problems by the computer-readable storage medium in the embodiments of the present application is similar to the above-mentioned face pinching data processing method in the embodiments of the present application, the implementation of the computer-readable storage medium can refer to the implementation of the method, and the repeated parts will not be elaborated.
[0223] A computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the processor executes the steps of the above-mentioned face pinching data processing method, specifically as follows:
[0224] In some embodiments, the processor further executes the following steps of training the target text-to-image generation model:
[0225] Obtain a target training set; the target training set includes first sample face pinching images conforming to the face pinching style of the target face pinching system, and sample text descriptions corresponding to the first sample face pinching images;
[0226] Train an initial text-to-image generation model based on the target training set to obtain a target text-to-image generation model corresponding to the target face pinching system.
[0227] In some embodiments, the sample text description corresponding to the first sample face pinching image includes descriptions in multiple different dimensions.
[0228] In some embodiments, the descriptions in multiple different dimensions include: face shape, facial features, makeup, character, type.
[0229] In some embodiments, when the processor executes the step of training the initial text-to-image generation model based on the target training set to obtain a target text-to-image generation model corresponding to the target face pinching system, it specifically executes the following steps:
[0230] Process the sample text description corresponding to the first sample face pinching image in the target training set to obtain multiple sample enhanced text descriptions corresponding to the first sample face pinching image; wherein, each first sample face pinching image corresponds to a sample text description and multiple sample enhanced text descriptions;
[0231] Train the initial text-to-image generation model based on the first sample face pinching image, the sample text description corresponding to the first sample face pinching image, and the multiple sample enhanced text descriptions to obtain a target text-to-image generation model corresponding to the target face pinching system.
[0232] In some embodiments, in the face pinching data processing device, when the training module processes the sample text description corresponding to the first sample face pinching image in the target training set to obtain multiple sample enhanced text descriptions corresponding to the first sample face pinching image, it is specifically used for:
[0233] Randomly extract the sample text descriptions corresponding to the first sample face pinching images and combine them to obtain multiple sample enhanced text descriptions corresponding to the first sample face pinching images;
[0234] and / or,
[0235] Extract the single-dimension text descriptions corresponding to each dimension in the sample text descriptions corresponding to the first sample face pinching images to obtain multiple sample enhanced text descriptions corresponding to the first sample face pinching images.
[0236] In some embodiments, when the processor executes the step of training the initial text-to-image model based on the first sample face pinching image, the sample text description corresponding to the first sample face pinching image, and multiple sample enhanced text descriptions to obtain the target text-to-image model corresponding to the target face pinching system, the following steps are specifically executed:
[0237] Obtain the first sample face pinching images in the target training set and the single-dimension text descriptions corresponding to the first sample face pinching images;
[0238] In some training rounds of training the initial text-to-image model, train based on the first sample face pinching images in the target training set and the single-dimension texts corresponding to the first sample face pinching images.
[0239] In some embodiments, the human faces in at least some of the first sample face pinching images in the target training set include makeup, and the sample text descriptions corresponding to the first sample face pinching images include descriptions corresponding to the makeup.
[0240] In some embodiments, when the processor executes the step of obtaining the target training set, the following steps are specifically executed:
[0241] Obtain the first sample face pinching images rendered by the target face pinching system;
[0242] Determine the sample text descriptions corresponding to each first sample face pinching image;
[0243] Based on the first sample face pinching images and the sample text descriptions, construct the target training set.
[0244] In some embodiments, when the processor executes the step of determining the sample text descriptions corresponding to each first sample face pinching image, the following steps are specifically executed:
[0245] Input the first sample face pinching image into the trained image-to-text model for processing to obtain the content text description of the first sample face pinching image;
[0246] Based on the adjustment instructions for the content text description of the first sample face pinching image, adjust the content text description to obtain the sample text description of the first sample face pinching image.
[0247] In some embodiments, the target parameter translator is trained based on the paired face pinching parameters and the second sample face pinching images of the target face pinching system.
[0248] In some embodiments, when the processor executes the step of obtaining the target face pinching text description of the user for the target face pinching system, the following steps are specifically executed:
[0249] Obtain the target face pinching text description of the text type directly input based on the terminal device; or,
[0250] Obtain the target face pinching text description converted from other types of face pinching descriptions.
[0251] In some embodiments, the other types of face pinching descriptions include at least one of the following:
[0252] Face pinching descriptions of the voice type, face pinching descriptions of the picture type, face pinching descriptions of the cross - language type.
[0253] In some embodiments, after the processor executes the step of sending the target face pinching parameters to the target face pinching system, so that the target face pinching system generates a target face pinching result that conforms to the face pinching text description based on the target face pinching parameters, the following steps are further executed:
[0254] Receive a parameter adjustment instruction for the parameter controller of the target face pinching system;
[0255] Send the face adjustment instruction to the target face pinching system, so that the parameter controller of the target face pinching system adjusts the generated target face pinching result based on the parameter adjustment instruction to obtain an adjusted target face pinching result.
[0256] In some embodiments, the processor further executes the following steps:
[0257] In response to a selection instruction for multiple face pinching styles of the target face pinching system, determine the target face pinching style of the target face pinching system; different face pinching styles of the target face pinching system correspond to different text - to - image models and parameter translators;
[0258] Determine the target text - to - image model and the target parameter translator corresponding to the target face pinching style of the target face pinching system.
[0259] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the method embodiments, and will not be elaborated herein. In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some communication interfaces. The indirect coupling or communication connection of the devices or modules can be in electrical, mechanical, or other forms.
[0260] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0261] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0262] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a platform server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs, etc., which can store program codes.
[0263] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A face pinching data processing method, characterized in that: The method comprises: Get the text description of the target face pinching input by the user; Processing the target face pinching text description through a pre-trained target text-based graph model to obtain a target face pinching image that matches the target face pinching text description and matches the face pinching style of the target face pinching system; Inputting the target face pinching image into the target parameter translator matched with the target Wensheng graph model, and processing the target face pinching image through the target parameter translator to determine the target face pinching parameters for the target face pinching system; The target face pinching parameters are sent to the target face pinching system, so that the target face pinching system generates a target face pinching result that conforms to the face pinching text description based on the target face pinching parameters.
2. The face pinching data processing method according to claim 1, characterized in that: The target text graph model is trained based on the following method: Acquire a target training set; the target training set includes a first sample face pinching image that conforms to the face pinching style of the target face pinching system and a sample text description corresponding to the first sample face pinching image; An initial Wensheng graph model is trained based on the target training set to obtain a target Wensheng graph model corresponding to the target face pinching system.
3. The face pinching data processing method according to claim 2, characterized in that: The sample text description corresponding to the first sample face-pinching image includes descriptions of multiple different dimensions; the descriptions of the multiple different dimensions include: face shape, facial features, makeup, role, and type.
4. The face pinching data processing method according to claim 2 or 3, characterized in that: The step of training the initial Wensheng graph model based on the target training set to obtain a target Wensheng graph model corresponding to the target face pinching system includes: Processing the sample text description corresponding to the first sample face pinching image in the target training set to obtain a plurality of sample enhanced text descriptions corresponding to the first sample face pinching image; wherein each first sample face pinching image corresponds to a sample text description and a plurality of sample enhanced text descriptions; An initial text-based graph model is trained based on a first sample face-pinching image, a sample text description corresponding to the first sample face-pinching image, and a plurality of sample enhanced text descriptions to obtain a target text-based graph model corresponding to a target face-pinching system.
5. The face pinching data processing method according to claim 4, characterized in that: The processing of the sample text description corresponding to the first sample face pinching image in the target training set to obtain a plurality of sample enhanced text descriptions corresponding to the first sample face pinching image includes: Randomly extracting sample text descriptions corresponding to the first sample face pinching image and combining them to obtain a plurality of sample enhanced text descriptions corresponding to the first sample face pinching image; and / or, A single-dimensional text description corresponding to each dimension in the sample text description corresponding to the first sample face-pinching image is extracted to obtain a plurality of sample enhanced text descriptions corresponding to the first sample face-pinching image.
6. The face pinching data processing method according to claim 5, characterized in that: An initial text-based graph model is trained based on a first sample face-pinching image, a sample text description corresponding to the first sample face-pinching image, and a plurality of sample enhanced text descriptions to obtain a target text-based graph model corresponding to a target face-pinching system, including: Obtain a first sample face pinching image in a target training set and a single-dimensional text description corresponding to the first sample face pinching image; In some training rounds of training the initial text-image model, training is performed based on the first sample face-pinching image in the target training set and the single-dimensional text corresponding to the first sample face-pinching image.
7. The face pinching data processing method according to claim 2, characterized in that: The faces of at least some of the first sample pinched face images in the target training set include makeup, and the sample text descriptions corresponding to the first sample pinched face images include descriptions corresponding to the makeup.
8. The face pinching data processing method according to claim 2, characterized in that: The step of obtaining a target training set includes: Obtain a first sample face pinching image rendered by the target face pinching system; Determine a sample text description corresponding to each first sample face pinching image; The target training set is constructed based on the first sample face pinching image and sample text description.
9. The face pinching data processing method according to claim 8, characterized in that: The determining of the sample text description corresponding to each first sample face pinching image includes: Inputting the first sample face-pinching image into the trained image-to-text model for processing to obtain a text description of the content of the first sample face-pinching image; Based on the adjustment instruction for the content text description of the first sample face pinching image, the content text description is adjusted to obtain a sample text description of the first sample face pinching image.
10. The face pinching data processing method according to claim 1, characterized in that: The target parameter translator is trained based on the paired face pinching parameters of the target face pinching system and the second sample face pinching image.
11. The face pinching data processing method according to claim 1, characterized in that: The step of obtaining a text description of a target face pinching input by a user includes: Get the target face-pinching text description based on the text type directly input by the terminal device; or, Get the target face pinching text description converted from other types of face pinching descriptions.
12. The face pinching data processing method according to claim 11, characterized in that: The other types of face-pinching descriptions include at least one of the following: There are voice type face pinching description, picture type face pinching description and cross-language type face pinching description.
13. The face pinching data processing method according to claim 1, characterized in that: After sending the target face pinching parameters to the target face pinching system so that the target face pinching system generates a target face pinching result that conforms to the face pinching text description based on the target face pinching parameters, the method further includes: Receive parameter adjustment instructions for a parameter controller of a target face pinching system; The face adjustment instruction is sent to the target face pinching system, so that the parameter controller of the target face pinching system adjusts the generated target face pinching result based on the parameter adjustment instruction to obtain the adjusted target face pinching result.
14. The face pinching data processing method according to claim 1, characterized in that: The method further comprises: In response to a selection instruction for a plurality of face pinching styles of a target face pinching system, determining a target face pinching style of the target face pinching system; different face pinching styles of the target face pinching system correspond to different Vincent graph models and parameter translators; Determine a target text graph model and a target parameter translator corresponding to a target face pinching style of a target face pinching system.
15. A face pinching data processing device, characterized in that: The device comprises: An acquisition module is used to acquire a target face pinching text description of a user for a target face pinching system; the target face pinching system corresponds to a pre-trained target text graph model and a target parameter translator; A first processing module is used to process the target face pinching text description through a pre-trained target text image model to obtain a target face pinching image that matches the target face pinching text description and matches the face pinching style of the target face pinching system; A second processing module is used to input the target face pinching image into the target parameter translator matched with the target Wensheng graph model, and process the target face pinching image through the target parameter translator to determine the target face pinching parameters for the target face pinching system; A sending module is used to send the target face pinching parameters to the target face pinching system, so that the target face pinching system generates a target face pinching result that conforms to the face pinching text description based on the target face pinching parameters.
16. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory communicate via the bus. When the machine-readable instructions are executed by the processor, the steps of the face-pinching data processing method as described in any one of claims 1 to 14 are performed.
17. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the face pinching data processing method as described in any one of claims 1 to 14 are executed.