Virtual image generation method and device, equipment and storage medium
By obtaining fine-grained attribute information to adjust the facial area of the virtual avatar, the problem of low matching degree between the virtual avatar and the human face image is solved, achieving high similarity and distinguishability of the virtual avatar and improving the display effect.
Patent Information
- Application Number
- CN202110203806.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-23
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2041-07-13
AI Technical Summary
The virtual avatars generated by existing technologies do not match human facial images well, and are prone to facial distortion. Furthermore, the similarity between different virtual avatars is insufficient, resulting in poor display effects.
By acquiring fine-grained attribute information of the target face, the facial area of the initial virtual avatar is adjusted to generate the target virtual avatar, thereby improving the matching accuracy.
It enhances the similarity between the target virtual image and the target human face, making it easier to distinguish virtual images generated from different face images and enriching the display effect of the virtual image.
Smart Images

Figure CN113569614B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device and storage medium for generating virtual images. Background Technology
[0002] With the development of artificial intelligence technology and the increasing entertainment needs of users, a function that generates virtual avatars based on facial images has emerged and is widely used in various scenarios, such as games, photo editing, and film and television. Users can choose their favorite virtual character and merge their own facial image with the virtual character to obtain a virtual avatar that closely resembles their real appearance.
[0003] The common approach to generating virtual avatars using related technologies is as follows: extract two-dimensional facial features from a face image, mine the extracted two-dimensional facial features to obtain corresponding attribute information, and then match the mined attribute information with the virtual character selected by the user to obtain a virtual avatar that closely matches the user's actual appearance.
[0004] The virtual avatars generated using the above method often have significantly different facial features from real human faces and are prone to facial distortion. In addition, the virtual avatars generated using the above method are also similar to each other based on different human face images, meaning that the distinction between different virtual avatars is not high, which leads to poor display results. Summary of the Invention
[0005] This application provides a method, apparatus, device, and storage medium for generating virtual avatars, which can enrich the display effects of virtual avatars. The technical solution is as follows:
[0006] On the one hand, a method for generating virtual avatars is provided, which includes:
[0007] Obtain a target face image, which includes the target face;
[0008] Based on the target face image and the selected target virtual character, an initial virtual avatar is generated;
[0009] Obtain fine-grained attribute information of the target face, which is used to identify at least one of the facial features, hairstyle, or accessory categories of the target face;
[0010] Based on this fine-grained attribute information, the facial region of the initial virtual image is adjusted in detail to generate the target virtual image; wherein, the second matching degree is greater than the first matching degree, the first matching degree refers to the matching degree between the facial region of the initial virtual image and the target face, and the second matching degree refers to the matching degree between the facial region of the target virtual image and the target face.
[0011] On the other hand, a virtual avatar generation device is provided, the device comprising:
[0012] The first acquisition module is used to acquire a target face image, which includes the target face;
[0013] The first generation module is used to generate an initial virtual image based on the target face image and the selected target virtual character;
[0014] The second acquisition module is used to acquire fine-grained attribute information of the target face, which is used to identify at least one of the facial features, hairstyle, or accessory categories of the target face.
[0015] The second generation module is used to make detailed adjustments to the facial region of the initial virtual image based on the fine-grained attribute information to generate the target virtual image; wherein, the second matching degree is greater than the first matching degree, the first matching degree refers to the matching degree between the facial region of the initial virtual image and the target face, and the second matching degree refers to the matching degree between the facial region of the target virtual image and the target face.
[0016] In one alternative implementation, the first generation module includes:
[0017] The acquisition unit is used to acquire the three-dimensional face parameters of the target face based on the target face image; wherein the three-dimensional face parameters include contour shape parameters and facial feature shape parameters that match the target face;
[0018] The generation unit is used to determine the facial material that matches the target face from the facial material library of the target virtual character, and to render the three-dimensional face parameters based on the determined facial material to generate the initial virtual image.
[0019] In one alternative implementation, the acquisition unit is used for:
[0020] Obtain the three-dimensional facial features of the target face from the target face image;
[0021] The 3D facial features are compared with the target comparison library to obtain the 3D facial parameters of the target face. The target comparison library is used to provide the correspondence between the 3D facial features and the 3D facial parameters.
[0022] In an alternative implementation, the acquisition unit is further configured to:
[0023] Feature extraction is performed on the target face image to obtain the two-dimensional facial features of the target face;
[0024] Basic attribute analysis is performed on the target face image to obtain the basic attribute information of the target face. The classification accuracy of the basic attribute information is lower than that of the fine-grained attribute information.
[0025] Based on the target face image, the target face is reconstructed in three dimensions to obtain a three-dimensional reconstruction model of the target face;
[0026] Based on the two-dimensional facial features, the basic attribute information, and the three-dimensional reconstruction model, the intermediate features of the target face are obtained, and the feature dimension of the intermediate features is greater than a preset threshold.
[0027] The intermediate features are weighted and fused to obtain the three-dimensional facial features of the target face.
[0028] In one alternative implementation, the device further includes:
[0029] The recognition module is used to perform facial key point recognition on the target face image according to the number of second key points, and obtain the facial feature proportion parameters of the target face. The number of second key points is less than the number of first key points.
[0030] The facial features adjustment module is used to adjust the proportions of the facial features of the initial virtual avatar based on the facial feature proportion parameters.
[0031] In one alternative implementation, the device further includes:
[0032] The art style adjustment module is used to adjust the art style of the facial area of the target virtual character according to the style of the virtual scene in which the target virtual character is located.
[0033] In one alternative implementation, the device further includes:
[0034] The first display module is used to display a first page in response to receiving a virtual avatar creation instruction. The first page includes an image selection window or a viewfinder window.
[0035] The third acquisition module is used to acquire the target face image obtained through the first page;
[0036] The second display module is used to display a second page after the target virtual image is generated, on which the target virtual image is displayed.
[0037] In one alternative implementation, the second page includes M adjustment options, where M is a positive integer, for adjusting the facial display of the target virtual avatar. The device also includes:
[0038] The third display module is used to display at least one target operation item corresponding to any of the adjustment options on the second page in response to a first trigger operation of any of the adjustment options on the second page;
[0039] The fourth acquisition module is used to acquire the adjustment parameters corresponding to the second trigger operation in response to any second trigger operation of the target operation item; and adjust the facial display of the target virtual image according to the adjustment method indicated by the target operation item and the adjustment parameters.
[0040] On the other hand, a computer device is provided, which includes a processor and a memory for storing at least one computer program, which is loaded and executed by the processor to perform the operations performed in the virtual image generation method in the embodiments of this application.
[0041] On the other hand, a computer-readable storage medium is provided that stores at least one computer program, which is loaded and executed by a processor to perform the operations performed in the virtual avatar generation method of the embodiments of this application.
[0042] On the other hand, a computer program product or computer program is provided, which includes computer program code stored in a computer-readable storage medium. A processor of a computer device reads the computer program code from the computer-readable storage medium and executes the computer program code, causing the computer device to perform the virtual avatar generation method provided in the various optional implementations described above.
[0043] The beneficial effects of the technical solutions provided in this application are:
[0044] In this embodiment, when generating a virtual image that resembles a target face based on a user-selected virtual character, an initial virtual image with a high degree of matching with the target face is first generated based on the selected target virtual character and the target face image. Then, through fine-grained attribute information of the target face, the facial area of the initial virtual image is adjusted in detail to ensure the matching degree between the final target virtual image and the target face. This method of adjusting the initial virtual image based on fine-grained attribute information greatly increases the similarity between the target virtual image and the target face, making it easy to distinguish between multiple virtual images generated from different face images and enriching the display effect of the virtual image. Attached Figure Description
[0045] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0046] Figure 1 This is a schematic diagram of the implementation environment of a virtual image generation method provided in the embodiments of this application;
[0047] Figure 2 This is a flowchart of a virtual avatar generation method provided according to an embodiment of this application;
[0048] Figure 3 This is a flowchart of another virtual image generation method provided according to an embodiment of this application;
[0049] Figure 4 This is a schematic diagram illustrating an embodiment of obtaining a three-dimensional reconstruction model according to an embodiment of this application;
[0050] Figure 5 This is a schematic diagram illustrating the acquisition of dense recognition results according to an embodiment of this application;
[0051] Figure 6 This is a schematic diagram of fine-grained attribute information provided according to an embodiment of this application;
[0052] Figure 7 This is a schematic diagram of a fine-grained attribute classification model provided according to an embodiment of this application;
[0053] Figure 8 This is a schematic diagram illustrating the adjustment of facial proportions according to an embodiment of this application;
[0054] Figure 9 This is a schematic diagram illustrating the generation of a target virtual image according to an embodiment of this application;
[0055] Figure 10 This is a flowchart of another virtual image generation method provided according to an embodiment of this application;
[0056] Figure 11 This is a schematic diagram of a virtual character selection page provided according to an embodiment of this application;
[0057] Figure 12 This is a schematic diagram of a virtual character customization page provided according to an embodiment of this application;
[0058] Figure 13 This is a schematic diagram of a prompt page provided according to an embodiment of this application;
[0059] Figure 14 This is a schematic diagram of a first page provided according to an embodiment of this application;
[0060] Figure 15 This is a schematic diagram of a second page provided according to an embodiment of this application;
[0061] Figure 16 This is a flowchart of another virtual image generation method provided according to an embodiment of this application;
[0062] Figure 17 This is a flowchart of a local detail optimization provided according to an embodiment of this application;
[0063] Figure 18 This is a schematic diagram of the structure of a virtual image generation device according to an embodiment of this application;
[0064] Figure 19 This is a schematic diagram of the structure of a terminal according to an embodiment of this application. Detailed Implementation
[0065] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0066] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0067] In this application, the terms "first," "second," etc., are used to distinguish identical or similar items that have essentially the same function. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor does it limit the quantity or execution order. It should also be understood that although the following description uses the terms "first," "second," etc., to describe various elements, these elements should not be limited by the terms.
[0068] These terms are simply used to distinguish one element from another. For example, without departing from the various examples, the first image can be referred to as the second image, and similarly, the second image can be referred to as the first image. Both the first image and the second image can be images, and in some cases, they can be separate and distinct images.
[0069] "At least one" refers to one or more images. For example, at least one image can be one image, two images, three images, or any integer number of images greater than or equal to one. "Multiple" refers to two or more images. For example, multiple images can be two images, three images, or any integer number of images greater than or equal to two.
[0070] The following describes the technologies that may be used in the virtual avatar generation scheme provided in the embodiments of this application.
[0071] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0072] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0073] Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instruction-based learning.
[0074] Computer vision (CV) is a science that studies how to enable machines to "see." More specifically, it refers to machine vision, which uses cameras and computers to replace human eyes in recognizing, tracking, and measuring targets, and then performs image processing to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, object-contextual representations (OCR), video processing, video semantic understanding, video content / behavior recognition, 3D reconstruction, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), and common biometric recognition technologies such as facial recognition and fingerprint recognition.
[0075] The following describes the key terms or abbreviations that may be used in the virtual avatar generation scheme provided in the embodiments of this application.
[0076] Multi-Task Convolutional Neural Network (MTCNN): A deep learning-based method for face detection and facial landmark recognition, capable of simultaneously performing face detection and facial landmark recognition tasks.
[0077] 3D Morphable Models (3DMMs): A type of model used to reconstruct the three-dimensional shape of a two-dimensional face image. It uses a fixed number of points to represent the face. The core idea of this method is that faces can be matched one-to-one in three-dimensional space and can be obtained by weighted linear summation of many other orthogonal basis images of faces.
[0078] Face registration: also known as facial landmark recognition. Based on face detection, it automatically locates key facial feature points, such as eyes, nose tip, corners of mouth, eyebrows, and contour points of various facial components, according to the input face image.
[0079] Classification task: Used to classify a certain type of object (such as an eye shape) and predict it as one of the multiple categories in the classification task.
[0080] The normalized exponential function (Softmax) is a generalization of the logistic function. It can "compress" a K-dimensional vector containing arbitrary real numbers into another K-dimensional real vector, such that each element is in the range [0, 1], and the sum of all elements is 1.
[0081] The following describes the implementation environment of the virtual avatar generation method provided in the embodiments of this application.
[0082] Figure 1 This is a schematic diagram of the implementation environment of the virtual avatar generation method provided in the embodiments of this application. The implementation environment includes: terminal 101 and server 102.
[0083] Terminal 101 and server 102 can be connected directly or indirectly via wired or wireless communication, and this application does not impose any limitations on this connection. Optionally, terminal 101 can be a smartphone, tablet, laptop, desktop computer, etc., but is not limited to these. Terminal 101 can install and run applications. Optionally, the application can be a game application, a social application, a camera application, or an image processing application, etc. Illustratively, terminal 101 is a terminal used by a user, and the user's account is logged into the application running on terminal 101.
[0084] Server 102 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. Server 102 is used to provide background services for the applications running on terminal 101.
[0085] Optionally, during the generation of the virtual avatar, server 102 undertakes the main computing work and terminal 101 undertakes the secondary computing work; or, server 102 undertakes the secondary computing work and terminal 101 undertakes the main computing work; or, server 102 or terminal 101 can each undertake computing work independently.
[0086] Optionally, terminal 101 generally refers to one of multiple terminals; this embodiment only uses terminal 101 as an example. Those skilled in the art will understand that the number of terminals 101 can be greater. For example, there may be dozens or hundreds, or even more, terminals 101. In this case, the implementation environment of the virtual avatar generation method also includes other terminals. This application embodiment does not limit the number of terminals or the type of device.
[0087] Optionally, the aforementioned wireless or wired networks use standard communication technologies and / or protocols. The network is typically the Internet, but can be any network, including but not limited to Local Area Networks (LANs), Metropolitan Area Networks (MANs), Wide Area Networks (WANs), mobile, wired or wireless networks, private networks, or any combination of virtual private networks. In some embodiments, technologies and / or formats, including Hyper Text Markup Language (HTML), Extensible Markup Language (XML), etc., are used to represent data exchanged over the network. Furthermore, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Networks (VPNs), and Internet Protocol Security (IPsec) can be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can be used to replace or supplement the aforementioned data communication technologies.
[0088] Indicatively, the application scenarios of the virtual avatar generation method provided in this application embodiment include, but are not limited to, the following example scenarios:
[0089] Scene 1: Game Scene
[0090] Many game applications offer a variety of game characters for users to choose from. After selecting a game character, users can upload their own facial image to blend their actual appearance with the game character, resulting in a game character image that closely matches their real appearance and enriching the display effect of the game character image.
[0091] Scenario 2: Emoji Scenarios
[0092] With the rise of emoji culture, many apps have added emoji creation features, allowing users to express their emotions and feelings through emojis. In some situations, users want to blend their facial images with various virtual avatars to create virtual characters that closely resemble their real-life appearance. Based on this, they create emojis, enriching their emoji display.
[0093] Scene 3: Movie and TV series scenes
[0094] In the process of filming movies and TV series, sometimes it is necessary to merge live-action and animated characters. Film and TV series producers can merge the actors' facial images with animated characters to obtain animated characters that are closer to the actors' actual appearance, thus enriching the display effect of animated characters in movies and TV series.
[0095] Figure 2 This is a flowchart of a virtual avatar generation method provided according to an embodiment of this application, such as... Figure 2 As shown, this embodiment of the application uses a terminal as an example for explanation. The virtual avatar generation method includes the following steps:
[0096] 201. Obtain the target face image, which includes the target face.
[0097] The target face image refers to the image containing the target face input by the user.
[0098] 202. Based on the target face image and the selected target virtual character, generate an initial virtual image.
[0099] The target virtual character refers to the virtual character selected by the user. The initial virtual image is a virtual image resulting from combining the target human face with the target virtual character.
[0100] Optionally, the target virtual character is a virtual character. For example, the target virtual character is a male character in ancient costume; or, for another example, the target virtual character is a female character from an anime, and so on.
[0101] Optionally, the target virtual character is a virtual animal character. For example, the target virtual character is a sphinx-like animal character. This application does not limit the specific type of the target virtual character.
[0102] 203. Obtain fine-grained attribute information of the target face, which is used to identify at least one of the facial features, hairstyle, or accessory categories of the target face.
[0103] In the embodiments of this application, fine-grained attribute information provides higher accuracy in face classification.
[0104] Among these categories, facial features refer to the type of facial features of the target person, such as almond-shaped eyes, a hooked nose, etc. Hairstyle refers to the type of hair on the target person, such as a ponytail, a bun, or a very short haircut, etc. Accessory refers to the type of ornaments on the target person, such as tassel earrings, square-framed glasses, etc. This application does not limit the specific forms of the above-mentioned facial features, hairstyle, and accessory categories.
[0105] 204. Based on the fine-grained attribute information, the facial region of the initial virtual image is adjusted in detail to generate the target virtual image; wherein, the second matching degree is greater than the first matching degree, the first matching degree refers to the matching degree between the facial region of the initial virtual image and the target face, and the second matching degree refers to the matching degree between the facial region of the target virtual image and the target face.
[0106] Among them, the matching degree is used to indicate the degree of similarity between the virtual image and the target human face, and can also be called similarity.
[0107] In this embodiment, when generating a virtual image that resembles a target face based on a user-selected virtual character, an initial virtual image with a high degree of matching with the target face is first generated based on the selected target virtual character and the target face image. Then, through fine-grained attribute information of the target face, the facial area of the initial virtual image is adjusted in detail to ensure the matching degree between the final target virtual image and the target face. This method of adjusting the initial virtual image based on fine-grained attribute information greatly increases the similarity between the target virtual image and the target face, making it easy to distinguish between multiple virtual images generated from different face images and enriching the display effect of the virtual image.
[0108] The above Figure 2 The diagram shown is only the basic process of this application. The following is a further explanation of the virtual image generation scheme provided by this application based on a specific implementation method.
[0109] Figure 3 This is a flowchart of another virtual avatar generation method provided according to an embodiment of this application, such as... Figure 3 As shown, this application embodiment uses a terminal as an example for illustration. The method includes the following steps:
[0110] 301. Obtain the target face image, which includes the target face.
[0111] In this embodiment, the terminal provides a virtual avatar creation function. The user triggers a creation command by performing a creation operation on the terminal. Upon receiving the creation command, the terminal displays an image upload page. Then, in response to the user's image input operation on the image upload page, the terminal acquires a target face image. Optionally, the target face image is a locally stored image on the terminal, or it is obtained by capturing a target face using the terminal's camera. This application does not limit the source of the target face image.
[0112] 302. Based on the target face image, obtain the three-dimensional face parameters of the target face; wherein, the three-dimensional face parameters include contour shape parameters and facial feature shape parameters that match the target face.
[0113] In this embodiment, the three-dimensional face parameters are used to indicate the contour shape and facial feature shape that match the target face. Optionally, the specific implementation of the terminal acquiring the three-dimensional face parameters of the target face includes the following steps 3021 and 3022:
[0114] 3021. Obtain the three-dimensional facial features of the target face from the target face image.
[0115] The three-dimensional facial features are obtained by extracting features from the target face image. These three-dimensional facial features are used to indicate the shape of the target face and the shape of its facial features. Optionally, the three-dimensional facial features can be represented in the form of a vector or a matrix. Optionally, the terminal obtains the three-dimensional facial features of the target face by performing feature extraction, basic attribute analysis, and three-dimensional reconstruction on the target face image. This optional implementation method is described in detail below, including steps 3021-1 to 3021-5.
[0116] 3021-1. Extract features from the target face image to obtain the two-dimensional facial features of the target face.
[0117] The terminal performs face detection on the target face image. After detecting the target face, it extracts features from the target face to obtain two-dimensional facial features. Optionally, these two-dimensional facial features include geometric relationships between facial features such as eyes, nose, and mouth, such as distance, area, and angle. Optionally, these two-dimensional facial features include histogram features obtained based on the grayscale information of the target face image. This application does not limit the method of obtaining the two-dimensional facial features.
[0118] 3021-2. Perform basic attribute analysis on the target face image to obtain the basic attribute information of the target face. The classification accuracy of the basic attribute information is lower than that of the fine-grained attribute information.
[0119] The basic attribute information is used to identify the facial features of the target face. Optionally, facial features include the target face's age, gender, and skin color. Optionally, the terminal performs facial landmark recognition on the target face image to obtain multiple landmarks of the target face. Based on these landmarks, basic attribute analysis is performed on the target face image to obtain the basic attribute information of the target face. This application does not limit the method of obtaining the basic attribute information.
[0120] 3021-3. Based on the target face image, perform three-dimensional reconstruction of the target face to obtain the three-dimensional reconstruction model of the target face.
[0121] For an illustrative example, the specific implementation method of the terminal performing 3D reconstruction of the target face can be found in [reference needed]. Figure 4 , Figure 4 This is a schematic diagram illustrating an embodiment of obtaining a three-dimensional reconstructed model provided in this application. For example... Figure 4 As shown, the terminal performs facial key point recognition on the target face image 1, which is also called face registration, and obtains the registration result 2 for the target face image. Based on the registration result 2, the 3DMM reconstruction method is used to perform pose projection 3, face shape acquisition 4, and expression shape acquisition 5 on the target face, and finally obtain the three-dimensional reconstruction model 6 of the target face.
[0122] It should be noted that in practical applications, the terminal can also obtain the 3D reconstruction model of the target face through other methods, such as the Shape From Shading (SFS) method or stereo matching method. This application does not limit the method of obtaining the 3D reconstruction model of the target face.
[0123] 3021-4. Based on the two-dimensional facial features, the basic attribute information, and the three-dimensional reconstruction model, obtain the intermediate features of the target face, wherein the feature dimension of the intermediate features is greater than a preset threshold.
[0124] The intermediate feature is a high-dimensional feature of the target face. This intermediate feature is used to represent the feature information of the target face from multiple dimensions. For example, these multiple dimensions include contour shape, facial feature shape, facial feature proportions, age, skin color, gender, and expression shape, etc. Optionally, the preset threshold is set to 128 dimensions, or 256 dimensions, etc. This application does not limit the preset threshold for this feature dimension.
[0125] 3021-5. Perform weighted fusion on the intermediate features to obtain the three-dimensional facial features of the target face.
[0126] Optionally, the terminal performs weighted fusion of the intermediate feature according to the importance of each dimension of the intermediate feature. Optionally, the terminal performs weighted fusion of the intermediate feature according to the degree of influence of each dimension of the intermediate feature on the target face. This application does not limit the specific implementation method of weighted fusion.
[0127] 3022. Compare the three-dimensional facial features of the target face with the target comparison library to obtain the three-dimensional facial parameters of the target face. The target comparison library is used to provide the correspondence between the three-dimensional facial features and the three-dimensional facial parameters.
[0128] The target comparison library contains multiple sets of correspondences between 3D facial features and 3D facial parameters. The terminal inputs the 3D facial features of the target face into the target comparison library for feature comparison to obtain contour shape parameters and facial feature shape parameters similar to the target face, which are the 3D facial parameters of the target face.
[0129] 303. From the facial material library of the target virtual character, determine the facial material that matches the target face, and render the three-dimensional face parameters based on the determined facial material to generate the initial virtual image.
[0130] In this embodiment, different virtual characters correspond to different facial image libraries. The terminal obtains the facial image library of the target virtual character, determines the facial image that matches the target face, and based on this, calls the renderer to render the 3D facial parameters to obtain the initial virtual image.
[0131] Optionally, taking a game scene as an example, the terminal loads the 3D face parameters into the game engine. The game engine obtains the face material library of the target virtual character, then determines the face material that matches the target face. Based on this, it calls the renderer in the game engine to render the 3D face parameters and obtain the initial virtual image.
[0132] It should be noted that, after steps 301 to 303 above, the terminal generates an initial virtual avatar that matches the target face based on the target face image and the selected target virtual character. This initial virtual avatar has a similar outline shape and facial features to the target face.
[0133] 304. Obtain fine-grained attribute information of the target face, which is used to identify at least one of the facial features, hairstyle, or accessory categories of the target face.
[0134] In this embodiment, the terminal performs facial landmark recognition on the target face image to obtain a dense recognition result of the target face, and then obtains fine-grained attribute information of the target face based on the dense recognition result. The implementation of step 304 is described in detail below, including steps 3041 to 3042:
[0135] 3041. Based on the number of the first key points, perform facial key point recognition on the target face image to obtain the dense recognition result of the target face.
[0136] Indicatively, in this embodiment, the number of first key points can be 1000. In some embodiments, the number of first key points is greater than 1000, or less than 1000. In practical applications, the number of first key points can be set according to requirements, and this application does not limit this.
[0137] Optionally, the terminal obtains dense recognition results of the target face based on a facial landmark recognition model. Optionally, the facial landmark recognition model is a model built based on a deep convolutional neural network. For example, the facial landmark recognition model is an MTCNN model, which is used to complete the recognition of facial landmarks in the target face image. This application embodiment does not limit the specific type of facial landmark recognition model.
[0138] Indicatively, for reference Figure 5 , Figure 5 This is a schematic diagram illustrating an embodiment of this application for obtaining dense recognition results. For example... Figure 5 As shown, the terminal uses a facial landmark recognition model to set the number of facial landmarks as the first number of landmarks, and then performs facial landmark recognition on the target face image, which is also called face registration. This includes recognizing the facial features of the target face. After stabilizing and refining the positions of the facial landmarks, the terminal finally outputs the dense recognition result of the target face.
[0139] In some embodiments, in the facial landmark recognition model described above, facial landmarks are classified into M categories according to their localization difficulty, where M is an integer greater than 1. Optionally, M can be 3, meaning the M categories include contour points, fine facial feature points, and facial feature localization points. Contour points are facial landmarks used to construct the facial contour, fine facial feature points are facial landmarks used to construct the facial feature contour, and facial feature localization points are facial landmarks used to locate the facial features. For example, the number of facial landmarks is 1000, including 200 contour points, 600 fine facial feature points, and 200 facial feature localization points, with each facial landmark evenly distributed in its corresponding region.
[0140] In other embodiments, in the above-mentioned facial landmark recognition model, facial landmarks are divided into strong semantic points and weak semantic points according to their semantic intensity. Among them, strong semantic points are the vertices and corners in the facial structure, such as the corners of the eyes, the tip of the nose, and the corners of the mouth; weak semantic points are distributed on the strong texture edges of the face and are used to indicate the arcs in the facial structure, such as points at the facial contour, bridge of the nose, and eye sockets.
[0141] The embodiments of this application do not limit the specific distribution location and type of facial key points.
[0142] Optionally, the training process of the facial landmark recognition model includes the following steps 3041-1 to 3041-4:
[0143] 3041-1. Input the training samples into the constructed initial model to obtain the prediction density recognition results of the training samples.
[0144] The training samples include face images with labeled real-world locations of facial key points. Correspondingly, this labeling result is referred to as the standard dense recognition result in this paper. Based on the network parameters of this initial model, the terminal performs facial key point recognition on the training samples according to the first number of key points, obtaining the predicted dense recognition result of the training samples. The predicted dense recognition result is the predicted location of the facial key points in the training samples output by the initial model.
[0145] The number of training samples is usually large. Training the facial landmark recognition model with a large number of training samples can make the final trained facial landmark recognition model have good universality and robustness. In addition, the embodiments of this application do not limit the network structure of the initial model.
[0146] 3041-2. Based on the predicted dense recognition results and the standard dense recognition results of the training samples, calculate the first loss function of the initial model. This first loss function is used to monitor the contour localization accuracy of the initial model.
[0147] The first loss function can be constructed based on the predicted and actual locations of all facial key points in the training samples. For example, it can be represented by the Euclidean distance between the predicted and actual locations of all facial key points. Alternatively, the first loss function can be any loss function commonly used in model training, such as the absolute value loss function, cosine similarity loss function, squared loss function, cross-entropy loss function, etc., and this embodiment does not limit the specific loss function used.
[0148] 3041-3. Perform detail calibration on the predicted dense recognition results of the training samples to obtain the second loss function of the initial model. The second loss function is used to monitor the facial feature localization accuracy of the initial model.
[0149] Detail calibration refers to comparing the facial feature regions in the predicted dense recognition results of the training samples with the facial feature regions in the standard dense recognition results. Optionally, the second loss function can be constructed based on the predicted and actual locations of the facial key points corresponding to the facial feature regions in the training samples. For example, it can be represented by the Euclidean distance between the predicted and actual locations of the facial key points corresponding to the facial feature regions.
[0150] It should be noted that, after steps 3041-2 and 3041-3 above, the terminal compares the predicted dense recognition result of the training sample with the standard dense recognition result twice. The first comparison is an overall comparison of all facial key points of the training sample, and the second comparison is a detailed comparison of the facial key points corresponding to the facial features of the training sample. Two loss functions are obtained through these two comparisons.
[0151] Furthermore, in this embodiment, the terminal executes steps 3041-2 and 3041-3 in a forward-to-back order. In some embodiments, the terminal executes step 3041-3 first, then step 3041-2. In other embodiments, the terminal executes steps 3041-2 and 3041-3 simultaneously. This embodiment does not limit the scope of the embodiments.
[0152] 3041-4. When both the first loss function and the second loss function meet the target conditions, the training is completed, and a well-trained facial landmark recognition model is obtained.
[0153] The objective condition is that the loss value (also known as the error value) is less than a set threshold. This threshold can be set according to actual needs, such as based on the model's positioning accuracy; this application does not impose any restrictions on this. Furthermore, in response to either the first loss function or the second loss function failing to meet the objective condition, the network parameters of the current model are adjusted, and then the process starts again from step 3041-1 until both the first and second loss functions meet the objective condition, at which point training stops.
[0154] Optionally, step 3041-4 above can also be replaced by "summing the first loss function and the second loss function according to preset weight coefficients to obtain the total loss function; in response to the total loss function meeting the target conditions, completing the training to obtain a trained facial landmark recognition model." Here, the total loss function monitors both the contour localization accuracy and the facial feature localization accuracy of the model.
[0155] It should be noted that the training process of the aforementioned facial landmark recognition model may also include other steps or other optional implementation methods, which are not limited in this application. Furthermore, by adopting this optional implementation method, the terminal introduces a detail calibration constraint as a loss function when training the facial landmark recognition model. That is, the loss function is related to detail calibration, which ensures that the obtained dense recognition results are closer to the actual target face, avoiding the adverse effects of subtle differences in facial landmarks on the dense recognition results.
[0156] Optionally, the terminal may perform multiple facial landmark recognitions on the target face image. This optional implementation is described in detail below, including any of the following scenarios:
[0157] Scenario 1: The terminal sequentially inputs the target face image into the facial landmark recognition model according to a preset number of inputs, obtaining multiple dense recognition results for the target face. Then, the error value between any two dense recognition results is calculated, and the multiple error values are summed to obtain the total error value of the multiple dense recognition results. If the total error value is less than or equal to a first threshold, the target face image is input into the facial landmark recognition model again to obtain a dense recognition result for the target face. If the total error value is greater than the first threshold, the target face image is input into the facial landmark recognition model again according to the preset number of inputs, and this process is repeated until the total error value is less than or equal to the first threshold. At this point, the target face image is input into the facial landmark recognition model again to obtain a dense recognition result for the target face. In this embodiment, the preset number of inputs and the first threshold are not limited.
[0158] Scenario 2: After the terminal inputs the target face image into the facial landmark recognition model, it obtains the first dense recognition result of the target face. Then, the terminal inputs the target face image into the facial landmark recognition model again to obtain the second dense recognition result of the target face, and calculates the error value between the second dense recognition result and the first dense recognition result. If the error value is less than or equal to a second threshold, the second dense recognition result is taken as the dense recognition result of the target face. If the error value is greater than the second threshold, the terminal inputs the target face image into the facial landmark recognition model again to obtain the third dense recognition result of the target face, and calculates the error value between the third dense recognition result and the second dense recognition result. This error value is compared with the second threshold, and so on, until the error value obtained by the terminal is less than or equal to the second threshold. At this point, the last dense recognition result obtained is taken as the dense recognition result of the target face. In this embodiment, the setting of the second threshold is not limited.
[0159] Scenario 3: The facial landmark recognition model itself has the function of multiple recognitions. That is, after the terminal inputs the target face image into the facial landmark recognition model, the model performs multiple facial landmark recognitions on the target face image until the error value between the multiple dense recognition results is less than or equal to a third threshold, at which point it outputs the dense recognition result of the target face. In this embodiment, the setting of the third threshold is not limited. Furthermore, the method by which the facial landmark recognition model determines the error value between multiple dense recognition results can refer to Scenario 1 or Scenario 2 above. That is, the facial recognition model can either perform recognition according to a preset number of recognitions, calculate the total error value of multiple dense recognition results, and finally output the dense recognition result of the target face; or it can recognize sequentially, calculate the error value between multiple dense recognition results sequentially, and finally output the dense recognition result of the target face. This embodiment will not be elaborated further here.
[0160] It should be noted that the embodiments of this application do not limit the method by which the terminal performs multiple facial landmark recognitions. In addition, by adopting this optional implementation method, the terminal performs multiple facial landmark recognitions on the target face, which can ensure the stability of the final dense recognition result, and make the dense recognition results generated by different poses of the same face have a certain consistency.
[0161] In addition, in the embodiments of this application, the dense recognition result obtained based on the number of first key points can accurately depict the detailed shape of the facial features of the target face, which is convenient for subsequent fine-grained attribute analysis of the target face image.
[0162] 3042. Based on the dense recognition result, fine-grained attribute analysis is performed on the target face image to obtain the fine-grained attribute information of the target face.
[0163] Indicatively, for reference Figure 6 , Figure 6 This is a schematic diagram illustrating fine-grained attribute information provided in an embodiment of this application. For example... Figure 6 As shown in the embodiments of this application, fine-grained attribute analysis of the target face image can obtain fine-grained attribute information 7 of the target face, such as willow leaf eyebrows, phoenix eyes, cherry lips, hooked nose, and oval face. Through these concise terms used to describe shapes, the shapes of distinct facial features are expressed with extremely high recognizability. Optionally, continue to refer to... Figure 6 Fine-grained attribute analysis of the target face image can also yield fine-grained attribute information, such as hairstyle (short hair), hair color (black), bangs (with or without bangs), glasses (with or without glasses), skin color (yellow skin), and expression (smiling), etc. This application does not limit this.
[0164] It should be noted that the above Figure 6 The illustrations are merely illustrative. In this embodiment, fine-grained attribute information is used to identify at least one of facial feature categories, hairstyle categories, and accessory categories. In some embodiments, fine-grained attribute information is used to identify facial feature categories, hairstyle categories, and accessory categories. In practical applications, the types of categories identified by the fine-grained attribute information can be set as needed. This application does not limit this.
[0165] Optionally, the terminal can perform fine-grained attribute analysis on the target face image using a fine-grained attribute classification model. This optional implementation method is described in detail below, including steps 3042-1 to 3042-3.
[0166] 3042-1. Based on the dense recognition result, the target face image is locally cropped to obtain N sub-regions. The N sub-regions include at least one of the facial features, hairstyle, or accessories of the target face, where N is a positive integer.
[0167] Specifically, based on the dense recognition result of the target face obtained in step 3041, the terminal performs local cropping on the target face image to obtain N sub-regions of the target face image. Optionally, taking the facial features region as an example, the N sub-regions include the eyebrow region, eye region, nose region, mouth region, and ear region, etc. Taking the hairstyle region as an example, the N sub-regions include the bangs region and other regions, etc. Taking the accessories region as an example, the N sub-regions include the earring region, glasses region, and hair accessory region, etc.
[0168] 3042-2. Input the N sub-regions into the fine-grained attribute classification model, which is a neural network model based on multi-task classification.
[0169] Indicatively, for reference Figure 7 , Figure 7 This is a schematic diagram of a fine-grained attribute classification model provided in an embodiment of this application. For example... Figure 7 As shown, this fine-grained attribute classification model is a neural network model based on multi-task classification. The terminal inputs N locally cropped sub-regions into this fine-grained attribute classification model, which then passes through the backbone network and the Softmax function, finally outputting the attribute information corresponding to each sub-region.
[0170] Optionally, taking the eye region as an example, the training process of this fine-grained attribute classification model includes: obtaining eye region samples obtained by local cropping based on dense recognition results; inputting the eye region samples into the fine-grained attribute classification model; classifying based on the network parameters of the fine-grained attribute classification model to obtain the fine-grained attribute information corresponding to the eye region samples; determining the loss function of the fine-grained attribute classification model; adjusting the network parameters of the fine-grained attribute classification model according to the loss function; inputting new eye region samples again based on the adjusted fine-grained attribute classification model; iteratively adjusting the network parameters of the fine-grained attribute classification model until the loss function of the fine-grained attribute classification model meets the target conditions, thus completing the training.
[0171] It should be noted that the training process of the fine-grained attribute classification model may also include other steps or other optional implementation methods, which are not limited in this application.
[0172] 3042-3. Use the output of the fine-grained attribute recognition model as the fine-grained attribute information of the target face.
[0173] The first point to note is that, through steps 3042-1 to 3042-3 above, this fine-grained attribute classification model based on multi-task classification is used to perform fine-grained attribute analysis on the target face image, which can simultaneously obtain multiple fine-grained attribute information, greatly improving the efficiency of obtaining fine-grained attribute information.
[0174] The second point to note is that, in the embodiments of this application, steps 301 to 304 are executed in a forward-to-back order. In some embodiments, the terminal executes steps 302 and 304 simultaneously after executing step 301; that is, after acquiring the target face image, the terminal acquires the fine-grained attribute information of the target face while acquiring the three-dimensional face parameters. In other embodiments, the terminal executes step 304 first after executing step 301, and then executes steps 302 and 303; that is, after acquiring the target face image, the terminal first acquires the fine-grained attribute information of the target face, and then acquires the three-dimensional face parameters of the target face to generate an initial virtual image. This application does not limit the execution order of step 304.
[0175] 305. Perform facial landmark recognition on the target face image according to the number of second key points to obtain the facial feature ratio parameters of the target face. The number of second key points is less than the number of first key points.
[0176] Schematic illustration: In this embodiment, the number of second key points is 94. After the terminal performs facial key point recognition on the target face image, it performs line drawing processing based on the distribution of facial features of the target face to obtain the facial feature ratio parameters of the target face. Schematic illustration: Referring to... Figure 8 , Figure 8 This is a schematic diagram illustrating the adjustment of facial proportions according to an embodiment of this application. Figure 8 As shown, line drawing is performed based on the distribution of facial features of the target face to obtain the facial feature ratio parameters 9 of the target face.
[0177] 306. Adjust the facial proportions of the initial virtual character based on the facial proportion parameters.
[0178] For illustrative purposes, please continue to refer to Figure 8 ,like Figure 8 As shown, based on the facial feature proportion parameters of the target human face, the facial feature proportions of the initial virtual image are adjusted in both the horizontal and vertical directions.
[0179] The first point to note is that in this embodiment, the number of second key points is 94. In some embodiments, the number of second key points is 68 or 39, etc. In practical applications, the number of second key points can be set according to requirements, and this application does not limit this.
[0180] The second point to note is that, after steps 305 and 306 above, the terminal adjusts the facial proportions of the initial virtual image, improving the overall distribution of facial features. Facial proportions are an important standard for facial impressions and largely determine whether the facial area of the initial virtual image resembles the target face. Therefore, the above method can increase the similarity between the facial area of the initial virtual image and the target face.
[0181] The third point to note is that steps 305 and 306 described above are an optional implementation provided by the embodiments of this application. In other embodiments, the terminal executes step 304 above and then directly executes step 307 below. This application does not limit this.
[0182] The fourth point to note is that, in the embodiments of this application, the terminal executes steps 304 to 306 in a sequential order. In some embodiments, the terminal executes steps 305 and 306 simultaneously with step 304. In other embodiments, the terminal executes steps 305 and 306 first, and then executes step 304. This application does not limit the execution order of steps 305 and 306.
[0183] 307. Based on the fine-grained attribute information, the facial region of the initial virtual image is adjusted in detail to generate the target virtual image; wherein, the second matching degree is greater than the first matching degree, the first matching degree refers to the matching degree between the facial region of the initial virtual image and the target face, and the second matching degree refers to the matching degree between the facial region of the target virtual image and the target face.
[0184] In this embodiment, the terminal adjusts the parameters of the facial region of the initial virtual avatar based on the fine-grained attribute information of the target face to generate the target virtual avatar. For example, if the eyebrow shape of the facial region in the initial virtual avatar is the default eyebrow shape, and the eyebrow shape of the target face is determined to be willow-leaf eyebrows based on the fine-grained attribute information, then the terminal fine-tunes the eyebrow shape of the initial virtual avatar to generate willow-leaf eyebrows. As another example, if the pupil color of the facial region of the initial virtual avatar is black, and the pupil color of the target face is determined to be brown based on the fine-grained attribute information, then the terminal adjusts the pupil color of the initial virtual avatar to brown. This application does not limit this.
[0185] Indicatively, for reference Figure 9 , Figure 9 This is a schematic diagram illustrating a method for generating a target virtual image based on fine-grained attribute information, as provided in an embodiment of this application. For example... Figure 9As shown, the fine-grained attribute information of the target face indicates that the target face has straight eyebrows, small eyes, an upturned nose, and thick lips that are closed. The target virtual image generated based on these fine-grained attribute information has features similar to the target face.
[0186] It should be noted that, after step 307 above, the terminal makes fine adjustments to the facial region of the initial virtual image based on the fine-grained attribute information of the target face, so that the facial region of the final generated target virtual image is more similar to the target face.
[0187] 308. Adjust the art style of the face area of the target virtual character according to the style of the virtual scene in which the target virtual character is located.
[0188] In this application embodiment, a virtual scene style can support multiple target virtual characters. Optionally, different virtual characters may be situated in different virtual scene styles. This application does not impose any limitations on this.
[0189] Specifically, the terminal adjusts the art style of the target virtual character's facial features based on the style of the virtual scene in which the target virtual character is situated, ensuring that the style of the target virtual character matches the style of the virtual scene. For example, in a game scenario, taking an ancient costume fighting game as an example, if the virtual scene in which the target virtual character is situated has an ancient costume fantasy style, then the terminal will adjust the art style of the target virtual character's facial features to match the ancient costume fantasy style. As another example, in a picture-based scenario, taking a cartoon character as an example, if the virtual scene in which the target virtual character is situated has an anime cute style, then the terminal will adjust the art style of the target virtual character's facial features to match the anime cute style. This application does not impose any limitations on this aspect.
[0190] It should be noted that step 308 above is an optional implementation method provided in the embodiments of this application. After step 308, the terminal adjusts the style of the facial area of the target virtual image, making the target virtual image more consistent with the style of the virtual scene and improving the aesthetics of the target virtual image.
[0191] Of course, the different implementation methods described above can be combined to form different implementation schemes, and this application does not limit this.
[0192] In this embodiment, when generating a virtual image that resembles a target face based on a user-selected virtual character, an initial virtual image with a high degree of matching with the target face is first generated based on the selected target virtual character and the target face image. Then, through fine-grained attribute information of the target face, the facial area of the initial virtual image is adjusted in detail to ensure the matching degree between the final target virtual image and the target face. This method of adjusting the initial virtual image based on fine-grained attribute information greatly increases the similarity between the target virtual image and the target face, making it easy to distinguish between multiple virtual images generated from different face images and enriching the display effect of the virtual image.
[0193] The following is combined with Figure 10 Taking the virtual avatar generation method provided in this application as an example applied to a game scene, the method will be illustrated below. (Reference) Figure 10 , Figure 10 This is a flowchart of another virtual avatar generation method provided according to an embodiment of this application, such as... Figure 10 As shown, this embodiment of the application uses a terminal applied in a game scenario as an example for illustration. The method includes the following steps:
[0194] 1001. In response to receiving a virtual avatar creation instruction, a first page is displayed, which includes an image selection window or a viewfinder window.
[0195] The image selection window is used to select the target face image stored locally on the terminal, and the viewfinder window is used to capture the target face image through the terminal's camera.
[0196] Optionally, step 1001 includes steps 1001-1 to 1001-4:
[0197] 1001-1. In response to receiving a virtual avatar creation instruction, a game character selection page is displayed. The game character selection page includes at least one game character for the user to choose from. The game character selection page also includes a first option for customizing the target game character in detail.
[0198] Indicatively, for reference Figure 11 , Figure 11 This is a schematic diagram of a game character selection page provided in an embodiment of this application. For example... Figure 11 As shown, in response to the virtual avatar creation command, the terminal displays a game character selection page. This page includes at least one game character 10 for the user to choose from; for example, at least one game character 10 includes male, female, and teenage girls. The game character selection page also includes a first option 11, which is the "detail customization" option.
[0199] 1001-2. In response to the triggering operation of the first option, a game character customization page is displayed, which includes a second option for customizing the game character based on the target facial image.
[0200] Indicatively, for reference Figure 12 , Figure 12 This is a schematic diagram of a game character customization page provided in an embodiment of this application. For example... Figure 12 As shown, the terminal responds to... Figure 11 The first option 11 is triggered to display the game character customization page, which includes the second option 12, also known as the "personalization" option.
[0201] 1001-3. In response to the triggering operation of the second option, a prompt page is displayed to prompt the user on how to input the target face image. The prompt page includes a third option for obtaining the target face image.
[0202] Indicatively, for reference Figure 13 , Figure 13 This is a schematic diagram of a prompt page provided in an embodiment of this application. For example... Figure 13 As shown, the terminal responds to the... Figure 12 The second option 12 triggers a prompt page. This prompt page includes the third option 13, which is the "Upload Photo" option.
[0203] 1001-4. In response to the triggering of the third option, a first page is displayed, which includes an image selection window or a viewfinder window.
[0204] Indicatively, for reference Figure 14 , Figure 14 This is a schematic diagram of a first page provided in an embodiment of this application. For example... Figure 14 As shown, taking the viewfinder window on the first page as an example, the terminal responds to this... Figure 13 The third option 13 is triggered to display the first page, which includes the viewfinder window 14.
[0205] It should be noted that the above Figures 11 to 14 The page layout shown is for illustrative purposes only. In actual applications, the various pages and options can be configured according to actual needs, and this application does not impose any limitations on this.
[0206] 1002. Obtain the target face image obtained through the first page.
[0207] 1003. After generating the target virtual avatar, a second page is displayed, which shows the target virtual avatar.
[0208] In this process, after performing step 1002, the terminal acquires the target face image and then uses the above-mentioned... Figure 3 The virtual avatar generation method shown generates a target virtual avatar and displays a second page. This second page includes M adjustment options, where M is a positive integer, used to adjust the facial display of the target virtual avatar.
[0209] Indicatively, for reference Figure 15 , Figure 15 This is a schematic diagram of a second page provided in an embodiment of this application. For example... Figure 15 As shown, the second page includes M adjustment options 15. Taking the adjustment of the target virtual avatar's face display as an example, the second page includes multiple adjustment parts, and the user can select the parts they want to adjust. Each adjustment part corresponds to M adjustment options. For example, if the user selects to adjust the face shape, adjustment option 15 includes adjustment options such as "forehead," "cheeks," and "chin." This application does not limit this.
[0210] 1004. In response to a first triggering operation on any adjustment option on the second page, display at least one target action item corresponding to the adjustment option.
[0211] The first triggering operation can be a click operation, and this application does not limit it.
[0212] For illustrative purposes, please continue to refer to Figure 15 ,like Figure 15 As shown, in response to a first trigger operation on any adjustment option on the second page, the terminal displays at least one target operation item 16 corresponding to the adjustment option, namely, the "left / right", "up / down", and "forward / backward" operation items.
[0213] 1005. In response to a second trigger operation on any target operation item, obtain the adjustment parameters corresponding to the second trigger operation; adjust the facial display of the target virtual image according to the adjustment method indicated by the target operation item and the adjustment parameters.
[0214] The second triggering operation can be a click operation on the target operation item or a drag operation on the target operation item; this application does not limit this.
[0215] For illustrative purposes, please continue to refer to Figure 15 ,like Figure 15As shown, taking the adjustment option 15 as the "cheekbone" adjustment option, the target operation item 16 as the "left and right" operation item, the adjustment parameter as 10 pixels, and the adjustment direction as left as an example, the terminal responds to the second trigger operation of the target operation item 16, obtains the adjustment parameter 10 corresponding to the second trigger operation, and moves the cheekbone of the target object to the left by 10 pixels according to the adjustment direction indicated by the target operation item.
[0216] The first point to note is that the form of the above-mentioned adjustment parameters is only illustrative. In some embodiments, the adjustment parameters may also be preset adjustment ratios, etc. This application does not limit them.
[0217] The second point to note is that, in this embodiment, different virtual character types are not affected by the gender of the target face. For example, if the target face is male, but the user selects a female target virtual character, the terminal can still generate the corresponding target virtual image based on the various features of the target face. This greatly enriches the display effects of the virtual image.
[0218] It should be noted that, thirdly, the aforementioned second page also includes several other functional items to allow users to personalize the generated target virtual avatar. For example, the makeup function is used to adjust the makeup display of the target virtual avatar, the hair styling function is used to adjust the hairstyle display, and so on. Furthermore, the second page also includes a "personalization customization" option, which can be used to re-upload the target facial image. In practical applications, the functional items on the second page can be set according to needs, and this application does not limit this.
[0219] The virtual avatar generation method provided in this application increases the similarity between the target virtual avatar and the target face, making it possible to distinguish between multiple target virtual avatars generated from different target face images, thus enriching the display effect of the target virtual avatar.
[0220] The following is combined with Figure 16 Taking the virtual avatar generation method provided in this application as an example applied to a game scene, the method will be illustrated below. (Reference) Figure 16 , Figure 16 This is a flowchart of another virtual image generation method provided in the embodiments of this application, such as... Figure 16 As shown, the method includes steps 1601 to 1604:
[0221] 1601. Obtain the target face image.
[0222] 1602. Generate an initial virtual image based on the three-dimensional face parameters of the target face image.
[0223] The process involves the terminal extracting facial features, analyzing basic attributes, and reconstructing 3D from the target face image. Based on this, a high-dimensional feature representation matching the target face is extracted, which is the intermediate feature in steps 3021-4 above. Then, the high-dimensional feature representation is weighted and fused to obtain the 3D facial features of the target face. Further, this 3D facial feature is compared with a target comparison library to obtain the 3D facial parameters of the target face. Finally, these 3D facial parameters are loaded into the game engine to render the initial virtual avatar.
[0224] 1603. Perform fine and dense registration on the target face image.
[0225] The terminal uses a facial landmark recognition model to perform facial landmark recognition on the target face image according to the number of first landmarks. This is also known as face registration, which includes recognizing the facial features of the target face. After stabilizing and refining the position of the facial landmarks, the terminal finally outputs the dense recognition result of the target face.
[0226] 1604. Perform local detail optimization on the initial virtual image to generate the target virtual image.
[0227] The terminal adjusts the facial proportions of the initial virtual avatar based on facial proportion parameters, then obtains fine-grained facial attribute information based on dense recognition results, and further adjusts the details of the facial display of the initial virtual avatar to finally obtain the target virtual avatar.
[0228] Indicatively, for reference Figure 17 , Figure 17 This is a flowchart illustrating a local detail optimization provided in an embodiment of this application. For example... Figure 17 As shown, the initial virtual image obtained after step 1602 has a defect of distorted facial features. To address this defect, the proportions of the facial features of the initial virtual image are aligned and adjusted. Furthermore, 1000 dense recognition results are introduced to accurately depict the facial details of the target face. Then, fine-grained attribute information of the target face is introduced to make detailed adjustments to the facial display of the initial virtual image, ultimately obtaining the target virtual image.
[0229] In this embodiment, when generating a virtual image that resembles a target face based on a user-selected virtual character, an initial virtual image with a high degree of matching with the target face is first generated based on the selected target virtual character and the target face image. Then, through fine-grained attribute information of the target face, the facial area of the initial virtual image is adjusted in detail to ensure the matching degree between the final target virtual image and the target face. This method of adjusting the initial virtual image based on fine-grained attribute information greatly increases the similarity between the target virtual image and the target face, making it easy to distinguish between multiple virtual images generated from different face images and enriching the display effect of the virtual image.
[0230] Figure 18 This is a schematic diagram of a virtual avatar generation device according to an embodiment of this application. This virtual avatar generation device is used to perform the steps of the above-described virtual avatar generation method. (See attached diagram.) Figure 18 The virtual avatar generation device includes: a first acquisition module 1801, a first generation module 1802, a second acquisition module 1803, and a second generation module 1804.
[0231] The first acquisition module 1801 is used to acquire a target face image, which includes a target face;
[0232] The first generation module 1802 is used to generate an initial virtual image based on the target face image and the selected target virtual character.
[0233] The second acquisition module 1803 is used to acquire fine-grained attribute information of the target face, which is used to identify at least one of the facial features, hairstyle, or accessory categories of the target face.
[0234] The second generation module 1804 is used to make detailed adjustments to the facial region of the initial virtual image based on the fine-grained attribute information to generate a target virtual image; wherein, the second matching degree is greater than the first matching degree, the first matching degree refers to the matching degree between the facial region of the initial virtual image and the target face, and the second matching degree refers to the matching degree between the facial region of the target virtual image and the target face.
[0235] In one alternative implementation, the second acquisition module 1803 includes:
[0236] The recognition unit is used to perform facial key point recognition on the target face image according to the number of first key points, and obtain the dense recognition result of the target face;
[0237] The analysis unit is used to perform fine-grained attribute analysis on the target face image based on the dense recognition result, and obtain the fine-grained attribute information of the target face.
[0238] In one alternative implementation, the analysis unit is used for:
[0239] Based on the dense recognition result, the target face image is locally cropped to obtain N sub-regions. The N sub-regions include at least one of the facial features, hairstyle, or accessories of the target face, where N is a positive integer.
[0240] The N sub-regions are input into a fine-grained attribute classification model, which is a neural network model based on multi-task classification.
[0241] The output of this fine-grained attribute classification model is used as the fine-grained attribute information of the target face.
[0242] In one alternative implementation, the first generation module 1801 includes:
[0243] The acquisition unit is used to acquire the three-dimensional face parameters of the target face based on the target face image; wherein the three-dimensional face parameters include contour shape parameters and facial feature shape parameters that match the target face;
[0244] The generation unit is used to determine the facial material that matches the target face from the facial material library of the target virtual character, and to render the three-dimensional face parameters based on the determined facial material to generate the initial virtual image.
[0245] In one alternative implementation, the acquisition unit is used for:
[0246] Obtain the three-dimensional facial features of the target face from the target face image;
[0247] The 3D facial features are compared with the target comparison library to obtain the 3D facial parameters of the target face. The target comparison library is used to provide the correspondence between the 3D facial features and the 3D facial parameters.
[0248] In an alternative implementation, the acquisition unit is further configured to:
[0249] Feature extraction is performed on the target face image to obtain the two-dimensional facial features of the target face;
[0250] Basic attribute analysis is performed on the target face image to obtain the basic attribute information of the target face. The classification accuracy of the basic attribute information is lower than that of the fine-grained attribute information.
[0251] Based on the target face image, the target face is reconstructed in three dimensions to obtain a three-dimensional reconstruction model of the target face;
[0252] Based on the two-dimensional facial features, the basic attribute information, and the three-dimensional reconstruction model, the intermediate features of the target face are obtained, and the feature dimension of the intermediate features is greater than a preset threshold.
[0253] The intermediate features are weighted and fused to obtain the three-dimensional facial features of the target face.
[0254] In one alternative implementation, the device further includes:
[0255] The recognition module is used to perform facial key point recognition on the target face image according to the number of second key points, and obtain the facial feature proportion parameters of the target face. The number of second key points is less than the number of first key points.
[0256] The facial features adjustment module is used to adjust the proportions of the facial features of the initial virtual avatar based on the facial feature proportion parameters.
[0257] In one alternative implementation, the device further includes:
[0258] The art style adjustment module is used to adjust the art style of the facial area of the target virtual character according to the style of the virtual scene in which the target virtual character is located.
[0259] In one alternative implementation, the device further includes:
[0260] The first display module is used to display a first page in response to receiving a virtual avatar creation instruction. The first page includes an image selection window or a viewfinder window.
[0261] The third acquisition module is used to acquire the target face image obtained through the first page;
[0262] The second display module is used to display a second page after the target virtual image is generated, on which the target virtual image is displayed.
[0263] In one alternative implementation, the second page includes M adjustment options, where M is a positive integer, for adjusting the facial display of the target virtual avatar. The device also includes:
[0264] The third display module is used to display at least one target operation item corresponding to any of the adjustment options on the second page in response to a first trigger operation of any of the adjustment options on the second page;
[0265] The fourth acquisition module is used to acquire the adjustment parameters corresponding to the second trigger operation in response to any second trigger operation of the target operation item; and adjust the facial display of the target virtual image according to the adjustment method indicated by the target operation item and the adjustment parameters.
[0266] In the embodiments of the present application, when generating a virtual image similar to the target human face according to the selected virtual character, first, based on the selected target virtual character and the target human face image, an initial virtual image with a relatively high matching degree with the target human face is generated. Then, through the fine-grained attribute information of the target human face, the facial area of the initial virtual image is adjusted in detail to ensure the matching degree between the finally obtained target virtual image and the target human face. This method of adjusting the initial virtual image based on the fine-grained attribute information greatly increases the similarity between the target virtual image and the target human face, enabling obvious distinction between multiple virtual images generated according to different human face images and enriching the display effect of the virtual images.
[0267] It should be noted that: when the virtual image generation device provided in the above embodiment generates a virtual image, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the virtual image generation device provided in the above embodiment and the embodiment of the virtual image generation method belong to the same concept. For the specific implementation process, please refer to the method embodiment, which will not be elaborated here.
[0268] In an exemplary embodiment, a computer device is further provided. Taking the computer device as an example of a terminal, Figure 19 FIG. 1900 shows a schematic structural diagram of a terminal 1900 provided in an exemplary embodiment of the present application. The terminal 1900 may be: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a notebook computer, or a desktop computer. The terminal 1900 may also be referred to by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.
[0269] Generally, the terminal 1900 includes: a processor 1901 and a memory 1902.
[0270] Processor 1901 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1901 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1901 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1901 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 1901 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0271] The memory 1902 may include one or more computer-readable storage media, which may be non-transitory. The memory 1902 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1902 are used to store at least one program code, which is executed by the processor 1901 to implement the virtual avatar generation method provided in the method embodiments of this application.
[0272] In some embodiments, the terminal 1900 may also optionally include a peripheral device interface 1903 and at least one peripheral device. The processor 1901, memory 1902, and peripheral device interface 1903 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1903 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 1904, a display screen 1905, a camera assembly 1906, an audio circuit 1907, and a power supply 1909.
[0273] Peripheral interface 1903 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 1901 and memory 1902. In some embodiments, processor 1901, memory 1902 and peripheral interface 1903 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1901, memory 1902 and peripheral interface 1903 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0274] The radio frequency (RF) circuit 1904 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1904 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1904 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 1904 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 1904 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: metropolitan area networks (MANs), various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks (WLANs), and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1904 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.
[0275] Display screen 1905 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 1905 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1901 for processing. In this case, display screen 1905 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1905, disposed on the front panel of terminal 1900; in other embodiments, there may be at least two display screens, disposed on different surfaces of terminal 1900 or in a folded design; in still other embodiments, display screen 1905 may be a flexible display screen, disposed on a curved or folded surface of terminal 1900. Furthermore, display screen 1905 may also be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. The display screen 1905 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0276] The camera assembly 1906 is used to acquire images or videos. Optionally, the camera assembly 1906 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 1906 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.
[0277] The audio circuit 1907 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting them into electrical signals that are input to the processor 1901 for processing, or to the radio frequency circuit 1904 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each positioned at a different location on the terminal 1900. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 1901 or the radio frequency circuit 1904 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 1907 may also include a headphone jack.
[0278] The power supply 1909 is used to power the various components in the terminal 1900. The power supply 1909 can be AC power, DC power, a disposable battery, or a rechargeable battery. When the power supply 1909 includes a rechargeable battery, the rechargeable battery can support wired or wireless charging. The rechargeable battery can also be used to support fast charging technology.
[0279] In some embodiments, the terminal 1900 further includes one or more sensors 1910. The one or more sensors 1910 include, but are not limited to: an acceleration sensor 1911, a gyroscope sensor 1912, a pressure sensor 1913, an optical sensor 1915, and a proximity sensor 1916.
[0280] Accelerometer 1911 can detect the magnitude of acceleration along the three axes of a coordinate system established by terminal 1900. For example, accelerometer 1911 can be used to detect the components of gravitational acceleration along the three axes. Processor 1901 can control display screen 1905 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 1911. Accelerometer 1911 can also be used for games or for acquiring user motion data.
[0281] The gyroscope sensor 1912 can detect the orientation and rotation angle of the terminal 1900. The gyroscope sensor 1912, in conjunction with the accelerometer sensor 1911, can collect 3D motion data from the user on the terminal 1900. Based on the data collected by the gyroscope sensor 1912, the processor 1901 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.
[0282] The pressure sensor 1913 can be disposed on the side bezel of the terminal 1900 and / or on the lower layer of the display screen 1905. When the pressure sensor 1913 is disposed on the side bezel of the terminal 1900, it can detect the user's grip signal on the terminal 1900, and the processor 1901 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 1913. When the pressure sensor 1913 is disposed on the lower layer of the display screen 1905, the processor 1901 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 1905. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0283] An optical sensor 1915 is used to collect ambient light intensity. In one embodiment, the processor 1901 can control the display brightness of the display screen 1905 based on the ambient light intensity collected by the optical sensor 1915. Specifically, when the ambient light intensity is high, the display brightness of the display screen 1905 is increased; when the ambient light intensity is low, the display brightness of the display screen 1905 is decreased. In another embodiment, the processor 1901 can also dynamically adjust the shooting parameters of the camera assembly 1906 based on the ambient light intensity collected by the optical sensor 1915.
[0284] The proximity sensor 1916, also known as a distance sensor, is typically located on the front panel of the terminal 1900. The proximity sensor 1916 is used to detect the distance between the user and the front of the terminal 1900. In one embodiment, when the proximity sensor 1916 detects that the distance between the user and the front of the terminal 1900 is gradually decreasing, the processor 1901 controls the display screen 1905 to switch from a screen-on state to a screen-off state; when the proximity sensor 1916 detects that the distance between the user and the front of the terminal 1900 is gradually increasing, the processor 1901 controls the display screen 1905 to switch from a screen-off state to a screen-on state.
[0285] Those skilled in the art will understand that Figure 19 The structure shown does not constitute a limitation on terminal 1900 and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0286] This application also provides a computer-readable storage medium applied to a computer device. The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to implement the operations performed by the computer device in the virtual image generation method of the above embodiments.
[0287] This application also provides a computer program product or computer program, which includes computer program code stored in a computer-readable storage medium. A processor of a computer device reads the computer program code from the computer-readable storage medium and executes the computer program code, causing the computer device to perform the virtual avatar generation method provided in the various optional implementations described above.
[0288] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0289] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for generating virtual avatars, characterized in that, The method includes: Acquire a target face image, wherein the target face image includes the target face; Based on the target face image and the selected target virtual character, an initial virtual image is generated; Based on the first number of key points, the target face image is subjected to multiple facial key point recognitions to obtain multiple dense recognition results. A final dense recognition result for the target face is determined based on these results. Based on the final dense recognition result, fine-grained attribute analysis is performed on the target face image to obtain fine-grained attribute information for the target face. This fine-grained attribute information is used to identify at least one of the following categories: facial features, hairstyle, or accessories. Furthermore, based on the second number of key points, facial key point recognition is performed on the target face image. Line drawing is performed based on the distribution of facial features to obtain facial feature proportion parameters for the target face. The second number of key points is less than the first number of key points. Based on these facial feature proportion parameters, the facial feature proportions of the initial virtual image are adjusted in both the horizontal and vertical directions. Based on the fine-grained attribute information, the facial region of the initial virtual image is adjusted in detail to generate the target virtual image; the matching degree between the facial region of the target virtual image and the target face is greater than the matching degree between the facial region of the initial virtual image and the target face. The display includes the target virtual image and multiple adjustable parts, each adjustable part corresponds to multiple adjustment options, and each adjustment option corresponds to at least one target operation item; Based on the selection of the adjustment area, adjustment options, and target operation items, the adjustment method and adjustment parameters are determined; based on the adjustment method and adjustment parameters, the facial display of the target virtual image is adjusted.
2. The method according to claim 1, characterized in that, The step of performing fine-grained attribute analysis on the target face image based on the final dense recognition result includes: Based on the final dense recognition result, the target face image is locally cropped to obtain N sub-regions. The N sub-regions include at least one of the facial features, hairstyle, or accessories of the target face, where N is a positive integer. The N sub-regions are input into a fine-grained attribute classification model, which is a neural network model based on multi-task classification. The output of the fine-grained attribute classification model is used as the fine-grained attribute information of the target face.
3. The method according to claim 1, characterized in that, The step of generating an initial virtual avatar based on the target face image and the selected target virtual character includes: Based on the target face image, obtain the three-dimensional face parameters of the target face; wherein, the three-dimensional face parameters include contour shape parameters and facial feature shape parameters that match the target face; From the facial material library of the target virtual character, facial materials matching the target face are determined, and the three-dimensional facial parameters are rendered based on the determined facial materials to generate the initial virtual image.
4. The method according to claim 3, characterized in that, The step of obtaining the three-dimensional facial parameters of the target face based on the target face image includes: Obtain the three-dimensional facial features of the target face from the target face image; The three-dimensional facial features are compared with the target comparison library to obtain the three-dimensional facial parameters of the target face. The target comparison library is used to provide the correspondence between the three-dimensional facial features and the three-dimensional facial parameters.
5. The method according to claim 4, characterized in that, The step of obtaining the three-dimensional facial features of the target face in the target face image includes: Feature extraction is performed on the target face image to obtain the two-dimensional facial features of the target face; Basic attribute analysis is performed on the target face image to obtain basic attribute information of the target face. The classification accuracy of the basic attribute information is lower than that of the fine-grained attribute information. Based on the target face image, a three-dimensional reconstruction of the target face is performed to obtain a three-dimensional reconstruction model of the target face; Based on the two-dimensional facial features, the basic attribute information, and the three-dimensional reconstruction model, intermediate features of the target face are obtained, and the feature dimension of the intermediate features is greater than a preset threshold. The intermediate features are weighted and fused to obtain the three-dimensional facial features of the target face.
6. The method according to claim 1, characterized in that, The method further includes: Adjust the art style of the facial area of the target virtual character according to the style of the virtual scene in which the target virtual character is located.
7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: In response to receiving a virtual avatar creation instruction, a first page is displayed, which includes an image selection window or a viewfinder window; Obtain the target face image obtained through the first page; After the target virtual image is generated, a second page is displayed, on which the target virtual image is shown.
8. The method according to claim 7, characterized in that, The second page includes M adjustment options, where M is a positive integer. These adjustment options are used to adjust the facial display of the target virtual avatar. The method further includes: In response to a first trigger operation on any of the adjustment options on the second page, at least one target operation item corresponding to the adjustment option is displayed; In response to a second trigger operation on any of the target operation items, the adjustment parameters corresponding to the second trigger operation are obtained; the facial display of the target virtual image is adjusted according to the adjustment method indicated by the target operation item and the adjustment parameters.
9. A virtual avatar generation device, characterized in that, The device includes: The first acquisition module is used to acquire a target face image, wherein the target face image includes a target face; The first generation module is used to generate an initial virtual image based on the target face image and the selected target virtual character; The second acquisition module includes an identification unit and an analysis unit; The recognition unit is used to perform multiple facial key point recognitions on the target face image according to the first number of key points, to obtain multiple dense recognition results, and to determine the final dense recognition result of the target face based on the multiple dense recognition results. The analysis unit is configured to perform fine-grained attribute analysis on the target face image based on the final dense recognition result, to obtain fine-grained attribute information of the target face, wherein the fine-grained attribute information is used to identify at least one of the following categories of the target face: facial features, hairstyle, or accessories; and, The recognition module is used to perform facial key point recognition on the target face image according to the number of second key points, and to perform line drawing processing based on the distribution of facial features of the target face to obtain the facial feature ratio parameters of the target face, wherein the number of second key points is less than the number of first key points. The facial features adjustment module is used to adjust the proportions of the facial features of the initial virtual image in both the horizontal and vertical directions based on the facial features proportion parameters. The second generation module is used to make detailed adjustments to the facial region of the initial virtual image based on the fine-grained attribute information to generate a target virtual image; the matching degree between the facial region of the target virtual image and the target face is greater than the matching degree between the facial region of the initial virtual image and the target face. The second display module is used to display a second page including the target virtual image and multiple adjustment parts, each adjustment part corresponding to multiple adjustment options, and each adjustment option corresponding to at least one target operation item; The module is used to perform the following steps: determining the adjustment method and adjustment parameters based on the selection of the adjustment part, adjustment option, and target operation item; and adjusting the facial display of the target virtual image based on the adjustment method and adjustment parameters.
10. The apparatus according to claim 9, characterized in that, The analysis unit is used for: Based on the final dense recognition result, the target face image is locally cropped to obtain N sub-regions. The N sub-regions include at least one of the facial features, hairstyle, or accessories of the target face, where N is a positive integer. The N sub-regions are input into a fine-grained attribute classification model, which is a neural network model based on multi-task classification. The output of the fine-grained attribute classification model is used as the fine-grained attribute information of the target face.
11. The apparatus according to claim 10, characterized in that, The first generation module includes: The acquisition unit is used to acquire three-dimensional face parameters of the target face based on the target face image; wherein the three-dimensional face parameters include contour shape parameters and facial feature shape parameters that match the target face; The generation unit is used to determine facial materials that match the target face from the facial material library of the target virtual character, and to perform rendering processing on the three-dimensional face parameters based on the determined facial materials to generate the initial virtual image.
12. The apparatus according to claim 11, characterized in that, The acquisition unit is used for: Obtain the three-dimensional facial features of the target face from the target face image; The three-dimensional facial features are compared with the target comparison library to obtain the three-dimensional facial parameters of the target face. The target comparison library is used to provide the correspondence between the three-dimensional facial features and the three-dimensional facial parameters.
13. The apparatus according to claim 12, characterized in that, The acquisition unit is further configured to: Feature extraction is performed on the target face image to obtain the two-dimensional facial features of the target face; Basic attribute analysis is performed on the target face image to obtain basic attribute information of the target face. The classification accuracy of the basic attribute information is lower than that of the fine-grained attribute information. Based on the target face image, a three-dimensional reconstruction of the target face is performed to obtain a three-dimensional reconstruction model of the target face; Based on the two-dimensional facial features, the basic attribute information, and the three-dimensional reconstruction model, intermediate features of the target face are obtained, and the feature dimension of the intermediate features is greater than a preset threshold. The intermediate features are weighted and fused to obtain the three-dimensional facial features of the target face.
14. The apparatus according to claim 9, characterized in that, The device further includes: The art style adjustment module is used to adjust the art style of the facial area of the target virtual character according to the style of the virtual scene in which the target virtual character is located.
15. The apparatus according to any one of claims 9 to 14, characterized in that, The device further includes: The first display module is used to display a first page in response to receiving a virtual avatar creation instruction. The first page includes an image selection window or a viewfinder window. The third acquisition module is used to acquire the target face image obtained through the first page; The second display module is used to display a second page after the target virtual image is generated, and the target virtual image is displayed on the second page.
16. The apparatus according to claim 15, characterized in that, The second page includes M adjustment options, where M is a positive integer. These adjustment options are used to adjust the facial display of the target virtual avatar. The device also includes: The third display module is configured to display at least one target operation item corresponding to any of the adjustment options on the second page in response to a first trigger operation. The fourth acquisition module is used to acquire the adjustment parameters corresponding to the second trigger operation in response to the second trigger operation of any of the target operation items; and to adjust the facial display of the target virtual image according to the adjustment method indicated by the target operation item and the adjustment parameters.
17. A computer device, characterized in that, The computer device includes a processor and a memory, the memory being used to store at least one computer program, the at least one computer program being loaded by the processor and executed as the virtual image generation method as described in any one of claims 1 to 8.
18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to implement the virtual image generation method as described in any one of claims 1 to 8.
19. A computer program product comprising computer program code stored in a computer-readable storage medium, wherein a processor of a computer device reads the computer program code from the computer-readable storage medium, and the processor executes the computer program code to cause the computer device to perform a virtual avatar generation method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Virtual image generation method, device and equipment, and readable storage medium
CN108510437A
Virtual image generation method and device, electronic equipment and storage medium
CN110782515A
Image data processing method and device and computer readable storage medium
CN110784728A
Three-dimensional face reconstruction network training and virtual face image generation method and device
CN111354079A