Cartoon Avatar Generation Method, Device, Equipment and Medium Based on Facial Images
The method addresses the inefficiency of manual cartoon avatar creation by using keypoint detection and GAN-based rendering to generate personalized avatars from face images, enhancing personalization and reducing manual effort.
Patent Information
- Application Number
- CN202210057898.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-19
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-01-19
AI Technical Summary
In the prior art, due to the limited energy of designers, it is difficult to draw suitable cartoon avatars for everyone, resulting in large engineering volume and difficult to achieve personalized cartoon avatar generation.
By detecting key points and extracting key features on the face image, a cartoon avatar is generated using a generative adversarial network model, and drawing it in combination with the conditions entered by the user to generate a cartoon avatar corresponding to the face image.
The workload of manual mapping is reduced, the generated cartoon avatars meet user needs, solve the problem of difficulty in cross-modal generation, and improve the stability and controllability of the generation effect.
Smart Images

Figure CN114463472B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technologies, and particularly to a method, apparatus, device, and medium for generating cartoon avatars based on face images. Background Art
[0002] With the continuous development and progress of computer technologies, various fashionable elements derived from computer technologies have become increasingly rich. Cartoon avatar drawings are deeply loved by the majority of users for their vivid and lively features and are widely used in fields such as social networking, gaming, film and television, and entertainment. The image processing application that can create a personalized cartoon avatar based on one's own image has attracted many users. Currently, when obtaining corresponding cartoon avatars according to different face shapes of users, the commonly used processing method is mainly for designers to draw as many different types of face contours as possible to increase the options for users. Due to the limited energy of designers and the diverse face contours among people, it is extremely difficult and time-consuming to directly draw face contours suitable for each person. Therefore, a method for generating cartoon avatars based on face images is needed. Summary of the Invention
[0003] In view of this, embodiments of the present disclosure provide a method, apparatus, device, and medium for generating cartoon avatars based on face images to solve the problem of how to generate cartoon avatars based on face images in the prior art.
[0004] In a first aspect of the embodiments of the present disclosure, a method for generating a cartoon avatar based on a face image is provided, including: performing key point detection on the obtained target face image; in response to determining that the key point detection is successful, extracting key features from the target face image to obtain a key feature image; based on conditions input by a user, drawing the key feature image to obtain a drawn key feature image; and generating a cartoon avatar corresponding to the target face image based on the drawn key feature image.
[0005] In a second aspect of the embodiments of the present disclosure, a device for generating a cartoon avatar based on a face image is provided. The device includes: a key point detection unit configured to perform key point detection on the obtained target face image; a key feature extraction unit configured to, in response to determining that the key point detection is successful, extract key features from the target face image to obtain a key feature image; an image drawing unit configured to draw the key feature image based on conditions input by a user to obtain a drawn key feature image; and a cartoon avatar generation unit configured to generate a cartoon avatar corresponding to the target face image based on the drawn key feature image.
[0006] In a third aspect of the embodiments of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.
[0007] In a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0008] One embodiment among the above various embodiments of the present disclosure has the following beneficial effects: First, perform key point detection on the obtained target face image; then, when it is determined that the key point detection is successful, extract key features from the above target face image to obtain a key feature image; then, based on the conditions input by the user, draw the above key feature image to obtain a drawn key feature image; finally, based on the above drawn key feature image, generate a cartoon avatar corresponding to the above target face image. The method provided by the present disclosure can generate a cartoon avatar corresponding to the target face image according to the detection, key feature extraction, and drawing of the target face image. Drawing the key feature image according to the conditions input by the user makes the generated cartoon avatar more meet the user's needs. The method provided by the present disclosure reduces the workload of manual drawing and solves the problem of difficult cross-modal generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more obvious. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the elements and elements are not necessarily drawn to scale.
[0010] Figure 1 is a schematic diagram of an application scenario of a method for generating a cartoon avatar based on a face image according to some embodiments of the present disclosure;
[0011] Figure 2 is a schematic flowchart of a method for generating a cartoon avatar based on a face image according to some embodiments of the present disclosure;
[0012] Figure 3 is a schematic structural diagram of a device for generating a cartoon avatar based on a face image according to some embodiments of the present disclosure;
[0013] Figure 4 is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0014] In the following description, specific details such as specific system architectures and technologies are presented for purposes of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art should understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from obscuring the description of the present disclosure.
[0015] A method, apparatus, electronic device, and medium for generating a cartoon avatar based on a face image according to an embodiment of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0016] Figure 1 It is a schematic diagram of an application scenario of a method for generating a cartoon avatar based on a face image according to some embodiments of the present disclosure.
[0017] In Figure 1 the application scenario, first, the computing device 101 can perform key point detection on the acquired target face image 102, as indicated by reference numeral 103 in the accompanying drawings. Then, in response to determining that the above key point detection is successful, the computing device 101 can perform key feature extraction on the above target face image 102, as indicated by reference numeral 104 in the accompanying drawings. After obtaining the key feature image 105, based on the condition 106 input by the user, the computing device 101 can draw the above key feature image 105 to obtain the drawn key feature image 107. Finally, based on the above drawn key feature image 107, the computing device 101 can generate a cartoon avatar 108 corresponding to the above target face image 102.
[0018] It should be noted that the above computing device 101 can be hardware or software. When the computing device is hardware, it can be implemented as a distributed cluster composed of multiple servers or terminal devices, or as a single server or a single terminal device. When the computing device is embodied as software, it can be installed in the above-listed hardware devices. It can be implemented as, for example, multiple software or software modules for providing distributed services, or as a single software or software module. No specific limitation is made here.
[0019] It should be understood that Figure 1 the number of computing devices in
[0020] Figure 2 is only illustrative. According to the implementation requirements, there can be any number of computing devices. Figure 2 The method for generating a cartoon avatar based on a face image in Figure 1 can be executed by the Figure 2As shown, the method for generating a cartoon avatar based on a face image includes the following steps:
[0021] Step S201: Detect key points of the acquired target face image.
[0022] In some embodiments, the execution entity of the method for generating a cartoon avatar based on a face image (such as Figure 1 the computing device 101 shown) can detect key points of the acquired target face image. Here, the result of the key point detection is whether valid key points are detected. Specifically, the target face image can be a photo with a face uploaded by the user. As an example, the above-mentioned execution entity can be connected to an electronic device for inputting a face image through a wired connection method or a wireless connection method, and determine the face image input by the user as the above-mentioned target face image.
[0023] As an example, the above-mentioned execution entity can use the ASM (Active Shape Model) algorithm to detect key points of the above-mentioned target face image. The ASM algorithm is a face key point detection algorithm extracted based on the feature point distribution model (Point Distribution Model, PDM). The active shape model abstracts the target object through the shape model.
[0024] As another example, the above-mentioned execution entity can use the AAM (Active Appearance Model) algorithm to detect key points of the above-mentioned target face image. The AAM algorithm is an image segmentation algorithm based on the active appearance model, mainly an improvement on the ASM algorithm, which not only uses shape constraints but also adds texture features of the entire face area.
[0025] As another example, the above-mentioned execution entity can use the CLM (Constrained local model) algorithm to detect key points of the above-mentioned target face image. The CLM algorithm is a Constrained Local Model, which completes face point detection by initializing the position of the average face and then allowing the feature points on each average face to search and match in their neighborhood positions.
[0026] It should be noted that the above-mentioned wireless connection method can include but is not limited to 3G / 4G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other currently known or future-developed wireless connection methods.
[0027] Step S202: In response to determining that the key point detection is successful, extract key features from the above-mentioned target face image to obtain a key feature image.
[0028] In some embodiments, in response to determining that the above key point detection is successful, the above execution entity may perform key feature extraction on the above target face image to obtain a key feature image. Here, the key feature image at least includes: a face contour, the positions of facial features, and a hair contour. As an example, the above execution entity may use an edge extraction algorithm to perform key feature extraction on the above target face image to obtain a face contour mixed with a hair contour. The above execution entity may use a face key point detection algorithm to perform key feature extraction on the above target face image to obtain the positions of facial features. Here, the face key point detection algorithm may be an ASM algorithm, an AAM algorithm, or a CLM algorithm.
[0029] As an example, the above execution entity may use the Roberts operator to perform key feature extraction on the above target face image to obtain a face contour mixed with a hair contour. The Roberts operator is also known as the cross differential algorithm. It is a gradient algorithm based on cross differences and detects edge lines through local difference calculations. It is often used to process images with steep low noise. When the image edge is close to 45 degrees positive or negative, the algorithm has a better processing effect.
[0030] As another example, the above execution entity may use the Prewitt operator to perform key feature extraction on the above target face image to obtain a face contour mixed with a hair contour. The Prewitt operator is a differential operator for image edge detection, and its principle is to use the differences generated by pixel gray values in a specific area to achieve edge detection.
[0031] Step S203: Based on the conditions input by the user, draw the above key feature image to obtain a drawn key feature image.
[0032] In some embodiments, based on the conditions input by the user, the above execution entity may draw the above key feature image to obtain a drawn key feature image. Here, the above conditions at least include: skin color and hair color. The way for the user to input conditions may be to input RGB values as the skin color and / or hair color. The RGB color model is a color standard that obtains various colors through the changes of the three color channels of red (R), green (G), and blue (B) and their superposition with each other. RGB represents the colors of the three channels of red, green, and blue.
[0033] As an example, the conditions input by the user can be "Skin color: 250, 194, 110; Hair color: 125, 0, 0". Based on the above conditions, the above-mentioned execution entity can draw the color of the area enclosed by the face contour of the above-mentioned key feature image as "250, 194, 110", and the above-mentioned execution entity can draw the color of the area enclosed by the above-mentioned hair contour as "125, 0, 0" to obtain the drawn key feature image.
[0034] Step S204: Generate a cartoon avatar corresponding to the above-mentioned target face image based on the above-mentioned drawn key feature image.
[0035] In some optional implementation manners of some embodiments, the above-mentioned execution entity can input the above-mentioned drawn key feature image into a pre-trained cartoon avatar generation model to generate a cartoon avatar corresponding to the above-mentioned target face image. Among them, the above-mentioned cartoon avatar generation model is trained by using the sample face images in the training sample set as inputs and the sample cartoon avatars in the above-mentioned training samples as expected outputs. The above-mentioned cartoon avatar generation model is a generative adversarial network model. Generative Adversarial Networks (GAN) is a deep learning model and a neural network architecture of a generative model. A generative model refers to using a model to generate new cases based on existing samples. For example, generating a set of new photos that are similar to but slightly different from an existing photo set. GAN is a generative model trained by using two neural network models.
[0036] In some optional implementation manners of some embodiments, the above-mentioned training sample set is obtained through the following steps: screening the obtained face image set to obtain face images that meet the first preset condition as sample face images to obtain a sample face image set; screening the obtained cartoon avatar set to obtain cartoon avatars that meet the second preset condition as sample cartoon avatars to obtain a sample cartoon avatar set; combining the above-mentioned sample face image set and the above-mentioned sample cartoon avatar set to obtain the above-mentioned training sample set. Here, the above-mentioned obtained face images and the above-mentioned obtained cartoon avatars can be obtained by downloading a public data set or by web crawling. The above-mentioned first preset condition may include: being able to extract a face contour, being able to extract a hair contour, and being able to detect at least a preset number of face key points. The above-mentioned second preset condition may be that the similarity between the cartoon avatars in the above-mentioned cartoon avatar set exceeds a preset similarity threshold. Here, the above-mentioned similarity specifically represents the approximation degree of the styles of the cartoon avatars.
[0037] In some optional implementation manners of some embodiments, the training steps of the above-mentioned cartoon avatar generation model include:
[0038] First step, using the edge extraction algorithm and the facial key point detection algorithm, extract the key features from each sample facial image in the above sample facial image set to obtain the key features of the above sample facial image, so as to form the first key feature set. Here, the key features of the sample facial image at least include: facial contour, facial feature positions, and hair contour.
[0039] Second step, using the above edge extraction algorithm and the above facial key point detection algorithm, extract the key features from each cartoon avatar in the above sample cartoon avatar set to obtain the key features of the above cartoon avatar, so as to form the second key feature set. Here, the key features of the cartoon avatar at least include: facial contour, facial feature positions, and hair contour.
[0040] Third step, using the preset image feature extraction algorithm, extract the features from each cartoon avatar in the above sample cartoon avatar set to obtain the conditional noise features of the above sample cartoon avatar set, where the above conditional noise features include: skin color feature, hair color feature. As an example, the above preset image feature extraction algorithm can be the above edge extraction algorithm, can be the algorithm related to key point detection mentioned above, or can be the CLD (Color Layout Descriptor) algorithm. The full name is the color layout descriptor, which is an efficient local color feature description in the mpeg-7 multimedia content standard description, and has the advantages of low computational cost and fast matching calculation speed.
[0041] Fourth step, combine the above second key feature set and the above conditional noise feature set to obtain the third key feature set.
[0042] Fifth step, based on the above first key feature set and the above third key feature set, train the initial model to obtain the above cartoon avatar generation model.
[0043] The fifth step stated above includes the following sub-steps:
[0044] First sub-step, perform data augmentation processing on the second key features in the above third key feature set. Here, data augmentation, also called data amplification, mainly makes the limited data produce the value equivalent to more data without substantially increasing the data. As an example, since the facial image has a large amount of noise, some noise can be added to the key features of the cartoon avatar to make the two closer.
[0045] Second sub-step, in response to determining that the similarity between the second key features after the data augmentation processing in the above third key feature set and the key features in the above first key feature set is greater than the preset threshold, determine that the above data augmentation processing is completed to obtain the third key feature set after the data augmentation processing.
[0046] The third sub-step is to use the above-mentioned third key feature set and the above-mentioned sample cartoon avatar set to perform image translation training on the above-mentioned initial model to obtain the above-mentioned cartoon avatar generation model. Here, the pix2pix algorithm (Image-to-Image Translation) is adopted in the process of image translation training.
[0047] In some optional implementation manners of some embodiments, the above method further includes: transmitting the above-mentioned cartoon avatar to a target device with a display function, and controlling the target device to display the above-mentioned cartoon avatar.
[0048] One embodiment of each of the above embodiments of the present disclosure has the following beneficial effects: First, perform key point detection on the obtained target face image; then, when it is determined that the key point detection is successful, extract key features from the above-mentioned target face image to obtain a key feature image; then, based on the conditions input by the user, draw the above-mentioned key feature image to obtain a drawn key feature image; finally, based on the above-mentioned drawn key feature image, generate a cartoon avatar corresponding to the above-mentioned target face image. The method provided by the present disclosure can detect, extract key features, and draw a target face image to generate a cartoon avatar corresponding to the target face image, reducing the workload of manual drawing. According to the conditions input by the user, the drawing of the key feature image makes the generated cartoon avatar more meet the user's needs. The method provided by the present disclosure obtains common features of two different domains through an artificial mapping algorithm, solving the problem of difficult cross-modal generation. In addition, by using the extracted common key features (the training steps of the cartoon avatar generation model), a paired data set of key features and cartoon avatars is created, improving the cross-modal generation effect and stability of the cartoon avatar generation model (GAN model). Moreover, by extracting the common key features of the face image and the cartoon avatar (face contour, facial feature positions, hair contour), the similarity after avatar cartoonization is guaranteed to a certain extent, and through conditional noise features (skin color, hair color), partial attributes of the cartoon avatar are made controllable.
[0049] All the above optional technical solutions can be combined arbitrarily to form optional embodiments of the present application, which will not be elaborated here one by one.
[0050] The following is an embodiment of the device of the present disclosure, which can be used to execute the embodiment of the method of the present disclosure. For details not disclosed in the embodiment of the device of the present disclosure, please refer to the embodiment of the method of the present disclosure.
[0051] Figure 3 is a schematic diagram of a cartoon avatar generation device based on a face image provided by an embodiment of the present disclosure. As Figure 3As shown in the figure, the cartoon avatar generation device based on a face image includes: a key point detection unit 301, a key feature extraction unit 302, an image drawing unit 303, and a cartoon avatar generation unit 304. Among them, the key point detection unit 301 is configured to perform key point detection on the acquired target face image; the key feature extraction unit 302 is configured to, in response to determining that the above key point detection is successful, extract key features from the above target face image to obtain a key feature image; the image drawing unit 303 is configured to draw the above key feature image based on the conditions input by the user to obtain a drawn key feature image; the cartoon avatar generation unit 304 is configured to generate a cartoon avatar corresponding to the above target face image based on the above drawn key feature image.
[0052] In some optional implementation manners of some embodiments, the above key feature image at least includes: a face contour, the positions of facial features, and a hair contour; the above conditions at least include: skin color and hair color.
[0053] In some optional implementation manners of some embodiments, the cartoon avatar generation unit 304 of the cartoon avatar generation device based on a face image is further configured to: input the above drawn key feature image into a pre-trained cartoon avatar generation model to generate a cartoon avatar corresponding to the above target face image, where the above cartoon avatar generation model is obtained by using the sample face images in the training sample set as inputs and the sample cartoon avatars in the above training samples as expected outputs; among them, the above cartoon avatar generation model is a generative adversarial network model.
[0054] In some optional implementation manners of some embodiments, the above training sample set is obtained through the following steps: screening the acquired face image set to obtain face images that meet the first preset condition as sample face images to obtain a sample face image set; screening the acquired cartoon avatar set to obtain cartoon avatars that meet the second preset condition as sample cartoon avatars to obtain a sample cartoon avatar set; combining the above sample face image set and the above sample cartoon avatar set to obtain the above training sample set.
[0055] In some alternative implementation manners of some embodiments, the training steps of the above-mentioned cartoon avatar generation model include: using an edge extraction algorithm and a face key point detection algorithm to extract key features from each sample face image in the above-mentioned sample face image set, obtaining the key features of the above-mentioned sample face image to form a first key feature set; using the above-mentioned edge extraction algorithm and the above-mentioned face key point detection algorithm to extract key features from each cartoon avatar in the above-mentioned sample cartoon avatar set, obtaining the key features of the above-mentioned cartoon avatar to form a second key feature set; using a preset image feature extraction algorithm to extract features from each cartoon avatar in the above-mentioned sample cartoon avatar set, obtaining the conditional noise features of the above-mentioned cartoon avatar to form a conditional noise feature set of the above-mentioned sample cartoon avatar set, where the above-mentioned conditional noise features include: skin color features, hair color features; combining the above-mentioned second key feature set and the above-mentioned conditional noise feature set to obtain a third key feature set; training an initial model based on the above-mentioned first key feature set and the above-mentioned third key feature set to obtain the above-mentioned cartoon avatar generation model.
[0056] In some alternative implementation manners of some embodiments, the training of the initial model based on the above-mentioned first key feature set and the above-mentioned third key feature set to obtain the above-mentioned cartoon avatar generation model includes: performing data augmentation processing on the second key features in the above-mentioned third key feature set; in response to determining that the similarity between the data-augmented second key features in the above-mentioned third key feature set and the key features in the above-mentioned first key feature set is greater than a preset threshold, determining that the data augmentation processing is completed to obtain a data-augmented third key feature set; using the above-mentioned third key feature set and the above-mentioned sample cartoon avatar set to perform image translation training on the above-mentioned initial model to obtain the above-mentioned cartoon avatar generation model.
[0057] In some alternative implementation manners of some embodiments, the cartoon avatar generation device based on a face image is further configured to: transmit the above-mentioned cartoon avatar to a target device with a display function, and control the target device to display the above-mentioned cartoon avatar.
[0058] It can be understood that the various units described in the device 300 correspond to the respective steps in the method described with reference to Figure 2 Therefore, the operations, features, and beneficial effects described above for the method also apply to the device 300 and the units included therein, and will not be repeated here.
[0059] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present disclosure.
[0060] Figure 4 is a schematic diagram of the computer device 4 provided by an embodiment of the present disclosure. As Figure 4 shown, the computer device 4 of this embodiment includes: a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable on the processor 401. When the processor 401 executes the computer program 403, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor 401 executes the computer program 403, the functions of each module / unit in the above-mentioned device embodiments are implemented.
[0061] Exemplarily, the computer program 403 may be divided into one or more modules / units. The one or more modules / units are stored in the memory 402 and executed by the processor 401 to complete the present disclosure. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 403 in the computer device 4.
[0062] The computer device 4 may be a desktop computer, a notebook, a palm computer, a cloud server, or other computer devices. The computer device 4 may include, but is not limited to, the processor 401 and the memory 402. Those skilled in the art can understand that Figure 4 merely examples of the computer device 4, and do not constitute a limitation to the computer device 4. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the computer device may further include input / output devices, network access devices, buses, etc.
[0063] The processor 401 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.
[0064] The memory 402 can be an internal storage unit of the computer device 4. For example, it can be the hard disk or memory of the computer device 4. The memory 402 can also be an external storage device of the computer device 4. For example, it can be a plug-in hard disk equipped on the computer device 4, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 402 can also include both the internal storage unit and the external storage device of the computer device 4. The memory 402 is used to store computer programs and other programs and data required by the computer device. The memory 402 can also be used to temporarily store the data that has been output or will be output.
[0065] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.
[0066] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0067] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.
[0068] In the embodiments provided in the present disclosure, it should be understood that the disclosed device / computer equipment and method can be implemented in other ways. For example, the device / computer equipment embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. Multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0069] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0070] In addition, each functional unit in various embodiments of the present disclosure can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0071] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of the present disclosure, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. The computer program can include computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0072] The above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them; although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present disclosure, and should all be included within the protection scope of the present disclosure.
Claims
1. A method for generating a cartoon avatar based on a face image, characterized in that, Including: Performing key point detection on the obtained target face image; In response to determining that the key point detection is successful, extracting key features from the target face image to obtain a key feature image, where the key feature image at least includes: a face contour, the positions of facial features, and a hair contour; Based on the conditions input by the user, drawing the key feature image to obtain a completed key feature image, including inputting the completed key feature image into a pre-trained cartoon avatar generation model to generate a cartoon avatar corresponding to the target face image, where the conditions at least include: skin color and hair color; Generating a cartoon avatar corresponding to the target face image based on the completed key feature image; The training steps of the cartoon avatar generation model include: Using an edge extraction algorithm and a face key point detection algorithm to extract key features from each sample face image in the sample face image set to obtain the key features of the sample face image, so as to form a first key feature set; Using the edge extraction algorithm and the face key point detection algorithm to extract key features from each cartoon avatar in the sample cartoon avatar set to obtain the key features of the cartoon avatar, so as to form a second key feature set; Using a preset image feature extraction algorithm to extract features from each cartoon avatar in the sample cartoon avatar set to obtain the conditional noise features of the cartoon avatar, so as to form a conditional noise feature set of the sample cartoon avatar set, where the conditional noise features include: skin color features and hair color features; Combining the second key feature set and the conditional noise feature set to obtain a third key feature set; Training an initial model based on the first key feature set and the third key feature set to obtain the cartoon avatar generation model.
2. The method for generating a cartoon avatar based on a face image according to claim 1, characterized in that, The generating a cartoon avatar corresponding to the target face image based on the completed key feature image includes: The cartoon avatar generation model is trained by taking the sample face images of the training samples in the training sample set as inputs and taking the sample cartoon avatars in the training samples as expected outputs; where the cartoon avatar generation model is a generative adversarial network model.
3. A method for generating a cartoon avatar based on a face image according to claim 2, characterized in that, The training sample set is obtained through the following steps: Screening the obtained face image set to obtain face images that meet the first preset condition as sample face images to obtain a sample face image set; Screening the obtained cartoon avatar set to obtain cartoon avatars that meet the second preset condition as sample cartoon avatars to obtain a sample cartoon avatar set; Combining the sample face image set and the sample cartoon avatar set to obtain the training sample set.
4. A method for generating a cartoon avatar based on a face image according to claim 3, characterized in that, The training an initial model based on the first key feature set and the third key feature set to obtain the cartoon avatar generation model includes: Performing data augmentation processing on the second key features in the third key feature set; Upon determining that the similarity between the second key feature after data augmentation processing in the third key feature set and the key feature in the first key feature set is greater than a preset threshold, it is determined that the data augmentation processing is completed, and the third key feature set after data augmentation processing is obtained; Using the third key feature set and the sample cartoon avatar set, perform image translation training on the initial model to obtain the cartoon avatar generation model.
5. A method for generating a cartoon avatar based on a face image according to any one of claims 1-4, characterized in that, The method further includes: Transmitting the cartoon avatar to a target device with a display function, and controlling the target device to display the cartoon avatar.
6. A cartoon avatar generation device based on a face image, characterized in that, Including: A key point detection unit configured to perform key point detection on the acquired target face image; A key feature extraction unit configured to, in response to determining that the key point detection is successful, extract key features from the target face image to obtain a key feature image, where the key feature image at least includes: a face contour, the positions of facial features, and a hair contour; An image drawing unit configured to draw the key feature image based on conditions input by the user to obtain a completed key feature image; A cartoon avatar generation unit configured to generate a cartoon avatar corresponding to the target face image based on the completed key feature image, where the conditions at least include: skin color, hair color; Generate a cartoon avatar corresponding to the target face image based on the completed key feature image; The training steps of the cartoon avatar generation model include: Using an edge extraction algorithm and a face key point detection algorithm, extract key features from each sample face image in the sample face image set to obtain the key features of the sample face image, so as to form a first key feature set; Using the edge extraction algorithm and the face key point detection algorithm, extract key features from each cartoon avatar in the sample cartoon avatar set to obtain the key features of the cartoon avatar, so as to form a second key feature set; Using a preset image feature extraction algorithm, extract features from each cartoon avatar in the sample cartoon avatar set to obtain the conditional noise features of the cartoon avatar, so as to form the conditional noise feature set of the sample cartoon avatar set, where the conditional noise features include: skin color features, hair color features; Combine the second key feature set and the conditional noise feature set to obtain a third key feature set; Based on the first key feature set and the third key feature set, train the initial model to obtain the cartoon avatar generation model.
7. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 4.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Portrait cartooning method and device, robot and storage medium
CN112465936A
Online live broadcast method and device based on image cartoonization and electronic equipment
CN112561786A
Generative adversarial network training method, image face swapping method and apparatus, and video face swapping method and apparatus
WO2021258920A1