Key point detection model training method and device, electronic equipment and storage medium

By generating a cartoon character image set and converting it into a real character image set, combining the real face key point training model, the problem of low anime face detection accuracy is solved, and high-precision anime face key point detection is achieved.

CN120472512APending Publication Date: 2025-08-12AIBASHI JAPAN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510538967.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing key point detection methods have low accuracy in animation face detection, mainly due to the large difference in facial features between animation faces and real faces, scarce data, and differences in different animation styles, the lack of model generalization ability.

Method used

By generating an image set containing multiple anime characters, and converting them into a real person image set based on prompt words and image generation control parameters, the initial key point detection model is iteratively trained using real face key points to establish a mapping relationship between anime and real faces, and improve detection accuracy.

Benefits of technology

It improves the accuracy of key points detection of anime faces, especially when facing diverse anime styles and complex expressions, and solves the problems of scarce anime face detection data and insufficient generalization capabilities of anime faces.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472512A_ABST
    Figure CN120472512A_ABST
Patent Text Reader

Abstract

The invention provides a key point detection model training method and device, electronic equipment and a storage medium, and relates to the technical field of face recognition. In the application, in response to a model training request for an initial key point detection model, a cartoon character image set including a plurality of cartoon characters is generated; then, based on the first cue word set and a plurality of preset image generation control parameters, generating a real character image set corresponding to the cartoon character image set; and finally, based on at least one real face key point corresponding to a plurality of real figure images included in the cartoon figure image set and the real figure image set, carrying out iterative training on the initial key point detection model, and obtaining a target key point detection model of which the key point detection accuracy is greater than a preset accuracy threshold. By adopting the mode, the key point detection precision of the cartoon face is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of face recognition technology, and in particular to a training method, device, electronic device and storage medium for a key point detection model. Background Art

[0002] Anime is an artistic expression of human figures, often exaggerating certain features to create a compelling effect. While the facial features of anime faces appear realistic, their scale, position, and aspect ratio are often exaggerated and varied. Therefore, key point detection in anime faces has broad applications in areas such as anime detection and editing.

[0003] Currently, keypoint detection methods for anime faces (or anime character faces) include, but are not limited to, keypoint detection based on traditional computer vision, keypoint detection based on deep learning methods, keypoint detection based on synthetic data, and keypoint detection based on cross-domain transfer learning. However, the aforementioned keypoint detection methods often suffer from low keypoint detection accuracy due to the significant differences in facial features between anime faces and real faces, the scarcity of anime face data, and the significant differences in the styles of anime faces across different anime. Summary of the Invention

[0004] The embodiments of the present application provide a training method, device, electronic device, and storage medium for a key point detection model to improve the accuracy of key point detection of animated faces.

[0005] In a first aspect, an embodiment of the present application provides a method for training a key point detection model, the method comprising:

[0006] In response to a model training request for an initial key point detection model, generating an animated character image set comprising a plurality of animated characters; wherein the plurality of animated character images included in the animated character set have different animated facial features;

[0007] generating a set of real-person images corresponding to the set of animated character images based on a first prompt word set and a plurality of preset image generation control parameters; wherein different first prompt words included in the first prompt word set are used to indicate different real-person styles, and different image generation control parameters correspond to different image generation operations;

[0008] Based on at least one real facial key point corresponding to multiple real person images included in the cartoon character image set and the real person image set, the initial key point detection model is iteratively trained to obtain a target key point detection model with a key point detection accuracy greater than a preset accuracy threshold.

[0009] In an optional implementation, generating an animated character image set including a plurality of animated characters includes:

[0010] Acquire a second prompt word set; the second prompt word set includes different second prompt words for indicating different cartoon character styles;

[0011] Based on the plurality of second prompt words included in the second prompt word set and a preset cartoon character generation model, generating cartoon character image subsets corresponding to the plurality of cartoon characters respectively;

[0012] Animation character image sets are obtained based on multiple animation character image subsets.

[0013] In an optional implementation, generating a real person image set corresponding to the cartoon character image set based on the first prompt word set and a plurality of preset image generation control parameters includes:

[0014] performing edge detection on a plurality of cartoon character images included in the cartoon character image set based on at least one edge detection parameter among the plurality of image generation control parameters to obtain a plurality of edge detection results;

[0015] For multiple anime character images, perform the following operations:

[0016] generating a subset of real-person images corresponding to the first cartoon character image based on an edge detection result of the first cartoon character image, a first prompt word set, and at least one image generation control parameter other than at least one edge detection parameter among a plurality of image generation control parameters; wherein the first cartoon character image is any one of the plurality of cartoon character images;

[0017] Save the subset of real person images to the real person image set.

[0018] In an optional implementation, iteratively training the initial key point detection model based on at least one real human face key point corresponding to a plurality of real person images included in the cartoon character image set and the real person image set includes:

[0019] Inputting multiple real person images into a pre-trained real face detection model to obtain at least one real face key point corresponding to each of the multiple real person images;

[0020] At least one real face key point corresponding to each of the multiple real person images is used as an animation face key point;

[0021] The initial key point detection model is iteratively trained based on at least one cartoon face key point corresponding to the cartoon character image set and multiple real character images.

[0022] In an optional implementation, an initial key point detection model is iteratively trained based on at least one key point of an animated face corresponding to an animated character image set and multiple real-life character images, including:

[0023] During each training process of the initial keypoint detection model, the following operations are performed:

[0024] Inputting a second cartoon character image into the initial key point detection model to obtain at least one predicted facial key point of the second cartoon character image; wherein the second cartoon character image is any one of a plurality of cartoon character images included in the cartoon character image set;

[0025] Based on a loss value between at least one predicted facial key point and at least one animated facial key point corresponding to the second animated character image, various network parameters in the initial key point detection model are adjusted.

[0026] In an optional implementation, after obtaining a target key point detection model whose key point detection accuracy is greater than a set key point accuracy threshold, the method further includes:

[0027] Obtaining a target anime character image subset from an anime character image set, and determining a real person image subset corresponding to the target anime character image subset from real person images;

[0028] Based on at least one real human face key point corresponding to at least one real human image included in the real human image subset, data annotation is performed on the target cartoon character image subset to obtain the target cartoon character image subset after data annotation;

[0029] Fine-tune the target key point detection model based on the labeled target anime character image subset.

[0030] In an optional implementation, the method further includes:

[0031] An animated character image containing a target animated character is input into a target key point detection model to obtain multiple facial key points of the target animated character.

[0032] In a second aspect, an embodiment of the present application further provides a training device for a key point detection model, the device comprising:

[0033] A first generating module is configured to generate, in response to a model training request for an initial key point detection model, an anime character image set comprising a plurality of anime characters; wherein the anime character images included in the anime character set have different anime facial features;

[0034] a second generation module, configured to generate a set of real-person images corresponding to the set of cartoon character images based on the first prompt word set and a plurality of preset image generation control parameters; wherein different first prompt words included in the first prompt word set are used to indicate different real-person styles, and different image generation control parameters correspond to different image generation operations;

[0035] The model training module is used to iteratively train the initial key point detection model based on at least one real facial key point corresponding to multiple real person images included in the cartoon character image set and the real person image set, so as to obtain a target key point detection model with a key point detection accuracy greater than a preset accuracy threshold.

[0036] In an optional implementation, when generating an animated character image set including a plurality of animated characters, the first generating module is specifically configured to:

[0037] Acquire a second prompt word set; the second prompt word set includes different second prompt words for indicating different cartoon character styles;

[0038] Based on the plurality of second prompt words included in the second prompt word set and a preset cartoon character generation model, generating cartoon character image subsets corresponding to the plurality of cartoon characters respectively;

[0039] Animation character image sets are obtained based on multiple animation character image subsets.

[0040] In an optional implementation, when generating a real person image set corresponding to the cartoon character image set based on the first prompt word set and a plurality of preset image generation control parameters, the second generation module is specifically configured to:

[0041] performing edge detection on a plurality of cartoon character images included in the cartoon character image set based on at least one edge detection parameter among the plurality of image generation control parameters to obtain a plurality of edge detection results;

[0042] For multiple anime character images, perform the following operations:

[0043] generating a subset of real-person images corresponding to the first cartoon character image based on an edge detection result of the first cartoon character image, a first prompt word set, and at least one image generation control parameter other than at least one edge detection parameter among a plurality of image generation control parameters; wherein the first cartoon character image is any one of the plurality of cartoon character images;

[0044] Save the subset of real person images to the real person image set.

[0045] In an optional implementation, when iteratively training the initial key point detection model based on at least one real human face key point corresponding to a plurality of real person images included in the cartoon character image set and the real person image set, the model training module is specifically configured to:

[0046] Inputting multiple real person images into a pre-trained real face detection model to obtain at least one real face key point corresponding to each of the multiple real person images;

[0047] At least one real face key point corresponding to each of the multiple real person images is used as an animation face key point;

[0048] The initial key point detection model is iteratively trained based on at least one cartoon face key point corresponding to the cartoon character image set and multiple real character images.

[0049] In an optional implementation, when iteratively training the initial key point detection model based on the set of cartoon character images and at least one cartoon face key point corresponding to a plurality of real character images, the model training module is specifically configured to:

[0050] During each training process of the initial keypoint detection model, the following operations are performed:

[0051] Inputting a second cartoon character image into the initial key point detection model to obtain at least one predicted facial key point of the second cartoon character image; wherein the second cartoon character image is any one of a plurality of cartoon character images included in the cartoon character image set;

[0052] Based on a loss value between at least one predicted facial key point and at least one animated facial key point corresponding to the second animated character image, various network parameters in the initial key point detection model are adjusted.

[0053] In an optional implementation, after obtaining a target key point detection model whose key point detection accuracy is greater than a set key point accuracy threshold, the model training module is further configured to:

[0054] Obtaining a target anime character image subset from an anime character image set, and determining a real person image subset corresponding to the target anime character image subset from real person images;

[0055] Based on at least one real human face key point corresponding to at least one real human image included in the real human image subset, data annotation is performed on the target cartoon character image subset to obtain the target cartoon character image subset after data annotation;

[0056] Fine-tune the target key point detection model based on the labeled target anime character image subset.

[0057] In an optional implementation, the device further includes a face detection module, which is specifically configured to:

[0058] An animated character image containing a target animated character is input into a target key point detection model to obtain multiple facial key points of the target animated character.

[0059] In a third aspect, an embodiment of the present application further provides an electronic device, including:

[0060] processor; and

[0061] Memory for storing programs,

[0062] The program includes instructions, which, when executed by a processor, cause the processor to execute the training method for the key point detection model as described in the first aspect.

[0063] In a fourth aspect, an embodiment of the present application further provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the training method of the key point detection model as described in the first aspect.

[0064] In a fifth aspect, the present application provides a computer program product, which, when called by a computer, enables the computer to execute the training method steps of the key point detection model as described in the first aspect.

[0065] The beneficial effects of this application are as follows:

[0066] In the training method of the key point detection model provided in the embodiment of the present application, in response to a model training request for an initial key point detection model, a cartoon character image set including multiple cartoon characters is generated; wherein, the cartoon facial features in the multiple cartoon character images included in the cartoon character set are different; then, based on a first prompt word set and a plurality of preset image generation control parameters, a real person image set corresponding to the cartoon character image set is generated; wherein, the different first prompt words included in the first prompt word set are used to indicate different real person styles, and different image generation control parameters correspond to different image generation operations; finally, based on at least one real face key point corresponding to each of the multiple real person images included in the cartoon character image set and the real person image set, the initial key point detection model is iteratively trained to obtain a target key point detection model with a key point detection accuracy greater than a preset accuracy threshold.

[0067] In this way, the anime character image is converted into a real person image through a first prompt word set and a plurality of preset image generation control parameters, so that the converted face is closer to a real person face; and because at least one real face key point of the real person image can be accurately detected, the target key point detection model obtained by iteratively training the initial key point detection model based on at least one real face key point corresponding to each of the multiple real person images included in the anime character image set and the real person image set can improve the key point detection accuracy of the anime face. In other words, by establishing a mapping relationship between real faces and anime faces, the problem of low key point detection accuracy of anime faces due to the large differences in facial features of anime faces compared to real faces, the scarcity of anime face data, and the large differences in the styles of anime faces from different anime is avoided, thereby improving the key point detection accuracy of anime faces.

[0068] In addition, other features and advantages of the present application will be described in the following description, and in part will become apparent from the description, or may be understood by practicing the present application. The objectives and other advantages of the present application can be realized and obtained through the structures particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described here are used to provide a further understanding of the present application, constitute a part of the present application, and do not constitute an improper limitation of the present application. In the drawings:

[0070] Figure 1 A schematic diagram of an optional system architecture applicable to the embodiments of the present application;

[0071] Figure 2 A schematic diagram of an implementation flow of a training method for a key point detection model provided in an embodiment of the present application;

[0072] Figure 3 A schematic diagram of a specific scenario for training a key point detection model provided in an embodiment of the present application;

[0073] Figure 4 A logical diagram of a fine-tuning target key point detection model provided in an embodiment of the present application;

[0074] Figure 5 A schematic diagram of the structure of a training device for a key point detection model provided in an embodiment of the present application;

[0075] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0076] The following describes embodiments of the present application in more detail with reference to the accompanying drawings. Although certain embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments described herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present application. It should be understood that the drawings and embodiments of the present application are for illustrative purposes only and are not intended to limit the scope of protection of the present application.

[0077] It should be understood that the various steps described in the method embodiments of the present application can be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present application is not limited in this respect.

[0078] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc. mentioned in this application are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0079] It should be noted that the modifications of "one" and "multiple" mentioned in this application are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0080] The names of the messages or information exchanged between multiple devices in the embodiments of the present application are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0081] The following is a brief introduction to the design concept of the embodiment of this application:

[0082] At present, key point detection methods for anime faces (or called anime character faces) include but are not limited to: key point detection based on traditional computer vision, key point detection based on deep learning methods, key point detection based on synthetic data, and key point detection based on cross-domain transfer learning.

[0083] Keypoint detection based on traditional computer vision primarily relies on traditional computer vision techniques, such as Haar features, histogram of oriented gradients (HOG) features, active shape models (ASM), and active appearance models (AAM). These methods rely on hand-crafted feature extractors and matching algorithms, and typically require large amounts of labeled data for training. However, for complex anime-style facial images, especially those featuring anime characters of varying styles and expressions, detection accuracy is low.

[0084] Keypoint detection based on deep learning methods performs well when processing real-world facial images, but accuracy remains limited for anime-style facial features, especially those of exaggerated anime characters. Keypoint detection based on synthetic data addresses the scarcity of anime facial data by augmenting the training set with synthetic data. This involves using image generation techniques to generate anime facial images and then training a keypoint detection model based on existing real-world facial datasets. However, such methods often require addressing the differences between anime facial features and those of real faces.

[0085] Keypoint detection based on cross-domain transfer learning bridges the gap between anime facial images and real faces by transferring a realistic-style face detection model (i.e., a real-world face detection model) to anime-style images or anime-style face images. This approach improves the ability to detect keypoints on anime faces by annotating real-world face data or images and fine-tuning a pre-trained real-world face keypoint detection model.

[0086] However, the above-mentioned key point detection method has problems such as large style differences between real faces and animated faces, lack of training data for animated faces, poor generalization ability of existing key point detection models, strong dependence on model architecture or algorithm, and poor real-time performance, which leads to low key point detection accuracy for animated faces.

[0087] Specifically, anime characters often have exaggerated facial features, diverse expressions, and stylistic variations. However, existing keypoint detection methods are mostly designed for real human faces and cannot effectively adapt to anime-style facial features. For example, the shape and position of parts like the eyes, mouth, and nose of anime characters can differ significantly from those on real faces, making it difficult for models trained on real face data to accurately detect keypoints in anime characters. Compared to real face images, annotated data for anime characters' facial keypoints is relatively scarce, especially high-quality annotated data covering different anime styles. As a result, existing keypoint detection models often rely on limited annotated data for training, making them incapable of broad applicability across styles and works. Furthermore, due to the diversity of anime styles and expressions, existing keypoint detection models struggle to generalize well across different anime works or characters. In particular, when faced with complex expressions, head angles, and visual effects, existing keypoint detection models suffer from low detection accuracy and are prone to keypoint detection errors. Many existing keypoint detection methods rely on specialized deep learning models or traditional algorithms, which may not fully understand and capture the unique characteristics of anime styles when processing anime faces. In particular, model performance may decline when faced with complex expressions or scene changes. Furthermore, some existing keypoint detection methods are computationally intensive and struggle to meet the demands of real-time applications. This is especially true for applications such as animation and virtual reality, which require efficient and fast detection systems.

[0088] In order to solve or improve the above-mentioned problems, an embodiment of the present application proposes a training method for a key point detection model, which may specifically include: in response to a model training request for an initial key point detection model, generating an animated character image set including multiple animated characters; wherein the animated facial features in the multiple animated character images included in the animated character set are different; based on a first prompt word set and a plurality of preset image generation control parameters, generating a real person image set corresponding to the animated character image set; wherein the different first prompt words included in the first prompt word set are used to indicate different real person styles, and different image generation control parameters correspond to different image generation operations; based on at least one real face key point corresponding to each of the multiple real face images included in the animated character image set and the real person image set, iteratively training the initial key point detection model to obtain a target key point detection model with a key point detection accuracy greater than a preset accuracy threshold.

[0089] In particular, the preferred embodiments of the present application are described below in conjunction with the drawings in the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application and are not used to limit the present application. In addition, the embodiments of the present application and the features in the embodiments may be combined with each other if there is no conflict.

[0090] See Figure 1 As shown, it is a schematic diagram of a system architecture applicable to an embodiment of the present application, and the system architecture may include: a terminal device (101a, 101b) and a server 102. The terminal device (101a, 101b) and the server 102 can exchange information through a communication network, wherein the communication mode adopted by the communication network may include: a wireless communication mode and a wired communication mode. Exemplarily, the terminal device (101a, 101b) can access the network through cellular mobile communication technology and communicate with the server 102. The cellular mobile communication technology, for example, includes the fifth generation mobile communication (5th generation mobile networks, 5G) technology or the next generation mobile communication technology. Optionally, the terminal device (101a, 101b) can access the network through a short-range wireless communication mode and communicate with the server 102. The short-range wireless communication mode, for example, includes wireless fidelity (Wi-Fi) technology.

[0091] The embodiment of the present application does not impose any restrictions on the number of communication devices involved in the above system architecture. For example, the above system architecture may include more terminal devices, or fewer terminal devices, or other network devices. Figure 1 As shown, only the terminal devices (101a, 101b) and the server 102 are described as examples, and the above-mentioned communication devices and their respective functions are briefly introduced below.

[0092] The terminal device (101a, 101b) is a device that can provide voice and / or data connectivity to users, and can be a device that supports wired and / or wireless connection.

[0093] Exemplarily, the terminal devices (101a, 101b) may include, but are not limited to: mobile phones, tablet computers, laptop computers, PDAs, mobile internet devices (MIDs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminal devices in industrial control, wireless terminal devices in unmanned driving, wireless terminal devices in smart grids, wireless terminal devices in transportation safety, wireless terminal devices in smart cities, or wireless terminal devices in smart homes, etc.

[0094] In addition, the terminal device (101a, 101b) may be installed with a relevant client, which may be software, such as an application (APP), a browser, a short video software, etc., or a web page, a mini-program, etc. It should be noted that the terminal device (101a, 101b) in the embodiment of the present application may enable the client related to the training of the key point detection model to send a model training request for the initial key point detection model to the server 102, so as to subsequently perform the method steps such as model training for the initial key point detection model.

[0095] Server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0096] It is worth mentioning that in an embodiment of the present application, the server 102 can be used to generate an animated character image set including multiple animated characters in response to a model training request for an initial key point detection model; wherein the animated facial features of the multiple animated character images included in the animated character set are different; then, based on a first prompt word set and a plurality of preset image generation control parameters, a real person image set corresponding to the animated character image set is generated; wherein, different first prompt words included in the first prompt word set are used to indicate different real person styles, and different image generation control parameters correspond to different image generation operations; finally, based on at least one real face key point corresponding to each of the multiple real person images included in the animated character image set and the real person image set, the initial key point detection model is iteratively trained to obtain a target key point detection model with a key point detection accuracy greater than a preset accuracy threshold.

[0097] The following describes the training method of the key point detection model provided by the exemplary embodiment of the present application in combination with the above-mentioned system architecture and reference to the accompanying drawings. It should be noted that the above-mentioned system architecture is only shown to facilitate understanding of the spirit and principles of the present application, and the implementation of the present application is not limited in this respect.

[0098] S201: In response to a model training request for an initial key point detection model, generate an animated character image set including a plurality of animated characters.

[0099] The anime character set includes multiple anime character images having different facial features. Specifically, the anime character set includes one or more anime character images corresponding to multiple anime characters, and the facial features of the same anime character in different anime character images are different. For example, the facial features include, but are not limited to, the shape and positional distribution of the anime character's eyes, mouth, nose, and other parts.

[0100] In an optional implementation, when executing step S201, once the server receives a model training request for the initial keypoint detection model, it can obtain a second prompt word set. Based on the multiple second prompt words included in the second prompt word set and a preset cartoon character generation model, it can generate a plurality of cartoon character image subsets corresponding to the cartoon characters, and then obtain a cartoon character image set based on the multiple cartoon character image subsets. The different second prompt words included in the aforementioned second prompt word set are used to indicate different cartoon character styles and / or facial gestures, i.e., different second prompt words correspond to different cartoon style types. Furthermore, the multiple second prompt words included in the second prompt word set are generated by calling an application programming interface (API). For example, the second prompt words can be generated in batches by calling an API of a chat generative pre-trained transformer (Chat GPT). It is understood that the second prompt words can also be referred to as cartoon style type prompt words, and of course other names are also possible, which are not specifically limited in this embodiment of the present application.

[0101] In addition, the aforementioned preset cartoon character generation model may be a stable diffusion extra large (SDXL) model capable of generating cartoon characters, which is not limited in the embodiments of the present application.

[0102] For example, if the second prompt word set includes six second prompt words, the server can use these six second prompt words and the SDXL model capable of generating anime characters to construct anime character data of different anime styles. For example, using five anime characters as an example, the anime character data includes at least anime character images of the six anime styles corresponding to the five characters, i.e., the anime character data includes at least 30 anime character images.

[0103] It should be understood that, in order to implement the training of the key point detection model provided in the embodiment of the present application, each cartoon character image in the above-mentioned cartoon character data includes the face of the cartoon character.

[0104] S202: Generate a real person image set corresponding to the cartoon character image set based on the first prompt word set and a plurality of preset image generation control parameters.

[0105] The different first prompt words included in the first prompt word set are used to indicate different real-life character styles and / or facial gestures. The first prompt words may also be referred to as realistic style types or real-life character style prompt words, which are not specifically limited in this embodiment of the present application. The multiple first prompt words included in the first prompt word set may also be generated by calling an API. For example, the Chat GPT interface may be called to batch generate first prompt words.

[0106] Different image generation control parameters among the aforementioned multiple preset image generation control parameters correspond to different image generation operations. Optionally, the aforementioned multiple preset image generation control parameters may include, but are not limited to, basic control parameters, preprocessor parameters, and advanced parameters in the ControlNet model. The aforementioned basic control parameters may include control strength, entry timing, termination timing, and control mode. The control strength ranges from 0 to 2. A control strength value of 0 indicates that the generation of a realistic person image relies entirely on the first prompt word. A control strength value of 1 indicates that the structure of the conditional image (e.g., line drawing, pose, etc.) is strictly adhered to. A control strength value greater than 1 indicates that the details of the conditional image (e.g., edge sharpening, depth layering, etc.) are enhanced. The intervention timing refers to the entry of the control conditional image at the beginning of the generation process, and the termination timing refers to the exit of the control conditional image at the end of the generation process. Both values range from 0 to 1. For example, an entry timing of 0.2 indicates that the conditional image begins to be applied at 20% of the generation steps, and a termination timing of 0.8 indicates that the conditional image ceases to be applied at 80% of the generation steps. The control modes may include: a balanced mode (eg, balancing the influence of the conditional image and the first prompt word), a conditional image priority mode (eg, retaining the line drawing), and a prompt word priority mode (eg, the first prompt word covers the line drawing details).

[0107] The above-mentioned preprocessing parameters are mainly used for fine-grained control of conditional images, and may include edge detection parameters, depth and normal parameters, and posture and segmentation parameters. Edge detection parameters may include edge low threshold (canny lowthreshold) / edge high threshold (canny high threshold) and resolution. Among them, the edge low threshold determines the degree of retention of weak edges, and the edge high threshold determines the extraction accuracy of strong edges. Depth and normal parameters may include algorithm selection parameters and color channel inversion parameters. Posture and segmentation parameters may include human body detection range parameters and color mapping parameters (for example, red = person, blue = background). The above-mentioned advanced parameters may include video memory optimization parameters and super-resolution control parameters.

[0108] Therefore, in an optional implementation method, when executing step S202, after generating the cartoon character image set, the server can perform edge detection on the multiple cartoon character images included in the cartoon character image set based on at least one edge detection parameter among the above-mentioned multiple image generation control parameters to obtain multiple edge detection results, and then perform the following operations for any one of the multiple cartoon character images, that is, the first cartoon character image: based on the edge detection result of the first cartoon character image, the first prompt word set and at least one image generation control parameter among the multiple image generation control parameters except at least one edge detection parameter, generate a real person image subset corresponding to the first cartoon character image, and save the real person image subset to the real person image set.

[0109] Exemplarily, the server can use multiple image generation control parameters included in the ControlNet model, multiple first prompt words included in the first prompt word set, and an SDXL model capable of generating real people (i.e., a preset real person generation model) to convert the cartoon character image into a real person image.

[0110] S203: Based on at least one real face key point corresponding to multiple real person images included in the cartoon character image set and the real person image set, the initial key point detection model is iteratively trained to obtain a target key point detection model with a key point detection accuracy greater than a preset accuracy threshold.

[0111] Each part of a real face (e.g., eyes, nose, and ears) in each real person image may correspond to one or more key points. Optionally, if a part of a real face corresponds to multiple key points, characteristic information (e.g., center position and size) of the part may be determined based on the distribution or position information of the multiple key points.

[0112] Taking the above-mentioned preset accuracy threshold of 95% as an example, if the key point detection accuracy of the key point detection model obtained by iterative training of the initial key point detection model is 83%, it indicates that the initial key point detection model needs to be iteratively trained; conversely, if the key point detection accuracy of the key point detection model obtained by iterative training of the initial key point detection model is 96%, it indicates that the key point detection model obtained by iterative training at this time is the target key point detection model.

[0113] In an optional implementation, when executing step S203, after the server obtains the cartoon character image set and the real person image set, it can input multiple real person images into a pre-trained real face detection model to obtain at least one real face key point corresponding to the multiple real person images; then, the at least one real face key point corresponding to the multiple real person images is used as the cartoon face key point; finally, based on the cartoon character image set and the at least one cartoon face key point corresponding to the multiple real person images, the initial key point detection model is iteratively trained.

[0114] It is understood that the above-mentioned real face detection model can also be called an animation facial key point detection model, and of course it can have other names, which are not specifically limited in the embodiments of this application. In addition, the embodiments of this application do not limit the specific type of real face detection model, and it can be any high-precision real face key point detection model.

[0115] Based on this approach, since the pre-trained real-life face detection model can accurately obtain real-life facial keypoints in real-life images, using these real-life facial keypoints as anime face keypoints ensures their accuracy. Therefore, using at least one anime face keypoint corresponding to multiple real-life images as a target for training the initial keypoint detection model improves the reliability of model training.

[0116] Optionally, during each training process of the initial key point detection model based on at least one animated face key point corresponding to the animated character image set and multiple real character images, the server can perform the following operations: input a second animated character image into the initial key point detection model to obtain at least one predicted facial key point of the second animated character image, and thereby adjust the various network parameters in the initial key point detection model based on the loss value between the at least one predicted facial key point and the at least one animated face key point corresponding to the second animated character image. The second animated character image is any one of the multiple animated character images included in the animated character image set. In this way, by adjusting the various network parameters in the initial key point detection model based on the loss value obtained from each training of the initial key point detection model, a facial key point detection model that performs high-precision detection of animated character images can be obtained.

[0117] It can be understood that when a target key point detection model is obtained whose key point detection accuracy is greater than a preset accuracy threshold, the loss value of the target key point detection model satisfies a preset convergence condition.

[0118] Based on the training method of the key point detection model described in steps S201 to S203 above, refer to Figure 3As shown, it is a specific scenario diagram of a training key point detection model provided by an embodiment of the present application. Among them, the animation data construction module can use animation style prompt words (i.e., the second prompt words) and an animation character generation model (for example, an SDXL model that can generate animation character images) to construct animation character data of different animation style types. The animation character to realistic character module can use the ControlNet model, realistic style prompt words (i.e., the first prompt words) and a realistic character generation model (for example, an SDXL model that can generate real character images to convert animation character data into realistic character data. The realistic character key point detection module can detect the facial key points of the realistic character data through the face key point detection model, thereby constructing the facial key point data of the animation character, and the facial key point data of the realistic character to be obtained is used as the facial key point data of the animation character. The animation character facial key point detection training module can train the animation facial key point detection model (i.e., the initial key point detection model) through the animation character data and the aforementioned animation character facial key point data, and combine the loss value to obtain the target key point detection model.

[0119] This demonstrates that, when using the ControlNet model to convert anime-style faces into realistic-style faces, fine-tuning the control parameters (i.e., image generation control parameters) allows the converted faces to approximate real-world facial features (i.e., facial features of real faces) as closely as possible. Furthermore, using an existing realistic facial keypoint detection model to perform keypoint detection on the converted realistic-style faces, accurate realistic-style facial keypoints can be obtained. Next, using realistic-style facial keypoints as targets, a mapping relationship between anime and realistic styles is established to construct training data, and based on this data, an initial keypoint detection model is trained for keypoint detection on anime-style faces. Finally, the trained target keypoint detection model is used to perform keypoint detection on anime characters, significantly improving the keypoint detection model's adaptability to different styles and detection accuracy. Therefore, the use of the target key point detection model can achieve accurate and robust facial key point detection of anime characters, especially when facing diverse anime styles and complex facial expressions, while still maintaining high precision and consistency. It effectively solves the problems of lack of training data and poor model generalization ability in existing anime character facial key point detection technologies, thereby improving the accuracy of key point detection on anime characters' faces.

[0120] In order to further improve the accuracy of the target key point detection model for the key point detection of anime characters' faces (i.e., anime faces). Figure 4As shown, after obtaining a target key point detection model with a key point detection accuracy greater than a set key point accuracy threshold, the server can also obtain a target animated character image subset from the animated character image set and determine a real person image subset corresponding to the target animated character image subset from real person images; then, based on at least one real person face key point corresponding to at least one real person image included in the real person image subset, the target animated character image subset is data labeled to obtain a data-labeled target animated character image subset; finally, the target key point detection model is fine-tuned based on the data-labeled target animated character image subset. In this way, by data labeling a small number of target animated character images, the problem of scarce training data for animated character facial key points can be further improved, thereby ensuring that the trained target key point detection model can more efficiently and accurately detect animated character facial key points.

[0121] Furthermore, after obtaining the target key point detection model, the server may input the cartoon character image containing the target cartoon character into the target key point detection model, thereby obtaining multiple facial key points of the target cartoon character.

[0122] To sum up, in the training method of the key point detection model provided in the embodiment of the present application, in response to a model training request for an initial key point detection model, a cartoon character image set including multiple cartoon characters is generated; wherein, the cartoon facial features in the multiple cartoon character images included in the cartoon character set are different; then, based on a first prompt word set and a plurality of preset image generation control parameters, a real person image set corresponding to the cartoon character image set is generated; wherein, the different first prompt words included in the first prompt word set are used to indicate different real person styles, and different image generation control parameters correspond to different image generation operations; finally, based on at least one real face key point corresponding to the multiple real person images included in the cartoon character image set and the real person image set, the initial key point detection model is iteratively trained to obtain a target key point detection model with a key point detection accuracy greater than a preset accuracy threshold.

[0123] In this way, the anime character image is converted into a real person image through a first prompt word set and a plurality of preset image generation control parameters, so that the converted face is closer to a real person face; and because at least one real face key point of the real person image can be accurately detected, the target key point detection model obtained by iteratively training the initial key point detection model based on at least one real face key point corresponding to each of the multiple real person images included in the anime character image set and the real person image set can improve the key point detection accuracy of the anime face. In other words, by establishing a mapping relationship between real faces and anime faces, the problem of low key point detection accuracy of anime faces due to the large differences in facial features of anime faces compared to real faces, the scarcity of anime face data, and the large differences in the styles of anime faces from different anime is avoided, thereby improving the key point detection accuracy of anime faces.

[0124] Furthermore, based on the same technical concept, the present application embodiment provides a training device for a key point detection model, which is used to implement the above method flow of the present application embodiment. Figure 5 As shown, the key point detection model training device 500 may include: a first generation module 501, a second generation module 502, a model training module 503 and a face detection module 504, wherein:

[0125] A first generating module 501 is configured to generate, in response to a model training request for an initial key point detection model, an anime character image set comprising a plurality of anime characters, wherein the anime character images included in the anime character set have different anime facial features;

[0126] A second generation module 502 is configured to generate a set of real-person images corresponding to the set of animated character images based on the first prompt word set and a plurality of preset image generation control parameters; wherein different first prompt words included in the first prompt word set are used to indicate different real-person styles, and different image generation control parameters correspond to different image generation operations;

[0127] The model training module 503 is used to iteratively train the initial key point detection model based on at least one real face key point corresponding to multiple real person images included in the cartoon character image set and the real person image set, so as to obtain a target key point detection model with a key point detection accuracy greater than a preset accuracy threshold.

[0128] In an optional implementation, when generating an animated character image set including a plurality of animated characters, the first generating module 501 is specifically configured to:

[0129] Acquire a second prompt word set; the second prompt word set includes different second prompt words for indicating different cartoon character styles;

[0130] Based on the plurality of second prompt words included in the second prompt word set and a preset cartoon character generation model, generating cartoon character image subsets corresponding to the plurality of cartoon characters respectively;

[0131] Animation character image sets are obtained based on multiple animation character image subsets.

[0132] In an optional implementation, when generating a real person image set corresponding to the cartoon character image set based on the first prompt word set and a plurality of preset image generation control parameters, the second generation module 502 is specifically configured to:

[0133] performing edge detection on a plurality of cartoon character images included in the cartoon character image set based on at least one edge detection parameter among the plurality of image generation control parameters to obtain a plurality of edge detection results;

[0134] For multiple anime character images, perform the following operations:

[0135] generating a subset of real-person images corresponding to the first cartoon character image based on an edge detection result of the first cartoon character image, a first prompt word set, and at least one image generation control parameter other than at least one edge detection parameter among a plurality of image generation control parameters; wherein the first cartoon character image is any one of the plurality of cartoon character images;

[0136] Save the subset of real person images to the real person image set.

[0137] In an optional implementation, when iteratively training the initial key point detection model based on at least one real human face key point corresponding to a plurality of real person images included in the cartoon character image set and the real person image set, the model training module 503 is specifically configured to:

[0138] Inputting multiple real person images into a pre-trained real face detection model to obtain at least one real face key point corresponding to each of the multiple real person images;

[0139] At least one real face key point corresponding to each of the multiple real person images is used as an animation face key point;

[0140] The initial key point detection model is iteratively trained based on at least one cartoon face key point corresponding to the cartoon character image set and multiple real character images.

[0141] In an optional implementation, when the initial key point detection model is iteratively trained based on at least one key point of an animated character face corresponding to the animated character image set and multiple real character images, the model training module 503 is specifically configured to:

[0142] During each training process of the initial keypoint detection model, the following operations are performed:

[0143] Inputting a second cartoon character image into the initial key point detection model to obtain at least one predicted facial key point of the second cartoon character image; wherein the second cartoon character image is any one of a plurality of cartoon character images included in the cartoon character image set;

[0144] Based on a loss value between at least one predicted facial key point and at least one animated facial key point corresponding to the second animated character image, various network parameters in the initial key point detection model are adjusted.

[0145] In an optional implementation, after obtaining a target key point detection model whose key point detection accuracy is greater than a set key point accuracy threshold, the model training module 503 is further configured to:

[0146] Obtaining a target anime character image subset from an anime character image set, and determining a real person image subset corresponding to the target anime character image subset from real person images;

[0147] Based on at least one real human face key point corresponding to at least one real human image included in the real human image subset, data annotation is performed on the target cartoon character image subset to obtain the target cartoon character image subset after data annotation;

[0148] Fine-tune the target key point detection model based on the labeled target anime character image subset.

[0149] In an optional implementation, the face detection module 504 is specifically configured to:

[0150] An animated character image containing a target animated character is input into a target key point detection model to obtain multiple facial key points of the target animated character.

[0151] Based on the description of the above method embodiment and apparatus embodiment, the exemplary embodiments of the present invention further provide an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, and when executed by the at least one processor, the computer program causes the electronic device to perform a method according to an embodiment of the present invention.

[0152] An embodiment of the present application further provides a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to execute a method according to an embodiment of the present application.

[0153] An embodiment of the present application further provides a computer program product, including a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to perform a method according to an embodiment of the present application.

[0154] See Figure 6 As shown, the structural block diagram of the electronic device 600 that can be used as the server or client of the present application will now be described, which is an example of a hardware device that can be applied to various aspects of the present application. The electronic device is intended to represent various forms of digital electronic computer equipment, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.

[0155] like Figure 6 As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the device 600 can also be stored in the RAM 603. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0156] Multiple components within electronic device 600 are connected to I / O interface 605, including an input unit 606, an output unit 607, a storage unit 608, and a communication unit 609. Input unit 606 can be any type of device capable of inputting information into electronic device 600. Input unit 606 can receive input numeric or character information and generate key input signals related to user settings and / or function control of the electronic device. Output unit 607 can be any type of device capable of presenting information and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. Storage unit 608 may include, but is not limited to, a magnetic disk or an optical disk. Communication unit 609 allows electronic device 600 to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunication networks and may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth device, a WiFi device, a Worldwide Interoperability for Microwave Access (WiMax) device, a cellular communication device, and / or the like.

[0157] The computing unit 601 can be various general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above. For example, in some embodiments, the training method of the key point detection model described above can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 600 via the ROM 602 and / or the communication unit 609. In some embodiments, the computing unit 601 can be configured to perform the training method of the key point detection model described above by any other appropriate means (e.g., by means of firmware).

[0158] The program code for implementing the methods of the present application can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flow charts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0159] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a RAM, a ROM, an erasable programmable read-only memory (EPROM) or a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0160] As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0161] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0162] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0163] Computer systems may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client and server relationship arises through computer programs running on the respective computers and having a client-server relationship to each other.

[0164] Furthermore, it should be understood that what is disclosed above is merely a preferred embodiment of the present application and certainly cannot be used to limit the scope of rights of the present invention. Therefore, equivalent changes made according to the claims of the present invention are still within the scope covered by the present application.

Claims

1. A training method for a key point detection model, characterized in that: include: In response to a model training request for an initial key point detection model, generating an animated character image set comprising a plurality of animated characters; wherein the plurality of animated character images included in the animated character set have different animated facial features; generating a set of real-person images corresponding to the set of cartoon-person images based on a first prompt word set and a plurality of preset image generation control parameters; wherein different first prompt words included in the first prompt word set are used to indicate different real-person styles, and different image generation control parameters correspond to different image generation operations; Based on at least one real human face key point corresponding to each of the multiple real person images included in the cartoon character image set and the real person image set, the initial key point detection model is iteratively trained to obtain a target key point detection model with a key point detection accuracy greater than a preset accuracy threshold.

2. The method according to claim 1, wherein Generating an anime character image set containing a plurality of anime characters comprises: Acquire a second prompt word set; the second prompt word set includes different second prompt words for indicating different cartoon character styles; Based on the plurality of second prompt words included in the second prompt word set and a preset cartoon character generation model, generating cartoon character image subsets corresponding to the plurality of cartoon characters respectively; The cartoon character image set is obtained based on a plurality of cartoon character image subsets.

3. The method according to claim 1, wherein The step of generating a real person image set corresponding to the cartoon character image set based on the first prompt word set and a plurality of preset image generation control parameters includes: performing edge detection on a plurality of cartoon character images included in the cartoon character image set based on at least one edge detection parameter among the plurality of image generation control parameters to obtain a plurality of edge detection results; For the multiple cartoon character images, perform the following operations respectively: generating a subset of real-person images corresponding to the first cartoon character image based on an edge detection result of the first cartoon character image, the first prompt word set, and at least one image generation control parameter other than the at least one edge detection parameter among the plurality of image generation control parameters; wherein the first cartoon character image is any one of the plurality of cartoon character images; The subset of real person images is saved in the real person image set.

4. The method according to any one of claims 1 to 3, wherein The iterative training of the initial key point detection model based on at least one real human face key point corresponding to a plurality of real human images included in the cartoon character image set and the real human image set comprises: Inputting the plurality of real person images into a pre-trained real face detection model to obtain at least one real face key point corresponding to each of the plurality of real person images; Using at least one real face key point corresponding to each of the plurality of real person images as an animated face key point; The initial key point detection model is iteratively trained based on at least one cartoon face key point corresponding to the cartoon character image set and the multiple real character images.

5. The method according to claim 4, wherein The iterative training of the initial key point detection model based on the at least one cartoon face key point corresponding to the cartoon character image set and the plurality of real character images includes: During each training process of the initial key point detection model, the following operations are performed: Inputting a second cartoon character image into the initial key point detection model to obtain at least one predicted facial key point of the second cartoon character image; wherein the second cartoon character image is any one of the multiple cartoon character images included in the cartoon character image set; Based on the loss value between the at least one predicted facial key point and the at least one animated facial key point corresponding to the second animated character image, various network parameters in the initial key point detection model are adjusted.

6. The method according to any one of claims 1 to 3, wherein After obtaining the target key point detection model whose key point detection accuracy is greater than the set key point accuracy threshold, the method further includes: Acquire a target cartoon character image subset from the cartoon character image set, and determine a real person image subset corresponding to the target cartoon character image subset from the real person images; Based on at least one real human face key point corresponding to at least one real human image included in the real human image subset, data labeling is performed on the target cartoon character image subset to obtain a target cartoon character image subset after data labeling; The target key point detection model is fine-tuned based on the target cartoon character image subset after the data is annotated.

7. The method according to any one of claims 1 to 3, wherein The method further comprises: An animated character image containing a target animated character is input into the target key point detection model to obtain multiple facial key points of the target animated character.

8. A training device for a key point detection model, characterized in that: include: A first generating module is configured to generate, in response to a model training request for an initial key point detection model, an anime character image set comprising a plurality of anime characters; wherein the anime character images included in the anime character set have different anime facial features; a second generation module, configured to generate a set of real-person images corresponding to the set of cartoon character images based on a first prompt word set and a plurality of preset image generation control parameters; wherein different first prompt words included in the first prompt word set are used to indicate different real-person styles, and different image generation control parameters correspond to different image generation operations; The model training module is used to iteratively train the initial key point detection model based on at least one real human face key point corresponding to multiple real human images included in the cartoon character image set and the real human image set, so as to obtain a target key point detection model with a key point detection accuracy greater than a preset accuracy threshold.

9. An electronic device comprising: processor; as well as Memory for storing programs, The program includes instructions, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.