Database construction method and device, equipment, medium and program product
By categorizing and storing reference character images into multiple virtual character categories, the problem of disproportionate facial features in the existing database is solved, enabling more accurate virtual character generation.
Patent Information
- Application Number
- CN202511035447.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-11-07
AI Technical Summary
Existing databases use a uniform flat structure in face-shaping technology, which leads to disproportionate facial features or facial asymmetry, resulting in inaccurate matching results.
Reference character images are categorized and stored into multiple virtual character categories. After determining the category through semantic keyword recognition, parameter matching is performed to generate image parameters.
It improves the matching relevance and accuracy of virtual character generation, reduces false matches, and enhances generation quality.
Smart Images

Figure CN120910330A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present specification relate to the technical field of data management, and in particular to a database construction method and device, equipment, medium and program product. BACKGROUND
[0002] In current various network games, role individualization shaping has become an important part of player experience, and the face shaping system as a core function allows users to customize facial features, skin color and makeup, thereby creating a unique virtual image. In recent years, the "text face shaping" technology based on text input to generate role appearance parameters has gradually emerged, which extracts keywords such as facial features, skin color and makeup through semantic analysis of the text, and maps them to the corresponding face parameter set to drive the game engine to generate the role face. This process relies on a pre-constructed face shaping data benchmark library to support the matching of keywords and face parameters.
[0003] However, the existing database adopts a unified flat structure, and all face parameters are stored in the same parameter space, which is easy to cause mis-matching, and thus the problems of facial feature proportion imbalance or facial disharmony. Therefore, there is an urgent need for a database design scheme with more accurate matching effect to improve the generation quality of virtual roles. SUMMARY
[0004] Therefore, the embodiments of the present specification provide a database construction method. One or more embodiments of the present specification also relate to a database construction device, a computing device, a computer-readable storage medium and a computer program product to solve the technical defects in the prior art.
[0005] According to a first aspect of the embodiments of the present specification, a database construction method is provided, comprising:
[0006] obtaining a reference role graph, wherein the reference role graph includes a reference role;
[0007] generating an image parameter of the reference role under a target role category based on the reference role graph, wherein the target role category is determined from a plurality of preset virtual role categories;
[0008] generating a role image database based on the role information and the image parameter of the reference role.
[0009] According to a second aspect of the embodiments of the present specification, a database construction device is provided, comprising:
[0010] The obtaining module is configured to obtain a reference role graph, wherein the reference role graph includes a reference role;
[0011] The first generation module is configured to generate an image parameter of the reference role under a target role category based on the reference role graph, wherein the target role category is determined from a plurality of preset virtual role categories;
[0012] The second generation module is configured to generate a role image database based on the role information and the image parameter of the reference role.
[0013] According to a third aspect of an embodiment of the present specification, a computing device is provided, comprising:
[0014] a memory and a processor;
[0015] The memory is configured to store computer executable instructions, and the processor is configured to execute the computer executable instructions, which implement the steps of the above database construction method when executed by the processor.
[0016] According to a fourth aspect of an embodiment of the present specification, a computer readable storage medium is provided, which stores computer executable instructions, which implement the steps of the above database construction method when executed by the processor.
[0017] According to a fifth aspect of an embodiment of the present specification, a computer program product is provided, comprising computer programs / instructions, which implement the steps of the above database construction method when executed by the processor.
[0018] An embodiment of the present specification implements a database construction method, which comprises: obtaining a reference role graph, wherein the reference role graph comprises a reference role; generating an image parameter of the reference role under a target role category based on the reference role graph, wherein the target role category is determined from a plurality of preset virtual role categories; and generating a role image database based on the role information and the image parameter of the reference role. In the present scheme, the role image parameters are classified and stored according to the preset virtual role categories, avoiding the problem of all parameters being mixed in the same space and being matched without distinction in the traditional flat database structure. In this way, when a keyword input by a user is received, the role category to which the keyword belongs can be identified according to semantics first, and then parameter retrieval and matching are performed under the category, so as to narrow the matching range, reduce interference, and significantly improve the relevance and accuracy of image parameter matching, thereby improving the quality of the virtual role generation result. BRIEF DESCRIPTION OF DRAWINGS
[0019] Figure 1 is a flowchart of a database construction method provided by an embodiment of the present specification;
[0020] Figure 2 is a process flowchart of a database construction method provided by an embodiment of the present specification;
[0021] Figure 3 is a structural schematic diagram of a database construction apparatus provided by an embodiment of the present specification.
[0022] Figure 4 is a structural block diagram of a computing device provided by an embodiment of the present specification. DETAILED DESCRIPTION
[0023] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present specification. However, the present specification can be practiced without the specific details, other than in the examples described herein. Those skilled in the art, having the benefit of the present description, can readily apply the basic inventive concepts to modify or adapt other applications and use the present specification.
[0024] The terminology used in one or more embodiments of the present specification is for the purpose of describing particular embodiments only and is not intended to be limiting of one or more embodiments of the present specification. As used in one or more embodiments of the present specification and the accompanying claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in one or more embodiments of the present specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0025] It will be understood that, although the terms first, second, etc. can be used herein to describe various information, these terms are not intended to denote a temporal or chronological order. Rather, these terms are used solely to distinguish one from another only. For example, without departing from the scope of one or more embodiments of the present specification, first can be termed second, and similarly, second can be termed first. The word "if' as used herein means "when" or "upon" or "in response to a determination" depending on the context.
[0026] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of the present specification are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.
[0027] In the present specification, a database construction method is provided, and the present specification also relates to a database construction apparatus, a computing device, a computer readable storage medium, and a computer program product, which are described in detail one by one in the following embodiments.
[0028] Referring to Figure 1 , Figure 1 A flowchart of a database construction method according to one embodiment of the present specification is shown, which specifically includes the following steps.
[0029] Step 102: Obtain a reference role image, wherein the reference role image includes a reference role.
[0030] The reference role image refers to an image or a set of images, which is used as an example image for constructing a database. Each image includes at least one "reference role" that has been designed, is representative, and can be modeled. The role can be a person or other target object with unique image features. The image information of these reference roles will be used as the basis for generating image parameters and constructing a database.
[0031] The reference role image can be obtained in various ways, including but not limited to manual drawing, photography, model generation, or collecting existing image resources.
[0032] In an optional implementation of the present embodiment, the specific implementation process of obtaining the reference role image includes the following steps:
[0033] Obtain an image set, wherein the image set includes a plurality of object images; extract the object images from the image set, perform key point detection on the object images, and obtain key point information of the target object; based on the key point information, intercept the face image of the target object to obtain the reference role image.
[0034] The image set refers to a set including multiple images. The object image refers to an image containing an object to be identified or processed in the image.
[0035] In the present embodiment, an image set is first obtained, which contains multiple images, and each image contains one or more objects. These images can be obtained in various ways, such as being synthesized by an image generation model, being manually drawn, or being sourced from network materials. Network materials can include public network photos, network video frames, etc. The present embodiment does not limit the specific source of the images, as long as the images contain object information that can be used to extract image features, they can all be used as components of the image set.
[0036] By way of example, the above-mentioned image set can be composed of photos of multiple stars. For this purpose, a reference database containing the names of more than 900 stars can be maintained in advance as a target list for image collection. Based on this list, images corresponding to these star names can be automatically captured from public networks through crawler technology.
[0037] To ensure image quality and applicability, further processing of the image set is required to obtain the final image set used to constitute the reference role map. The following scheme is the further processing process of the image set.
[0038] Among them, extracting the object image from the image set means screening out the image containing the effective object from the image set.
[0039] Among them, key point detection is a computer vision technology mainly used for accurate positioning of the key structure area of the target object in the image. For example, in the task of face recognition or modeling, key point detection can be used to identify and mark important facial regions such as eyebrows, eyes, nose, mouth, and facial contours. The process usually outputs the position coordinates of multiple key points to represent the facial structure of the object. Key point information refers to the position set of the key structure of the target object obtained by key point detection.
[0040] Through the above key point detection process, not only the accurate positioning information of the facial key structure in the image can be obtained, but also a structured data basis is provided for subsequent face cropping, feature extraction, image generation, etc.
[0041] After obtaining the key point information, further optimization processing of the image set based on the key point information is performed to obtain the facial image of the target object, and the specific process is as follows.
[0042] First, key point detection is performed on the object image to obtain the positions of multiple typical structure points of the face (such as eyebrows, eyes, nose tip, corners of the mouth, chin contour, etc.). These key points provide the geometric positioning basis for the face region.
[0043] Next, in real-world applications, the symmetry and distribution of the key points or the number of key points are analyzed based on the key point information to determine whether the target object in the image is a frontal face. Only frontal face images can be used for subsequent processing.
[0044] For example, if the key points exhibit good symmetry (such as eyes on the horizontal line, nose in the middle, etc.), the image is determined to be a frontal face image.
[0045] For example, if 68 key points including eyes, nose tip, corners of the mouth, etc. can be successfully detected, the image is determined to be a frontal face image.
[0046] It's important to note that commonly used keypoint detection algorithms are trained on frontal facial images, which annotate the key structural regions of the target object in the image. These include areas such as the eye contour, eyebrows, bridge of the nose, nostrils, lip contour, and jawline. When the target object in the image is facing frontally, all these features are visible, allowing the algorithm to predict all points with high confidence. However, non-frontal orientations can cause occlusion or distortion. For example, in a profile view, an eye or corner of the mouth might be obscured, preventing the keypoint detection algorithm from predicting the corresponding point. Furthermore, when the head is tilted up or down, the tip of the nose or the jawline may shift or become blurred, causing some points to fail to locate. Additionally, hair, hands, hats, etc., obscuring the face can also affect the completeness of keypoint recognition. Therefore, the number of keypoints output by the model can be used as a criterion for determining whether the target object in the image is a frontal face. If the number of keypoints is insufficient, it can be inferred that the target object is not a standard frontal view; only when the entire face of the target object is visible, with minimal distortion and high symmetry can the algorithm identify complete keypoints (e.g., 68 keypoints) with high quality, which is a key characteristic of a frontal face.
[0047] Considering that in some cases, although some substandard images can be automatically filtered out based on key point information (such as failed face detection, insufficient number of facial key points, etc.), the automatic method still has certain misjudgments (such as mistakenly retaining partially occluded faces), manual secondary screening is required to ensure that the final image meets the following conditions: unobstructed face, no obvious shadows, and a pose close to frontal (no large angle of deviation).
[0048] Next, if the image is identified as a frontal face, a facial bounding box is automatically calculated based on the distribution of key points, and the facial region of the target object is cropped out. This is known as "facial image cropping." This step eliminates redundant content such as background information and half-body or full-body images, allowing the data to focus more on facial features.
[0049] Then, a reference character image is determined based on facial images cropped from the image set.
[0050] This embodiment does not limit the specific method for determining the reference character image based on facial images extracted from the image set. For example, one facial image may be selected as the reference character image; or multiple facial images may be aggregated, such as through image fusion, style averaging, feature weighted synthesis, etc., to generate a comprehensive and representative reference character image.
[0051] Following the example, after obtaining an image set composed of photos of multiple stars, first process each photo through a face key point detection algorithm to extract its facial key point information. If a complete 68 key points are successfully detected, the image is retained; if the detection result is less than 68 key points, it is determined as a non-frontal face image or poor quality, and is excluded. Subsequently, the image set after preliminary screening is screened again to further exclude images with obvious occlusion, strong shadow or excessive posture deflection, and only high-quality photos with frontal face orientation, no occlusion and uniform illumination are retained.
[0052] On this basis, using the corresponding key point information in each retained image, a face box region is automatically calculated and generated for accurately cropping the facial image region of the star. Finally, a high-quality image is selected from the facial region image of the star as the reference role graph, which is used as the standard image sample for constructing the face sculpting parameter database.
[0053] The embodiment of the present specification realizes a method for obtaining a reference role graph, which comprises: acquiring an image set, wherein the image set comprises a plurality of object images; extracting an object image from the image set, performing key point detection on the object image, and obtaining key point information of a target object; based on the key point information, cropping a facial image of the target object to obtain a reference role graph. In the present scheme, low-quality images with face posture deflection, facial occlusion or uneven illumination are effectively excluded through key point detection, ensuring that the retained images are high-quality images with frontal face orientation, no occlusion and no shadow, thereby improving the reference accuracy and reliability of subsequent image parameters. Secondly, the face region of the image with frontal face orientation is cropped to effectively remove background interference and redundant information, thereby effectively improving the efficiency of the entire database construction process and the final generation effect of the subsequent face sculpting process.
[0054] In another optional implementation of the present embodiment, the specific implementation process of obtaining a reference role graph comprises the following steps:
[0055] Obtain image description information; determine the target prompt word corresponding to the image description information based on the image description information; generate image data corresponding to the image description information based on the target prompt word; and obtain a reference role graph based on the image data.
[0056] First, obtain the image description information input by the user or automatically collected, which can come from user interaction, network query or feedback results generated by a large language model (such as ChatGPT).
[0057] Specifically, a plurality of Chinese keywords for describing the image can be obtained, and the following are some example word categories:
[0058] Words for describing eyebrows: messy eyebrows, willow leaf eyebrows, sword eyebrows, thick eyebrows, thick eyebrows.
[0059] Words for describing eyes: almond eyes, phoenix eyes, peach blossom eyes.
[0060] Words for describing mouth: cherry mouth, thick lips, thin lips.
[0061] Words for describing nose: high bridge, hawk nose, flat nose.
[0062] Words for describing skin: fair skin, dark skin, smooth skin.
[0063] Words for describing lip color: red lips, pale lip color, rosy lip color.
[0064] Words for describing face shape: oval face, square face, round face.
[0065] After obtaining the above-mentioned multiple Chinese image description words, the information is combined, and the description information obtained by the combination is used to generate a prompt word that describes the overall image in natural language, such as: "a girl, full face, peach blossom eyes, willow-leaf eyebrows, red lips, fair skin, oval face, Chinese style, centered, white background".
[0066] When connected to an English image generation system, it is converted into an English prompt word, for example: "1 girl, full face, peach blossom eyes, willow-leaf eyebrows, red lips, fair skin, oval face, Chinese style, centered, white background".
[0067] Among them, the image data refers to the visual image data generated or expressed according to the prompt word to represent the image characteristics. Specifically, the image data is usually presented in the form of an image.
[0068] After generating the target prompt word, the prompt word is input into an image generation engine, such as Stable Diffusion, white-xl model specialized in ancient style, DALL·E, Midjourney, or a custom multi-modal model. The model generates image data corresponding to the prompt word according to the prompt word.
[0069] In this embodiment, the image data is a set of images generated based on the prompt word. In order to determine the final reference role image from the image data, any of the following methods can be used for selection or processing:
[0070] Single-image selection: directly selecting an image that best meets the expectations from the image data as the reference role image. The image can be manually specified by the user, or automatically selected by the system according to image quality and other indicators.
[0071] The image aggregation generates a reference character image: multiple images are selected from the image data, and the multiple images are subjected to image fusion, style averaging or feature weighting synthesis processing to generate a reference character image with comprehensive representation.
[0072] Through any of the above methods, a representative reference character image can be obtained from the image data generated by the prompt word as a standard input for subsequent modeling, rendering or visual generation.
[0073] The embodiments of the present specification realize another method for obtaining a reference character image, which comprises: obtaining image description information; determining a target prompt word corresponding to the image description information based on the image description information; generating image data corresponding to the image description information based on the target prompt word; and obtaining a reference character image based on the image data. In this scheme, the user only needs to provide intuitive image description information (such as "willow leaf eyebrows" and "peach blossom eyes"), without the need for manual drawing or collecting sample images, and can quickly and batch generate reference character images that meet the requirements, greatly improving the efficiency of constructing reference character images.
[0074] It should be noted that the above embodiments provide two methods for obtaining a reference character image: one is based on an image set, and the reference character image is generated by key point detection and image extraction; the other is based on image description information, and the reference character image is generated by a prompt word driven image generation model. Although the two methods differ in implementation path, they are not mutually exclusive in actual application, but can be complementary and used together. In particular, in scenarios where a character image database with wide coverage and high diversity needs to be constructed, the combination of the two methods has obvious advantages. For example: some reference images are derived from real person photos, and high-quality front face images can be extracted by key point detection technology as reference character images; another part is generated by a text description driven diffusion model to construct virtual character images with specific styles or custom features.
[0075] This method helps to achieve an effective balance between data authenticity and expression flexibility, further improving the quality, size and application adaptability of the character image database.
[0076] Taking the functional requirement of "players generating virtual character faces in game scenes through text description" as the background, we designed the following two parts of image data construction scheme to support the benchmark image library of this function:
[0077] First part: star face shape image (real image source)
[0078] We maintain a database of 900+ celebrity names; use web crawlers to automatically scrape photos of the corresponding celebrities from public channels; apply face key point detection and image filtering algorithms to eliminate photos with side faces, obstructions, and abnormal lighting, and retain clear front-facing images; crop the face area of the photos and retain only the facial feature area, removing the background and irrelevant content; after manual review and screening, we finally obtain about 2000 high-quality celebrity face images as one of the components of the reference role image.
[0079] Part II: Customized Face Shape Images (Generated Image Sources)
[0080] For the facial features of "good-looking" in a general sense, we use two Diffusion models to generate images: the standard Stable Diffusion model and the specialized ancient style model white-xl. Based on network data and large model feedback, we have constructed 300+ facial feature and makeup description words (with Chinese-English translation), such as:
[0081] Messy eyebrows: unkempt eyebrows;
[0082] Almond-shaped eyes: almond-shaped eyes;
[0083] Red lips: red lips;
[0084] Through structured design of prompt templates, different facial feature words are arranged and combined to generate a variety of facial images.
[0085] Example prompts are as follows: 1 girl, half_body, full_face, round eyes, bold eyebrows, bigmouth, dumpy nose, red lips, fair skin, round face, looking at viewer, centered, white background, detail_face, Chinese-style.
[0086] The generated images are manually screened, and only high-quality samples are retained to be included in the reference face image library.
[0087] By combining the above two methods, the reference role graph dataset constructed has the following advantages: covering the full spectrum of data structures of real images and virtually generated images; meeting the dual needs of high restoration (such as star faces) and high customization (user features); significantly expanding the size of the face database, improving the flexibility and diversity of role construction; supporting the input needs of complex tasks such as face pinching, image synthesis, and style transfer in multiple scenarios.
[0088] Step 104: Based on the reference role graph, generate the image parameters of the reference role in the target role category, wherein the target role category is determined from a plurality of preset virtual role categories.
[0089] Among them, "preset virtual role category" can include but is not limited to: virtual roles with different body styles (such as adult men, loli, and shoujo); roles with different gender characteristics (such as male, female, and neutral style); roles with different artistic styles and styles (such as reality, cartoon, realism, and two-dimensional); in this embodiment, "preset virtual role category" is not specifically limited.
[0090] For example, based on the game setting, the following four body types are predefined: adult male, adult female, shoujo, and loli.
[0091] Among them, the target role category is any of the plurality of preset virtual role categories.
[0092] Among them, the image parameters are controllable numerical values used to represent the reference role in a certain virtual role category, which usually include:
[0093] The basic parameters corresponding to the basic attributes of the face: the geometric parameters used to describe the facial structure of the role. Usually include: the position and shape of the eyes, the position and shape of the mouth, the position and shape of the nose, the position and shape of the eyebrows, and the position and shape of the ears.
[0094] Among them, "shape" can be represented in two ways: using specific numerical shape parameters (such as curve control point coordinates); selecting the shape with the highest matching degree from a preset group of template shapes, and using the template identifier of the shape as the shape parameter. For example: for adult men, 5 eyebrow shapes are preset; based on the reference role graph, the best match is selected, and the identifier (such as the serial number) of the eyebrow shape is used as the shape parameter; for adult women, 10 eyebrow shapes are preset; similarly, select the highest matching one, and use its eyebrow shape identifier as the shape parameter.
[0095] Additional parameters corresponding to additional attributes of the face, used to describe appearance details and surface performance. For example, makeup parameters.
[0096] In an optional implementation of the embodiment, a specific implementation process of generating the image parameters of the reference role under the target role category based on the reference role graph includes the following steps.
[0097] At least one facial basic attribute corresponding to the reference role and at least one facial additional attribute corresponding to the reference role are extracted from the reference role graph; based on the at least one facial basic attribute, the basic parameters corresponding to each facial basic attribute of the reference role under the target role category are generated; based on the at least one facial additional attribute, the additional parameters of the reference role under the target role category are generated; and based on the basic parameters and the additional parameters, the image parameters of the reference role under the target role category are generated.
[0098] In a possible implementation, first, two independent autoencoders are trained, one image encoder encodes the reference graph visually, and the other parameter encoder encodes the virtual role parameters (facial features and makeup); second, a small amount of paired samples (image parameters) are used to train the mapping network to map the visual latent vector output by the image encoder to the parameter latent vector space; finally, during inference, the visual latent vector of the reference role graph is input, mapped to the parameter space, and the complete image parameters are directly output by the parameter decoder.
[0099] In another possible implementation, the reference role graph can be input into the "graph-based face modeling" model corresponding to the target role category. The "graph-based face modeling" model extracts at least one facial basic attribute corresponding to the reference role and at least one facial additional attribute corresponding to the reference role from the reference role graph; based on the at least one facial basic attribute, the basic parameters corresponding to each facial basic attribute of the reference role under the target role category are generated through the neural network corresponding to the target role category; based on the at least one facial additional attribute, the additional parameters of the reference role under the target role category are generated; and based on the basic parameters and the additional parameters, the image parameters of the reference role under the target role category are generated.
[0100] Exemplarily, the reference role image will be input into the "face kneading with image" model corresponding to the target role category. The "face kneading with image" model extracts the image area of the facial features (eyes, nose, mouth, ears, eyebrows) of the reference role from the reference role image; and extracts the skin color and lip color corresponding to the reference role from the reference role image. Based on the image area of the facial features of the reference role, the basic parameters (position and shape of eyes, position and shape of mouth, position and shape of nose, position and shape of eyebrows, position and shape of ears) corresponding to each facial basic attribute of the reference role under the target role category are generated through the neural network corresponding to the target role category. The skin color of the reference role is input into a classification network for skin color to obtain the color category of the skin color, and the lip color of the reference role is input into a classification network for lip color to obtain the color category of the lip color; then the color category of the skin color and the color category of the lip color are input into a classification network for predicting makeup to obtain the additional parameters about the predicted makeup. Based on the basic parameters and the additional parameters, the image parameters of the reference role under the target role category are generated.
[0101] Following the above example, the preset virtual role categories are adult male, adult female, shou tai and loli.
[0102] Reference role Figure 1 Part of the reference role image is composed of star face image group, and part of the reference role image is composed of model generated image group.
[0103] For the reference role image composed of star face image, the reference role image is first divided into male reference role image and female reference role image. Among them, the male reference role image is input into the "face kneading with image" model corresponding to the adult male, and the image parameters of the reference role under the adult male can be obtained; the male reference role image is input into the "face kneading with image" model corresponding to the shou tai, and the image parameters of the reference role under the shou tai can be obtained; the female reference role image is input into the "face kneading with image" model corresponding to the adult female, and the image parameters of the reference role under the adult female can be obtained; the female reference role image is input into the "face kneading with image" model corresponding to the loli, and the image parameters of the reference role under the loli can be obtained.
[0104] For the reference role image generated by the model, the customized face image is input into the "face kneading with image" model corresponding to the adult male, the "face kneading with image" model corresponding to the shou tai, the "face kneading with image" model corresponding to the adult female, and the "face kneading with image" model corresponding to the loli, respectively. The image parameters of the reference role under the adult male, the image parameters of the reference role under the shou tai, the image parameters of the reference role under the adult female, and the image parameters of the reference role under the loli can be obtained respectively.
[0105] The embodiment of the specification implements a generation method of an image parameter, which comprises: extracting at least one facial basic attribute corresponding to a reference role from a reference role graph, and at least one facial additional attribute corresponding to the reference role; generating a basic parameter corresponding to each facial basic attribute of the reference role under a target role category based on the at least one facial basic attribute; generating an additional parameter of the reference role under the target role category based on the at least one facial additional attribute; and generating an image parameter of the reference role under the target role category based on the basic parameter and the additional parameter. In the scheme, the image parameter is automatically generated without human intervention, and batch processing of multiple reference graphs is supported, which greatly improves the construction efficiency of the database. At the same time, the image parameter includes both the basic parameter corresponding to the facial basic attribute and the additional parameter corresponding to the facial additional attribute, so that the finally generated virtual role more accurately reflects the reference role graph.
[0106] In an optional embodiment of the present embodiment, after obtaining the basic parameter and the additional parameter of the reference role graph under the target role category in the above embodiment, the basic parameter and the additional parameter can be directly used as the image parameter of the reference role graph under the target role category.
[0107] Considering that in some cases, the basic parameter generated only by relying on the facial basic attribute extracted from the reference role graph may deviate from the standard facial structure (such as the distance between the eyebrows, nose and mouth being too close, which does not conform to the body type characteristics) that the target role category (for example, “male”) should have due to image quality, posture or expression, etc., therefore, before generating the image parameter of the reference role under the target role category based on the basic parameter and the additional parameter, the basic parameter should be corrected to ensure that it is consistent with the typical geometric characteristics of the target role category.
[0108] In another optional embodiment of the present embodiment, before generating the image parameter of the reference role under the target role category based on the basic parameter and the additional parameter, it further comprises:
[0109] generating a virtual role graph of the reference role under the target role category based on the basic parameter and a first virtual model corresponding to the target role category; performing key point detection on the virtual role graph to determine each facial key point of the virtual role graph; determining the distance between each facial basic attribute in the virtual role graph based on the facial key points; and correcting the basic parameter corresponding to each facial basic attribute of the reference role under the target role category based on the distance between each facial basic attribute, to obtain a corrected basic parameter.
[0110] First, the generated basic parameter is input into a game engine or a 3D renderer, and the game engine or the 3D renderer renders the basic parameter into a virtual role graph corresponding to the first virtual model based on the basic parameter and the first virtual model corresponding to the target role category.
[0111] For example, the target character category is a young man, and the generated basic parameters (such as the position and shape of the eyebrows, the position and shape of the eyes, the position and shape of the nose, the position and shape of the mouth, etc.) are set to the character driving system in the game engine. The engine reads the preset virtual model of the “young man” body type (the first virtual model), maps these basic parameters to the character geometry and facial morphology, and performs a traditional rendering pipeline to finally generate a virtual character graph of the “young man” body type and adapted to the changes in the input basic parameters, which is used for subsequent key point detection and basic parameter correction steps.
[0112] Next, key point detection is performed on the virtual character graph to automatically identify each facial key point of the virtual character graph.
[0113] Continuing with the above example, after rendering the virtual character graph of the “young man” body type, the system inputs the generated virtual character graph into a face key point detection algorithm. The algorithm automatically identifies and locates the precise coordinates of each facial key area in the virtual character graph, including the eyebrow area, eye area, nose area, and mouth area, and represents the geometric distribution of these areas with multiple well-defined key points (such as the eye corner, nose tip, lip peak, and eyebrow head / tail points).
[0114] Next, based on the detected key points, the main spatial distances between the facial basic attributes are calculated to quantify the relative structural proportions between the facial basic attributes.
[0115] Continuing with the above example, after obtaining the key points corresponding to the virtual character graph of the “young man” body type (including the key points of the eyebrows, eyes, nose, and mouth), the interocular distance (the distance between the centers of the left and right eyes), the eyebrow-eye distance, and the nose-mouth distance are calculated.
[0116] Next, the actually calculated distances are compared with the standard template distances of the “young man” body type. If they deviate from the preset range, it indicates that the basic parameters are deviated. Statistical shape models (such as Point Distribution Model) or regression models are applied to uniformly compensate each basic parameter to make it more consistent with the overall structural characteristics of the “young man”. Combined with the compensated parameter values, the final corrected basic parameters are generated to ensure that the corrected basic parameters not only retain the characteristics of the reference character but also meet the structural specifications of the target body type.
[0117] Finally, these corrected basic parameters are combined with the original additional parameters to output the image parameters that meet the target character category.
[0118] The embodiment of the specification implements a method for correcting a basic parameter, which comprises: generating a virtual character graph of a reference character in a target character category based on a first virtual model corresponding to the basic parameter and the target character category; performing key point detection on the virtual character graph to determine each facial key point of the virtual character graph; determining the distance between each facial basic attribute in the virtual character graph based on each facial key point; and correcting the basic parameter corresponding to each facial basic attribute of the reference character in the target character category based on the distance between each facial basic attribute, to obtain a corrected basic parameter. In this way, the automatic correction of facial features is realized, the basic parameter more consistent with the characteristics of the target character category is obtained, the accuracy and consistency of the virtual character generation are improved, the need for manual intervention is reduced, and the user experience is improved.
[0119] The above embodiment automatically generates a plurality of preset virtual character categories corresponding to the image parameters. For example, in the case of a plurality of preset virtual character categories being adult male, adult female, shou tai and loli, the image parameters corresponding to adult male, adult female, shou tai and loli are obtained. However, due to the limitations of the "graph-based face modeling" model for generating image parameters based on reference character graphs, etc., the generation effect of part of the image parameters may not meet the requirements. Therefore, in real application, the generated image parameters need to be further screened. Specifically, after all the image parameters are generated, the image parameters of all virtual character categories (such as adult male, adult female, shou tai and loli) are imported into the game engine, and the engine renders the virtual model of the corresponding category based on each set of image parameters to generate the corresponding virtual character graph. These virtual character graphs are used as visual references, and the evaluator evaluates each graph in terms of aesthetic degree, structural coordination, style consistency, etc. For each virtual character category, only one set of image parameters that pass the visual evaluation is finally retained as the image parameters of the category.
[0120] In some embodiments, after generating the image parameters for each virtual character category (e.g., male, female, shoujo, lolita), a plurality of emotion / style sub-categories (e.g., "sad", "lively", "pure" and the like) can be further defined. Since the number of original images for each main category is limited, the system cannot directly learn the subtle differences corresponding to these sub-categories from the existing images. Therefore, the system will perform style fine-tuning based on the main category image parameters to supplement the sub-category parameters. For example, to generate the "lively" sub-category, the system can fine-tune multiple existing male image parameters to obtain a version with higher eyebrow and eye arch, more outwardly expanded mouth corners, and higher skin brightness, thereby forming a "male-lively" sub-category. Similarly, by adjusting the eyebrow height, eyelid opening degree, and mouth line, the "sad" or "pure" sub-categories can be generated. This style-oriented parameter fine-tuning mechanism can generate rich and expression-oriented sub-category image parameters in the case of data scarcity, significantly enhancing the diversity and expressiveness of the system.
[0121] In an optional implementation of the present embodiment, before generating the character image database based on the character information and image parameters of the reference character, the method further includes: adjusting the image parameters to generate at least two image adjustment parameters; and determining the image parameters corresponding to at least two sub-categories under the target character category based on the at least two image adjustment parameters.
[0122] Correspondingly, the character image database is generated based on the character information of the reference character and the image parameters corresponding to at least two sub-categories under the target character category.
[0123] First, at least one facial basic attribute that affects the facial feeling and temperament is determined, as well as the adjustment direction of the at least one facial basic attribute. According to the adjustment direction of the at least one facial basic attribute, a plurality of adjustment strategies are combined.
[0124] Exemplarily, a facial basic attribute that affects the facial feeling and temperament, and the adjustment direction of the facial basic attribute are shown in Table 1.
[0125] Table 1
[0126]
[0127] First, according to the three influential facial basic attributes of eyebrows, eyes and mouth listed in Table 1, permutation and combination is performed. Since there are 3 adjustment methods for each facial basic attribute, plus the treatment of "maintaining unchanged", theoretically, 4*4*4=64 adjustment strategies can be combined. Subsequently, the system applies these 64 strategies to each set of image parameters in batches to generate 64 groups of different "image adjustment parameters". Since some of the combinations may not be logical in vision (such as eyebrows down and eyes extremely drooping at the same time, which looks dull or unnatural as a whole), it is necessary to filter these adjustment strategies and only keep the logical and visually coordinated strategies to finally form an optimized and usable image adjustment strategy library.
[0128] Exemplarily, Table 2 provides an optimized and usable image adjustment strategy library after screening.
[0129] Table 2
[0130] Number Eyebrow Eye Mouth v1 Overall Up Guaranteed Invariant Guaranteed Invariant v2 Level Out Guaranteed Invariant Guaranteed Invariant v3 Overall Down Guaranteed Invariant Guaranteed Invariant v4 Guaranteed Invariant Level Up Guaranteed Invariant v5 Guaranteed Invariant Level Guaranteed Invariant
[0131] For each adjustment strategy in the image adjustment strategy library (for example: eyebrows up + eyes horizontal + mouth drooping), the strategy is applied to all original image parameters to generate multiple sets of adjusted image parameters. These adjusted image parameters are input into the game engine to batch render corresponding virtual character images based on the virtual models corresponding to each virtual character type. Finally, a set of candidate image library is generated for each virtual character type, which includes all logically reasonable and stylistically different adjustment results for subsequent screening and application.
[0132] It should be noted that in this embodiment, for each image parameter adjustment method, we predefine the corresponding numerical coefficient for parameter transformation, for example: "overall up": multiply the original image parameter by a coefficient greater than 1 (such as x1.1) to achieve the overall lifting effect of the organ; "overall down": multiply the original parameter by a coefficient less than 1 (such as x0.9) to achieve the overall lowering. This adjustment parameter defined by multiplication coefficient realizes the explicit mapping between each adjustment strategy and the value, making the adjustment intuitive and measurable.
[0133] The embodiment of the specification realizes a method of refining the image parameters under each virtual role category into image parameters under at least two subcategories, which comprises: adjusting the image parameters to generate at least two image adjustment parameters; and determining the image parameters corresponding to the at least two subcategories under the target role category based on the at least two image adjustment parameters. In the scheme, the image parameters under each virtual role category are further refined to obtain the image parameters corresponding to the at least two subcategories under each virtual role category. Thus, more rich role styles (such as "pure" and "melancholy") are brought, and the adaptation ability of the database to the more variable and more detailed needs of the user is improved.
[0134] In an optional embodiment of the present embodiment, after a set of image adjustment parameters are generated and screened, the screened image adjustment parameters are first vectorized and input into a clustering algorithm (such as K-Means Clustering Algorithm, K-means clustering algorithm) to automatically generate multiple clusters. Then, the system selects at least one representative parameter (such as the parameter vector of the cluster center) from each cluster as the candidate image parameter corresponding to the cluster, and inputs it into the game engine to render a virtual role graph through the virtual model corresponding to the target role category preset in the engine. According to the visual style of the graph, the specific subcategory (such as "grown-up male-sunny" or "grown-up male-steady") to which the parameter belongs is determined. Thus, multiple subcategory image parameters with different style expressions can be automatically distinguished and constructed for each target role category.
[0135] In another optional embodiment of the present embodiment, determining the image parameters corresponding to the at least two subcategories under the target role category based on the at least two image adjustment parameters comprises: generating an image graph of the image adjustment parameters under the target role category based on the image adjustment parameters and the second virtual model corresponding to the target role category; determining the category description information of the target subcategory and the image description information corresponding to each image graph of the image adjustment parameters under the target role category, wherein the target subcategory is any one of the at least two subcategories; determining the target image graph matched with the target subcategory based on the category description information of the target subcategory and the image description information corresponding to each image graph; and determining the image parameters corresponding to the target subcategory based on the image adjustment parameters corresponding to the target image graph.
[0136] First, a set of image adjustment parameters under the "target role category" are generated based on the reference role graph. Each set of image adjustment parameters is input into the game engine, and the game engine renders the virtual model corresponding to the target role category according to the image adjustment parameters to obtain the virtual role graph corresponding to the target role category.
[0137] Then, an image description information such as "smile", "calm", "cold and stern" is added to each rendered virtual character graph, and the image description information can be generated by a pre-defined sub-category label or automatically recognized by a classification model. By comparing the image description information of the rendered virtual character graph, the virtual character graph most consistent with the target sub-category style is found by selecting the category description information of the target sub-category (such as "lively" or "melancholy"). The image adjustment parameters used by these virtual character graphs are selected as the image parameters of the target sub-category.
[0138] In some embodiments, the image adjustment parameters corresponding to the target image graph are directly determined as the image parameters corresponding to the target sub-category.
[0139] In some embodiments, the image adjustment parameters include makeup parameters, and the image parameters corresponding to the target sub-category are determined based on the image adjustment parameters corresponding to the target image graph, and further include: determining a makeup adjustment strategy corresponding to a target makeup matched with the target sub-category based on a preset makeup matching relationship; and adjusting the image adjustment parameters corresponding to the target image graph based on the makeup adjustment strategy to obtain the target image parameters corresponding to the target sub-category.
[0140] In some embodiments, the image adjustment parameters include makeup parameters, and the image parameters corresponding to the target sub-category are determined based on the image adjustment parameters corresponding to the target image graph, and further include: determining a makeup adjustment strategy corresponding to a target makeup matched with the target sub-category based on a preset makeup matching relationship; and adjusting the image adjustment parameters corresponding to the target image graph based on the makeup adjustment strategy to obtain the target image parameters corresponding to the target sub-category.
[0141] In the above example, the virtual character categories include male, female, loli, and shota. Taking the target character category as an example, the above scheme is described.
[0142] In the female category, there are 8 sub-categories, including: beautiful, lively, pure, girl, doll face, gentle, sissy, and intellectual.
[0143] First, based on the reference role figure, a set of basic image parameters under the "adult female" is generated; then, through the fine-tuning methods such as eyebrows, eyes, and corners of the mouth, a plurality of image adjustment parameter samples are formed. Each set of adjustment parameters is input into the game engine to render a virtual role figure for the "adult female" body type. For each subcategory, its category description information is defined: lively: eyebrows are lightly raised, eyes are bright, corners of the mouth are slightly raised; intellectual: eyebrows are straight, eyes are focused, smile is implicit … and so on. For each virtual role figure, the system (or artificial, auxiliary classification model) analyzes its visual features and extracts the corresponding "image description information". The category description information of each subcategory and the image description information of each figure are semantically matched, for example: if the image description information of the virtual role figure A is: "beautiful", "national beauty", "attitude is thick and meaning is far away, and the skin is delicate and the flesh is uniform", according to the semantic matching model, the virtual role figure can be matched with the "beautiful" subcategory; then, the image adjustment parameters corresponding to the virtual role figure matched by each subcategory are used as the image parameters of the subcategory.
[0144] Through the above steps, the corresponding image parameters of each subcategory under each virtual role category are generated, and a multi-level and multi-style role image database is constructed.
[0145] For example, under the "adult female" category, the image parameters of the "beautiful", "lively", "pure", "girl", "doll face", "gentle", "sister", and "intelligent" subcategories can be generated through the above method, thereby meeting the diversified needs of users for role styles.
[0146] The embodiment of the present application realizes a method for determining the image parameters corresponding to the subcategories under a virtual role category, which comprises: determining the image parameters corresponding to at least two subcategories under a target role category based on at least two image adjustment parameters, comprising: generating an image figure of the image adjustment parameters under the target role category based on the image adjustment parameters and a second virtual model corresponding to the target role category; determining the category description information of the target subcategory and the image description information corresponding to each image figure of the image adjustment parameters under the target role category, wherein the target subcategory is any one of the at least two subcategories; determining a target image figure matched with the target subcategory based on the category description information of the target subcategory and the image description information corresponding to each image figure; and determining the image parameters corresponding to the target subcategory based on the image adjustment parameters corresponding to the target image figure. In the present scheme, by semantically matching the visual features of the image figure with the category description information of the subcategory, the image figure can be more accurately classified into the corresponding subcategory, thereby improving the accuracy of classification. At the same time, the method can adapt to the diversity of different role categories and subcategories, flexibly extend to new role types and styles, and improve the adaptability of the system.
[0147] When building the character image database, ensuring the emotional inclination balance of subcategories under the virtual character category is crucial for improving user experience. If the number of subcategories with positive inclinations (such as "sunshine" and "dignified") and negative inclinations (such as "melancholy" and "dejected") is found to be imbalanced, measures should be taken to adjust. Here are some optimization suggestions:
[0148] Dig up new subcategory labels: Identify potential subcategory labels by analyzing text descriptions and gallery content. For example, new labels for positive inclinations may include "vibrant" and "sweet", while new labels for negative inclinations may include "cold" and "lonely".
[0149] Adjust gallery content: Increase or decrease the number of character images with specific emotional inclinations in the gallery to achieve visual balance. For example, increase the number of "vibrant" style character images and decrease the number of "cold" style character images.
[0150] Through the above measures, the emotional inclination of the character subcategory can be effectively balanced, and the diversity of the character database and the personalized experience of the user can be improved.
[0151] Finally, manual screening is performed to remove unqualified image parameters due to not meeting the subcategory, not being aesthetically pleasing, or changing after makeup, etc., to ensure the quality of the final image parameters for each subcategory under each virtual character category.
[0152] Step 106: Based on the character information and image parameters of the reference character, generate a character image database.
[0153] In an optional implementation of this embodiment, based on the character information and target image parameters of the reference character, a character image database is generated, including:
[0154] Construct a keyword text library, wherein the keyword text library includes first type text and second type text, the first type text is the category description information corresponding to at least two subcategories under the target character category, and the second type text is the character information of the reference character;
[0155] Construct the corresponding relationship between the first type text and the image parameters, and the corresponding relationship between the second type text and the image parameters, and generate a character image database.
[0156] In this embodiment, we manually maintain a keyword text library for each image parameter and divide it into two types of text to build a character image database:
[0157] The first type of text is a description related to the characteristics of the subcategory, including idioms, poems, long text descriptions, etc., which are related to the characteristics of the subcategory.
[0158] The second type of text is the role information of the reference role, which can include star names, nicknames, film / literature / historical characters, etc.
[0159] In constructing the role image database, first, the first type of text is sorted and classified, and a corresponding relationship is established with the image parameters corresponding to the subcategories described above. For example, under the "cheng female" body type, we establish a corresponding relationship between the descriptive text in the text library and the image parameters corresponding to the subcategory of the matching subcategory. For example, the descriptive text "cheng female pure", and the image parameters corresponding to the "pure" subcategory under the "cheng female" are established.
[0160] For the second type of text, we directly correspond the text with a clear corresponding image such as star names to the image parameters of the corresponding star; for texts such as film characters, historical characters, etc. without a clear image, we first consider the star who plays the role in the film, and then try to match a star with a similar temperament, and then correspond to the image parameters of the star.
[0161] For the second type of text, the following optimization strategies are adopted to ensure that the text and the image parameters can be accurately corresponded:
[0162] For texts such as star names or nicknames that have a real and fixed image, the system can directly establish a one-to-one mapping between the text and the image parameters of the corresponding star;
[0163] For texts of film, literature or historical characters (whose standard image may not be clear), we use a hierarchical mapping strategy:
[0164] Priority mapping: if the role is played by a specific star in a film, the image parameters of the star in the film are directly mapped as the mapping target;
[0165] Secondary mapping: if the role has no specific actor, but can be matched with a star with similar temperament and style, the image parameters of the reference star are mapped as the mapping target.
[0166] Through this "priority-secondary" mapping mechanism, it is ensured that all second type of texts (including roles without a clear prototype image) can be stably associated with available image parameters, thereby enriching the coverage and expressiveness of the role image database.
[0167] In the embodiments of the present specification, by constructing a keyword text library and accurately corresponding it to the image parameters, significant advantages are brought in many aspects: first, by establishing a direct mapping relationship between the first type of text (i.e. subcategory style tags such as "pure", "sunshine", "sister") and the image parameters of the corresponding subcategory, the system can realize the semantic Accurate correspondence of image parameters. Such parameter mapping mechanism ensures that the generated character image not only meets the expected style semantically, but also is highly consistent in visual performance, thereby greatly improving the accuracy and efficiency of the style refinement process. Secondly, cross-modal retrieval can be achieved through the second type of text (star name, film / historical character name), supporting users to quickly obtain the corresponding face shaping results with only keywords or familiar characters, enhancing the interactive experience.
[0168] In an optional implementation of the embodiment, before generating the character image database, the method further includes:
[0169] determining at least one key element corresponding to the preset virtual character, and description text used to describe each key element;
[0170] combining the description text corresponding to the at least one key element to obtain at least one combined text; and constructing a correspondence between each combined text and an image parameter.
[0171] Among them, the key element refers to the facial feature element that cannot be replaced flexibly later, such as "eyes", "nose", "face shape", and a set of representative feature description texts (such as "big eyes", "fleshy nose", "round face", etc.) are maintained for each element.
[0172] Combined description text generation: randomly or according to requirements, combine these description texts to form a combined text such as "big eyes, fleshy nose, round face".
[0173] Combined text and image parameter mapping relationship construction: select and correspond a set of image parameters for each combined text to form a mapping relationship.
[0174] In runtime, when the player inputs the key element combination description (such as "big eyes + round face"), the system does not need to modify each attribute one by one, but directly extracts the corresponding image parameters from the corresponding database as the output, ensuring the overall style harmony and beauty.
[0175] In some embodiments, the scheme also maintains a synonym mapping table for facial features, which is used to improve the accuracy of natural language input analysis by the system. For example, non-standard expressions such as "big eyes" are mapped to the standard word "big eyes". This method is equivalent to normalizing the text preprocessing step, automatically replacing based on a dictionary or dictionary mapping table, improving the consistency and robustness of text analysis. This not only simplifies the synonym processing process, but also ensures that user input can be accurately identified as a facial feature keyword already in the system, thereby ensuring high precision and stability of subsequent parameter mapping.
[0176] In the embodiments of the present specification, by introducing the "key element + combined text + parameter mapping" mechanism before generating the role image database, high-fidelity control of core facial structures such as eyes, nose and face shape is achieved. When the player proposes a demand such as "big eyes + round face", the system can directly find the corresponding preset combined text and randomly select the face shaping parameters that have been verified for aesthetics and structure coordination, thus avoiding the imbalance of the face caused by later replacement, and maintaining the consistency of response speed and experience. In addition, with the help of synonym processing functions, such as "big eyes" automatically mapping to "big eyes", the robustness of natural language analysis is enhanced. This mechanism does not require complex program intervention, and only needs to maintain the text and parameter mapping library, which can efficiently support new combinations and small batch updates, ensuring that the virtual role library has multiple advantages such as structural coordination, aesthetic guarantee and flexible expansion.
[0177] In an optional implementation of the present embodiment, before generating the role image database, the following steps are further included:
[0178] Obtaining at least one image description text; for each image description text, performing feature extraction on the image description text to obtain the text features corresponding to the image description text; and constructing a corresponding relationship between the text features corresponding to each image description text and the image parameters.
[0179] In the present embodiment, in order to cope with the situation of insufficient player input information, a role image parameter library storing the corresponding relationship between text features and image parameters is constructed. First, some typical image description texts are collected, and a pre-trained sentence vector model (such as Sentence-BERT, Sentence Bidirectional Encoder Representations from Transformers) is used to convert these image description texts into text features. Then, these text features are screened to ensure that they have sufficient difference in the semantic space, so as to ensure the representativeness of each image description text in semantics. Subsequently, each image description text is associated with the image parameters that meet its description.
[0180] In the embodiments of the present specification, in order to cope with the situation of insufficient player input information or inability to extract specific information, a role image parameter library storing the corresponding relationship between text features and image parameters is constructed. When the player input cannot be directly parsed, the system can select the text feature closest to the input semantics from the library and recommend the corresponding image parameters. In this way, by introducing a backup solution, the system can still provide reasonable output when facing ambiguous or missing input, enhancing the stability and reliability of the system.
[0181] The following is described in conjunction with the accompanying drawings Figure 2With the application of the database construction method provided in the specification in the game as an example, the database construction method is further described. Among them, Figure 2 A flow chart of the processing process of a database construction method provided by an embodiment of the specification is shown, which specifically includes the following steps.
[0182] Step 202: Obtain a first image set composed of star images.
[0183] Specifically, a reference database containing more than 900 star names is maintained in advance as a target list for image collection. Based on this list, image materials corresponding to these star names can be automatically captured from public networks through crawler technology. After obtaining an image set composed of photos of multiple stars, first, the face key point detection algorithm is used to process each photo to extract its facial key point information. If 68 complete key points are successfully detected, the image is retained; if the detection result is less than 68 key points, it is determined as a non-frontal face image or poor quality, and is excluded. Subsequently, the image set after preliminary screening is screened again to further exclude images with obvious occlusion, strong shadow or excessive posture deflection, and only high-quality photos with frontal face orientation, no occlusion and uniform illumination are retained. On this basis, the corresponding key point information in each retained image is used to automatically calculate and generate a face box region for accurately cropping the star's face image region. After cropping each retained image, the first image set is obtained.
[0184] Step 204: Obtain a second image set, the images in the second image set are generated based on a model.
[0185] Specifically, through network query and large model feedback, 300+ words and English descriptions of facial features are obtained, such as unkempt eyebrows: unkempt eyebrows, almond-shaped eyes: almond-shaped eyes, etc. Then, various facial features are arranged and combined, and based on the arranged and combined facial features, a prompt word is generated. The prompt word is input into the standard stablediffusion model to generate corresponding image data, and the prompt word is input into the white-xl model specialized in ancient style to generate corresponding image data, and the second image set is composed of these two parts of image data.
[0186] Step 206: Based on the first image set and the second image set, obtain a target image set for constructing a role image database.
[0187] Step 208: Input the target image set into the "image-based face modeling" model to obtain the image parameters under each virtual role category.
[0188] Specifically, there are four virtual role categories in the game scene: adult male, adult female, shoujo, and loli, so we need to generate the image parameters of the four virtual role categories respectively by face modeling.
[0189] For the first image set, first, the first image set is divided into a female image set and a male image set according to gender. For the female image set, the female image set is input into the face modeling by image model corresponding to the adult female to generate the image parameters of the adult female category; and the female image set is input into the face modeling by image model corresponding to the loli to generate the image parameters of the loli category. For the male image set, the male image set is input into the face modeling by image model corresponding to the adult male to generate the image parameters of the adult male category; and the male image set is input into the face modeling by image model corresponding to the shoujo to generate the image parameters of the shoujo category.
[0190] For the second image set, the second image set is input into the face modeling by image model corresponding to the adult female to generate the image parameters of the adult female category; the second image set is input into the face modeling by image model corresponding to the loli to generate the image parameters of the loli category; the second image set is input into the face modeling by image model corresponding to the adult male to generate the image parameters of the adult male category; and the second image set is input into the face modeling by image model corresponding to the shoujo to generate the image parameters of the shoujo category.
[0191] In this way, the image parameters of the adult male category, the adult female category, the shoujo category, and the loli category are obtained by integrating the image parameters of the above two parts.
[0192] The image parameters in the embodiment include facial feature parameters and makeup parameters.
[0193] The process of generating the image parameters of the adult female category by the face modeling by image model corresponding to the adult female is described as an example. First, the image regions of the facial features of the female images in the female image set are extracted, as well as the skin color and the lip color of the characters; the facial feature parameters of the facial features are predicted by the neural network corresponding to the adult female category; the skin color is input into a classification network for skin color to obtain the color category of the skin color, and the lip color is input into a classification network for lip color to obtain the color category of the lip color; and then the color category of the skin color and the color category of the lip color are input into a classification network for predicting makeup to obtain the makeup parameters.
[0194] In addition, the generated five facial feature parameters need to be corrected by the face modeling by image. The correction process is: inputting the generated five facial feature parameters into a game engine or a 3D renderer, rendering the five facial feature parameters into a virtual character corresponding to the virtual model based on the five facial feature parameters and the virtual model corresponding to the female; performing key point detection on the virtual character image to automatically identify each facial key point of the virtual character image; calculating the spatial distance between the five facial features according to the detected key points; comparing the actually calculated distance with the standard template distance of the female body type, and if the distance deviates from the standard template distance, applying a statistical shape model (such as a Point Distribution Model) or a regression model to compensate for the five facial feature parameters to obtain corrected five facial feature parameters.
[0195] After the above steps, the image parameters corresponding to the male, the image parameters corresponding to the female, the image parameters corresponding to the boy, and the image parameters corresponding to the girl are obtained. In real applications, the generated image parameters need to be further screened. Specifically, after all the image parameters are generated, the image parameters of all virtual character categories (male, female, boy, and girl) are imported into a game engine, and the engine renders a virtual model based on each set of image parameters to generate a corresponding virtual character image. These virtual character images are used as visual references, and a selector evaluates each image in terms of aesthetic degree, structural coordination, style consistency, and the like. For each virtual character category, one set of image parameters that is relatively optimal in visual evaluation is finally retained as the image parameters under the category.
[0196] Step 210: For each virtual character category, adjusting the image parameters under the virtual character category to generate image adjustment parameters under the virtual character category; and determining the image parameters corresponding to the subcategory under the virtual character category based on the image adjustment parameters.
[0197] The specific process of adjusting the image parameters and the specific process of determining the image parameters of each subcategory under each virtual character category based on the adjusted image parameters can be referred to the foregoing description of the database construction method.
[0198] Finally, manual screening is performed to remove image parameters that are unqualified due to not meeting the subcategory, not being aesthetically pleasing, or changing in effect after makeup, and the like, to ensure the quality of the final image parameters of each subcategory under each virtual character category.
[0199] Corresponding to the method embodiments described above, the present specification also provides embodiments of a database construction apparatus, Figure 3 A structural schematic diagram of a database construction apparatus provided by one embodiment of the present specification is shown. As shown in the figure, Figure 3 The apparatus includes:
[0200] The acquisition module 302 is configured to acquire a reference role graph, wherein the reference role graph comprises a reference role;
[0201] The first generation module 304 is configured to generate an image parameter of the reference role under a target role category based on the reference role graph, wherein the target role category is determined from a plurality of preset virtual role categories;
[0202] The second generation module 306 is configured to generate a role image database based on role information and the image parameter of the reference role.
[0203] Optionally, before acquiring the reference role graph, the apparatus further comprises a first obtaining module configured to:
[0204] acquire an image set, wherein the image set comprises a plurality of object images; extract an object image from the image set, perform key point detection on the object image, and obtain key point information of a target object; and based on the key point information, intercept a face image of the target object to obtain the reference role graph.
[0205] Optionally, before acquiring the reference role graph, the apparatus further comprises a second obtaining module configured to:
[0206] acquire image description information; determine a target prompt word corresponding to the image description information based on the image description information; generate image data corresponding to the image description information based on the target prompt word; and obtain the reference role graph based on the image data.
[0207] Optionally, the first generation module 304 is further configured to:
[0208] extract at least one face basic attribute corresponding to the reference role and at least one face additional attribute corresponding to the reference role from the reference role graph; generate a basic parameter corresponding to each face basic attribute of the reference role under the target role category based on the at least one face basic attribute; generate an additional parameter of the reference role under the target role category based on the at least one face additional attribute; and generate the image parameter of the reference role under the target role category based on the basic parameter and the additional parameter.
[0209] Optionally, the first generation module 304 is further configured to, before generating the image parameter of the reference role under the target role category based on the basic parameter and the additional parameter, generate a virtual role graph of the reference role under the target role category based on the basic parameter and a first virtual model corresponding to the target role category; perform key point detection on the virtual role graph to determine each face key point of the virtual role graph; determine distances between each face basic attribute in the virtual role graph based on the each face key point; and correct the basic parameter corresponding to each face basic attribute of the reference role under the target role category based on the distances between the each face basic attribute to obtain a corrected basic parameter.
[0210] Optionally, the first generation module 304 is further configured to: before generating the role image database based on the role information of the reference role and the image parameters, adjusting the image parameters to generate at least two image adjustment parameters; and determining the image parameters corresponding to at least two sub-categories of the target role category based on the at least two image adjustment parameters.
[0211] Optionally, the first generation module 304 is further configured to:
[0212] generate an image graph of the image adjustment parameters in the target role category based on the image adjustment parameters and the second virtual role corresponding to the target role category; determine the category description information of the target sub-category and the image description information corresponding to each image graph of the image adjustment parameters in the target role category, wherein the target sub-category is any one of the at least two sub-categories; determine a target image graph matched with the target sub-category based on the category description information of the target sub-category and the image description information corresponding to each image graph; and determine the image parameters corresponding to the target sub-category based on the image adjustment parameters corresponding to the target image graph.
[0213] Optionally, the second generation module 306 is further configured to:
[0214] construct a keyword text library, wherein the keyword text library includes first texts and second texts, the first texts are category description information corresponding to at least two sub-categories of the target role category, and the second texts are role information of the reference role; construct a corresponding relationship between the first texts and the image parameters, and a corresponding relationship between the second texts and the image parameters, and generate the role image database.
[0215] Optionally, the second generation module 306 is further configured to: before generating the role image database, determine at least one key element corresponding to a preset virtual role and description text used to describe each key element; combine the description text corresponding to the at least one key element to obtain at least one combined text; and construct a corresponding relationship between each combined text and the image parameters.
[0216] Optionally, the second generation module 306 is further configured to: before generating the role image database, obtain at least one image description text; for each image description text, perform feature extraction on the image description text to obtain text features corresponding to the image description text; and construct a corresponding relationship between the text features corresponding to each image description text and the image parameters.
[0217] The above is a schematic solution of the database construction apparatus of the embodiment. It should be noted that the technical solution of the database construction apparatus and the technical solution of the database construction method described above belong to the same concept, and the details of the technical solution of the database construction apparatus which are not described in detail can be referred to the description of the technical solution of the database construction method.
[0218] Figure 4 A structural block diagram of a computing device 400 according to one embodiment of the present specification is shown. The components of the computing device 400 include, but are not limited to, a memory 410 and a processor 420. The processor 420 is connected to the memory 410 through a bus 430, and a database 450 is used to save data.
[0219] The computing device 400 also includes an access device 440, which enables the computing device 400 to communicate via one or more networks 460. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 440 can include one or more of any type of network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC).
[0220] In one embodiment of the present specification, the above-mentioned components of the computing device 400 and other components not shown in the Figure 4 may be connected to each other, for example, through a bus. It should be understood that Figure 4 The structural block diagram of the computing device shown is only for the purpose of example, and is not a limitation on the scope of the present specification. Those skilled in the art can add or replace other components as needed.
[0221] The computing device 400 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other type of mobile device, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 400 can also be a mobile or stationary server.
[0222] The processor 420 is configured to execute computer-executable instructions to implement the steps of the database construction method described above.
[0223] The above is a schematic solution of the computing device of the embodiment. It should be noted that the technical solution of the computing device and the technical solution of the database construction method described above belong to the same concept, and the details of the technical solution of the computing device that are not described in detail can be referred to the description of the technical solution of the database construction method.
[0224] An embodiment of the present specification further provides a computer-readable storage medium, which stores computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the steps of the database construction method described above.
[0225] The above is a schematic solution of the computer-readable storage medium of the embodiment. It should be noted that the technical solution of the storage medium and the technical solution of the database construction method described above belong to the same concept, and the details of the technical solution of the storage medium that are not described in detail can be referred to the description of the technical solution of the database construction method.
[0226] An embodiment of the present specification further provides a computer program, and the computer program causes a computer to perform the steps of the database construction method described above when the computer program is executed in the computer.
[0227] The above is a schematic solution of the computer program of the embodiment. It should be noted that the technical solution of the computer program and the technical solution of the database construction method described above belong to the same concept, and the details of the technical solution of the computer program that are not described in detail can be referred to the description of the technical solution of the database construction method.
[0228] The above-described embodiments of the application have several aspects, no single one of which is solely responsible for the application's desirable attributes. Without limiting the scope of the application as expressed by the claims which follow, some further embodiments make these aspects even more useful. Other embodiments can result in less desirable attributes.
[0229] The computer readable medium can include any entity or apparatus capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, software distribution medium, etc. It should be noted that the computer readable medium can include appropriate contents according to the requirements of patent practice, for example, according to the patent practice in some regions, the computer readable medium does not include electrical carrier signals and telecommunication signals.
[0230] It should be noted that for the foregoing method embodiments, the purposes of brief description are to express them as a combination of a series of acts, but those skilled in the art should know that the embodiments of the present specification are not limited by the order of the described acts, because according to the embodiments of the present specification, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the acts and modules involved are not necessarily all the embodiments of the present specification.
[0231] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0232] The preferred embodiments of the present specification disclosed above are only used to help explain the present specification. The alternative embodiments do not describe all the details and limit the application to the specific embodiments described. Obviously, according to the content of the embodiments of the present specification, many modifications and changes can be made. The present specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present specification, so that those skilled in the art can well understand and use the present specification. The present specification is limited by the claims and their entire scope and equivalents.
Claims
1. A database construction method characterized by comprising: The method comprises: obtaining a reference role graph, wherein the reference role graph comprises a reference role; generating an image parameter of the reference role under a target role category based on the reference role graph, wherein the target role category is determined from a plurality of preset virtual role categories; generating a role image database based on role information of the reference role and the image parameter.
2. The method of claim 1, wherein, Before the reference role graph is obtained, the method further comprises: obtaining an image set, wherein the image set comprises a plurality of object images; extracting the object images from the image set, performing key point detection on the object images, and obtaining key point information of a target object; based on the key point information, the face image of the target object is intercepted to obtain a reference role graph.
3. The method of claim 1, wherein, Before the reference role graph is obtained, the method further comprises: obtaining image description information; based on the image description information, determining a target prompt word corresponding to the image description information; based on the target prompt word, generating image data corresponding to the image description information; based on the image data, obtaining a reference role graph.
4. The method of claim 1, wherein, The method further comprises: extracting at least one face basic attribute corresponding to the reference role and at least one face additional attribute corresponding to the reference role from the reference role graph; based on the at least one face basic attribute, generating a basic parameter corresponding to each face basic attribute of the reference role under the target role category; based on the at least one face additional attribute, generating an additional parameter of the reference role under the target role category; based on the basic parameter and the additional parameter, generating an image parameter of the reference role under the target role category.
5. The method of claim 4, wherein, Before the basic parameter and the additional parameter are generated, the method further comprises: based on the basic parameter and a first virtual model corresponding to the target role category, generating a virtual role graph of the reference role under the target role category; performing key point detection on the virtual role graph to determine each face key point of the virtual role graph; based on the each face key point, determining the distance between each face basic attribute in the virtual role graph; based on the distance between each face basic attribute, correcting the basic parameter corresponding to each face basic attribute of the reference role under the target role category to obtain a corrected basic parameter.
6. The method of claim 1, wherein, Before the role image database is generated, the method further comprises: adjusting the image parameter to generate at least two image adjustment parameters; based on the at least two image adjustment parameters, determining an image parameter corresponding to at least two subcategories under the target role category.
7. The method of claim 6, wherein, The method further comprises: based on the image adjustment parameter and a second virtual role corresponding to the target role category, generating an image graph of the image adjustment parameter under the target role category; determine category description information of a target sub-category, and image description information corresponding to the image of each image adjustment parameter under the target character category, wherein the target sub-category is any one of the at least two sub-categories; determine a target image corresponding to the target sub-category based on the category description information of the target sub-category and the image description information corresponding to each image; determine the image parameter corresponding to the target sub-category based on the image adjustment parameter corresponding to the target image.
8. The method of claim 6, wherein, The generating of the character image database based on the character information of the reference character and the target image parameter includes: constructing a keyword text library, wherein the keyword text library includes first texts and second texts, the first texts are category description information corresponding to the at least two sub-categories under the target character category, and the second texts are character information of the reference character; constructing a corresponding relationship between the first texts and the image parameters, and a corresponding relationship between the second texts and the image parameters, and generating the character image database.
9. The method of claim 8, wherein, Before the generating of the character image database, the method further includes: determining at least one key element corresponding to a preset virtual character, and description texts for describing each key element; combining the description texts corresponding to the at least one key element to obtain at least one combined text; constructing a corresponding relationship between each combined text and the image parameter.
10. The method according to claim 8 or 9, characterized in that, Before the generating of the character image database, the method further includes: obtaining at least one image description text; performing feature extraction on the image description text to obtain a text feature corresponding to the image description text for each image description text; constructing a corresponding relationship between the text feature corresponding to each image description text and the image parameter.
11. A database construction apparatus characterized by comprising: includes: an acquisition module configured to acquire a reference character graph, wherein the reference character graph includes a reference character; a first generation module configured to generate an image parameter of the reference character under a target character category based on the reference character graph, wherein the target character category is determined from a plurality of preset virtual character categories; a second generation module configured to generate a character image database based on character information of the reference character and the image parameter.
12. A computing device, comprising: includes: a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions, and the computer executable instructions, when executed by the processor, realize the steps of the database construction method of any one of claims 1-10.
13. A computer-readable storage medium, characterized in that, which stores computer executable instructions, and the computer executable instructions, when executed by the processor, realize the steps of the database construction method of any one of claims 1-10.
14. A computer program product, characterised in that, includes computer programs / instructions, and the computer programs / instructions, when executed by the processor, realize the steps of the database construction method of any one of claims 1-10.