Virtual image generation method, computer terminal, storage medium and product

By obtaining input text data and posture templates, and using adapter models and control models to generate virtual images, the problem in the existing technology of being difficult to generate custom avatars conveniently and quickly while saving shooting costs is solved, and the generation of custom avatars is realized.

CN120599089APending Publication Date: 2025-09-05ALIBABA (CHINA) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410252024.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-05
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

It is difficult with existing technologies to generate customized avatars quickly and conveniently while saving shooting costs.

Method used

By obtaining input text data and posture templates, a virtual image is generated using an adapter model, a control model, and an image generation model, including determining the adapter model and the control model to match the user's style and posture requirements, and using a large-scale deep learning model for image generation.

Benefits of technology

It can save shooting costs while generating customized avatars conveniently and quickly, thus meeting the personalized needs of users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599089A_ABST
    Figure CN120599089A_ABST
Patent Text Reader

Abstract

The invention discloses a virtual image generation method, a computer terminal, a storage medium and a product, and relates to the field of large model technology and image generation. The method comprises the steps that input text data and a posture template are obtained, the text data at least comprise prompt information having an incidence relation with generation of a virtual image, and the posture template is used for representing posture information of the virtual image; based on first text data in the input text data, determining an adapter model, the first text data being used for representing prompt information having an association relationship with the style and the attribute of the virtual image; a control model matched with the posture template is determined, and the control model is used for controlling the posture of the virtual image; and generating a virtual image based on second text data in the input text data by using the adapter model, the control model and the image generation model, the second text data being used for representing prompt information having an association relationship with the content of the virtual image. According to the method and the device, the technical problem that the shooting cost is difficult to save while the user-defined head portrait is generated conveniently and quickly in the related technology is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of large-scale model technology and image generation, and more specifically, to a method for generating a virtual image, a computer terminal, a storage medium, and a product. Background Art

[0002] On social media, a user profile picture with a precise image and exquisite style has become a basic need for communication and interaction. Therefore, how to automatically generate AI profile pictures with customized styles and postures quickly and conveniently while saving users' shooting costs is an urgent problem to be solved. Summary of the Invention

[0003] The embodiments of the present application provide a method for generating a virtual image, a computer terminal, a storage medium, and a product to at least solve the technical problem in the related art that it is difficult to generate a customized avatar conveniently and quickly while saving shooting costs.

[0004] According to one aspect of an embodiment of the present application, a method for generating a virtual image is provided, comprising: obtaining input text data and a posture template, wherein the text data includes at least prompt information associated with the generation of the virtual image, and the posture template is used to represent the posture information of the virtual image; determining an adapter model based on first text data in the input text data, wherein the first text data is used to represent prompt information associated with the style and attributes of the virtual image; determining a control model that matches the posture template, wherein the control model is used to control the posture of the virtual image; and generating a virtual image based on second text data in the input text data using the adapter model, the control model, and the image generation model, wherein the second text data is used to represent prompt information associated with the content of the virtual image.

[0005] According to another aspect of an embodiment of the present application, a avatar generation method is also provided, including: in response to an input operation on an operation interface, displaying input text data and a posture template on the operation interface, wherein the text data at least includes prompt information associated with the generation of a virtual image, and the posture template is used to represent the posture information of the virtual image; in response to the avatar generation operation on the operation interface, displaying the virtual image on the operation interface, wherein the virtual image is generated based on second text data in the input text data using an adapter model, a control model matching the posture template, and an image generation model, the control model is used to control the posture of the virtual image, the adapter model is a model determined based on first text data in the input text data, the first text data is used to represent prompt information associated with the style and attributes of the virtual image, and the second text data is used to represent prompt information associated with the content of the virtual image.

[0006] According to another aspect of an embodiment of the present application, a avatar generation method is also provided, including: obtaining input text data and a posture template by calling a first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter includes text data and a posture template, the text data includes prompt information associated with the generation of a virtual image, and the posture template is used to characterize the posture information of the virtual image; determining an adapter model based on the first text data in the input text data, wherein the first text data is used to characterize prompt information associated with the style and attributes of the virtual image; determining a control model that matches the posture template, wherein the control model is used to control the posture of the virtual image; using the adapter model, the control model and the image generation model, generating a virtual image based on the second text data in the input text data, wherein the second text data is used to characterize prompt information associated with the content of the virtual image; outputting the virtual image by calling a second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the virtual image.

[0007] According to one aspect of an embodiment of the present application, a device for generating a virtual image is provided, including: an acquisition module for acquiring input text data and a posture template, wherein the text data at least includes prompt information associated with the generation of the virtual image, and the posture template is used to represent the posture information of the virtual image; a first determination module for determining an adapter model based on first text data in the input text data, wherein the first text data is used to represent prompt information associated with the style and attributes of the virtual image; a second determination module for determining a control model that matches the posture template, wherein the control model is used to control the posture of the virtual image; a generation module for generating a virtual image based on second text data in the input text data using the adapter model, the control model and the image generation model, wherein the second text data is used to represent prompt information associated with the content of the virtual image.

[0008] According to another aspect of an embodiment of the present application, an avatar generation device is provided, including: a first display module for responding to an input operation on an operation interface, and displaying input text data and a posture template on the operation interface, wherein the text data at least includes prompt information associated with the generation of a virtual image, and the posture template is used to represent the posture information of the virtual image; a second display module for responding to an avatar generation operation on the operation interface, and displaying the virtual image on the operation interface, wherein the virtual image is generated based on second text data in the input text data using an adapter model, a control model matching the posture template, and an image generation model, the control model is used to control the posture of the virtual image, the adapter model is a model determined based on first text data in the input text data, the first text data is used to represent prompt information associated with the style and attributes of the virtual image, and the second text data is used to represent prompt information associated with the content of the virtual image.

[0009] According to another aspect of an embodiment of the present application, an avatar generation device is also provided, including: an acquisition module for acquiring input text data and a posture template by calling a first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter includes text data and a posture template, the text data includes prompt information associated with the generation of a virtual image, and the posture template is used to represent the posture information of the virtual image; a first determination module for determining an adapter model based on the first text data in the input text data, wherein the first text data is used to represent prompt information associated with the style and attributes of the virtual image; a second determination module for determining a control model matching the posture template, wherein the control model is used to control the posture of the virtual image; a generation module for generating a virtual image based on the second text data in the input text data using the adapter model, the control model and the image generation model, wherein the second text data is used to represent prompt information associated with the content of the virtual image; an output module for outputting the virtual image by calling a second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the virtual image.

[0010] In an embodiment of the present application, input text data and a posture template are obtained, wherein the text data at least includes prompt information associated with the generation of a virtual image, and the posture template is used to represent the posture information of the virtual image; an adapter model is determined based on first text data in the input text data, wherein the first text data is used to represent prompt information associated with the style and attributes of the virtual image; a control model matching the posture template is determined, wherein the control model is used to control the posture of the virtual image; and a virtual image is generated based on second text data in the input text data using the adapter model, the control model, and the image generation model, wherein the second text data is used to represent prompt information associated with the content of the virtual image. It is easy to notice that the adapter model can be determined based on the first text data in the input text data, and the control model that matches the posture template can be determined, so that the adapter model, the control model and the image generation model can be used to generate a virtual image based on the second text data in the input text data. Since the virtual image is generated based on the input text data and the posture template, the shooting cost can be saved. At the same time, since the input text data contains prompt information that is associated with the generation of the virtual image, that is, the generated virtual image can be customized based on the text data, so that the customized avatar can be generated conveniently and quickly while saving shooting costs, thereby solving the technical problem in the related art that it is difficult to generate customized avatars conveniently and quickly while saving shooting costs.

[0011] It is easy to notice that the above general description and the following detailed description are merely for the purpose of exemplifying and explaining the present application, and do not constitute a limitation of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0013] Figure 1 is a schematic diagram of an application scenario of a method for generating a virtual image according to an embodiment of the present application;

[0014] Figure 2 is a flowchart of a method for generating a virtual image according to Example 1 of the present application;

[0015] Figure 3 is a schematic diagram of a data-enhanced image generation method according to an embodiment of the present application;

[0016] Figure 4 is a flowchart of the avatar generation method according to Example 2 of the present application;

[0017] Figure 5 is a flowchart of the avatar generation method according to Example 3 of the present application;

[0018] Figure 6 is a schematic diagram of a device for generating a virtual image according to Example 4 of the present application;

[0019] Figure 7 is a schematic diagram of an avatar generation device according to Example 5 of the present application;

[0020] Figure 8 is a schematic diagram of an avatar generation device according to Example 6 of the present application;

[0021] Figure 9 This is a structural block diagram of a computer terminal according to an embodiment of the present application. DETAILED DESCRIPTION

[0022] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0023] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0024] The technical solution provided in this application is mainly implemented using large-scale model technology. The large model here refers to a deep learning model with large-scale model parameters, which can usually contain hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. The large model can also be called a cornerstone model / foundation model (Foundation Model). The large model is pre-trained by large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks and has good generalization capabilities, such as large-scale language models (LLMs) and multi-modal pre-training models.

[0025] It should be noted that when the large model is actually used, the pre-trained model can be fine-tuned through a small number of samples, so that the large model can be applied to different tasks. For example, the large model can be widely used in natural language processing (NLP), computer vision, speech processing and other fields. Specifically, it can be applied to computer vision tasks such as visual question answering (VQA), image caption (IC), image generation, etc. It can also be widely used in natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. Therefore, the main application scenarios of the large model include but are not limited to digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc. In the embodiment of the present application, the data processing is performed by a stable diffusion model (also called a stable diffusion model) in the image generation scenario as an example for explanation.

[0026] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following interpretations:

[0027] Long-term evolution model: also known as the Long Range Low Power model, abbreviated as LoRa model, is a low-power, long-distance communication wireless transmission technology suitable for application scenarios that require long-distance transmission and low power consumption, such as Internet of Things devices, sensor networks, radio scheduling, etc.

[0028] Adapter model: is a software design pattern used to convert the interface of a class into another interface expected by the client. In the image processing scenario, the adapter pattern can be used to connect different image processing libraries or tools so that they can be compatible with each other and work together. The adapter pattern can convert the interface of one image processing library into the interface required by another image processing library, so that they can work together seamlessly.

[0029] Diffusion control model: also known as the ControlNet model, is a network model used for industrial automation control systems. It is a real-time communication network. The ControlNet network uses a high-speed transmission communication method to achieve high-performance data transmission and control in industrial environments.

[0030] Stable diffusion model: also known as stable diffusion model, is a mathematical model used in image processing. It can be used to simulate the diffusion process of stable substances or energy in images. This model is often used in applications such as image denoising, edge detection, and image enhancement.

[0031] Example 1

[0032] According to an embodiment of the present application, a method for generating an avatar is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0033] Considering the huge number of model parameters of large models and the limited computing resources of mobile terminals, Figure 1 This is a schematic diagram of an application scenario of a method for generating a virtual image according to an embodiment of the present application. The above-mentioned avatar generation method provided in the embodiment of the present application can be applied to Figure 1 The application scenarios shown are not limited to this. Figure 1 In the illustrated application scenario, the large model is deployed on a server 10. The server 10 can be connected to one or more client devices 20 via a local area network, a wide area network, the Internet, or other types of data networks. The client devices 20 herein may include, but are not limited to, smartphones, tablet computers, laptops, PDAs, personal computers, smart home devices, and in-vehicle devices. The client devices 20 can interact with users via a graphical user interface to access the large model and thereby implement the methods provided in the embodiments of the present application.

[0034] In an embodiment of the present application, a system composed of a client device and a server can perform the following steps: the client device executes to receive input instructions on a graphical user interface, and the server executes to obtain input text data and a posture template, wherein the text data at least includes prompt information associated with the generation of a virtual image, and the posture template is used to represent the posture information of the virtual image; based on the first text data in the input text data, an adapter model is determined, wherein the first text data is used to represent prompt information associated with the style and attributes of the virtual image; a control model that matches the posture template is determined, wherein the control model is used to control the posture of the virtual image; and using the adapter model, the control model and the image generation model, a virtual image is generated based on the second text data in the input text data, wherein the second text data is used to represent prompt information associated with the content of the virtual image. It should be noted that if the operating resources of the client device can meet the deployment and operation conditions of the large model, the embodiment of the present application can be performed in the client device.

[0035] Under the above operating environment, this application provides Figure 2 The method for generating the virtual image shown. Figure 2 is a flow chart of the method for generating a virtual image according to Example 1 of the present application, such as Figure 2 As shown, the method may include the following steps:

[0036] Step S202: obtaining input text data and a posture template, wherein the text data at least includes prompt information associated with the generation of the avatar, and the posture template is used to represent the posture information of the avatar.

[0037] The above-mentioned input text data can be data for generating a virtual image, wherein the virtual image can be data of an artificial intelligence image (also called an Artificial Intelligence image, referred to as an AI image). Optionally, the above-mentioned AI image can be a person image, an animal image, a cartoon image, etc. In this application, the AI ​​image is taken as a person image as an example for illustration. Specifically, the above-mentioned text data may include but is not limited to: style adapter model trigger words (for example, style lora model trigger words), face lora model trigger words, and text image prompt words, wherein the style lora model trigger words can be used to enhance the style of the generated AI image, the face lora model trigger words can be used to enhance the ability to retain the character ID corresponding to the AI ​​image, the character ID can represent the identity of the character, and the text image prompt words are used to represent the specific content of the generated AI avatar, for example, it can be a girl with long hair, etc.

[0038] The above-mentioned posture template can be an AI avatar posture template image, which can be used to generate a virtual image with a specific posture, for example, an AI avatar.

[0039] In an optional embodiment, the user can input text data in any form, for example, by giving a text description of the virtual image to be generated, that is, input prompt information associated with the virtual image to be generated, for example, the size of the virtual image, the hair color of the virtual image, the expression of the virtual image, etc., and input a posture template in any form, that is, input an AI avatar posture template image. Optionally, the present application does not impose any specific restrictions on the type of the input AI avatar posture template image. Specifically, it can be a smiling AI avatar posture template, a focused thinking AI avatar posture template, a welcoming posture AI avatar posture template, a confident and upright AI avatar posture template, a friendly greeting AI avatar posture template, a surprised expression AI avatar posture template, a professional work AI avatar posture template, a thinking AI avatar posture template, a straight AI avatar posture template, a lively and cute AI avatar posture template, etc.

[0040] Step S204: determining an adapter model based on first text data in the input text data, wherein the first text data is used to represent prompt information associated with the style and attributes of the virtual image.

[0041] The above-mentioned first text data can be a prompt word for determining the adapter model that the user wants to use. Optionally, the first text data can include at least one of the following types of prompt words: prompt words for the style lora model, prompt words for the face lora model. The prompt words for the style lora model can be words of a specific style, for example, cartoon style, doomsday style, Chinese style, etc. The prompt words for the face lora model can be words of a specific user ID, for example, oval face, melon seed face, square face, or directly indicate that it is a certain person's face, etc.

[0042] The above-mentioned adapter model can be at least one of a face lora model and a style lora model, wherein the face lora model is a model for identifying and analyzing facial features, which is usually used in face recognition systems. The model can identify various features of the face, such as eyes, nose, mouth, etc., thereby helping to identify different individuals. The style lora model is a model for analyzing and identifying artistic styles, which is usually used for the classification and style recognition of artworks. The style lora model can identify different artistic styles, such as impressionism, realism, abstract art, etc.

[0043] In an optional embodiment, in order to meet the user's customized needs for different styles and different faces, multiple different types of adapter models can be pre-configured. For example, multiple different style LoRa models and multiple different face LoRa models can be pre-configured. Therefore, after the user enters text data, the adapter model the user wants to use can be directly determined based on the first text data in the input text data. Optionally, assuming that the content contained in the first text data indicates that an oval face is desired to be generated, an adapter model that can generate an oval face can be found from multiple adapter models.

[0044] In another optional embodiment, when determining the adapter model based on the first text data in the input text data, the following steps may be performed:

[0045] First, it is necessary to train the adapter model, using the collected prompt words and the corresponding data set, that is, the first text data, to train the adapter model so that it can generate results similar to the face lora model based on the prompt words.

[0046] Secondly, the model parameters need to be adjusted. That is, during the training of the adapter model, the model parameters and hyperparameters may need to be continuously adjusted to obtain better performance.

[0047] Finally, evaluate the model effect. That is, you need to evaluate the effect of the trained adapter model to check whether it can generate expected results based on the prompt words. If the effect is not good, you may need to readjust the model or replace it with another model for training.

[0048] Step S206: Determine a control model that matches the posture template, wherein the control model is used to control the posture of the virtual image.

[0049] The above-mentioned control model can be a controlnet model based on human posture (also called human pose) and depth (also called depth). Common control models include proportional integral derivative controller, fuzzy controller, neural network controller, etc.

[0050] In an optional embodiment, in order to meet the different needs of users, different control models can be pre-configured, for example, different controlnet models can be pre-configured, so that after determining the type of posture template, the control model that matches the posture template can be directly determined from the pre-configured control model. Optionally, assuming that the posture template is a posture template for generating cartoon images, the control model that matches the posture template for generating cartoon images can be determined. Assuming that the posture template is for generating real-life images, the control model that matches the posture template for generating real-life images can be determined, so that the posture of the virtual image can be controlled based on the control model.

[0051] In another optional embodiment, when determining a control model that matches a posture template, the following steps may be followed:

[0052] First, it is necessary to determine the posture template to be matched, including the description of the target posture, the expected motion trajectory, etc.

[0053] Secondly, the appropriate control model can be selected according to the characteristics and requirements of the posture template.

[0054] Furthermore, according to the actual situation and system characteristics, the parameters of the selected control model are adjusted so that it can better match the posture template.

[0055] Thirdly, simulation software or actual systems can be used for simulation and verification to check whether the control model can effectively match the posture template.

[0056] Afterwards, the control model can be optimized and adjusted based on the simulation and verification results to make it better suited to the requirements of the posture template.

[0057] Finally, the optimized control model can be applied to the actual system for actual control and adjustment, and the control model can be continuously optimized and improved to better match the posture template.

[0058] Step S208: Generate a virtual image based on second text data in the input text data using the adapter model, the control model, and the image generation model, wherein the second text data is used to represent prompt information associated with the content of the virtual image.

[0059] The image generation model described above is an artificial intelligence model that can generate realistic facial avatars based on input features and parameters. This model is typically based on deep learning technology and is trained using large facial datasets to learn facial features and expressions. The avatar generation model can generate avatars with various features and styles, including those of different ages, genders, and races. It can also generate avatars with specific expressions, postures, and clothing as needed. Optionally, in this application, the avatar generation model is described using the stable diffusion model as an example.

[0060] In an optional embodiment, after determining the adapter model and the control model, the control model can be used to control the operation of the adapter model and the avatar generation model. Specifically, the control model controls the adapter model and the avatar generation model to generate an avatar based on the second text data in the input text data. Alternatively, the adapter model's weights can be integrated into the avatar generation model, and the avatar generation model can be controlled by the control model to achieve the purpose of generating an avatar based on the second text data in the input text data.

[0061] In another optional embodiment, to generate a virtual image using the adapter model, the control model, and the avatar generation model, the following steps may be followed:

[0062] First, an adapter model is used to extract key information from the prompt information associated with the image generation model. The adapter model can help convert the prompt information into a format suitable for the input of the avatar generation model for subsequent processing.

[0063] Next, the extracted key information can be input into the control model, which can adjust the parameters of the image generation model based on the prompt information to ensure that the generated avatar meets the user's expectations and requirements. The control model can also fine-tune the avatar during the generation process to obtain better results.

[0064] Finally, the prompt information processed by the adapter model and the control model can be input into the image generation model to generate a virtual image. The avatar generation model can use deep learning technology to generate an avatar that meets the requirements based on the input prompt information.

[0065] Optionally, in this way, by utilizing the adapter model, the control model and the avatar generation model, a virtual image that meets the user's requirements can be generated according to the prompt information associated with the avatar generation model.

[0066] In an embodiment of the present application, input text data and a posture template are obtained, wherein the text data at least includes prompt information associated with the generation of a virtual image, and the posture template is used to represent the posture information of the virtual image; an adapter model is determined based on first text data in the input text data, wherein the first text data is used to represent prompt information associated with the style and attributes of the virtual image; a control model matching the posture template is determined, wherein the control model is used to control the posture of the virtual image; and a virtual image is generated based on second text data in the input text data using the adapter model, the control model, and the image generation model, wherein the second text data is used to represent prompt information associated with the content of the virtual image. It is easy to notice that the adapter model can be determined based on the first text data in the input text data, and the control model that matches the posture template can be determined, so that the adapter model, the control model and the image generation model can be used to generate a virtual image based on the second text data in the input text data. Since the virtual image is generated based on the input text data and the posture template, the shooting cost can be saved. At the same time, since the input text data contains prompt information that is associated with the generation of the virtual image, that is, the generated virtual image can be customized based on the text data, so that the customized avatar can be generated conveniently and quickly while saving shooting costs, thereby solving the technical problem in the related art that it is difficult to generate customized avatars conveniently and quickly while saving shooting costs.

[0067] In the above embodiment of the present application, the adapter model includes an attribute adapter model and a style adapter model; based on the first text data in the input text data, the adapter model is determined, including: based on the first sub-data in the first text data, the attribute adapter model is determined, wherein the first sub-data is used to represent the object attribute information that has an associated relationship with the virtual image; based on the second sub-data in the first text data, the style adapter model is determined, wherein the second sub-data is used to represent the style information that has an associated relationship with the virtual image.

[0068] The above-mentioned first sub-data is used to represent object attribute information that is associated with the attribute adapter model. Specifically, it may include facial feature data, facial recognition algorithm, facial database, facial adaptation parameters, facial recognition results, etc., wherein the above-mentioned object attribute information may be user ID (Identification), user name, and other information that can indicate the user's identity.

[0069] The above-mentioned attribute adapter model can be a model for adapting and matching different facial data. This model is usually based on deep learning and computer vision technology, and can identify and compare the features of different faces, and then adapt and match them, thereby realizing functions such as face recognition, face verification and face comparison.

[0070] Optionally, the attribute adapter model typically includes the following main components:

[0071] Face detector: used to detect the face area in the image and perform preliminary face positioning and extraction.

[0072] Feature extractor: used to extract feature vectors from facial images, usually based on deep learning technology, and can convert facial images into recognizable feature vectors.

[0073] Feature matcher: used to compare and match the feature vectors of different faces, thereby realizing face recognition and comparison functions.

[0074] Adapter: used to adapt and match different facial data, usually including image preprocessing, feature extraction and feature matching functions.

[0075] The second sub-data is used to represent style information associated with the style adapter model, and specifically may include data representing the clothing style of the person in the picture, data representing the background style of the person in the picture, and the like.

[0076] The above-mentioned style adapter model can be a model for adapting and matching different style data. It can identify and compare the features of different styles based on deep learning and computer vision technology, and then adapt and match them.

[0077] In an optional embodiment, different types of attribute adapter models and different types of style adapter models can be pre-configured, for example, different types of face lora models and different types of style lora models, so that when the adapter model needs to be determined, the object attribute information associated with the attribute adapter model can be directly determined from the pre-configured different types of attribute adapter models, that is, the attribute adapter model corresponding to the first sub-data, and the style information associated with the style adapter model can be directly determined from the pre-configured different types of style adapter models, that is, the style adapter model corresponding to the second sub-data. Optionally, assuming that the object attribute information contained in the first sub-data indicates that a melon-seed face needs to be generated, the attribute adapter model that can generate a melon-seed face can be determined. assuming that the style information contained in the second sub-data indicates that a cartoon-style avatar needs to be generated, the style adapter model that can generate a cartoon-style avatar can be directly determined.

[0078] Optionally, to determine the corresponding attribute adapter model based on the object attribute information that has an associated relationship with the attribute adapter model, an attribute matching algorithm can be used to identify and match specific object attribute information, and then find the corresponding attribute adapter model based on the matching results. A common method is to use a feature extraction algorithm to extract feature vectors of the object attribute information, and then use a similarity matching algorithm to compare these feature vectors with the feature vectors of the attribute adapter model to find the most similar attribute adapter model.

[0079] Furthermore, to determine the corresponding style adapter model using style information associated with the style adapter model, a style matching algorithm can be used to identify and match specific style information, and then find the corresponding style adapter model based on the matching results. A common method is to use a feature extraction algorithm to extract feature vectors of the style information, and then use a similarity matching algorithm to compare these feature vectors with the feature vectors of the style adapter model to find the most similar style adapter model.

[0080] In the above embodiment of the present application, the method also includes: obtaining a first original image containing a preset image; rotating the first original image to obtain a rotated image, wherein the preset image in the rotated image faces a preset direction; segmenting and adjusting parameters of the rotated image to obtain a first training image; annotating the first training image to obtain first annotated data of the first training image; and adjusting the image generation model using the first training image and the first annotated data to obtain an attribute adapter model.

[0081] The first original image may be an unprocessed image containing a face region. Optionally, the first original image may be obtained from an image database or uploaded by a user.

[0082] The above-mentioned preset direction can be a unified rotation direction set by a technician in this field. Optionally, since the directions of the faces in the first original image acquired may be different, it is difficult to obtain a neat and unified face area. Therefore, it is necessary to rotate the first original image so that the faces in the first original image are oriented in the same direction. Optionally, the preset direction is not specifically limited in this application. In this application, the preset direction is taken as an example for explanation as the forward direction.

[0083] The first annotated data may be data with refined labels, wherein the refined labels may be used to indicate the image content in the first training image, for example, a man wearing glasses, a smiling woman, and the like.

[0084] In an optional embodiment, after obtaining a first original image containing a face area, the first original image can be rotated using an image rotation model based on orientation judgment and a face refinement rotation method based on face detection and a key point model to rotate the first original image to a positive direction. Furthermore, the rotated image can be segmented to segment the face area from the first original image, and the skin parameters of the face area can be adjusted, that is, the skin of the face area is beautified to obtain a first training image. Optionally, the first training image can be labeled, that is, the first training image is marked with a refined label of the training image to obtain first labeled data. Optionally, in the process of labeling the first training image, it can be labeled manually or by a related computational model. In this application, there is no specific limitation on the labeling method. After obtaining the first labeled data, the first training image and the first labeled data can be used to adjust the avatar generation model, that is, the avatar generation model is trained, and the parameters of the avatar generation model are adjusted to optimize the performance of the avatar generation model to obtain an attribute adapter model.

[0085] In the above embodiment of the present application, the first annotation data includes first content annotation data and object attribute annotation data, wherein the first content annotation data is used to represent the content of the preset image, and the object attribute annotation data is used to represent the attributes of the preset image.

[0086] The first content annotation data may be data carrying image tags, for example, a beautiful woman wearing sunglasses, smiling, etc. Optionally, the first content annotation data may be obtained by annotating the content of the first training image using a relevant text annotation model.

[0087] The above-mentioned object attribute annotation data can be data used to represent facial attribute information such as gender, age, expression, facial accessories, etc. of the person contained in the first training image. Optionally, the object attribute annotation data can be obtained by processing the first training image using a facial attribute model.

[0088] In an optional embodiment, the content of the first training image can be annotated using a relevant text annotation model to obtain first content annotation data, and the first training image can be processed using a relevant face attribute model to obtain object attribute annotation data, wherein the above-mentioned text annotation model is a model for annotating or classifying specific content in the text. It can identify entities, emotions, events and other content in the text and annotate them for subsequent analysis and application. The above-mentioned face attribute model can be a model for identifying and analyzing facial features. It can identify facial attributes such as age, gender, race, expression, glasses, beard, etc. This model is usually based on deep learning and computer vision technology. It learns the characteristics of facial attributes by training a large amount of face image data, so that it can accurately identify and analyze facial attributes. Furthermore, the first content annotation data and the object attribute annotation data can be summarized to obtain first annotation data.

[0089] In the above embodiment of the present application, the method also includes: obtaining a second original image of a preset style; performing data augmentation on the second original image to obtain a data augmented image; cropping the data augmented image to obtain a second training image; annotating the second training image to obtain second annotated data of the second training image; and adjusting the image generation model using the second training image and the second annotated data to obtain a style adapter model.

[0090] The second original image of the preset style may be a photo of a given style, wherein the type of photo may be set by those skilled in the art according to needs, for example, a certain number of portraits with the same or similar clothing and background.

[0091] In an optional embodiment, a second original image of a preset style may be given by a person skilled in the art, or a second original image of a preset style may be obtained by other means, and data enhancement may be performed on the second original image. Optionally, the above-mentioned data enhancement may include enhancing the brightness, clarity, and pixel value of the second original image to obtain a data-enhanced image. Further, the data-enhanced image may be cropped to obtain a second training image, and the second training image may be annotated to obtain second annotated data for the second training image, that is, the image style, type and other features of the second training image are annotated on the basis of the second training image. For example, "wearing sunglasses" and "laughing" may be annotated. Optionally, in the process of annotating the second training image, manual annotation or annotation may be performed through a related computational model. In this application, no specific limitation is imposed on the annotation method. After obtaining the second annotated data, the second training image and the second annotated data may be used to adjust the image generation model, that is, the performance of the image generation model may be optimized to obtain a style adapter model.

[0092] In the above embodiment of the present application, data enhancement is performed on the second original image to obtain a data-enhanced image, including: obtaining training text data corresponding to the second original image, wherein the training text data is used to describe the style of the second original image; extracting the original human body posture of the second original image; using a first training control model and an image generation model that match the original human body posture to generate a target generated image based on the training text data; extracting image key points of the target generated image; and using a second training control model and an image generation model that match the image key points to generate a data-enhanced image based on the training text data, the image key points, and the second original image.

[0093] The above-mentioned training text data may be text prompt words used to describe the style of the second original image, for example, the training text data may be Japanese-style girl, wearing Chinese-style clothes, etc.

[0094] The above-mentioned original human body posture can be the skeleton posture of the second original image, that is, the skeleton posture of the character contained in the second original image, which is used to represent the action posture of the character contained in the second original image at this time. Optionally, assuming that the original human body posture shows that the human body's knee bones are curled up, and the thigh bones are flush with the calf bones, it can be inferred that the current posture of the human body is sitting. Therefore, the first training control model and image generation model matching "sitting" can be determined, and the target generation image can be generated based on the training text data.

[0095] The first training control model mentioned above can be a controlnet model that matches the skeleton posture of the second original image.

[0096] The target generated image may be a result in which the facial position and overall posture conform to the input image.

[0097] The above-mentioned image key points can be facial key points. Optionally, the number of image key points is not specifically limited in this application. In this application, an example of 68 image key points is used for illustration.

[0098] The second training control model mentioned above can be a controlnet model that matches the image key points.

[0099] In an optional embodiment, when performing data enhancement on the second original image to obtain a data enhanced image, the text prompt word corresponding to the second original image can be first obtained, and the original human body posture of the second original image can be extracted using a related feature extraction model, that is, the skeleton posture of the second original image is extracted, and the controlnet model and image generation model that match the skeleton posture of the second original image are used to generate a corresponding target generated image based on the training text data. Furthermore, the image key points of the target generated image can be extracted, and the controlnet model and image generation model that match the image key points are used to generate a data enhanced image based on the training text data, image key points and the second original image, that is, an image that can maintain the overall content and style of the input image and has the facial shape and facial features of the corresponding ID is generated.

[0100] Figure 3 is a schematic diagram of a data enhanced image generation method according to an embodiment of the present application, such as Figure 3 As shown in , when determining the data enhancement image, it can be divided into two processing steps, such as Figure 3 As shown in , the thinner arrow represents the first processing process, and the thicker arrow represents the second processing process. The original human body posture can be extracted from the second original image, that is, the posture detection is performed on the second original image, and the diffusion control model and image generation model that match the original human body posture are used to generate a target generated image based on the training text data. Further, the image key points of the target generated image can be extracted, that is, facial recognition is performed on the facial landmarks of the second original image, so that the diffusion control model and the image generation model can be used to generate a data enhanced image. Optionally, the obtained data enhanced image can be annotated to obtain second annotated data, and the second training image can be obtained by cropping the data enhanced image. The image generation model can be adjusted using the second training image and the second annotated data, so as to determine the style adapter model from the adapter model. Furthermore, the adapter model can be fused with the image generation model based on the weight to obtain a weight fusion model, and the weight fusion model can be controlled by the control model, and a virtual image can be generated based on the second text data.

[0101] In the above embodiment of the present application, a first training control model and an image generation model that match the original human body posture are used to generate a target generated image based on the training text data, including: using the first training control model and the image generation model to generate an initial generated image based on the training text data; fusing the preset image image with the initial generated image to obtain the target generated image.

[0102] The aforementioned preset image may be a face-attached image that is finally obtained by selecting from training images using a face quality assessment model.

[0103] In an optional embodiment, when generating a target generated image, the first training control model and the image generation model can be used to generate an initial generated image based on the training text data. Furthermore, the face-attached image finally obtained by selecting the training image through the face quality assessment model can be fused with the initial generated image to obtain the target generated image. Optionally, after obtaining the initial generated image, the face-attached image is fused with the initial generated image through the face quality assessment model. That is, after two image generation processes, the final target generated image can be made more accurate and more in line with user needs.

[0104] In the above embodiment of the present application, the preset image is an image obtained by screening the training face image using the face quality assessment model.

[0105] A face quality assessment model can be a model used to evaluate the quality of facial images, and is typically used in applications such as face recognition, face detection, and face verification. The model can evaluate the quality of facial images by analyzing factors such as clarity, lighting, posture, and occlusion, and provide corresponding scores or recommendations.

[0106] In an optional embodiment, the face quality assessment model may be used to screen the training face images, thereby screening out an image with higher quality from the training face images, and determining the image as the preset face image.

[0107] In the above embodiment of the present application, an adapter model, a control model and an image generation model are used to generate a virtual image based on the second text data in the input text data, including: fusing the weight of the adapter model with the image generation model to obtain a weight fusion model; and using the control model to control the weight fusion model to generate a virtual image based on the second text data.

[0108] The weights of the above-mentioned adapter model can be parameters of the adapter model, wherein the weights of the adapter model usually depend on the specific application scenario and data set, and are usually determined by training the model. During the training process, the weights are continuously adjusted to maximize the performance and accuracy of the model. These weights include weights and bias terms connecting different layers, which determine how the model converts input data into output. During the training process, the weights are adjusted according to the gradient of the loss function to make the model's prediction results closer to the actual results. Therefore, the weights of the adapter model change dynamically and need to be determined through training.

[0109] In an optional embodiment, when using a control model to control the adapter model and the image generation model to generate a virtual image based on the second text data in the input text data, the weight of the adapter model can be first fused with the image generation model to obtain a weight fusion model. Further, the control model can be used to control the weight fusion model to generate a virtual image based on the second text data.

[0110] In another optional embodiment, to fuse the weights of the adapter model with the image generation model, the following steps may be used:

[0111] First, ensure that the architectures of the adapter model and the image generation model are similar so that their weights can be fused.

[0112] Next, we need to load the weights of the adapter model and the image generation model.

[0113] Again, the weight of each model is adjusted as needed, for example, normalizing the weight or weighting different models.

[0114] Furthermore, the weights of the two models need to be fused, and a weighted average method can be used, that is, the weights of the two models are weighted averaged to obtain a new fusion weight.

[0115] Afterwards, the fused weights can be used to construct a new weight fusion model, thus obtaining a weight fusion model of the adapter model and the image generation model.

[0116] Finally, the new weight fusion model can be evaluated and tested to ensure that the fused model performs well in the actual task.

[0117] Optionally, after obtaining the weight fusion model, the control model may be used to control the weight fusion model to generate a virtual image based on the second text data.

[0118] In the above embodiment of the present application, a control model is used to control a weight fusion model to generate a virtual image based on the second text data, including: using a control model to control a weight fusion model to generate an initial avatar based on the second text data; fusing a preset image with the initial avatar to obtain a virtual image.

[0119] In an optional embodiment, when the weight fusion model is controlled by the control model and the initial avatar is generated based on the second text data, the following steps may be adopted:

[0120] Data preprocessing: That is, the second text data needs to be preprocessed, including text cleaning, word segmentation, part-of-speech tagging, etc., in order to convert the text into a data format that can be processed by the model.

[0121] Constructing a weight fusion model, that is, using a deep learning framework or a custom algorithm to construct a weight fusion model, which can receive multiple inputs, including the second text data and other related information (such as the feature vector of the virtual image, etc.), and output the generated virtual image.

[0122] Designing a control model, that is, designing a control model that can receive the second text data as input and output a set of control parameters for controlling the behavior of the weight fusion model. These control parameters may include weight adjustment, feature selection, generator output, etc.

[0123] Training and optimizing the model, that is, using the labeled dataset to train and optimize the weight fusion model and the control model so that they can effectively generate the initial avatar.

[0124] Generate an initial avatar, that is, use the trained weight fusion model and control model, input the second text data, obtain corresponding control parameters through the control model, and then use the weight fusion model to generate the initial avatar.

[0125] Optionally, after the initial image is generated, the preset image may be fused with the initial avatar to obtain a virtual image.

[0126] In the above embodiment of the present application, obtaining an input posture template includes: outputting multiple posture templates; and in response to a selection instruction of a user to select multiple posture templates, determining the posture template corresponding to the selection instruction as the posture template.

[0127] In an optional embodiment, when obtaining the input posture template, multiple different posture templates can be displayed on the display interface of the client, and the user can issue a selection instruction in any form to determine the posture template the user wants to select from multiple different posture templates.

[0128] In the above embodiment of the present application, the method also includes: outputting an adapter model; obtaining modified text data in response to a user's modification instruction to modify the first text data; determining the adapter model based on the modified text data, and generating a virtual image based on the second text data using the new adapter model, control model and image generation model.

[0129] In an optional embodiment, the user can adjust the adapter model, that is, the user can give modification instructions in any form to instruct the modification of the first text data, so that the adapter model the user wants can be determined based on the modified text data, and the new adapter model, control model and image generation model can be used to generate a virtual image based on the second text data.

[0130] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0131] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.

[0132] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0133] Example 2

[0134] According to another aspect of the embodiment of the present application, a method for generating an avatar is further provided. Figure 4 This is a flow chart of the avatar generation method according to Example 2 of the present application. Figure 4 As shown, the method may include the following steps:

[0135] Step S402: In response to an input operation on the operation interface, input text data and a posture template are displayed on the operation interface, wherein the text data at least includes prompt information associated with the generation of the virtual image, and the posture template is used to represent the posture information of the virtual image.

[0136] Step S404: In response to the avatar generation operation performed on the operation interface, a virtual image is displayed on the operation interface, wherein the virtual image is generated based on the second text data in the input text data using an adapter model, a control model that matches the posture template, and an image generation model. The control model is used to control the posture of the virtual image, the adapter model is a model determined based on the first text data in the input text data, the first text data is used to represent prompt information that is associated with the style and attributes of the virtual image, and the second text data is used to represent prompt information that is associated with the content of the virtual image.

[0137] In an optional embodiment, the user can give input instructions in any form on the graphical user interface of the client 40, that is, on the operation interface, wherein the client 40 is connected to the server 41 via a network, and after the server receives the input instructions, it can control the display of the input text data and the posture template on the operation interface. Optionally, when the user gives an avatar generation instruction in any form on the graphical user interface of the client 40, the server 41 can use the adapter model, the control model matching the posture template and the image generation model to generate a virtual image based on the second text data in the input text data, and control the display of the virtual image on the operation interface, wherein the control model is used to control the posture of the virtual image, the adapter model is a model determined based on the first text data in the input text data, the first text data is used to represent prompt information associated with the adapter model, and the second text data is used to represent prompt information associated with the image generation model.

[0138] Example 3

[0139] According to another aspect of the embodiment of the present application, a method for generating an avatar is further provided. Figure 5 This is a flow chart of the avatar generation method according to Example 3 of the present application. Figure 5 As shown, the method may include the following steps:

[0140] Step S502: Obtain input text data and a posture template by calling a first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter includes text data and a posture template, the text data includes prompt information associated with the generation of a virtual image, and the posture template is used to represent the posture information of the virtual image.

[0141] Step S504: determining an adapter model based on first text data in the input text data, wherein the first text data is used to represent prompt information associated with the style and attributes of the virtual image.

[0142] Step S506: Determine a control model that matches the posture template, wherein the control model is used to control the posture of the virtual image.

[0143] Step S508: Generate a virtual image based on second text data in the input text data using the adapter model, the control model, and the image generation model, wherein the second text data is used to represent prompt information associated with the content of the virtual image.

[0144] Step S510: Outputting the virtual image by calling the second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the virtual image.

[0145] In an optional embodiment, a first interface on the client 50 can be called to obtain input text data and a posture template, and a second interface on the client 50 can be called to output a virtual image, wherein the client 50 and the server 51 are connected via a network, and the server 51 can determine an adapter model based on the first text data in the input text data, wherein the first text data is used to represent prompt information associated with the adapter model, and determine a control model that matches the posture template, wherein the control model is used to control the posture of the virtual image. Furthermore, the adapter model, the control model and the image generation model can be used to generate a virtual image based on the second text data in the input text data, wherein the second text data is used to represent prompt information associated with the image generation model.

[0146] Example 4

[0147] According to an embodiment of the present application, a device for implementing the above-mentioned method for generating a virtual image is also provided. Figure 6 is a schematic diagram of a device for generating a virtual image according to Example 4 of the present application, as shown in FIG. Figure 6 As shown, the device includes: an acquisition module 602 , a first determination module 604 , a second determination module 606 , and a generation module 608 .

[0148] Among them, the acquisition module 602 is used to obtain input text data and a posture template, wherein the text data at least includes prompt information associated with the generation of the virtual image, and the posture template is used to represent the posture information of the virtual image; the first determination module 604 is used to determine the adapter model based on the first text data in the input text data, wherein the first text data is used to represent the prompt information associated with the style and attributes of the virtual image; the second determination module 606 is used to determine the control model that matches the posture template, wherein the control model is used to control the posture of the virtual image; the generation module 608 is used to generate the virtual image based on the second text data in the input text data using the adapter model, the control model and the image generation model, wherein the second text data is used to represent the prompt information associated with the content of the virtual image.

[0149] In the above embodiment of the present application, the first determination module 604 includes: a first determination unit, used to determine the attribute adapter model based on the first sub-data in the first text data, wherein the first sub-data is used to represent the object attribute information that has an associated relationship with the virtual image; a second determination unit, used to determine the style adapter model based on the second sub-data in the first text data, wherein the second sub-data is used to represent the style information that has an associated relationship with the virtual image.

[0150] In the above embodiment of the present application, the device also includes: a second acquisition module, used to acquire a first original image containing a preset image; a rotation module, used to rotate the first original image to obtain a rotated image, wherein the preset image in the rotated image faces a preset direction; segmenting and parameter adjustment of the rotated image to obtain a first training image; a first annotation module, used to annotate the first training image to obtain first annotation data of the first training image; and a first adjustment module, used to adjust the image generation model using the first training image and the first annotation data to obtain an attribute adapter model.

[0151] In the above embodiment of the present application, the device also includes: a third acquisition module, used to obtain a second original image of a preset style; an enhancement module, used to perform data enhancement on the second original image to obtain a data-enhanced image; a cropping module, used to crop the data-enhanced image to obtain a second training image; a second labeling module, used to label the second training image to obtain second labeling data of the second training image; and a second adjustment module, used to adjust the image generation model using the second training image and the second labeling data to obtain a style adapter model.

[0152] In the above embodiment of the present application, the enhancement module includes: an acquisition unit for acquiring training text data corresponding to the second original image, wherein the training text data is used to describe the style of the second original image; a first extraction unit for extracting the original human body posture of the second original image; a first generation unit for generating a target generated image based on the training text data using a first training control model and an image generation model that match the original human body posture; a second extraction unit for extracting image key points of the target generated image; and a second generation unit for generating a data-enhanced image based on the training text data, image key points, and the second original image using a second training control model and an image generation model that match the image key points.

[0153] In the above embodiment of the present application, the first generation unit includes: a generation subunit, which is used to generate an initial generated image based on training text data using a first training control model and an image generation model; and a fusion subunit, which is used to fuse the preset image image with the initial generated image to obtain a target generated image.

[0154] In the above embodiment of the present application, the generation module 608 includes: a fusion unit for fusing the weight of the adapter model with the image generation model to obtain a weight fusion model; a control unit for controlling the weight fusion model using a control model to generate a virtual image based on the second text data.

[0155] In the above embodiment of the present application, the control unit includes: a first fusion subunit, which is used to control the weight fusion model using the control model to generate an initial avatar based on the second text data; and a first fusion subunit, which is used to fuse the preset image with the initial avatar to obtain a virtual image.

[0156] In the above embodiment of the present application, the acquisition module 602 includes: an output unit for outputting multiple posture templates; a third determination unit for responding to a user's selection instruction for selecting multiple posture templates, and determining that the posture template corresponding to the selection instruction is the posture template.

[0157] In the above embodiment of the present application, the device further includes: an output module for outputting the adapter model;

[0158] The modification module is used to obtain modified text data in response to a user's modification instruction to modify the first text data; the third determination module is used to determine a new adapter model based on the modified text data; and the second generation module is used to generate a virtual image based on the second text data using the new adapter model, control model and image generation model.

[0159] It should be noted that the acquisition module 602, the first determination module 604, the second determination module 606, and the generation module 608 correspond to steps S202 to S208 in Example 1. The examples and application scenarios implemented by the modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 1. It should be noted that the modules or units can be hardware components or software components stored in a memory and processed by one or more processors. The modules can also be part of the device and can run in the server 10 provided in Example 1.

[0160] It should be noted that the preferred implementation scheme involved in the above embodiments of this application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0161] Example 5

[0162] According to an embodiment of the present application, a device for implementing the above-mentioned avatar generation method is also provided. Figure 7 is a schematic diagram of an avatar generation device according to Example 5 of the present application, as shown in FIG. Figure 7 As shown, the device includes: a first display module 702 and a second display module 704.

[0163] Among them, the first display module 702 is used to respond to the input operation on the operation interface and display the input text data and posture template on the operation interface, wherein the text data at least includes prompt information associated with the generation of the virtual image, and the posture template is used to represent the posture information of the virtual image; the second display module 704 is used to respond to the avatar generation operation on the operation interface and display the virtual image on the operation interface, wherein the virtual image is generated based on the second text data in the input text data using an adapter model, a control model matching the posture template, and an image generation model, the control model is used to control the posture of the virtual image, the adapter model is a model determined based on the first text data in the input text data, the first text data is used to represent prompt information associated with the style and attributes of the virtual image, and the second text data is used to represent prompt information associated with the content of the virtual image.

[0164] It should be noted that the first display module 702 and the second display module 704 correspond to steps S402 to S404 in Example 2. The examples and application scenarios implemented by the modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 1. It should be noted that the modules or units can be hardware components or software components stored in a memory and processed by one or more processors. The modules can also be part of the device and can run in the server 10 provided in Example 1.

[0165] It should be noted that the preferred implementation scheme involved in the above embodiments of this application is the same as the scheme provided in Example 2, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 2.

[0166] Example 6

[0167] According to an embodiment of the present application, a device for implementing the above-mentioned avatar generation method is also provided. Figure 8 is a schematic diagram of an avatar generation device according to Example 6 of the present application, as shown in FIG. Figure 8 As shown, the apparatus includes: an acquisition module 802 , a first determination module 804 , a second determination module 806 , a generation module 808 , and an output module 810 .

[0168] Among them, the acquisition module 802 is used to obtain input text data and a posture template by calling a first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter includes text data and a posture template, the text data includes prompt information associated with the generation of the virtual image, and the posture template is used to represent the posture information of the virtual image; the first determination module 804 is used to determine the adapter model based on the first text data in the input text data, wherein the first text data is used to represent the prompt information associated with the style and attributes of the virtual image; the second determination module 806 is used to determine the control model that matches the posture template, wherein the control model is used to control the posture of the virtual image; the generation module 808 is used to generate the virtual image based on the second text data in the input text data using the adapter model, the control model and the image generation model, wherein the second text data is used to represent the prompt information associated with the content of the virtual image; the output module 810 is used to output the virtual image by calling a second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the virtual image.

[0169] It should be noted that the acquisition module 802, the first determination module 804, the second determination module 806, the generation module 808, and the output module 810 correspond to steps S502 to S510 in Example 3. The examples and application scenarios implemented by the modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned Example 1. It should be noted that the above-mentioned modules or units can be hardware components or software components stored in a memory and processed by one or more processors. The above-mentioned modules can also be run as part of the device in the server 10 provided in Example 1.

[0170] It should be noted that the preferred implementation scheme involved in the above embodiments of this application is the same as the scheme provided in Example 3, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 3.

[0171] Example 7

[0172] The embodiment of the present application can provide a computer terminal, which can be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the computer terminal can also be replaced by a terminal device such as a mobile terminal.

[0173] Optionally, in this embodiment, the computer terminal may be located in at least one network device among a plurality of network devices of a computer network.

[0174] In this embodiment, the above-mentioned computer terminal can execute the program code of the following steps in the avatar generation method: obtaining input text data and a posture template, wherein the text data at least includes prompt information associated with the generation of the virtual image, and the posture template is used to represent the posture information of the virtual image; determining an adapter model based on the first text data in the input text data, wherein the first text data is used to represent the prompt information associated with the style and attributes of the virtual image; determining a control model that matches the posture template, wherein the control model is used to control the posture of the virtual image; using the adapter model, the control model and the image generation model, generate a virtual image based on the second text data in the input text data, wherein the second text data is used to represent the prompt information associated with the content of the virtual image.

[0175] Optionally, Figure 9 This is a structural block diagram of a computer terminal according to an embodiment of the present application. Figure 9 As shown, the computer terminal A may include: one or more (only one is shown in the figure) processors 902, a memory 904, a storage controller, and a peripheral interface, wherein the peripheral interface is connected to a radio frequency module, an audio module and a display.

[0176] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the avatar generation method and device in the embodiments of the present application. The processor executes the software programs and modules stored in the memory to perform various functional applications and data processing, thereby implementing the above-mentioned avatar generation method. The memory can include high-speed random access memory and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory can further include memory remotely located relative to the processor, and these remote memories can be connected to terminal A via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0177] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtain input text data and a posture template, wherein the text data at least includes prompt information associated with the generation of the virtual image, and the posture template is used to represent the posture information of the virtual image; determine the adapter model based on the first text data in the input text data, wherein the first text data is used to represent the prompt information associated with the style and attributes of the virtual image; determine the control model that matches the posture template, wherein the control model is used to control the posture of the virtual image; use the adapter model, the control model and the image generation model to generate the virtual image based on the second text data in the input text data, wherein the second text data is used to represent the prompt information associated with the content of the virtual image.

[0178] Optionally, the processor may also execute the program code of the following steps: determining an attribute adapter model based on the first sub-data in the first text data, wherein the first sub-data is used to represent object attribute information associated with the virtual image; and determining a style adapter model based on the second sub-data in the first text data, wherein the second sub-data is used to represent style information associated with the virtual image.

[0179] Optionally, the processor may also execute the program code of the following steps: obtaining a first original image containing a preset image; rotating the first original image to obtain a rotated image, wherein the preset image in the rotated image faces a preset direction; segmenting and adjusting parameters of the rotated image to obtain a first training image; annotating the first training image to obtain first annotated data of the first training image; and adjusting the image generation model using the first training image and the first annotated data to obtain an attribute adapter model.

[0180] Optionally, the processor may also execute the program code of the following steps: obtaining a second original image of a preset style; performing data augmentation on the second original image to obtain a data augmented image; cropping the data augmented image to obtain a second training image; annotating the second training image to obtain second annotated data of the second training image; and adjusting the image generation model using the second training image and the second annotated data to obtain a style adapter model.

[0181] Optionally, the processor may also execute the program code of the following steps: obtaining training text data corresponding to the second original image, wherein the training text data is used to describe the style of the second original image; extracting the original human body posture of the second original image; generating a target generated image based on the training text data using a first training control model and an image generation model that match the original human body posture; extracting image key points of the target generated image; generating a data enhanced image based on the training text data, image key points, and the second original image using a second training control model and an image generation model that match the image key points.

[0182] Optionally, the processor may also execute the program code of the following steps: generating an initial generated image based on training text data using the first training control model and the image generation model; and fusing the preset image with the initial generated image to obtain a target generated image.

[0183] Optionally, the processor may further execute program code of the following steps: fusing the weight of the adapter model with the image generation model to obtain a weight fusion model; controlling the weight fusion model using the control model to generate a virtual image based on the second text data.

[0184] Optionally, the processor may further execute program code of the following steps: using the control model to control the weight fusion model to generate an initial avatar based on the second text data; and fusing the preset image with the initial avatar to obtain a virtual image.

[0185] Optionally, the processor may further execute program code of the following steps: outputting a plurality of posture templates; and in response to a selection instruction from a user to select the plurality of posture templates, determining the posture template corresponding to the selection instruction as the posture template.

[0186] Optionally, the processor may also execute the program code for the following steps: outputting an adapter model; obtaining modified text data in response to a user's modification instruction to modify the first text data; determining a new adapter model based on the modified text data; and generating a virtual image based on the second text data using the new adapter model, control model, and image generation model.

[0187] In an embodiment of the present application, input text data and a posture template are obtained, wherein the text data at least includes prompt information associated with the generation of a virtual image, and the posture template is used to represent the posture information of the virtual image; an adapter model is determined based on first text data in the input text data, wherein the first text data is used to represent prompt information associated with the style and attributes of the virtual image; a control model matching the posture template is determined, wherein the control model is used to control the posture of the virtual image; and a virtual image is generated based on second text data in the input text data using the adapter model, the control model, and the image generation model, wherein the second text data is used to represent prompt information associated with the content of the virtual image. It is easy to notice that the adapter model can be determined based on the first text data in the input text data, and the control model that matches the posture template can be determined, so that the adapter model, the control model and the image generation model can be used to generate a virtual image based on the second text data in the input text data. Since the virtual image is generated based on the input text data and the posture template, the shooting cost can be saved. At the same time, since the input text data contains prompt information that is associated with the generation of the virtual image, that is, the generated virtual image can be customized based on the text data, so that the customized avatar can be generated conveniently and quickly while saving shooting costs, thereby solving the technical problem in the related art that it is difficult to generate customized avatars conveniently and quickly while saving shooting costs.

[0188] Those skilled in the art will appreciate that the structure shown in the figure is for illustration only, and the computer terminal may also be a smart phone (such as an Android phone, iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, or other terminal devices. Figure 9 It does not limit the structure of the above electronic device. For example, the computer terminal A may also include Figure 9 More or fewer components (such as network interfaces, display devices, etc.) shown in, or with Figure 9 Different configurations shown.

[0189] A person skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0190] Example 8

[0191] The embodiment of the present application further provides a storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the avatar generation method provided in the first embodiment.

[0192] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.

[0193] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: obtaining input text data and a posture template, wherein the text data includes at least prompt information associated with the generation of a virtual image, and the posture template is used to represent the posture information of the virtual image; determining an adapter model based on first text data in the input text data, wherein the first text data is used to represent prompt information associated with the style and attributes of the virtual image; determining a control model that matches the posture template, wherein the control model is used to control the posture of the virtual image; and generating a virtual image based on second text data in the input text data using the adapter model, the control model, and the image generation model, wherein the second text data is used to represent prompt information associated with the content of the virtual image.

[0194] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0195] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0196] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0197] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0198] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0199] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0200] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A method for generating a virtual image, characterized in that: include: Acquiring input text data and a posture template, wherein the text data includes prompt information associated with the generation of the virtual image, and the posture template is used to represent the posture information of the virtual image; determining an adapter model based on first text data in the input text data, wherein the first text data is used to represent prompt information associated with the style and attributes of the virtual image; determining a control model that matches the posture template, wherein the control model is used to control the posture of the avatar; The virtual image is generated based on second text data in the input text data using the adapter model, the control model and the image generation model, wherein the second text data is used to represent prompt information associated with the content of the virtual image.

2. The method according to claim 1, characterized in that The adapter model includes an attribute adapter model and a style adapter model; and determining the adapter model based on the first text data in the input text data includes: determining the attribute adapter model based on first sub-data in the first text data, wherein the first sub-data is used to represent attribute information of an object associated with the virtual image; The style adapter model is determined based on second sub-data in the first text data, wherein the second sub-data is used to represent style information associated with the virtual image.

3. The method according to claim 2, characterized in that The method further comprises: Acquire a first original image containing a preset image; Rotating the first original image to obtain a rotated image, wherein the preset image in the rotated image faces a preset direction; Segmenting and adjusting parameters of the rotated image to obtain a first training image; Annotating the first training image to obtain first annotated data of the first training image; The image generation model is adjusted using the first training image and the first annotated data to obtain the attribute adapter model.

4. The method according to claim 3, characterized in that The first annotation data includes first content annotation data and object attribute annotation data, wherein the first content annotation data is used to represent the content of the preset image, and the object attribute annotation data is used to represent the attributes of the preset image.

5. The method according to claim 2, characterized in that The method further comprises: Obtain a second original image of a preset style; performing data enhancement on the second original image to obtain a data enhanced image; Cropping the data augmented image to obtain a second training image; annotating the second training image to obtain second annotated data of the second training image; The image generation model is adjusted using the second training image and the second annotated data to obtain the style adapter model.

6. The method according to claim 5, characterized in that The performing data enhancement on the second original image to obtain a data enhanced image includes: Acquiring training text data corresponding to the second original image, wherein the training text data is used to describe the style of the second original image; extracting an original human body posture from the second original image; Generate a target generated image based on the training text data using a first training control model that matches the original human body posture and the image generation model; Extracting image key points of the target generated image; The data-augmented image is generated based on the training text data, the image key points and the second original image using a second training control model that matches the image key points and the image generation model.

7. The method according to claim 6, characterized in that The method of generating a target generated image based on the training text data by using the first training control model and the image generation model that match the original human body posture includes: generating an initial generated image based on the training text data using the first training control model and the image generation model; The preset image is fused with the initial generated image to obtain the target generated image.

8. The method according to any one of claims 1 to 7, characterized in that The step of generating the virtual image based on the second text data in the input text data by using the adapter model, the control model, and the image generation model includes: fusing the weight of the adapter model with the image generation model to obtain a weight fusion model; The weight fusion model is controlled by the control model to generate the virtual image based on the second text data.

9. The method according to claim 8, characterized in that The step of controlling the weight fusion model by using the control model to generate the virtual image based on the second text data includes: Using the control model to control the weight fusion model, and generating an initial avatar based on the second text data; The preset image is merged with the initial avatar to obtain the virtual image.

10. The method according to any one of claims 1 to 7, characterized in that The step of obtaining an input posture template includes: Output multiple posture templates; In response to a selection instruction from a user to select the plurality of posture templates, the posture template corresponding to the selection instruction is determined to be the posture template.

11. The method according to any one of claims 1 to 7, characterized in that The method further comprises: outputting the adapter model; In response to a modification instruction from a user to modify the first text data, obtaining modified text data; determining a new adapter model based on the modified text data; The virtual image is generated based on the second text data using the new adapter model, the control model and the image generation model.

12. A method for generating an avatar, characterized in that: include: In response to an input operation performed on an operation interface, displaying input text data and a posture template on the operation interface, wherein the text data at least includes prompt information associated with the generation of an avatar, and the posture template is used to represent posture information of the avatar; In response to an avatar generation operation performed on the operation interface, the virtual image is displayed on the operation interface, wherein the virtual image is generated based on second text data in the input text data using an adapter model, a control model matching the posture template, and an image generation model, the control model being used to control the posture of the virtual image, the adapter model being a model determined based on first text data in the input text data, the first text data being used to represent prompt information associated with the style and attributes of the virtual image, and the second text data being used to represent prompt information associated with the content of the virtual image.

13. A method for generating an avatar, characterized in that: include: Acquiring input text data and a posture template by calling a first interface, wherein the first interface includes a first parameter, a parameter value of the first parameter includes the text data and the posture template, the text data includes prompt information associated with generation of an avatar, and the posture template is used to represent posture information of the avatar; determining an adapter model based on first text data in the input text data, wherein the first text data is used to represent prompt information associated with the style and attributes of the virtual image; determining a control model that matches the posture template, wherein the control model is used to control the posture of the avatar; generating the virtual image based on second text data in the input text data using the adapter model, the control model, and the image generation model, wherein the second text data is used to represent prompt information associated with the content of the virtual image; The virtual image is output by calling a second interface, wherein the second interface includes a second parameter, and a parameter value of the second parameter includes the virtual image.

14. A computer terminal, characterized in that: include: a memory storing an executable program; A processor, configured to run the program, wherein the program executes the method according to any one of claims 1 to 13 when running.

15. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored executable program, wherein when the executable program is run, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 13.

16. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 13.