Virtual image generation method and device, storage medium and program product
By extracting the appearance features in the input image and injecting the generation model, the problem of users in the prior art requiring detailed feature descriptions is solved, and the stability, consistency and diversity generation of virtual images are achieved.
Patent Information
- Application Number
- CN202411766286.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-03
- Publication Date
- 2025-05-02
AI Technical Summary
The existing virtual image generation method requires the user to pre-configure detailed feature descriptions as prompt words, resulting in limited user text expression ability, and the generated virtual image is difficult to match the user's needs, and image deformation or distortion is prone to occur during the generation process.
By extracting the appearance features of the reference object in the input image and injecting these features into the pre-trained generative model, the model is guided to generate anthropomorphic virtual avatars with specific appearance features.
The appearance characteristics of the virtual image are accurately determined through input images, which reduces the user's text expression requirements, improves the stability and consistency of automatic generation of virtual images, and enriches the diversity of virtual images.
Smart Images

Figure CN119919547A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information processing technology, and in particular to a method, device, storage medium and program product for generating a virtual image. Background Art
[0002] With the rapid development of artificial intelligence in the field of image generation, more and more application scenarios require the generation of anthropomorphic animal images. Especially in the fields of entertainment, advertising, games, etc., there is a huge demand for cute pet images.
[0003] Traditional methods generally use manual design of cute pet images, but the traditional manual design method is inefficient. In order to solve the problems in the traditional method, some virtual image generation platforms based on artificial intelligence technology have emerged, which can automatically generate animal images. However, the current virtual image generation method requires users to pre-configure detailed feature descriptions as prompts, which places high demands on users' text expression ability and often makes it difficult for the generated virtual images to be consistent with user needs. Summary of the invention
[0004] The main purpose of the embodiments of the present application is to provide a virtual image generation method, device, storage medium and program product, which can achieve accurate customization of the appearance characteristics of the virtual image through input images, which can not only reduce the user's text expression requirements, but also improve the stability and consistency of the automatically generated virtual image, and enrich the diversity of the virtual image.
[0005] In a first aspect, an embodiment of the present application provides a method for generating a virtual image, comprising: in response to a virtual image generation instruction, obtaining an input image, wherein the input image includes a reference object; extracting appearance features of the reference object in the input image; and injecting the appearance features into a preset generation model, so that the preset generation model generates an anthropomorphic virtual image having the appearance features.
[0006] In a second aspect, an embodiment of the present application provides a method for training a generative model, comprising: obtaining a sample image, wherein the sample image includes a reference object; fine-tuning a pre-trained generative model through a preset adapter according to the sample image to obtain a preset generative model; wherein the adapter is used to extract the appearance features of the reference object in the sample image, and inject the appearance features into the generative model, so that the fine-tuned preset generative model generates an anthropomorphic virtual image having the appearance features of the reference object in the input image.
[0007] In a third aspect, an embodiment of the present application provides a navigation method based on a virtual image, comprising: in response to a map navigation request, displaying a preset anthropomorphic virtual image on a map interaction interface, wherein the anthropomorphic virtual image is generated by the virtual image generation method described in any of the above aspects; and driving the anthropomorphic virtual image to indicate navigation information in the map interaction interface.
[0008] In a fourth aspect, an embodiment of the present application provides a virtual image generation device, comprising:
[0009] An acquisition module, configured to acquire an input image in response to a virtual image generation instruction, wherein the input image includes a reference object;
[0010] An extraction module, used for extracting appearance features of the reference object in the input image;
[0011] A generation module is used to inject the appearance features into a preset generation model so that the preset generation model generates an anthropomorphic virtual image with the appearance features.
[0012] In a fifth aspect, an embodiment of the present application provides an electronic device, including:
[0013] at least one processor; and
[0014] a memory communicatively coupled to the at least one processor;
[0015] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the electronic device to execute the method described in any one of the above aspects.
[0016] In a sixth aspect, an embodiment of the present application provides a cloud device, including:
[0017] at least one processor; and
[0018] a memory communicatively coupled to the at least one processor;
[0019] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the cloud device to execute the method described in any one of the above aspects.
[0020] In a seventh aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the method described in any one of the above aspects is implemented.
[0021] In an eighth aspect, an embodiment of the present application provides a computer program product, including a computer program, which implements the method described in any of the above aspects when executed by a processor.
[0022] The virtual image generation method, device, storage medium and program product provided in the embodiments of the present application can extract the appearance features of the reference object in the input image and inject the extracted appearance features into the pre-trained preset generation model. It can guide the preset generation model to generate an anthropomorphic virtual image through the appearance features of the reference object, so that the generated anthropomorphic virtual image has the appearance features of the reference object in the input image. In this way, the appearance characteristics of the virtual image can be accurately customized through the input image, which not only reduces the user's text expression requirements, but also improves the stability and consistency of the automatically generated virtual image and enriches the diversity of the virtual image. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the drawings described below are some embodiments of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative labor.
[0024] Figure 1 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;
[0025] Figure 2 A schematic diagram of an application scenario of a virtual image generation system provided in an embodiment of the present application;
[0026] Figure 3 A schematic diagram of a flow chart of a method for generating a virtual image provided in an embodiment of the present application;
[0027] Figure 4 A flowchart of a training method for generating a model provided in an embodiment of the present application;
[0028] Figure 5 A schematic diagram of an adapter module provided in an embodiment of the present application performing fine-tuning training on a pre-trained diffusion model;
[0029] Figure 6 A schematic diagram of stylized training of a pre-trained generative model based on LoRA technology provided in an embodiment of the present application;
[0030] Figure 7 A schematic diagram of a pre-trained generative model fine-tuned by a DPO method provided in an embodiment of the present application;
[0031] Figure 8A flowchart of a navigation method based on a virtual image provided in an embodiment of the present application;
[0032] Fig. 9 A schematic diagram of the structure of a virtual image generating device provided in an embodiment of the present application;
[0033] Fig.10 A schematic diagram of the structure of a cloud device provided in an embodiment of the present application.
[0034] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0035] Here, exemplary embodiments are described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application.
[0036] The term "and / or" in this article is used to describe the association relationship of associated objects, specifically indicating that there may be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0037] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0038] In order to clearly describe the technical solution of the embodiment of the present application, the terms involved in the present application are first defined:
[0039] Avatar: A graphic or three-dimensional model in a digital or virtual environment. Avatars are widely used in various digital platforms and technologies.
[0040] LLM: Large Language Model, large language model, hereinafter referred to as "large model".
[0041] CLIP: Contrastive Language-Image Pre-Training, contrastive language-image pre-training.
[0042] AI: Artificial Intelligence.
[0043] 3D:Three-Dimensional, three-dimensional.
[0044] 2D: Two-Dimensional, two-dimensional.
[0045] Diffusion Models: Diffusion models are a type of generative model that generates data by simulating the process of gradually denoising from noise. Diffusion models were originally used in particle diffusion simulations in the field of physics, but in deep learning, diffusion models are used to generate images, audio, and other forms of data. The model training process involves gradually adding noise from the data distribution and learning how to denoise inversely to reconstruct the original data. This method can generate high-quality and high-diversity samples.
[0046] SD: Stable Diffusion is a variant of the diffusion model designed for efficient image generation. Unlike traditional diffusion models, Stable Diffusion optimizes the diffusion process and dimensionality reduction techniques to significantly reduce the consumption of computing resources while maintaining the quality of generated images. This makes it more practical in practical applications, especially in the task of generating images, and is therefore widely used in AI painting, image restoration and other fields.
[0047] GAN: Generative Adversarial Network, a generative model, consists of two neural networks: the generator and the discriminator. The generator is responsible for generating data, while the discriminator determines whether the generated data is real or forged. The two compete with each other during the training process, with the generator continuously optimized to generate more realistic data, while the discriminator continuously improves its ability to identify forged data. GAN has been widely used in image generation, style transfer, data enhancement and other fields.
[0048] cGAN: Conditional Generative Adversarial Network, a variant of Generative Adversarial Network (GAN), in which both the generator and the discriminator accept additional conditional information as input. This conditional information can be labels, images, or any other form of data.
[0049] Text-to-Image Generation: Text-to-Image Generation refers to the technology of generating corresponding images from text descriptions, which is usually implemented using deep learning models. The goal of this type of technology is to extract visual features from natural language descriptions and generate images that match the descriptions. Typical models include architectures based on diffusion models (such as Stable Diffusion) or generative adversarial networks (GANs). Text-to-Image Generation technology has a wide range of applications in creative design, advertising generation, image editing and other fields.
[0050] Image-to-Image Translation: Image-to-Image Translation is a technique that converts an input image into a target image, usually keeping the basic structure of the input image but changing some of its properties or style. This type of technology is widely used in image style transfer, image restoration, image enhancement, and image editing. Image-to-Image models are usually implemented using Generative Adversarial Networks (GANs), Conditional Generative Adversarial Networks (cGANs), or diffusion models.
[0051] LoRA: Low-Rank Adaptation is a technique for fine-tuning a pre-trained model by introducing a low-rank matrix to adjust the model's weights so that the model can adapt to new tasks or new data distributions. The main advantage of LoRA is that it can quickly and efficiently adjust large pre-trained models to specific application scenarios without significantly increasing computational costs. This technology is particularly suitable for generative models, language models, and other fields when rapid transfer learning is required.
[0052] Adapter: Adapter is a lightweight model fine-tuning technology designed to reduce the cost and time of training large models. Adapter modules are usually smaller trainable modules inserted between certain layers of the model, allowing the model to be fine-tuned without changing the original model parameters. This method is particularly suitable for multi-task learning or scenarios where the same basic model needs to be applied in multiple fields, and can effectively improve the generalization ability and adaptability of the model.
[0053] DPO: Direct Preference Optimization, is an optimization method that is often used to train models to directly optimize user preferences or reward functions. Compared with traditional policy optimization methods, DPO directly targets user preferences and adjusts the output of the model to make it more in line with user needs or goals. DPO has significant advantages in recommendation systems, personalized advertising, and other tasks involving user preferences.
[0054] VAE: Variational Autoencoder, a generative model that generates data through an encoding-decoding process. The core of VAE is to model data using probabilistic methods, mapping the input data to a latent continuous space (usually a Gaussian distribution), and then mapping the points in the latent space back to the original data distribution through the decoder. The advantage of VAE is that it can not only generate new data, but also interpolate in the latent space to generate smooth transition samples.
[0055] The virtual image generation method of the embodiment of the present application can be applied to any field that requires a personified virtual image.
[0056] Artificial Intelligence (AI) is the study of using computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.), which can enable computers to achieve higher-level applications.
[0057] With the rapid development of artificial intelligence in the field of image generation, more and more application scenarios require the generation of anthropomorphic animal images. Especially in the fields of entertainment, advertising, games, etc., there is a huge demand for cute pet images.
[0058] The traditional method generally uses manual design of cute pet images, but the traditional manual design method is inefficient. In order to solve the problems in the traditional method, some virtual image generation platforms based on artificial intelligence technology have emerged, which can automatically generate animal images. However, the current virtual image generation method requires users to pre-configure detailed feature descriptions as prompt words, which places high demands on the user's text expression ability. The model often cannot understand the concept of anthropomorphism well, and the generated anthropomorphic animal image may not meet expectations in posture and limb structure. Especially in the case of diversified input, it often makes it difficult for the generated virtual image to be consistent with user needs, and it is difficult to ensure the consistency of batch-generated virtual images and their similarity to the specified image, and the user's customization ability is limited. In addition, the existing solutions are prone to image deformation or distortion during the generation process, and it is difficult to maintain the stability of the generated results.
[0059] In order to solve at least one of the above problems, an embodiment of the present application provides a virtual image generation solution, which extracts the appearance features of a reference object in an input image and injects the extracted appearance features into a pre-trained preset generation model. The appearance features of the reference object can be used to guide the preset generation model to generate an anthropomorphic virtual image, so that the generated anthropomorphic virtual image has the appearance features of the reference object in the input image. In this way, the appearance characteristics of the virtual image can be accurately customized through the input image, which not only reduces the user's text expression requirements, but also improves the stability and consistency of the automatically generated virtual image, enriching the diversity of the virtual image.
[0060] Some embodiments of the present application are described in detail below in conjunction with the accompanying drawings. In the case where there is no conflict between the embodiments, the following embodiments and the features in the embodiments can be combined with each other. In addition, the step sequence in the following method embodiments is only an example and not a strict limitation.
[0061] like Figure 1 As shown, this embodiment provides an electronic device 1, including: at least one processor 11 and a memory 12, Figure 1 A processor is taken as an example. The processor 11 and the memory 12 are connected via a bus 10. The memory 12 stores instructions that can be executed by the processor 11, and the instructions are executed by the processor 11 so that the electronic device 1 can execute all or part of the process of the method in the following embodiment to achieve accurate customization of the appearance characteristics of the virtual image through the input image, which can not only reduce the user's text expression requirements, but also improve the stability and consistency of the automatically generated virtual image, and enrich the diversity of the virtual image.
[0062] In one embodiment, the electronic device 1 may be a mobile phone, a tablet computer, a laptop computer, a desktop computer, or a large computing system composed of multiple computers.
[0063] Figure 2 Schematic diagram of a virtual image generation system application scenario 200 provided in an embodiment of the present application. Figure 2 As shown, the system includes: a server 210 and a terminal 220, wherein:
[0064] The server 210 may be a data platform that provides virtual image generation services, such as an electronic map service platform. In actual scenarios, an electronic map service platform may have multiple servers 210. Figure 2 In the figure, one server 210 is taken as an example.
[0065] The terminal 220 may be a mobile device for logging into the electronic map service platform, such as a computer, a mobile phone, a tablet, etc. There may also be multiple terminals 220. Figure 2 Two terminals 220 are taken as an example for illustration.
[0066] The terminal 220 and the server 210 can transmit information via the Internet, so that the terminal 220 can access the data on the server 210. The terminal 220 and / or the server 210 can be implemented by the electronic device 1.
[0067] The virtual image generation solution of the embodiment of the present application can be deployed on the server 210, can also be deployed on the terminal 220, or can be partially deployed on the server 210 and partially deployed on the terminal 220. In actual scenarios, it can be selected based on actual needs, and this embodiment does not limit it.
[0068] When the virtual image generation solution is fully or partially deployed on the server 210 , a calling interface may be opened to the terminal 220 to provide algorithm support to the terminal 220 .
[0069] The method provided in the embodiment of the present application can be implemented by executing the corresponding software code by the electronic device 1, and can be implemented by exchanging data with the server. The electronic device 1 can be a local terminal device. When the method is run on the server, the method can be implemented and executed based on the cloud interaction system, wherein the cloud interaction system includes a server and a client device.
[0070] In a possible implementation, the method provided in the embodiment of the present application provides a graphical user interface through a terminal device, wherein the terminal device can be the local terminal device mentioned above, or can be the client device in the cloud interaction system mentioned above.
[0071] Please see Figure 3 , which is a method for generating a virtual image according to an embodiment of the present application, the method can be Figure 1 The electronic device 1 shown is used to perform and can be applied to Figure 2 In the virtual image generation application scenario shown in , the appearance characteristics of the virtual image can be accurately customized by inputting an image, which can not only reduce the user's text expression requirements, but also improve the stability and consistency of the automatically generated virtual image, and enrich the diversity of the virtual image. In this embodiment, the terminal 220 is used as an example of the execution end, and the method includes the following steps:
[0072] Step 301: In response to a virtual image generation instruction, an input image is acquired, where the input image includes a reference object.
[0073] In this step, the reference object is used to indicate that the generated anthropomorphic virtual image is similar to it. Assuming that the user wants to generate an anthropomorphic virtual image similar to kitten A, a photo of kitten A can be provided as an input image. The reference object can be a physical object, such as a physical animal object. In this case, the input image can be a photo of a physical animal, such as a photo of a kitten or a puppy. The reference object can also be a virtual object, such as a virtual animal object designed by computer software. In this case, the input image can be an image containing the virtual animal object. The generation instruction can be actively triggered by the user or automatically triggered by the system. The input image can be an image uploaded by the user or an image automatically loaded by the system. The input image at least includes the head image of the reference object, so as to provide an accurate data source for subsequent appearance feature extraction.
[0074] Step 302: Extract appearance features of the reference object in the input image.
[0075] In this step, the appearance features of the reference object are used to characterize the reference object's appearance and determine the reference object's visual features. Appearance features include but are not limited to the reference object's head features, facial features, and limb features. Detailed appearance features can provide accurate guidance information for the subsequent generation of a virtual image.
[0076] In one embodiment, taking the reference object as an animal object as an example, the appearance features of the reference object include: one or more of the head features, facial features, and limb features of the animal.
[0077] In this embodiment, taking an animal object as an example, the appearance characteristics of an animal can be described according to different classification standards and observation angles. For example, the appearance characteristics of an animal may include: body shape and weight characteristics, body structure characteristics (such as thin and long, round and fat), skin and hair characteristics, head contour characteristics (such as head shape structure, head color, head texture and other image information), facial features (including but not limited to facial structure, color, texture and other features, such as eyes, ears, nose, mouth structure, color, texture and other image information), limb characteristics (such as limb types may include hooves, claws, fins, wings, and the number of limbs may include two feet, four feet, multiple feet and other features), tail characteristics, color and pattern characteristics and special features (such as the presence or absence of horns, the size and shape of the horns, the presence or absence of fins, the size and shape of the fins, the presence or absence of wings, the size and shape of the wings, etc.). Appearance characteristics can help identify and classify different animal images and provide detailed feature references for the virtual image generation process.
[0078] In one embodiment, step 302 may specifically include: extracting image features of the input image by using a preset image encoder, and determining the image features as appearance features of the reference object.
[0079] In this embodiment, by processing the input image using a preset image encoder, the feature information of the image can be accurately extracted. For example, the pre-trained CLIP image encoder model can be used to extract image features from image prompts. These image features can fully reflect the appearance characteristics of the reference object and can capture the subtle appearance details of the reference object, such as facial features, hair, expressions, head structure, etc., so as to ensure that the generated virtual image can highly restore the real appearance of the reference object. This provides a reliable basis for subsequent image analysis, recognition and processing. This method not only improves the accuracy and efficiency of appearance feature extraction, but also reduces the need for manual intervention, and has a high degree of automation and practical value.
[0080] In an optional embodiment, the image features of the input image may also be obtained by a convolutional neural network or other feature extraction methods.
[0081] Step 303: injecting the appearance features into the pre-trained preset generation model so that the preset generation model generates an anthropomorphic virtual image with the appearance features.
[0082] In this step, by extracting the appearance features of the reference object, the generated virtual image can be highly personalized, accurately reflecting the unique appearance features of the reference object, and meeting the user's demand for personalized virtual images. Specifically, the appearance features extracted in step 302 include but are not limited to the features of the reference object's facial structure, skin color, hairstyle, eye shape, nose and mouth, which are important parameters for describing the appearance of the reference object. By injecting the appearance features of the reference object in the input image into the pre-trained preset generation model, so that the appearance features of the reference object are used as the basic parameters of the preset generation model when generating the virtual image, it is ensured that the generated virtual image can reflect the appearance features of the reference object in the input image. The virtual image generated by this scheme is not only similar to the reference object in appearance, but also can meet the personalized needs of the user.
[0083] In addition, the use of pre-trained preset generation models can automatically generate virtual images, reducing manual intervention and design time, and improving generation efficiency. The pre-trained model has been trained with a large amount of data to ensure the consistency and accuracy of the generated virtual image in appearance features, avoiding the errors that may be caused by manual design. In addition, this method can adapt to different types of reference objects, whether human or other anthropomorphic objects, and can generate corresponding virtual images by extracting their appearance features, which has a wide range of application scenarios. By generating virtual images that are highly similar to reference objects, users can get a better interactive experience and emotional resonance, which enhances the user's sense of identity and satisfaction with the virtual image.
[0084] Optionally, the preset generative model can also be a pre-trained graph-based model, which can apply the style of an input image (such as an oil painting style) to another image while keeping the content unchanged. For example, by inputting a partially missing image, a complete image is generated to achieve image restoration. Or a low-resolution image is converted to a high-resolution image to achieve image enhancement. Graph-based technology requires that the model can not only understand the content of the input image, but also generate the expected output image according to the target task, maintaining the coherence and quality of the image.
[0085] Optionally, the preset generation model can be a pre-trained text-based image model, which can accurately extract semantic information from the prompt word text through natural language understanding, map it to visual features, and generate high-quality images that conform to the text description based on the extracted visual features. For example, complex scenes, character designs, or product models can be generated by using simple text descriptions as input prompt words.
[0086] In one embodiment, step 303 may specifically include: extracting text features of the input prompt word of the preset generation model. Mapping the appearance features and the text features to the same feature space. Calculating the first attention of the appearance features and the second attention of the text features respectively. Fusing the appearance features and the text features according to the first attention and the second attention to generate a fusion feature of the appearance features and the text features. The preset generation model is used to generate an anthropomorphic virtual image with appearance features according to the fusion feature.
[0087] In this embodiment, the preset generation model can be a Wensheng graph model, such as a Wensheng graph model that can be implemented based on the SD diffusion model, and the corresponding image can be generated by the preset input prompt word, such as the input prompt word can be "an anthropomorphic animal image". In order to ensure the accuracy of the generated virtual image, the preset generation model can be guided in combination with the prompt word and the input image at the same time, so that the generated virtual image has both the features described by the prompt word and the appearance features of the reference object in the input image. Specifically, firstly, the text features of the input prompt word are extracted, and the appearance features and text features in the input image are mapped to the same feature space, then, the first attention of the appearance features and the second attention of the text features are calculated respectively, and the appearance features and text features are fused according to these two attentions to generate fusion features. Finally, the preset generation model generates an anthropomorphic virtual image with appearance features according to the fusion features. In this way, the appearance features extracted from the input image and the text features of the prompt word are fused through the decoupled cross attention mechanism, so that the fusion features simultaneously contain the features described by the prompt word and the features of the reference object in the image, which ensures the consistency between different features, makes the generated virtual image more coordinated and natural, and thus improves the accuracy of the anthropomorphic virtual image. This embodiment can not only process multiple input forms (such as images and text), but also customize and generate diversified virtual images according to different input prompt words and input images, with high flexibility and adaptability. It can significantly enhance the user's interactive experience in the virtual environment and meet the user's personalized needs.
[0088] In one embodiment, the appearance features and text features are fused according to the first attention and the second attention to generate fusion features of the appearance features and the text features, including: determining a preset ratio between the appearance features and the text features. The first attention and the second attention are weighted and summed according to the preset ratio to generate the fusion attention of the appearance features and the text features. The appearance features and the text features are weighted and summed according to the fusion attention to generate the fusion features of the appearance features and the text features.
[0089] In this embodiment, during the inference stage of the preset generation model, the proportion of attention between the input image and the prompt word can be adjusted according to actual needs, so as to adjust the proportion of the appearance features of the reference object in the input image in the final fusion feature. For example, if the input image is an animal, and the prompt word indicates to generate "an anthropomorphic animal image", the proportion of the appearance features in the input image can be adjusted to adjust whether the final generated virtual image is more like the animal in the input image or more like the anthropomorphic features described in the prompt word, thereby improving the flexibility of the virtual image.
[0090] Specifically, the preset ratio between appearance features and text features is first determined. Here, the user can customize the preset ratio between appearance features and text features according to actual needs. The system performs weighted summation of the first attention and the second attention according to the preset ratio to generate a fused attention. Finally, the fused attention performs weighted summation of the appearance features and the text features to generate a fused feature. The attention mechanism makes the feature fusion process adaptive, and can dynamically adjust the fusion strategy according to the feature weights of different inputs, thereby generating a fused feature that better meets the actual needs of the user. Improve the personalization of the anthropomorphic virtual image.
[0091] In one embodiment, step 303 may specifically include: injecting the appearance features into the cross-attention layer of the preset generation model, so that the preset generation model generates an anthropomorphic virtual image under the guidance of the input prompt words and the appearance features.
[0092] In this embodiment, the preset generation model can be a model with an attention mechanism. Taking the preset generation model based on the diffusion model as an example, the fusion between the injected appearance features and other features can be processed by a cross-attention layer. The attention mechanism was originally introduced in the field of natural language processing to enhance the ability of the model in processing sequence data. Its core idea is to dynamically focus on the most relevant information by calculating the correlation (attention weight) between each element in the input sequence. Cross-attention is an extension of the attention mechanism for processing multimodal data. It fuses information by calculating the correlation between different modalities. For example, in an image-text task, the cross-attention module can calculate the correlation between image features and text features to generate a richer representation. In this embodiment, by injecting the appearance features of the reference object in the input image into the cross-attention layer of the preset generation model, the cross-attention layer can pay attention to the input prompt words and the appearance features at the same time, so that the model can dynamically adjust the attention to the input prompt words and the appearance features, so that the preset generation model takes the input prompt words as the target, and generates an anthropomorphic virtual image with appearance features under the joint action of the input prompt words and the appearance features, thereby flexibly guiding the appearance features of the virtual image during the generation process, so that the generation result is more in line with expectations. In addition, by injecting the appearance features of the reference object in the input image into the cross-attention layer, the model can better adapt to different combinations of input prompt words and appearance features, generate a variety of anthropomorphic virtual images, and has high adaptability and flexibility.
[0093] In one embodiment, the preset generation model is obtained by fine-tuning a pre-trained generation model through a stylized sample set, the stylized sample set includes multiple text-image pairs, the text-image pairs include sample prompt words for indicating the generation of a virtual image and stylized sample images, and the stylized sample images include an anthropomorphic object with human body features.
[0094] In this embodiment, the pre-trained generative model refers to a model that is preliminarily trained on a large and diverse data set. Its main purpose is to learn the basic features and structure of the data so that it can be fine-tuned and optimized faster and more effectively in subsequent specific tasks. Pre-trained generative models are widely used in various generative tasks, such as image generation, text generation, image-to-image conversion (such as style transfer), text-to-image generation, etc. By fine-tuning on specific tasks, the pre-trained generative model can generate high-quality output that meets specific needs. For example, by fine-tuning the pre-trained generative model using multiple text-image pairs in the stylized sample set, the quality and consistency of the generated anthropomorphic virtual image can be significantly improved. Specifically, the text-image pairs in the stylized sample set include sample prompt words and anthropomorphic object images with human body features, which enables the generative model to more accurately understand and generate virtual images that meet the expected style. In this way, the generative model not only retains the basic capabilities of the pre-trained model, but also enhances its expressiveness and detail processing capabilities in a specific style, so that the generated virtual image is more vivid, realistic, and has unique style characteristics.
[0095] Optionally, the LoRA technology can be applied to perform stylized training on the pre-trained generative model to achieve fine-tuning of the generative model so that the generative model learns the human limb features in the stylized sample set, such as learning the standing posture of the human body, so that the anthropomorphic virtual image stands up like a human. The stylized sample images in the stylized sample set can be images of anthropomorphic animals, such as anthropomorphic animal images with animal head features (such as a kitten's head) and human limb features, so that the pre-trained generative model can learn the posture, limb structure, expression and other features of the anthropomorphic animals in the stylized sample images, ensuring that the final generated virtual image meets the anthropomorphic requirements.
[0096] In one embodiment, the preset generation model is obtained by fine-tuning a pre-trained generation model through a user preference sample set, the user preference sample set includes at least one text image pair and a user preference image, the text image pair includes a sample prompt word for indicating the generation of a virtual image and an image of a specified reference object, the user preference image includes the user's preference information for marking the output image, and the output image is generated by the generation model using the text image pair as input.
[0097] In this embodiment, in order to improve the personalization of the preset generation model, the preference alignment optimization task can be achieved by fine-tuning the pre-trained generation model. The text image pair in the user preference sample set includes a sample prompt word and an image of a specified reference object, wherein the specified reference object is used to indicate that the anthropomorphic virtual image generated in the output image is similar to it. For details, please refer to the aforementioned interpretation of the reference object in step 301, which will not be repeated here. Assuming that the sample prompt word is "an anthropomorphic animal image", the text image pair in the user preference sample set is first input into the pre-trained generation model for reasoning to obtain the output image of the pre-trained generation model. At this time, the user can mark the output image with preference information. For example, the output image can contain multiple, and the user can mark each output image with preference information to obtain the corresponding user preference image. Then the user preference image is fed back to the generation model to guide the generation model to make further adjustments and optimizations, so that the generation model can understand the user's basic preference needs and significantly improve the personalization and user satisfaction of the final generated virtual image.
[0098] Optionally, the pre-trained generative model can be fine-tuned by the DPO method to fix unsatisfactory details generated by the model and achieve the model's preference alignment optimization task.
[0099] By combining the user's preferred sample set, the preset generation model not only retains the basic capabilities of the pre-trained model, but can also better capture and reflect the user's personalized needs and preferences, so that the generated avatar is more in line with the user's expectations. Ultimately, the generated avatar is closer to the user's preferences in terms of details, style and overall effect, significantly improving the user experience and satisfaction.
[0100] Please see Figure 4 , which is a training method for generating a model according to an embodiment of the present application, and the method can be Figure 1 The electronic device 1 shown is used to perform and can be applied to Figure 2 In the virtual image generation application scenario shown in , the appearance characteristics of the virtual image can be accurately customized by inputting an image, which can not only reduce the user's text expression requirements, but also improve the stability and consistency of the automatically generated virtual image, and enrich the diversity of the virtual image. In this embodiment, the terminal 220 is used as an example of the execution end, and the method includes the following steps:
[0101] Step 401: Acquire a sample image, where the sample image includes a reference object.
[0102] Step 402: fine-tune the pre-trained generation model through a preset adapter according to the sample image to obtain a preset generation model.
[0103] The adapter is used to extract the appearance features of the reference object in the sample image and inject the appearance features into the generation model so that the fine-tuned preset generation model generates an anthropomorphic virtual image with the appearance features of the reference object in the input image.
[0104] In this embodiment, the pre-trained generative model refers to a model that is preliminarily trained on a large and diverse dataset. Its main purpose is to learn the basic features and structure of the data so that it can be fine-tuned and optimized faster and more effectively in subsequent specific tasks. By fine-tuning on specific tasks, the pre-trained generative model can generate high-quality output that meets specific needs. Large Model Adapter is a technology for fine-tuning or adapting specific tasks on a large pre-trained model. Its main purpose is to quickly and efficiently adjust the model to meet specific task or domain requirements without completely retraining the entire large model. In this embodiment, the pre-trained generative model can be fine-tuned based on the preset adapter to obtain the preset generative model. Specifically, a sample image containing a reference object is first obtained to ensure the diversity and authenticity of the training data, which provides a solid foundation for subsequent model fine-tuning. Then the pre-trained generative model is fine-tuned using the preset adapter. The design of the adapter enables it to efficiently extract the appearance features of the reference object in the sample image and inject these features into the generative model, which not only improves the adaptability of the model, but also enhances the ability of the generative model to capture specific appearance features. Through the above-mentioned fine-tuned preset generation model, an anthropomorphic virtual image with the appearance characteristics of the reference object in the input image can be generated. This generation method ensures that the appearance characteristics of the virtual image are highly consistent with the reference object, thereby improving the authenticity and personalization of the generation result.
[0105] Optionally, taking anthropomorphic animal images as an example, the Adapter can be applied to perform similarity training on the pre-trained generative model, and the image features of the animal with a specified ID can be learned. For example, the pre-trained diffusion model can be fine-tuned through the Adapter module, and the animal appearance features in the sample image are injected into the diffusion model to learn the appearance of the animal with a specified ID, so that the generated anthropomorphic animal image is highly similar to the appearance of the animal object in the input image, and the anthropomorphic virtual image can be flexibly customized according to the input image.
[0106] Optionally, the Adapter may be applied to perform human posture learning training on the pre-trained generative model, so that the generative model learns the "standing" posture of the human body during the training process, thereby ensuring that the generated anthropomorphic virtual image can stand up like a human.
[0107] like Figure 5As shown, it is a schematic diagram of an adapter module provided in an embodiment of the present application for fine-tuning the pre-trained diffusion model. Assuming that the diffusion model (Diffusion Models) is used as the pre-trained generative model, the adapter training sample can be a sample image containing an anthropomorphic reference object. The reference object in the sample image can be an anthropomorphic animal image with animal head features (such as a kitten's head). In addition, the sample prompt word "an anthropomorphic animal image" can be added. During the training process, only the adapter Adapter can be optimized while keeping the parameters of the pre-trained diffusion model unchanged. Using the sample image (such as a kitten's head image) as the input image, a pre-trained image encoder (Image Encoder) can be used to extract image features from the input image prompt, and the image features are injected into the diffusion model through the adapter Adapter. So that the diffusion model learns the appearance of the animal in the specified input image.
[0108] In one embodiment, the method further includes: obtaining a stylized sample set, the stylized sample set including a plurality of text-image pairs, the text-image pairs including sample prompt words for instructing to generate a virtual image and stylized sample images, the stylized sample images including an anthropomorphic object with human body features, and fine-tuning a pre-trained generative model according to the stylized sample set so that the fine-tuned generative model learns stylized information.
[0109] In this embodiment, the text image pairs in the stylized sample set include sample prompt words and anthropomorphic object images with human body features, which enables the generative model to more accurately understand and generate virtual images that meet the expected style. The pre-trained generative model is fine-tuned according to the stylized sample set so that the fine-tuned generative model learns stylized information. The fine-tuned generative model not only retains the basic capabilities of the pre-trained model, but also enhances its expressiveness and detail processing capabilities under a specific style, such as learning human body features, so that the generated virtual image is more vivid and lifelike, and has unique human body style features.
[0110] Optionally, the LoRA technology can be applied to perform stylized training on the pre-trained generative model to achieve fine-tuning of the generative model so that the generative model learns the human limb features in the stylized sample set, such as learning the standing posture of the human body, so that the anthropomorphic virtual image can stand up like a human.
[0111] like Figure 6As shown, it is a schematic diagram of stylized training of a pre-trained generative model based on LoRA technology provided by an embodiment of the present application. It is assumed that a diffusion model is used as a pre-trained generative model, and the stylized sample set includes sample prompt words and stylized sample images, wherein the sample prompt word can be "an anthropomorphic animal image". The stylized sample images can be images of various anthropomorphic animals, such as anthropomorphic animal images with animal head features (such as a kitten's head, a puppy's head) and human limb features. The stylized sample images should cover a variety of postures, limb structures and expressions to ensure that the model can learn rich features. The stylized sample images are input into LoRA Models (LoRA models), and the LoRA model extracts the basic features of the reference object in the stylized sample image, such as features similar to human limbs. Then the LoRA model fuses the extracted features into the diffusion model, so that the diffusion model learns the posture, limb structure, expression and other features of the anthropomorphic animal in the stylized sample image with the sample prompt word as the purpose, to ensure that the final generated virtual image meets the anthropomorphic requirements.
[0112] In one embodiment, the method further includes: obtaining at least one text-image pair, the text-image pair including a sample prompt word for indicating the generation of a virtual image and an image of a designated reference object. Inputting the text-image pair into a pre-trained generative model, and obtaining an output image of the pre-trained generative model. Obtaining a user preference image corresponding to the output image, the user preference image including preference information marked by the user on the output image. Fine-tuning the pre-trained generative model according to the user preference image, so that the fine-tuned generative model learns the user's preference information on the output image.
[0113] In this embodiment, in order to enhance the personalization of the preset generation model during the model training process, the preference alignment optimization task can be achieved by fine-tuning the pre-trained generation model with user preferences. First, a user preference sample set is prepared, which includes at least a text image pair, and the text image pair includes a sample prompt word and an image of a specified reference object, wherein the sample prompt word can be "an anthropomorphic animal image". The specified reference object is used to indicate that the anthropomorphic virtual image generated in the output image is similar to it. For details, please refer to the aforementioned interpretation of the reference object in step 301, which will not be repeated here. Then the text image pair is input into the pre-trained generation model for reasoning. Here, the pre-trained generation model can be a model that has been fine-tuned by stylized training using the LoRA technology of the aforementioned embodiment, and the output image after the pre-trained generation model reasoning is obtained. At this time, the user can mark the output image with preference information. For example, the output image can contain multiple, and the user can mark each output image with preference information to obtain the corresponding user preference image. The preference information can include positive examples and / or negative examples. The positive example indicates that the output image meets expectations, and the negative example indicates that the output image does not meet expectations. The user preference image is then fed back to the pre-trained generative model to guide the pre-trained generative model to make further adjustments and optimizations, so that the final fine-tuned generative model can understand the user's basic preference needs and significantly improve the personalization and user satisfaction of the final generated virtual image.
[0114] Optionally, the pre-trained generative model can be fine-tuned by the DPO method to fix unsatisfactory details generated by the model and achieve the model's preference alignment optimization task.
[0115] like Figure 7As shown, it is a schematic diagram of a pre-trained generative model fine-tuned by the DPO method provided by an embodiment of the present application. Assuming that a diffusion model is used as a pre-trained generative model, the sample prompt word in the text image pair is "an anthropomorphic husky image", and the image of the designated reference object in the text image pair is a puppy image. The puppy image is used as an input image. First, the image features of the puppy image are extracted through an image encoder, and then the image features of the puppy image are injected into the diffusion model through an adapter, so that the diffusion model fuses the text features of the sample prompt word with the image features of the puppy image to generate an output image. Assuming that the output image includes two anthropomorphic puppy images, the user can mark the output image with preference information, such as marking the output image of the anthropomorphic puppy standing on two legs as a positive example, and marking the output image of the anthropomorphic puppy standing on four legs as a negative example, and then feeding back the output image (positive example) and the output image (negative example) with preference marking information to the diffusion model through DPO Training (data preference optimization training). LoRA technology can also be used here to achieve fine-tuning of preference information. For example, through LoRAModels (LoRA model), the preference tag information in the output image (positive example) and the output image (negative example) is merged into the diffusion model to guide the model for further adjustment and optimization, so that the final generation model can understand the user's basic preference needs and significantly improve the personalization and user satisfaction of the final generated virtual image.
[0116] Through preference alignment optimization, the preset generation model not only retains the basic capabilities of the pre-trained model, but also better captures and reflects the user's personalized needs and preferences, so that the generated avatar is more in line with the user's expectations. Ultimately, the generated avatar is closer to the user's preferences in terms of details, style and overall effect, significantly improving the user experience and satisfaction.
[0117] Optionally, the pre-trained generative model can be implemented by combining a generative adversarial network (GAN) and a variational autoencoder (VAE) to address the stability issue of the generated avatar.
[0118] Optionally, the pre-trained generation model can be an image generation model based on a deep residual network, which uses the deep residual network to generate high-quality images and enhance the stability and consistency of the generated virtual image.
[0119] Optionally, the pre-trained generation model may be an image generation model based on an autoregressive network, which uses the autoregressive network to generate images of virtual images.
[0120] In actual use, you can choose a suitable network model according to actual needs.
[0121] The above model training method proposes a three-stage training method that combines stylized training, consistency training, and preference alignment optimization to ensure the stability and consistency of generation. The pre-trained generative model is fine-tuned through the adapter module to learn the characteristics of the input image and ensure the similarity between the generated virtual image and the animal in the input image, so that users can customize anthropomorphic virtual images that meet their needs. The concept of anthropomorphic animals can also be learned through LoRA technology to ensure that the generated anthropomorphic virtual image meets the expected requirements in terms of posture, limb structure, expression, etc. During the model training process, the pre-trained generative model can be trained in three stages using stylized LoRA learning, Adapter module, and preference alignment optimization to ensure that the generative model has stable generation capabilities and avoids image deformation or distortion.
[0122] Please see Figure 8 , which is a navigation method based on a virtual image in an embodiment of the present application, the method can be Figure 1 Compared with the above-mentioned embodiments, this embodiment takes the application of an anthropomorphic virtual image in map navigation as an example. The method includes the following steps:
[0123] Step 801: In response to a map navigation request, a preset anthropomorphic virtual image is displayed on a map interaction interface, where the anthropomorphic virtual image is generated using a virtual image generation method as described in any of the aforementioned embodiments.
[0124] Step 802: driving the anthropomorphic virtual image to indicate navigation information in the map interaction interface.
[0125] In an embodiment of the present application, the method for generating a virtual image can be applied to a navigation scenario of a digital map, and the generated anthropomorphic virtual image is used as a digital navigator image. When a user triggers a map navigation request, an anthropomorphic virtual image matching the user is displayed on the map interaction interface, and when displaying navigation information, the anthropomorphic virtual image is driven to indicate the navigation information in the interaction interface. For example, in scenarios such as autonomous driving vehicles and intelligent voice assistants, anthropomorphic animal images are displayed to indicate navigation information, allowing users to feel immersed in the navigation of cute pet images and improving the interactive experience.
[0126] In some optional embodiments, the method for generating a virtual image of the present application may also be applied in the following scenarios:
[0127] 1. Games and virtual reality: For example, character generation systems that require high-quality anthropomorphic animal images are needed, especially in games or virtual worlds that require a large number of customized characters.
[0128] 2. Social media and digital content creation tools: used to generate personalized anthropomorphic animal images, suitable for avatar generators, emoticon package design, etc.
[0129] 3. Advertising and marketing: Help brands quickly generate anthropomorphic animal images that match their brand image for use in advertising, animation and other content creation.
[0130] For details of each step of the above method, please refer to the relevant description of the above embodiment, which will not be repeated here.
[0131] Please see Fig. 9 , which is a virtual image generation device 900 of an embodiment of the present application, the device can be applied to a terminal and can be applied to Figure 2 In the virtual image generation application scenario shown in , the appearance characteristics of the virtual image can be accurately customized by inputting an image, which can not only reduce the user's text expression requirements, but also improve the stability and consistency of the automatically generated virtual image, and enrich the diversity of the virtual image. The device includes: an acquisition module 901, an extraction module 902 and a generation module 903, and the functional principles of each module are as follows:
[0132] The acquisition module 901 is used to acquire an input image in response to a virtual image generation instruction, where the input image includes a reference object.
[0133] The extraction module 902 is used to extract the appearance features of the reference object in the input image.
[0134] The generation module 903 is used to inject the appearance features into the preset generation model so that the preset generation model generates an anthropomorphic virtual image with the appearance features.
[0135] In one embodiment, the extraction module 902 is used to extract image features of the input image through a preset image encoder to determine the image features as appearance features of the reference object.
[0136] In one embodiment, the reference object is an animal object. The appearance features of the reference object include: one or more of the head features, facial features, and limb features of the animal.
[0137] In one embodiment, the generation module 903 is used to extract the text features of the input prompt words of the preset generation model. The appearance features and the text features are mapped to the same feature space. The first attention of the appearance features and the second attention of the text features are calculated respectively. The appearance features and the text features are fused according to the first attention and the second attention to generate fusion features of the appearance features and the text features. The preset generation model is used to generate an anthropomorphic virtual image with appearance features according to the fusion features.
[0138] In one embodiment, the generation module 903 is used to determine a preset ratio between the appearance feature and the text feature. The first attention and the second attention are weighted and summed according to the preset ratio to generate a fusion attention of the appearance feature and the text feature. The appearance feature and the text feature are weighted and summed according to the fusion attention to generate a fusion feature of the appearance feature and the text feature.
[0139] In one embodiment, the generation module 903 is used to inject the appearance features into the cross-attention layer of the preset generation model, so that the preset generation model generates an anthropomorphic virtual image under the guidance of the input prompt words and the appearance features.
[0140] In one embodiment, the preset generation model is obtained by fine-tuning a pre-trained generation model through a stylized sample set, the stylized sample set includes multiple text-image pairs, the text-image pairs include sample prompt words for indicating the generation of a virtual image and stylized sample images, and the stylized sample images include an anthropomorphic object with human body features.
[0141] In one embodiment, the preset generation model is obtained by fine-tuning the pre-trained generation model through a user preference sample set, the user preference sample set includes at least one text image pair and a user preference image, the text image pair includes a sample prompt word for indicating the generation of a virtual image and an image of a specified reference object, the user preference image includes the user's preference information for marking the output image, and the output image is generated by the pre-trained generation model with the text image pair as input.
[0142] For a detailed description of the above-mentioned virtual image generating device 900, please refer to the description of the relevant method steps in the above-mentioned embodiment. Its implementation principle and technical effects are similar, and will not be repeated here in this embodiment.
[0143] Fig.10 The following is a schematic diagram of a cloud device 100 provided in an exemplary embodiment of the present application. The cloud device 100 can be used to run the method provided in any of the above embodiments. Fig.10 As shown, the cloud device 100 may include: a memory 1004 and at least one processor 1005, Fig.10 A processor is taken as an example.
[0144] The memory 1004 is used to store computer programs and can be configured to store various other data to support operations on the cloud device 100. The memory 1004 can be an object storage service (OSS).
[0145] Memory 1004 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0146] The processor 1005 is coupled to the memory 1004 and is used to execute the computer program in the memory 1004 to implement the solution provided by any of the above method embodiments. The specific functions and technical effects that can be achieved are not repeated here.
[0147] Furthermore, if Fig.10 The cloud device also includes: a firewall 1001, a load balancer 1002, a communication component 1006, a power supply component 1003 and other components. Fig.10 Only some components are shown schematically, which does not mean that the cloud device only includes Fig.10 Components shown.
[0148] In one embodiment, the above Fig.10 The communication component 1006 in is configured to facilitate wired or wireless communication between the device where the communication component 1006 is located and other devices. The device where the communication component 1006 is located can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G, LTE (Long Term Evolution, Long Term Evolution, referred to as LTE), 5G and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component 1006 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1006 also includes a near field communication (Near Field Communication, referred to as NFC) module to facilitate short-range communication. For example, the NFC module can be based on Radio Frequency Identification (Radio Frequency Identification, referred to as RFID) technology, Infrared Data Association (Infrared Data Association, referred to as IrDA) technology, Ultra Wide Band (UWB) technology, Bluetooth (bluetooth, referred to as BT) technology and other technologies to achieve.
[0149] In one embodiment, the above Fig.10 The power supply component 1003 provides power to various components of the device where the power supply component 1003 is located. The power supply component 1003 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device where the power supply component is located.
[0150] An embodiment of the present application further provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the method of any of the aforementioned embodiments is implemented.
[0151] An embodiment of the present application also provides a computer program product, including a computer program, which implements the method of any of the aforementioned embodiments when executed by a processor.
[0152] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of modules is only a logical function division, and there may be other division methods in actual implementation, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.
[0153] The above-mentioned integrated module implemented in the form of a software function module can be stored in a computer-readable storage medium. The above-mentioned software function module is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform some steps of the methods of various embodiments of the present application.
[0154] It should be understood that the above-mentioned processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The memory may include a high-speed RAM (Random Access Memory) memory, and may also include non-volatile storage NVM (Nonvolatile memory, NVM for short), such as at least one disk storage, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a disk or an optical disk, etc.
[0155] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The storage medium can be any available medium that can be accessed by a general or special computer.
[0156] An exemplary storage medium is coupled to a processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic device or a main control device.
[0157] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, clothing or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, clothing or device. In the absence of more restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, clothing or device including the element.
[0158] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0159] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a disk, or an optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods of each embodiment of the present application.
[0160] In the technical solution of this application, the collection, storage, use, processing, transmission, provision and disclosure of user data and other information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0161] The above are only preferred embodiments of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A method for generating a virtual image, characterized in that: include: In response to a virtual image generation instruction, acquiring an input image, wherein the input image includes a reference object; Extracting appearance features of the reference object in the input image; The appearance features are injected into a preset generation model so that the preset generation model generates an anthropomorphic virtual image with the appearance features.
2. The method according to claim 1, characterized in that The extracting the appearance features of the reference object in the input image comprises: Extracting image features of the input image by using a preset image encoder, and determining the image features as appearance features of the reference object; and / or, The reference object is an animal object; the appearance features of the reference object include: one or more of the head features, facial features, and limb features of the animal.
3. The method according to claim 1 or 2, characterized in that: The step of injecting the appearance features into a preset generation model so that the preset generation model generates an anthropomorphic virtual image having the appearance features comprises: Extracting text features of the input prompt words of the preset generation model; Mapping the appearance feature and the text feature to the same feature space; Calculating the first attention of the appearance feature and the second attention of the text feature respectively; fusing the appearance feature and the text feature according to the first attention and the second attention to generate a fusion feature of the appearance feature and the text feature; The preset generation model is used to generate an anthropomorphic virtual image with the appearance features according to the fusion features.
4. The method according to claim 3, characterized in that The step of fusing the appearance feature and the text feature according to the first attention and the second attention to generate a fusion feature of the appearance feature and the text feature includes: Determining a preset ratio between the appearance feature and the text feature; Performing a weighted summation of the first attention and the second attention according to the preset proportion to generate a fusion attention of the appearance feature and the text feature; The appearance feature and the text feature are weightedly summed according to the fused attention to generate a fused feature of the appearance feature and the text feature.
5. The method according to claim 1, characterized in that: The step of injecting the appearance feature into a preset generation model so that the preset generation model generates a virtual image having the appearance feature comprises: The appearance features are injected into the cross-attention layer of the preset generation model, so that the preset generation model generates the anthropomorphic virtual image under the guidance of the input prompt words and the appearance features.
6. The method according to claim 1, characterized in that The preset generation model is obtained by fine-tuning a pre-trained generation model through a stylized sample set, wherein the stylized sample set includes a plurality of text-image pairs, wherein the text-image pairs include sample prompt words for indicating generation of a virtual image and stylized sample images, wherein the stylized sample images include an anthropomorphic object having human body features.
7. A training method for generating a model, characterized in that: include: Acquire a sample image, wherein the sample image includes a reference object; Fine-tuning the pre-trained generative model through a preset adapter according to the sample image to obtain a preset generative model; The adapter is used to extract the appearance features of the reference object in the sample image and inject the appearance features into the generation model so that the fine-tuned preset generation model generates an anthropomorphic virtual image with the appearance features of the reference object in the input image.
8. A navigation method based on a virtual image, characterized in that: include: In response to a map navigation request, displaying a preset anthropomorphic virtual image on a map interaction interface, wherein the anthropomorphic virtual image is generated by the method according to any one of claims 1 to 6; The anthropomorphic virtual image is driven to indicate navigation information in the map interaction interface.
9. A virtual image generating device, characterized in that: include: An acquisition module, configured to acquire an input image in response to a virtual image generation instruction, wherein the input image includes a reference object; An extraction module, used for extracting appearance features of the reference object in the input image; A generation module is used to inject the appearance features into a preset generation model so that the preset generation model generates an anthropomorphic virtual image with the appearance features.
10. A computer program product, characterized in that The method comprises a computer program, which implements the method according to any one of claims 1 to 8 when being executed by a processor.