Method, device, equipment and storage medium for obtaining virtual image

By fusing the first image generation model and the second image generation model, the problem of difficulty in obtaining sample images that retain the object's ontological features and have specific attributes is solved, and the quality and efficiency of virtual images are improved.

CN113705316BActive Publication Date: 2025-10-03TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Application Number
CN202110394129.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-13
Publication Date
2025-10-03
Estimated Expiration
2041-04-13

AI Technical Summary

Technical Problem

In the existing technology, it is difficult to obtain sample images that retain the object's ontological features and have specific attributes, resulting in low efficiency in training image generation models, which in turn affects the efficiency and quality of virtual image acquisition.

Method used

By fusing the first image generation model and the second image generation model, the target image generation model trained separately can focus on both the object ontology features and the target attributes, thereby improving the quality and acquisition efficiency of the virtual image.

Benefits of technology

While ensuring the quality of virtual images, the time consumption of acquiring the target image generation model is shortened and the efficiency of acquiring virtual images is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113705316B_ABST
    Figure CN113705316B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, device and storage medium for acquiring virtual images, and belongs to the field of artificial intelligence technology. The method includes: acquiring a target image generation model, the target image generation model is obtained by fusing a first image generation model and a second image generation model, the first image generation model is trained based on a sample original image that retains the object ontology characteristics of the sample object, and the second image generation model is trained based on a sample virtual image with target attributes; based on the target image generation model, acquiring a target virtual image corresponding to the original image of the target object. In the above manner, the target image generation model can pay attention to both the object ontology characteristics and the target attributes at the same time, which is conducive to ensuring the quality of the acquired virtual image. Both the sample original image and the sample virtual image are relatively easy to acquire, which is conducive to improving the efficiency of acquiring virtual images. It is possible to improve the efficiency of acquiring virtual images while ensuring the quality of the acquired virtual images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and storage medium for acquiring a virtual image. Background Art

[0002] With the development of artificial intelligence technology, more and more application scenarios require obtaining virtual images corresponding to the original image of an object, wherein the virtual image corresponding to the original image of an object retains the object's ontological characteristics and has certain specific attributes (such as anime style, cartoon style, etc.).

[0003] In related technologies, sample images that retain the object's ontological features and have specific attributes are obtained, and then an image generation model is trained using the obtained sample images. Then, a virtual image corresponding to the original image of a certain object is obtained based on the obtained image generation model.

[0004] In this approach, ensuring good quality in the virtual images obtained requires obtaining a sufficient number of sample images that retain the object's ontological features and possess specific attributes. In many cases, obtaining sample images that retain the object's ontological features and possess specific attributes is difficult, resulting in low efficiency in training the image generation model and, in turn, low efficiency in obtaining virtual images. If fewer sample images are obtained to improve virtual image acquisition efficiency, the quality of the virtual images generated by the image generation model will be poor. Summary of the Invention

[0005] The embodiments of the present application provide a method, apparatus, device, and storage medium for acquiring virtual images, which can be used to improve the efficiency of acquiring virtual images while ensuring the quality of the acquired virtual images. The technical solution is as follows:

[0006] In one aspect, an embodiment of the present application provides a method for acquiring a virtual image, the method comprising:

[0007] Obtaining a target image generation model, where the target image generation model is obtained by fusing a first image generation model and a second image generation model, wherein the first image generation model is trained based on a sample original image retaining object ontology features of the sample object, and the second image generation model is trained based on a sample virtual image having target attributes;

[0008] Based on the target image generation model, a target virtual image corresponding to the original image of the target object is obtained, where the target virtual image retains the object ontology features of the target object and has the target attributes.

[0009] In another aspect, a device for acquiring a virtual image is provided, the device comprising:

[0010] a first acquisition unit, configured to acquire a target image generation model, the target image generation model being obtained by fusing a first image generation model and a second image generation model, the first image generation model being trained based on a sample original image retaining object ontology features of the sample object, and the second image generation model being trained based on a sample virtual image having target attributes;

[0011] A second acquisition unit is configured to acquire a target virtual image corresponding to the original image of the target object based on the target image generation model, where the target virtual image retains the object ontology features of the target object and has the target attributes.

[0012] In one possible implementation, the apparatus further includes:

[0013] a third acquisition unit, configured to acquire the sample original image and the sample virtual image; train the first image generation model based on the sample original image; and train the second image generation model based on the sample virtual image;

[0014] A fusion unit is used to fuse the first image generation model and the second image generation model to obtain the target image generation model.

[0015] In one possible implementation, the fusion unit is used to determine at least one fusion method, any fusion method is used to indicate a method of fusing the first image generation model and the second image generation model; based on the at least one fusion method, at least one candidate image generation model is obtained, and any candidate image generation model is obtained by fusing the first image generation model and the second image generation model based on any fusion method; in response to the image generation function of the third image generation model meeting the reference condition, the third image generation model is used as the target image generation model, and the third image generation model is a candidate image generation model that meets the selection condition among the at least one candidate image generation model.

[0016] In one possible implementation, the first image generation model and the second image generation model both include a reference number layer network, and any fusion method includes a method for determining the target network parameters corresponding to the reference number layer networks respectively; the fusion unit is also used to determine the target network parameters corresponding to the reference number layer networks respectively based on the method for determining the target network parameters corresponding to the reference number layer networks respectively included in any fusion method; based on the target network parameters corresponding to the reference number layer networks respectively, the parameters of the reference number layer networks in the specified image generation model are adjusted, and the image generation model obtained after adjustment is used as any candidate image generation model; wherein, the method for determining the target network parameters corresponding to any layer network included in any fusion method is used to indicate the relationship between the target network parameters corresponding to any layer network and at least one of the first network parameter and the second network parameter; the first network parameter is the parameter of any layer network in the first image generation model, and the second network parameter is the parameter of any layer network in the second image generation model.

[0017] In one possible implementation, the apparatus further includes:

[0018] a fourth acquiring unit, configured to acquire a supplementary image in response to the image generation function of the third image generation model not satisfying the reference condition, wherein the supplementary image is used to supplement the sample virtual image;

[0019] The third acquisition unit is further configured to train the second image generation model based on the supplementary image and the sample virtual image to obtain an updated second image generation model;

[0020] The fusion unit is further used to fuse the first image generation model and the updated second image generation model to obtain a fourth image generation model; in response to the image generation function of the fourth image generation model satisfying the reference condition, the fourth image generation model is used as the target image generation model.

[0021] In one possible implementation, the fourth acquisition unit is used to acquire enhanced image features, where the enhanced image features are used to enhance the image generation function of the second image generation model; the second image generation model is called to process the enhanced image features to obtain the supplementary image.

[0022] In one possible implementation, the fourth acquisition unit is used to obtain an image driving model, which is trained based on a sample video having the target attributes; calling the image driving model to drive the sample virtual image to obtain an enhanced video corresponding to the sample virtual image; and extracting a video frame from the enhanced video corresponding to the sample virtual image as the supplementary image.

[0023] In one possible implementation, the second acquisition unit is used to obtain original image features corresponding to the original image of the target object; obtain target image features based on the original image features; call the target image generation model to process the target image features to obtain a target virtual image corresponding to the original image of the target object.

[0024] In a possible implementation, the second acquisition unit is further configured to convert the original image features using an image feature conversion method corresponding to the target image generation model, and use the image features obtained after the conversion as the target image features.

[0025] In one possible implementation, the second acquisition unit is configured to call the target image generation model, process the candidate image features, and obtain a candidate virtual image corresponding to the candidate image features; obtain a target image translation model based on the candidate original image corresponding to the candidate image features and the candidate virtual image corresponding to the candidate image features, where the candidate original image corresponding to the candidate image features is an image identified by the candidate image features that retains the object's ontological features; and call the target image translation model to process the original image of the target object to obtain a target virtual image corresponding to the original image of the target object.

[0026] In one possible implementation, the second acquisition unit is further used to call the initial image translation model, process the candidate original image corresponding to the candidate image feature, and obtain a candidate predicted image corresponding to the candidate image feature; determine a loss function based on the difference between the candidate predicted image corresponding to the candidate image feature and the candidate virtual image corresponding to the candidate image feature; and use the loss function to train the initial image translation model to obtain the target image translation model.

[0027] In one possible implementation, the apparatus further includes:

[0028] an acquisition unit, configured to acquire an original image of the interactive object in response to a virtual image display instruction for the target attribute;

[0029] A display unit is used to call a target image prediction model, process the original image of the interactive object, and obtain a virtual image corresponding to the original image of the interactive object, wherein the target image prediction model is trained based on the original image of the target object and the target virtual image corresponding to the original image of the target object; and display the virtual image corresponding to the original image of the interactive object.

[0030] On the other hand, a computer device is provided, comprising a processor and a memory, wherein the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor so that the computer device implements any of the above-mentioned methods for obtaining a virtual image.

[0031] On the other hand, a computer-readable storage medium is provided, in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor to enable a computer to implement any of the above-mentioned methods for obtaining a virtual image.

[0032] In another aspect, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform any of the aforementioned methods for acquiring a virtual image.

[0033] The technical solutions provided by the embodiments of the present application bring at least the following beneficial effects:

[0034] In an embodiment of the present application, a target image generation model for acquiring a virtual image is obtained by fusing a first image generation model and a second image generation model, wherein the first image generation model can focus on the object ontology features, and the second image generation model can focus on the target attributes. Therefore, the target image generation model can focus on both the object ontology features and the target attributes at the same time, which is beneficial to ensuring the quality of the acquired virtual image. In addition, compared to sample images that retain the object ontology features and have target attributes, the sample original images based on which the first image generation model is trained and the sample virtual images based on which the second image generation model is trained are both easier to obtain, which is beneficial to shortening the acquisition time of the target image generation model, thereby improving the efficiency of acquiring virtual images. In other words, the method for acquiring virtual images provided in an embodiment of the present application can improve the efficiency of acquiring virtual images while ensuring the quality of the acquired virtual images. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0036] Figure 1 is a schematic diagram of an implementation environment of a method for obtaining a virtual image provided in an embodiment of the present application;

[0037] Figure 2 is a flow chart of a method for obtaining a virtual image provided by an embodiment of the present application;

[0038] Figure 3 This is a flowchart of a process of fusing a first image generation model and a second image generation model to obtain a target image generation model, provided by an embodiment of the present application;

[0039] Figure 4 is a schematic diagram of an image generated by at least one image generation model after performing an image generation test on at least one candidate image generation model, provided by an embodiment of the present application;

[0040] Figure 5 is a flowchart of a process for obtaining a fourth image generation model provided by an embodiment of the present application;

[0041] Figure 6 is a schematic diagram of a process for obtaining a target image generation model provided by an embodiment of the present application;

[0042] Figure 7 This is a schematic diagram of a process of displaying a virtual image that retains object characteristics and has target attributes, provided by an embodiment of the present application;

[0043] Figure 8 Schematic diagram of an original image of an interactive object and a virtual image that retains the object ontology features of the interactive object and has target attributes, provided by an embodiment of the present application;

[0044] Figure 9 is a schematic diagram of a device for acquiring a virtual image provided by an embodiment of the present application;

[0045] Figure 10 is a schematic diagram of a device for acquiring a virtual image provided by an embodiment of the present application;

[0046] Figure 11 This is a schematic diagram of the structure of a server provided in an embodiment of the present application;

[0047] Figure 12This is a schematic diagram of the structure of a terminal provided in an embodiment of the present application. DETAILED DESCRIPTION

[0048] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0049] It should be noted that the terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments of the application described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. Instead, they are merely examples of devices and methods consistent with certain aspects of the present application as detailed in the appended claims.

[0050] In an exemplary embodiment, the method for obtaining a virtual image provided in the embodiment of the present application can be applied to the field of artificial intelligence technology. Next, artificial intelligence technology is introduced.

[0051] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0052] Artificial intelligence technology is a comprehensive discipline covering a wide range of fields, encompassing both hardware and software technologies. Basic AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. Artificial intelligence software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning. The methods for acquiring virtual images provided in the embodiments of this application involve computer vision and machine learning technologies.

[0053] Computer vision (CV) technology is the study of how machines can "see." Specifically, it refers to the use of cameras and computers to replace the human eye in identifying and measuring objects, and further image processing to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, aiming to build artificial intelligence systems that can extract information from images or multidimensional data. Computer vision technology generally includes virtual image acquisition, image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / action recognition, three-dimensional object reconstruction, 3D (three-dimensional) technology, virtual reality, augmented reality, and map construction. It also includes common biometric recognition technologies such as facial recognition and fingerprint recognition.

[0054] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning through demonstration.

[0055] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, automatic driving, drones, robots, smart medical care, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0056] In an exemplary embodiment, the method for obtaining a virtual image provided in the embodiment of the present application is implemented in a blockchain system. The images (sample original image, sample virtual image, original image of the target object, target virtual image, etc.) and models (target image generation model, first image generation model, second image generation model, etc.) involved in the method for obtaining a virtual image provided in the embodiment of the present application are all stored on the blockchain in the blockchain system, and the security and reliability of the images and models are relatively high.

[0057] Figure 1A schematic diagram of an implementation environment of the method for obtaining a virtual image provided by an embodiment of the present application is shown. The implementation environment may include: a terminal 11 and a server 12.

[0058] The method for obtaining a virtual image provided in the embodiment of the present application can be executed by the terminal 11, can be executed by the server 12, or can be executed jointly by the terminal 11 and the server 12, and the embodiment of the present application does not limit this. In the case where the method for obtaining a virtual image provided in the embodiment of the present application is jointly executed by the terminal 11 and the server 12, the server 12 undertakes the primary computing work and the terminal 11 undertakes the secondary computing work; or, the server 12 undertakes the secondary computing work and the terminal 11 undertakes the primary computing work; or, the server 12 and the terminal 11 adopt a distributed computing architecture to perform collaborative computing.

[0059] In one possible implementation, the terminal 11 may be any electronic product that can interact with a user through one or more methods such as a keyboard, touchpad, touch screen, remote control, voice interaction, or handwriting device, such as a PC (Personal Computer), a mobile phone, a smart phone, a PDA (Personal Digital Assistant), a wearable device, a PPC (Pocket PC), a tablet computer, a smart car computer, a smart TV, a smart speaker, etc. The server 12 may be a single server, a server cluster consisting of multiple servers, or a cloud computing service center. The terminal 11 establishes a communication connection with the server 12 via a wired or wireless network.

[0060] Those skilled in the art should understand that the above-mentioned terminal 11 and server 12 are only examples. Other existing or future terminals or servers that are applicable to this application should also be included in the scope of protection of this application and are included here by reference.

[0061] Based on the above Figure 1 In the implementation environment shown, the embodiment of the present application provides a method for obtaining a virtual image. Taking the method applied to a computer device as an example, the computer device can be a server or a terminal, and the embodiment of the present application does not limit this. Figure 2 As shown, the method for obtaining a virtual image provided in an embodiment of the present application includes the following steps 201 and 202.

[0062] In step 201, a target image generation model is obtained, which is obtained by fusing a first image generation model and a second image generation model. The first image generation model is trained based on a sample original image that retains the object ontology features of the sample object, and the second image generation model is trained based on a sample virtual image with target attributes.

[0063] The target image generation model is a model based on which a virtual image that retains the object ontology features and has the target attributes is obtained. In an embodiment of the present application, the target image generation model is obtained by fusing a first image generation model and a second image generation model. The first image generation model is trained based on a sample original image that retains the object ontology features of the sample object, and the second image generation model is trained based on a sample virtual image with the target attributes. Compared with the sample images that retain the object ontology features and have the target attributes in the related art, the sample original images based on which the first image generation model is trained and the sample virtual images based on which the second image generation model is trained are both easier to obtain, and the acquisition time of the first image generation model and the second image generation model are both shorter, which is conducive to shortening the acquisition time of the target image generation model, thereby improving the efficiency of obtaining virtual images.

[0064] In addition, the first image generation model trained based on the sample original image that retains the object ontology features of the sample object can focus on the object ontology features, and the second image generation model trained based on the sample virtual image with target attributes can be related to the target attributes. Therefore, the target image generation model obtained by fusing the first image generation model and the second image generation model can focus on both the object ontology features and the target attributes at the same time, which is conducive to ensuring the quality of the acquired virtual image.

[0065] A sample object refers to any entity object. The embodiments of the present application do not limit the type of entity object. Exemplarily, the type of entity object is a human face, an animal face, or an object. The sample original image retains the object ontology features of the sample object. The object ontology features of the sample object are used to uniquely identify the sample object. Exemplarily, in the case where the type of the sample object is a human face, the face of the sample object can be known based on the object ontology features of the sample object. Preserving the object ontology features can be achieved by retaining the face shape, texture, sub-parts, etc. of the face. Exemplarily, the sample original image refers to the real image obtained after image acquisition of the sample object. For example, in the case where the type of the sample object is a human face, the sample original image refers to a real face image.

[0066] The target attribute refers to an attribute of an image. The target attribute is set according to actual needs or flexibly adjusted according to the application scenario. The embodiments of the present application do not limit this. For example, the target attribute refers to a specific style (such as anime style, cartoon style, etc.); or, the target attribute refers to a specific special effect (such as European and American children's special effects, Disney children's special effects, etc.). The sample virtual image has a target attribute. In an exemplary embodiment, the sample virtual image includes objects of the same type as the sample object to ensure the acquisition quality of the target image generation model. The embodiments of the present application refer to the objects included in the sample virtual image as virtual objects for easy distinction. It should be noted that the virtual objects included in the sample virtual images are objects obtained by virtualizing real objects. The sample virtual images may retain the object ontology characteristics of the real objects corresponding to the virtual objects, or may not retain the object ontology characteristics of the real objects corresponding to the virtual objects. The embodiments of the present application do not limit this.

[0067] Acquiring the target image generation model in step 201 may refer to directly extracting a pre-stored target image generation model, or may refer to fusing the first image generation model with the second image generation model to obtain the target image generation model, which is not limited in the present embodiment. In one possible implementation, the process of acquiring the target image generation model is as follows: acquiring a sample original image and a sample virtual image; training a first image generation model based on the sample original image; training a second image generation model based on the sample virtual image; and fusing the first image generation model and the second image generation model to obtain the target image generation model.

[0068] The embodiment of the present application does not limit the method of obtaining the sample original images and the sample virtual images. For example, the sample original images can be obtained by crawling from the Internet, or directly extracted from the original image library, or uploaded by the staff; the sample virtual images can be obtained by crawling from the Internet, or extracted from the virtual image library, or uploaded by the staff. It should be noted that the embodiment of the present application does not limit the number of sample original images and the number of sample virtual images, which can be flexibly set according to actual conditions. For example, the number of sample original images is tens of thousands; the number of sample virtual images is dozens to hundreds. The time required to obtain dozens to hundreds of sample virtual images with target attributes is relatively short, which is conducive to improving the efficiency of obtaining the second image generation model.

[0069] The first image generation model has the function of generating an image that retains the ontological features of the object. For example, taking the type of the sample object as a human face, the first image generation model has the function of generating a real person face image that retains the face identity information. In an exemplary embodiment, based on the sample original image, the method of training the first image generation model is as follows: obtaining a first basic image generation model; based on the sample original image, training the first basic image generation model to obtain the first image generation model. The first basic image generation model can refer to an image generation model that has not undergone any training, or it can refer to an image generation model that has been pre-trained, and the embodiments of the present application are not limited to this.

[0070] Exemplarily, the first image generation model is a generative adversarial model including a generative model and a discriminative model. The first image generation model is obtained by training the first basic image generation model based on the sample original image using a generative adversarial training method. The embodiment of the present application does not limit the model structure of the first image generation model. Exemplarily, the first image generation model is a StyleGAN2 (Style Generative Adversarial Networks 2, second-generation style generative adversarial network) model, a Progressive GAN (progressive generative adversarial network) model, etc.

[0071] In the process of training the first basic image generation model based on the sample original image using a generative adversarial training method, the training goal of the discriminant model in the first basic image generation model is to accurately determine whether an image is a false image generated by the image generation model or a true image in the sample original image. The training goal of the generative model in the first basic image generation model is to generate an image that is as close to the sample original image as possible, so that the discriminant model cannot determine whether the generated image is a false image or a true image.

[0072] Exemplarily, during the training of the first basic image generation model based on the sample original image, image enhancement is performed on the sample original image, and then the first basic image generation model is trained based on the enhanced sample original image. Exemplarily, methods for enhancing the sample original image include, but are not limited to, rotation, noise addition, and cropping. In this manner, applying image enhancement online to the training of the image generation model is beneficial for improving the quality of images generated by the image generation model.

[0073] Exemplarily, the number of first image generation models is one. During the process of training the first basic image generation model based on the sample original image, the training effect is continuously monitored, and the image generation model obtained when the training effect meets the conditions is used as the first image generation model. Exemplarily, the method of monitoring the training effect is to test the image generation effect of the image generation model obtained during the training process. The quality of the image generation effect of the image generation model can be judged by professionals or by a computer device according to a judgment rule, which is not limited in the embodiment of the present application. The judgment rule can be set by a professional and uploaded to the computer device.

[0074] The second image generation model has the function of generating an image with target attributes. For example, taking the target attribute as a cartoon style, the second image generation model can generate cartoon-style images. In one possible implementation, the second image generation model is trained based on the sample virtual image by: obtaining a second base image generation model; and training the second base image generation model based on the sample virtual image to obtain the second image generation model.

[0075] The second basic image generation model can refer to an untrained image generation model, a pre-trained image generation model, or a first image generation model trained based on the sample original image, and this is not limited in the present embodiment. If the second basic image generation model refers to the first image generation model trained based on the sample original image, the process of obtaining the second image generation model can be regarded as the process of fine-tuning the first image generation model using the sample virtual object.

[0076] In an exemplary embodiment, the second basic image generation model refers to the first image generation model obtained by training based on the sample original image, and the training of the second basic image generation model is stopped while ensuring that the model generates target attributes with sufficient detail and does not produce a collapsed image.

[0077] In an exemplary embodiment, the second image generation model is a generative adversarial model comprising a generative model and a discriminative model. The second image generation model is obtained by training the second basic image generation model using a generative adversarial training approach based on sample virtual images. This embodiment of the application does not limit the model structure of the second image generation model; illustratively, the model structure of the second image generation model is the same as that of the first image generation model.

[0078] In the process of training the second basic image generation model based on the sample virtual image using the generative adversarial training method, the training goal of the discriminant model in the second basic image generation model is to accurately determine whether an image is a false image generated by the image generation model or a true image in the sample virtual image, and the training goal of the generative model in the second basic image generation model is to generate an image as close as possible to the sample virtual image, so that the discriminant model cannot determine whether the generated image is a false image or a true image.

[0079] In an exemplary embodiment, the number of second image generation models is one or more, which is limited by the embodiments of the present application. In the case where the number of second image generation models is one, in the process of training the second basic image generation model based on the sample virtual image, the training effect is continuously monitored, and the image generation model obtained when the training effect meets the conditions is used as the second image generation model. In the case where the number of second image generation models is multiple, different second image generation models are image generation models obtained under different training durations in the process of training the second basic image generation model based on the sample virtual image. The process of obtaining multiple second image generation models can be regarded as a process of controlling the learning ability of the fine-tuning model by supervising the intermediate training results of the model.

[0080] In an exemplary embodiment, the process of obtaining the second image generation model can be regarded as a process of fine-tuning the first image generation model using a sample virtual object and in the case where the number of second image generation models is multiple, the second image generation model obtained under a shorter training time can generate images that retain more object ontology features and have a weaker degree of target attribution; the second image generation model obtained under a longer training time can generate images that retain fewer object ontology features and have a stronger degree of target attribution.

[0081] After obtaining the first image generation model and the second image generation model, the first image generation model and the second image generation model are fused to obtain a target image generation model. The target image generation model has the function of generating an image that retains the object ontology features and has the target attributes. In one possible implementation, see Figure 3 The process of fusing the first image generation model and the second image generation model to obtain the target image generation model includes the following steps 301 to 304.

[0082] Step 301: Determine at least one fusion method, where any fusion method is used to indicate a method of fusing a first image generation model and a second image generation model.

[0083] The fusion method is used to indicate the method for fusing the first image generation model and the second image generation model. Based on the fusion method, it is possible to determine how to fuse the first image generation model and the second image generation model. In one possible implementation, the process of determining at least one fusion method is: displaying the model parameters of the first image generation model and the model parameters of the second image generation model to a professional for review, and obtaining at least one fusion method uploaded by the professional after analyzing the parameters of the first image generation model and the parameters of the second image generation model. In another possible implementation, the at least one fusion method is obtained by a computer device after analyzing the parameters of the first image generation model and the parameters of the second image generation model according to parameter analysis rules.

[0084] In one possible implementation, both the first image generation model and the second image generation model include a reference quantity layer network, that is, the model structure of the first image generation model and the second image generation model is the same. In this case, the candidate image generation model to be obtained based on the fusion method also includes a reference quantity layer network. The embodiments of the present application do not limit the specific value of the reference quantity, which can be determined based on the model structure of the image generation model and can also be flexibly adjusted according to the actual application scenario.

[0085] In an exemplary embodiment, the image generation model is a generative adversarial model, and different layers of the network in the image generation model are different resolution layers. Different resolution layers focus on features at different levels. For example, when the type of sample object is a face, the low-resolution layer in the image generation model focuses on features such as the posture and face shape of the face in the generated image; the high-resolution layer in the image generation model focuses on features such as light and texture in the generated image.

[0086] In an exemplary embodiment, when both the first image generation model and the second image generation model include reference number layer networks, each fusion method includes a method for determining target network parameters corresponding to each reference number layer network. Different fusion methods include different methods for determining target network parameters corresponding to each reference number layer network. It should be noted that the different methods for determining target network parameters corresponding to each reference number layer network included in the two fusion methods may indicate that the methods for determining target network parameters corresponding to a certain layer or layers of networks included in the two fusion methods are different.

[0087] The present application embodiment uses a fusion method as an example for description. The method for determining the target network parameters corresponding to any layer of the network included in any fusion method is used to indicate the relationship between the target network parameters corresponding to any layer of the network and at least one of a first network parameter and a second network parameter. The first network parameter is a parameter of the network at any layer in the first image generation model, and the second network parameter is a parameter of the network at any layer in the second image generation model.

[0088] The relationship between the target network parameter corresponding to any layer of the network and at least one of the first network parameter and the second network parameter is used to clarify how the target network parameter corresponding to any layer of the network is determined based on at least one of the first network parameter and the second network parameter. For example, if there are multiple second image generation models, there are also multiple second network parameters, and each second image generation model corresponds to one second network parameter.

[0089] Exemplarily, the relationship between the target network parameter corresponding to any layer of network and at least one of the first network parameter and the second network parameter may refer to the relationship between the target network parameter corresponding to any layer of network and the first network parameter, or may refer to the relationship between the target network parameter corresponding to any layer of network and one or more second network parameters, or may refer to the relationship between the target network parameter corresponding to any layer of network and the first network parameter and one or more second network parameters, which is not limited in this embodiment of the present application. Based on the relationship between the target network parameter corresponding to any layer of network and at least one of the first network parameter and the second network parameter, the target network parameter corresponding to any layer of network can be clearly determined.

[0090] It should be noted that the methods for determining the target network parameters corresponding to different network layers included in the same fusion method may be the same or different, and this embodiment of the present application does not limit this. The methods for determining the target network parameters corresponding to the same layer network included in different fusion methods may be the same or different, and this embodiment of the present application also does not limit this.

[0091] For example, assume there is one first image generation model, denoted as Model A; and five second image generation models, denoted as Model B1, Model B2, Model B3, Model B4, and Model B5. Both the first and second image generation models include 18 networks. The target network parameters for the first network layer included in fusion method 1 are determined by using the parameters of the first network layer in Model A as the target network parameters for the first network layer. The target network parameters for the seventh network layer included in fusion method 1 are determined by weighting the parameters of the seventh network layer in Model A, the seventh network layer in Model B1, the seventh network layer in Model B2, the seventh network layer in Model B3, the seventh network layer in Model B4, and the seventh network layer in Model B5, using specified weights, and using the resulting weighted sum as the target network parameters for the seventh network layer. The specified weights are given by the method for determining the target network parameters for the seventh network layer included in fusion method 1. The target network parameters corresponding to the 18th layer network in fusion mode 1 are determined by using the parameters of the 18th layer network in model B5 as the target network parameters corresponding to the 18th layer network.

[0092] For example, based on the above example, the target network parameters for the first layer of the network included in fusion method 2 are determined by using the parameters of the first layer of the network in model A as the target network parameters for the first layer of the network. The target network parameters for the seventh layer of the network included in fusion method 2 are determined by using the average parameters of the seventh layer of the network in model A and the seventh layer of the network in model B4 as the target parameters for the seventh layer of the network. The target network parameters for the eighteenth layer of the network included in fusion method 2 are determined by using the parameters of the eighteenth layer of the network in model B3 as the target network parameters for the eighteenth layer of the network.

[0093] In an exemplary embodiment, when there is one first image generation model and multiple second image generation models, the determination methods included in different fusion methods all involve the first image generation model and at least one second image generation model. The second image generation models involved in the different fusion methods may be the same or different. For example, the second image generation models involved in the determination methods included in fusion method 1 are Model B1, Model B2, Model B3, Model B4, and Model B5; and the second image generation models involved in the determination methods included in fusion method 2 are Model B3 and Model B4.

[0094] In an exemplary embodiment, for the case where the second image generation model involved in the determination method included in different fusion methods is the same, the different fusion methods can be regarded as fusion methods corresponding to different parameter fusion schemes; for the case where the second image generation models involved in the determination method included in different fusion methods are different, since the training time corresponding to the second image generation model is different, the different fusion methods can be regarded as fusion methods corresponding to different training time.

[0095] It should be noted that the method of determining the target network parameters corresponding to certain layer networks in the above fusion method is only an example, and the embodiments of the present application are not limited thereto and can be flexibly set according to actual circumstances.

[0096] Step 302: Based on at least one fusion method, obtain at least one candidate image generation model, where any candidate image generation model is obtained by fusing the first image generation model and the second image generation model based on any fusion method.

[0097] After determining at least one fusion method, a candidate image generation model is obtained based on each fusion method. The number of candidate image generation models is the same as the number of fusion methods. The embodiment of the present application is described by taking the example of obtaining any candidate image generation model based on any fusion method. Any candidate image generation model is obtained by fusing the first image generation model and the second image generation model based on any fusion method.

[0098] In one possible implementation, based on any fusion method, the process of obtaining any candidate image generation model is: based on the method of determining the target network parameters corresponding to the reference number layer networks included in any fusion method, determine the target network parameters corresponding to the reference number layer networks; based on the target network parameters corresponding to the reference number layer networks, adjust the parameters of the reference number layer networks in the specified image generation model, and use the image generation model obtained after adjustment as any candidate image generation model.

[0099] As described in step 301 regarding the method for determining the target network parameters corresponding to any layer of network included in any fusion method, the target network parameters corresponding to any layer of network can be determined based on the method for determining the target network parameters corresponding to any layer of network included in any fusion method. Therefore, based on the method for determining the target network parameters corresponding to the reference number of layer networks included in any fusion method, the target network parameters corresponding to the reference number of layer networks can be determined. It should be noted that the target network parameters corresponding to the reference number of layer networks here are determined based on any fusion method, and the target network parameters corresponding to the reference number of layer networks determined based on different fusion methods may differ.

[0100] It should be noted that the target network parameters corresponding to the reference number layer networks determined according to different fusion methods are different, which means that the target network parameters corresponding to a certain layer network or some layer networks determined according to different fusion methods are different.

[0101] After determining the target network parameters corresponding to the reference number layer networks according to any fusion method, the parameters of the reference number layer networks in the specified image generation model are adjusted based on the determined target network parameters corresponding to the reference number layers, and the image generation model obtained after the adjustment is used as any candidate image generation model. The specified image generation model is the image generation model corresponding to any fusion method that requires parameter adjustment. The specified image generation models that require parameter adjustment corresponding to different fusion methods can be the same or different, and this is not limited in this embodiment of the present application.

[0102] The embodiments of the present application do not limit the method for determining the designated image generation model. For example, the designated image generation model is the first image generation model, or the designated image generation model is any second image generation model, or the designated image generation model is any model that includes a reference number of layer networks. Regardless of the determination method, the determined designated image generation model includes the reference number of layer networks.

[0103] The purpose of adjusting the parameters of the reference number of layer networks in the specified image generation model based on the target network parameters corresponding to the reference number of layer networks is to make the parameters of any layer network in the image generation model obtained after the adjustment be the target network parameters corresponding to the any layer network. Exemplarily, in the process of adjusting the parameters of the reference number of layer networks in the specified image generation model based on the target network parameters corresponding to the reference number of layer networks, if the parameters of a certain layer network in the specified image generation model are the same as the target network parameters corresponding to the layer network, then the parameters of the layer network in the specified image generation model remain unchanged; if the parameters of a certain layer network in the specified image generation model are different from the target network parameters corresponding to the layer network, then the parameters of the layer network in the specified image generation model are replaced with the target network parameters corresponding to the layer network. After the parameters of the reference number of layer networks in the specified image generation model are adjusted, the image generation model obtained after the adjustment is used as any candidate image generation model.

[0104] The above description uses the example of obtaining any candidate image generation model based on any fusion method. Using the methods described above, a candidate image generation model can be obtained based on each fusion method, and then at least one candidate image generation model can be obtained. For example, the process of obtaining at least one candidate image generation model can be considered a model fusion process.

[0105] In one possible implementation, after obtaining at least one candidate image generation model, a candidate image generation model that meets a selection condition is determined from the at least one candidate image generation model, and the candidate image generation model that meets the selection condition is used as the third image generation model.

[0106] The candidate image generation model that meets the selection criteria is set based on experience or flexibly adjusted according to the actual application scenario, and the embodiments of the present application do not limit this. Exemplarily, the candidate image generation model that meets the selection criteria is a candidate image generation model whose image generation effect meets the specified conditions in at least one candidate image generation model. The image generation effect of the candidate image generation model can be determined by performing an image generation test on the candidate image generation model and then analyzing the image generated by the image generation model. The image generation effect meeting the specified conditions means that the image generation effect is closest to the image generation effect required by the actual application scenario. Which image generation effect is closest to the actual application scenario is set by professionals based on experience.

[0107] Exemplarily, after performing an image generation test on at least one candidate image generation model, the image generated by at least one image generation model is as follows: Figure 4 As shown. Figure 4 In the first column, the real person original image is used. The process of image testing at least one candidate image generation model is achieved by inputting the image features of the real person original image into at least one candidate image generation model. Figure 4 In the images shown, the target attribute is the American comic style. In addition to the original real-life images, the images from left to right represent the images generated by the candidate image generation model determined by the fusion method corresponding to different parameter fusion schemes. The further to the right, the stronger the American comic stylization; the images from top to bottom represent the images generated by the candidate image generation model determined by the fusion method corresponding to different training times (from short to long training time). The longer the training time, the stronger the American comic stylization. Figure 4 Among the images generated by the other six candidate image generation models except the real-life image shown, the one whose effect is closer to the image generation effect required by the actual application scenario will be used as the candidate image generation model that meets the selection conditions, and the candidate image generation model that generates the image will be used as the third image generation model.

[0108] In an exemplary embodiment, a candidate image generation model with better image generation effect is a model that selects appropriate parameters of a second image generation model at a specific high-resolution layer and selects appropriate parameters of a first image generation model at a specific low-resolution layer, so that the generated virtual image can have more textures and details of the target attributes and retain more object ontology features.

[0109] After determining the third image generation model, a determination is made as to whether the image generation function of the third image generation model satisfies the reference condition. If the image generation function of the third image generation model satisfies the reference condition, step 303 is executed; if the image generation function of the third image generation model does not satisfy the reference condition, step 304 is executed. Whether the image generation function of a particular image generation model satisfies the reference condition is determined based on empirical experience or can be flexibly adjusted based on the application scenario, and this is not limited in the present embodiment. For example, satisfying the reference condition for the image generation function of a particular image generation model means that the image generation model is capable of generating a desired virtual image.

[0110] Exemplarily, the image generation function of a certain image generation model meeting the reference conditions means that the image generation model is capable of generating a virtual image of a virtual object with a specified pose. The specified pose is set as needed or flexibly adjusted based on the actual application scenario. Exemplarily, for a virtual object of a human face, the specified poses include, but are not limited to, frontal, low-angle profile, high-angle profile, open mouth, closed mouth, open eyes, and closed eyes.

[0111] Step 303: In response to the image generation function of the third image generation model satisfying the reference condition, the third image generation model is used as the target image generation model.

[0112] When the image generation function of the third image generation model satisfies the reference condition, it indicates that there is no need to continue obtaining other image generation models, and the third image generation model is directly used as the target image generation model. The third image generation model is a candidate image generation model that satisfies the selection condition among the at least one candidate image generation model.

[0113] Step 304: In response to the image generation function of the third image generation model not satisfying the reference condition, acquiring a fourth image generation model; in response to the image generation function of the fourth image generation model satisfying the reference condition, using the fourth image generation model as the target image generation model.

[0114] When the image generation function of the third image generation model does not meet the reference condition, it indicates that a fourth image generation model needs to be further acquired to improve the image generation function of the image generation model. Compared with the third image generation model, the image generation function of the fourth image generation model is closer to the reference condition.

[0115] In one possible implementation, see Figure 5 The process of obtaining the fourth image generation model includes the following steps 501 to 503.

[0116] Step 501: Acquire a supplementary image, where the supplementary image is used to supplement the sample virtual image.

[0117] In an exemplary embodiment, training a relatively stable image generation model usually requires tens of thousands of data. However, in the actual research and development process, it is often impossible to collect enough data, especially for relatively good sample virtual images with target attributes, which may generally only have a few hundred or even dozens of data. Not only is the data quality not high, but the data coverage is also very low. For example, when the type of virtual object is a human face, sample virtual images with the virtual object's postures of closed eyes, open mouth, and large-angle side faces are very scarce. However, in actual scenarios, these are common facial postures. In other words, in the embodiment of the present application, it is believed that the reason why the image generation function of the third image generation model does not meet the reference conditions is that the quality of the sample virtual images is not high or the coverage of the sample virtual images is low.

[0118] When the quality of the sample virtual image is not high or the coverage of the sample virtual image is low, although the third image generation model obtained by the model fusion method can generate an image that retains the object's ontological features and has target attributes, the generation effect of the image for the posture not covered by the sample virtual image is poor. For example, when the type of virtual object is a human face, the sample virtual image covers the posture of the front face, but does not cover postures such as closed eyes and wide-angle side faces. The third image generation model has a better generation effect for the image of the virtual object including the front face, but has a poor generation effect for the image of the virtual object including closed eyes and wide-angle side faces.

[0119] Supplementary images are used to supplement the sample virtual images, thereby enriching the initially collected sample virtual images. Supplementary images are images with target attributes not covered by the sample virtual images. The process of acquiring supplementary images can be considered as data augmentation for the sample virtual images. Through data augmentation, insufficient sample data can be generated when data is lacking, significantly reducing the reliance on sample virtual images. Data augmentation can effectively reduce the number of sample virtual images with target attributes that need to be collected, thereby improving the acquisition efficiency of the second image generation model.

[0120] In one possible implementation, methods for obtaining the supplementary image include but are not limited to the following two methods:

[0121] Method 1: Obtain enhanced image features, which are used to enhance the image generation function of the second image generation model; call the second image generation model to process the enhanced image features to obtain a supplementary image.

[0122] In this manner, a supplemental image is acquired using enhanced image features, and the enhanced image features are used to enhance the image generation function of the second image generation model. The supplemental image is acquired using the enhanced image features that enhance the image generation function of the second image generation model, and the acquired supplemental image is then used to enhance the image generation function of the second image generation model.

[0123] The enhanced image features are determined based on the image generation function that needs to be enhanced in the second image generation model, which is not limited in the embodiments of the present application. The image generation function of the second image generation model is used to indicate what kind of virtual image the second image generation model can generate. In one possible implementation, the enhanced image features are determined by a professional based on the image generation function that the second image generation model already has and uploaded to a computer device. In another implementation, the enhanced image features are automatically obtained by the computer device by analyzing the image generation function that the second image generation model already has according to image generation function analysis rules.

[0124] In an exemplary embodiment, the enhanced image features are obtained by analyzing the image generation function already possessed by the second image generation model to determine the various types of images that the second image generation model can generate; comparing the various types of images that the second image generation model can generate with the images of the required type to determine which types of images the second image generation model cannot generate; and using the image features used to indicate the types of images that the second image generation model cannot generate as enhanced image features. Exemplarily, the number of enhanced image features may be one or more, and this is not limited in the embodiments of the present application. The required type is set based on experience or flexibly adjusted according to the actual application scenario.

[0125] For example, the image features used to indicate the type of image that the second image generation model cannot generate can be obtained by editing randomly generated image features, or by editing image features of certain real images, which is not limited in this embodiment of the present application. For example, the image features can also be called latent vectors.

[0126] The process of acquiring enhanced image features can be viewed as a process of gradually expanding training data through image feature editing. For example, if a certain image feature is a feature of an image of a virtual object with eyes open, then by editing this image feature toward eyes closed, an image feature indicating an image of a virtual object with eyes closed can be obtained. Using the image feature indicating an image of a virtual object with eyes closed, an image of a virtual object with eyes closed can be obtained.

[0127] Exemplarily, the process of editing an image feature is a process of editing the image feature using a direction vector corresponding to a sub-feature in the image feature that needs to be adjusted. The sub-feature in the image feature that needs to be adjusted is determined based on actual conditions, and the embodiments of the present application do not limit this. Exemplarily, if the sub-feature in the image feature that needs to be adjusted is a facial angle sub-feature, the image feature is edited using a direction vector corresponding to the facial angle sub-feature, thereby editing the image feature of the image of a virtual object including a frontal face into the image feature of an image of a virtual object including a side face.

[0128] Exemplarily, the direction vector corresponding to the sub-feature in the image feature is used to edit the sub-feature, and the direction vector corresponding to the sub-feature is set based on experience, or flexibly adjusted according to the method of obtaining the image feature, which is not limited in the embodiments of the present application. Exemplarily, in the process of editing the image feature, attention should be paid to controlling the degree of editing, which cannot be too large, otherwise the generated image will collapse due to the lack of corresponding data. Exemplarily, a certain enhanced image feature can refer to an image feature that is edited once, or it can refer to an image feature that is edited multiple times continuously. In the case of an enhanced image feature that is edited multiple times continuously to obtain a certain image feature, each time the editing is continued on the basis of the image feature obtained after the previous edit until the final enhanced image feature is obtained.

[0129] For example, assuming that a certain enhanced image feature is a feature of an image of a virtual object including a large-angle side face, the enhanced image feature is obtained by continuously editing multiple times based on a feature of an image of a virtual object including a front face, and each edit edits the angle of the face to a smaller angle based on the previous result until a higher-quality enhanced image feature for indicating an image of a virtual object including a large-angle side face is generated.

[0130] After acquiring the enhanced image features, the second image generation model is invoked to process the enhanced image features to obtain a supplementary image. The second image generation model is capable of generating an image based on the input image features. The enhanced image features are input into the second image generation model, which processes the enhanced image features and outputs a generated image. The image output by the second image generation model is used as the supplementary image.

[0131] Exemplarily, the process of invoking the second image generation model to process the enhanced image features is an internal processing process of the second image generation model. The specific processing method is related to the model structure of the second image generation model and is not limited in this embodiment of the present application. Exemplarily, the second image generation model is a generative adversarial model, and the process of the second image generation model processing the enhanced image features is the process of the generative model within the second image generation model processing the enhanced image features.

[0132] Method 2: Obtain an image-driven model, which is trained based on a sample video with target attributes; call the image-driven model to drive the sample virtual image to obtain an enhanced video corresponding to the sample virtual image; extract video frames from the enhanced video corresponding to the sample virtual image as supplementary images.

[0133] In this second approach, an image-driven model is used to acquire a supplementary image for supplementing a sample virtual image. The image-driven model has the function of driving an image of a virtual object having target attributes and a certain posture into an enhanced video having the target attributes. The enhanced video includes video frames of virtual objects having multiple postures. The postures of the virtual objects included in different video frames may be the same or different. For example, in the enhanced video, the virtual objects included in different video frames have similar appearances.

[0134] Acquiring the image driving model in this second method may refer to extracting a pre-stored image driving model, or may refer to obtaining an image driving model based on training of a sample video with target attributes, which is not limited in this embodiment of the present application.

[0135] In an exemplary embodiment, sample videos with target attributes are acquired by professionals and uploaded to a computer device, or automatically crawled online by the computer device. For example, sample video frames within the sample videos with target attributes include the same virtual object, and the virtual object has various poses. For example, if the virtual object is a human face, the poses of the virtual object include, but are not limited to, open mouth, closed mouth, open eyes, closed eyes, wide-angle profile, narrow-angle profile, and full-face.

[0136] In one possible implementation, the process of obtaining an image driving model based on training of a sample video with target attributes is as follows: randomly select two sample video frames from the sample video, use one of the sample video frames as the original image, and use the other sample video frame as the target image; obtain motion information of the posture of the virtual object in the target image relative to the posture of the virtual object in the original image; input the original image and the motion information into the initial driving model to obtain a predicted image output by the initial driving model; based on the difference between the target image and the predicted image, obtain a driving loss function, and use the driving loss function to train the initial driving model to obtain an image driving model. The embodiment of the present application does not limit the model structure of the initial driving model. Exemplarily, the model structure of the initial driving model is a convolutional neural network model, for example, a VGG (Visual Geometry Group) model.

[0137] Exemplarily, the virtual object in the target image has a similar appearance to the virtual object in the original image, and motion information about the posture of the virtual object in the target image relative to the posture of the virtual object in the original image is used to indicate motion information required to adjust the posture of the virtual object in the original image to the posture of the virtual object in the target image. Exemplarily, the motion information about the posture of the virtual object in the target image relative to the posture of the virtual object in the original image is obtained by decoupling appearance information and posture information of the target image and the original image, and then comparing the posture information of the target image and the original image.

[0138] The above-mentioned training method for obtaining an image-driven model is a supervised training method. The image-driven model obtained by training using the above-mentioned method can learn the posture of virtual objects in each sample video frame in the sample video, and can obtain a video composed of video frames of virtual objects with the learned posture based on an image of a virtual object with a certain posture.

[0139] After acquiring the image-driven model, a sample virtual image is input into the image-driven model, which then drives the sample virtual image to obtain an enhanced video corresponding to the sample virtual image. This process of the image-driven model driving the sample virtual image is an internal processing step of the image-driven model. The specific processing method depends on the model structure of the image-driven model and is not limited in this embodiment of the present application.

[0140] The enhanced video corresponding to the sample virtual image is a video composed of video frames of virtual objects in multiple postures, and the multiple postures are the postures of the virtual objects in the sample video frames in the sample video. In an exemplary embodiment, if there are multiple sample virtual images, the image driving model can be called to drive all the sample virtual images separately, or the image driving model can be called to drive some of the sample virtual images. The embodiment of the present application does not limit this. By calling the image driving model to drive each sample virtual object, an enhanced video can be obtained. The enhanced videos corresponding to different sample virtual images may be the same or different.

[0141] After obtaining the enhanced video corresponding to the sample virtual image, video frames are extracted from the enhanced video corresponding to the sample virtual image as supplementary images. Exemplarily, the base video frame or frames in the enhanced video corresponding to the sample virtual image that need to be extracted are set according to requirements. Exemplarily, the video frames in the enhanced video corresponding to the sample virtual image that need to be extracted are video frames that include a virtual object in a specified posture among the base video frames in the enhanced video corresponding to the sample virtual image. Exemplarily, the specified posture refers to a posture that the virtual object in the sample virtual image lacks. For example, if the posture of the virtual object in the sample virtual image is open eyes, the specified posture is the closed eyes posture that the virtual object in the sample virtual image lacks.

[0142] For example, the process of obtaining supplementary images using this second method can be considered as a process of data augmenting the sample virtual image using image-driven technology. An image-driven model is pre-trained using image-driven technology. When supplementary images are needed, the image-driven model is called upon to drive the sample virtual image to obtain an enhanced video, from which video frames are extracted as supplementary images. For virtual objects such as human faces, this second method utilizes the image-driven model to create images of the virtual object in relatively uncommon poses, such as profile faces, closed eyes, and open mouths. This can supplement the images of the virtual object in poses such as wide-angle profile faces and specific expressions that are lacking in the originally collected sample virtual images.

[0143] Step 502: Based on the supplementary image and the sample virtual image, the second image generation model is trained to obtain an updated second image generation model.

[0144] Since the supplementary image is used to supplement the sample virtual image, after acquiring the supplementary image, the second image generation model is trained based on the supplementary image and the sample virtual image to obtain an updated second image generation model. The updated second image generation model has better image generation capabilities than the second image generation model before the update.

[0145] Exemplarily, the second image generation model is trained based on the supplementary image and the sample virtual image by mixing the supplementary image and the sample virtual image to form a new sample virtual image; and then training the second image generation model based on the new sample virtual image to obtain an updated second image generation model. Exemplarily, if there are multiple second image generation models, each second image generation model is trained separately based on the new sample virtual image to obtain multiple updated second image generation models.

[0146] Exemplarily, the second image generation model is a generative adversarial model, and the updated second image generation model is obtained by training the second image generation model based on a new sample virtual image using a generative adversarial training method.

[0147] Step 503: Fuse the first image generation model and the updated second image generation model to obtain a fourth image generation model.

[0148] After obtaining the updated second image generation model, the first image generation model and the updated second image generation model are fused to obtain a fourth image generation model. Compared with the third image generation model, the fourth image generation model has better image generation function.

[0149] In one possible implementation, the process of fusing the first image generation model and the updated second image generation model to obtain the fourth image generation model includes: determining at least one second fusion method, where any second fusion method indicates a method for fusing the first image generation model and the updated second image generation model; obtaining at least one second candidate image generation model based on the at least one second fusion method, where any second candidate image generation model is obtained by fusing the first image generation model and the updated second image generation model using any second fusion method; and selecting a second candidate image generation model that meets a selection condition from the at least one second candidate image generation model as the fourth image generation model. The implementation of this process is described in steps 301 and 302 and is not further described here.

[0150] It should be noted that the method for obtaining the fourth image generation model described in steps 501 to 503 above is merely an exemplary implementation, and the embodiments of the present application are not limited thereto. Exemplarily, the process for obtaining the fourth image generation model is as follows: calling the third image generation model to generate a second generated image; obtaining second enhanced image features, where the second image enhanced features are used to enhance the image generation function of the third image generation model; calling the third image generation model to process the second enhanced image features to obtain a third generated image; and training the third image generation model based on the second and third generated images to obtain the fourth image generation model.

[0151] The second generated image is an image that can be generated using the image generation function already possessed by the third image generation model. The third generated image is an image generated using the second image enhancement feature used to enhance the image generation function of the third image generation model. The third generated image can complement the second generated image. Therefore, by training the third image generation model based on the second and third generated images, a fourth image generation model with better image generation function than the third image generation model can be obtained.

[0152] After obtaining the fourth image generation model, determine whether the image generation function of the fourth image generation model meets the reference conditions. The method for determining whether the image generation function of the fourth image generation model meets the reference conditions refers to the method for determining whether the third image generation model meets the reference conditions, which will not be repeated here.

[0153] If the image generation function of the fourth image generation model satisfies the reference condition, the fourth image generation model is used as the target image generation model. If the image generation function of the fourth image generation model does not satisfy the reference condition, the fifth image generation model is continuously acquired until an image generation model with an image generation function that satisfies the reference condition is acquired, and the image generation model with an image generation function that satisfies the reference condition is used as the target image generation model. The process for acquiring the fifth image generation model is similar to the process for acquiring the fourth image generation model and is not further described here.

[0154] For example, the process of obtaining the target image generation model is as follows: Figure 6 As shown. First, sample virtual images are collected, and then the sample virtual images are used to obtain a second image generation model; the first image generation model and the second image generation model are fused to obtain a third image generation model; and it is determined whether the image generation function of the third image generation model meets the reference conditions. If the image generation function of the third image generation model meets the reference conditions, the third image generation model is used as the target image generation model. If the image generation function of the third image generation model does not meet the reference conditions, a supplementary image is obtained; the supplementary image and the sample virtual image are used to obtain an updated second image generation model, and so on, until an image generation model with an image generation function that meets the reference conditions is obtained, and the image generation model with an image generation function that meets the reference conditions is used as the target image generation model.

[0155] In step 202 , based on the target image generation model, a target virtual image corresponding to the original image of the target object is obtained. The target virtual image retains the object ontology features of the target object and has target attributes.

[0156] After acquiring the target image generation model, a target virtual image corresponding to the original image of the target object is obtained based on the target image generation model. Since the target image generation model can simultaneously focus on the object ontology features and target attributes, the target virtual image retains the object ontology features of the target object and has the target attributes.

[0157] The target object is a physical object of the same type as the sample object. The target object may have one or more original images, and the source of the original images of the target object is diverse. For example, the original images of the target object may be obtained from an original image library; or, the original images of the target object may be crawled from the Internet; or, the original images of the target object may be uploaded to a computer device by a professional, although this embodiment of the present application is not limited thereto.

[0158] In one possible implementation, based on the target image generation model, methods for obtaining the target virtual image corresponding to the original image of the target object include but are not limited to the following two:

[0159] Method 1: Obtain original image features corresponding to the original image of the target object; obtain target image features based on the original image features; call the target image generation model, process the target image features, and obtain a target virtual image corresponding to the original image of the target object.

[0160] The original image features corresponding to the original image of the target object are obtained by performing feature extraction on the original image of the target object. The embodiment of the present application does not limit the manner in which the features of the original image of the target object are extracted. For example, the feature extraction model is called to perform feature extraction on the original image of the target object. After obtaining the original image features, the target image features for inputting into the target image generation model are obtained based on the original image features. The target image generation model processes the target image features, and the virtual image generated by the target image generation model according to the input target image features is used as the target virtual image corresponding to the original image of the target object. The processing process of the target image generation model processing the target image features is an internal processing process of the target image generation model. The specific processing method is related to the model structure of the target image generation model, and the embodiment of the present application does not limit this.

[0161] Target image features are image features derived from original image features and required to be input into the target image generation model. In an exemplary embodiment, the target image features are derived based on the original image features by directly using the original image features as the target image features. In this manner, the target image generation model is directly invoked to process the original image features to obtain the target virtual image, resulting in a more efficient acquisition of the target virtual image.

[0162] In one possible implementation, the target image features are obtained based on the original image features by converting the original image features using an image feature conversion method corresponding to the target image generation model, and using the converted image features as the target image features.

[0163] The image feature conversion method corresponding to the target image generation model is used to indicate the method in which the original image features need to be converted. Exemplarily, the image feature conversion method is set based on experience, or is flexibly adjusted according to the image generation effect of the target image generation model. The purpose of setting the image feature conversion method is to enable the virtual image generated by the target image generation model to take into account both the object's ontological characteristics and the target attributes. Exemplarily, in the case where the object type is a human face, the purpose of setting the image feature conversion method is to enable the virtual image generated by the target image generation model to maintain the similarity of the real person's face shape and facial features, while having more details of the target attributes. Exemplarily, if the image generation effect of the target image generation model is an effect with a weak target attribute, the image feature conversion method is used to indicate that the original image features are converted in the direction of enhancing the target attribute. According to the image features converted using the image feature conversion method, it is beneficial to improve the quality of the generated virtual image.

[0164] In the case where the target image generation model corresponds to an image feature conversion method, after obtaining the original image features corresponding to the original image of the target object, the original image features are converted using the image feature conversion method to obtain converted image features. The converted image features are then input into the target image generation model, which processes the converted image features. The virtual image generated by the target image generation model based on the input converted image features is used as the target virtual image corresponding to the original image of the target object. Compared to the target virtual image obtained by directly calling the target image generation model to process the original image features, the target virtual image obtained by calling the target image generation model to process the converted image features has higher quality.

[0165] Method 2: Call the target image generation model, process the candidate image features, and obtain the candidate virtual images corresponding to the candidate image features; obtain the target image translation model based on the candidate original image corresponding to the candidate image features and the candidate virtual images corresponding to the candidate image features, where the candidate original image corresponding to the candidate image features is the image identified by the candidate image features that retains the object's ontological features; call the target image translation model, process the original image of the target object, and obtain the target virtual image corresponding to the original image of the target object.

[0166] In this method 2, a model is first generated based on the target image to obtain a target image translation model, and then the target image translation model is called to obtain a target virtual image. The target image translation model has the function of outputting a virtual image that retains the same object ontology features and has target attributes based on the input original image that retains the object ontology features. The target image translation model is obtained based on the candidate original image corresponding to the candidate image features and the candidate virtual image corresponding to the candidate image features. Exemplarily, the candidate image features are randomly generated image features, or image features obtained by feature extraction of an image that retains the object ontology features, which is not limited in the embodiments of the present application.

[0167] It should be noted that the number of candidate image features is one or more, and each candidate image feature identifies an image that retains the object's ontological features. In the embodiment of the present application, the image that retains the object's ontological features and is identified by the candidate image feature is referred to as the candidate original image corresponding to the candidate image feature. In the case where the candidate image feature is a randomly generated image feature, the candidate original image corresponding to the candidate image feature is a false image; in the case where the candidate image feature is an image feature obtained by feature extraction of an image that retains the object's ontological features, the candidate original image corresponding to the candidate image feature is a real image. Image features correspond one-to-one to images, and the candidate original image corresponding to the candidate image feature can be obtained by performing image restoration on the candidate image feature.

[0168] The candidate virtual image corresponding to the candidate image feature is obtained by calling the target image generation model, and the candidate image feature is input into the target image generation model. The target image generation model processes the candidate image feature to obtain the candidate virtual image corresponding to the candidate image feature. The candidate virtual image corresponding to the candidate image feature retains the object ontology features retained by the candidate original image identified by the candidate image feature and has target attributes. The processing process of the candidate image feature by the target image generation model is an internal processing process of the target image generation model. The specific processing method is related to the model structure of the target image generation model, and the embodiments of the present application are not limited to this.

[0169] The candidate original images corresponding to the candidate image features and the candidate virtual images corresponding to the candidate image features serve as the training data required to obtain the target image translation model. After obtaining the candidate original images and candidate virtual images corresponding to the candidate image features, the target image translation model can be obtained through training. In one possible implementation, the process of obtaining the target image translation model based on the candidate original images and candidate virtual images corresponding to the candidate image features is as follows: calling the initial image translation model, processing the candidate original images corresponding to the candidate image features, and obtaining candidate predicted images corresponding to the candidate image features; determining a loss function based on the difference between the candidate predicted images and the candidate virtual images corresponding to the candidate image features; and using the loss function to train the initial image translation model to obtain the target image translation model.

[0170] The initial image translation model refers to the image translation model to be trained. While the present embodiments do not limit the model structure of the initial image translation model, the initial image translation model exemplarily comprises an encoder-decoder structure. The process of training the target image translation model based on candidate original images corresponding to candidate image features and candidate virtual images corresponding to candidate image features is a supervised training process and will not be further elaborated upon here.

[0171] Since the target image translation model is trained using a supervised training method based on the candidate original image corresponding to the candidate image features and the candidate virtual image corresponding to the candidate image features, wherein the candidate virtual image is obtained by calling the target image generation model, and the candidate virtual image can retain the object ontology features and have the target attributes, the target image translation model has the function of outputting a virtual image that retains the object ontology features and has the target attributes based on the input original image.

[0172] After obtaining the target image translation model, the original image of the target object is input into the target image translation model. The target image translation model processes the original image of the target object to obtain a target virtual image corresponding to the original image of the target object. The processing process of the target image translation model on the original image of the target object is an internal processing process of the target image translation model. The specific processing method is related to the model structure of the target image translation model and is not limited to this embodiment of the present application.

[0173] Exemplarily, the purpose of obtaining the target virtual image corresponding to the original image of the target object is to use the original image-target virtual image data pair of the target object as training data for the image prediction model corresponding to the target attribute. In this case, under the condition that there are only dozens to hundreds of sample virtual images, through the method provided in the embodiment of the present application, it is possible to obtain tens of thousands of virtual images corresponding to high-quality original images, obtain tens of thousands of original image-virtual image data pairs, and then use the obtained original image-virtual image data pairs to train the image prediction model corresponding to the target attribute. Exemplarily, the image prediction model corresponding to the target attribute is the model required by the product, and the purpose of training the image prediction model using the obtained original image-virtual image data pairs is to enable the image prediction model to meet the product launch standard.

[0174] Exemplarily, in the case where the purpose of obtaining a target virtual image corresponding to an original image of a target object is to use the target object's original image-target virtual image data pair as training data for an image prediction model corresponding to a target attribute, after obtaining the target virtual image corresponding to the original image of the target object based on the target image generation model, the method further includes: inputting the original image of the target object into an initial image prediction model to obtain a predicted image corresponding to the original image of the target object predicted by the initial image prediction model; determining a loss function based on the difference between the predicted image and the target virtual image; and training the initial image prediction model using the loss function to obtain a target image prediction model. In other words, the target image prediction model is trained based on the original image of the target object and the target virtual image corresponding to the original image of the target object.

[0175] The present embodiment does not limit the model structure of the initial image prediction model. For example, the model structure of the initial image prediction model is an encoder-decoder structure. The process of training the target image prediction model using the original image of the target object and the target virtual image is a supervised training process and will not be further described here.

[0176] For example, the target image prediction model and the target image translation model may have the same or different model structures. For example, when the target image prediction model and the target image translation model have the same model structures, the number of model parameters of the target image prediction model is smaller than the number of model parameters of the target image translation model, so as to conserve resources of the product deployment device.

[0177] For example, after obtaining the target image prediction model, the target image prediction model can be applied to the product, so that the gameplay of displaying a virtual image that retains the object's ontological features and has the target attributes can be launched online. For example, the process of launching the gameplay of displaying a virtual image that retains the object's ontological features and has the target attributes can be as follows: Figure 7As shown. Data is collected, including but not limited to sample original images that retain object ontology features of the sample object and sample virtual images with target attributes; a first image generation model and a second image generation model are trained using the collected data, and model fusion and data enhancement are performed on the first image generation model and the second image generation model to obtain a target image generation model; training data required for training a target image prediction model is obtained based on the target image generation model, and the obtained training data is used to perform on-end model training to obtain a target image prediction model; after obtaining the target image prediction model, the target image prediction model is applied to the product, so that a gameplay that displays a virtual image that retains the object ontology features and has the target attributes is launched online.

[0178] exist Figure 7 In the process shown, model fusion and data augmentation play a connecting role. They are responsible for learning the target attributes from dozens to hundreds of collected sample virtual images and migrating the target attributes to tens of thousands of original images. This provides high-quality training data for on-device model training and helps launch gameplay.

[0179] In one implementation, after the gameplay of displaying a virtual image that retains the object's intrinsic features and has target attributes is launched, the game further includes: capturing an original image of the interactive object in response to a virtual image display instruction for the target attribute; invoking a target image prediction model to process the original image of the interactive object to obtain a virtual image corresponding to the original image of the interactive object; and displaying the virtual image corresponding to the original image of the interactive object. The target image prediction model is trained based on the original image of the target object and the target virtual image corresponding to the original image of the target object.

[0180] The virtual image display instruction for the target attribute is used to trigger the capture of the original image of the interactive object. The embodiments of this application do not limit the method for acquiring the virtual image display instruction for the target attribute. For example, a display interface displays multiple selectable attributes. In response to detecting a trigger operation for the target attribute, a virtual image display instruction for the target attribute is acquired. After acquiring the virtual image display instruction for the target attribute, the original image of the interactive object is captured. For example, the interactive object is the face of the user who generated the virtual image display instruction for the target attribute.

[0181] Since the target image prediction model is trained based on the original image-virtual image data pair that retains the object ontology features and has the target attributes, the original image of the interactive object is input into the target image prediction model, and the target image prediction model processes the original image of the interactive object to obtain a virtual image corresponding to the original image of the interactive object. The virtual image corresponding to the original image of the interactive object retains the object ontology features of the interactive object and has the target attributes. The processing process of the target image prediction model on the original image of the interactive object is the internal processing process of the target image prediction model. The specific processing method is related to the model structure of the target image prediction model, and the embodiment of the present application does not limit this. After obtaining the virtual image corresponding to the original image of the interactive object, the virtual image corresponding to the original image of the interactive object is displayed in the display interface.

[0182] For example, after displaying the virtual image corresponding to the original image of the interactive object, new virtual images may be continuously displayed with reference to new original images of the interactive object. Different original images of the interactive object may have different facial expressions, and thus different displayed virtual images may also have different facial expressions. The facial expression of the displayed virtual image is the same as the facial expression of the original image from which the virtual image was obtained.

[0183] For example, the original image of the interactive object and the virtual image that retains the object ontology features of the interactive object and has target attributes are as follows: Figure 8 shown. Figure 8 (1) is the original image obtained based on the face of the first user (interaction object), Figure 8 (2) is based on Figure 8 The original image shown in (1) is used to obtain a virtual image that retains the facial features of the first user and has the target attribute of the first special effect. Figure 8 (3) is a virtual image obtained based on the original image with transformed facial expression, which retains the facial features of the first user and has the target attribute of the first special effect.

[0184] Figure 8 (4) is the original image obtained based on the face of the second user (interaction object), Figure 8 (5) is based on Figure 8 The original image shown in (4) obtains a virtual image that retains the facial features of the second user and has the target attribute of the second special effect. Figure 8 (6) is a virtual image obtained based on the original image with transformed facial expression, which retains the facial features of the second user and has the target attribute of the second special effect.

[0185] In an embodiment of the present application, by means of model fusion and data enhancement, the number of sample virtual images required to obtain the target image generation model can be reduced to dozens, and a stable real-time effect can be trained through dozens of sample virtual images, reducing the dependence of the image generation model on large-scale data, greatly improving the scope of application of the image generation model, and expanding the application scenario of the image generation model. Based on the method provided in the embodiment of the present application, large-scale data is not required, which reduces the annotation cost and R&D cost. In addition, the data enhancement method quickly supplements images at a low cost, greatly speeds up the acquisition speed of training data, can speed up the R&D cycle, and help the gameplay on the product to be quickly implemented. Based on the target image generation model, a virtual image that retains good object ontology features and has good-looking target attributes can be obtained, which can achieve good visual effects and enhance user experience.

[0186] The method provided in the embodiment of the present application is applied to the field of virtual human faces. When the amount of data is limited, the real-person model (i.e., the first image generation model) and the virtual style model (i.e., the second image generation model) are fused through model fusion, and a fusion model with the function of generating images that retain real-person textures and have virtual styles can be obtained; through data enhancement, the expressions and postures of virtual objects in the virtual images generated by the fusion model can be enriched.

[0187] In an embodiment of the present application, a target image generation model for acquiring a virtual image is obtained by fusing a first image generation model and a second image generation model, wherein the first image generation model can focus on the object ontology features, and the second image generation model can focus on the target attributes. Therefore, the target image generation model can focus on both the object ontology features and the target attributes at the same time, which is beneficial to ensuring the quality of the acquired virtual image. In addition, compared to sample images that retain the object ontology features and have target attributes, the sample original images based on which the first image generation model is trained and the sample virtual images based on which the second image generation model is trained are both easier to obtain, which is beneficial to shortening the acquisition time of the target image generation model, thereby improving the efficiency of acquiring virtual images. In other words, the method for acquiring virtual images provided in an embodiment of the present application can improve the efficiency of acquiring virtual images while ensuring the quality of the acquired virtual images.

[0188] See also Figure 9 , an embodiment of the present application provides a device for acquiring a virtual image, the device comprising:

[0189] A first acquisition unit 901 is configured to acquire a target image generation model, where the target image generation model is obtained by fusing a first image generation model with a second image generation model, wherein the first image generation model is trained based on a sample original image that retains the object ontology features of the sample object, and the second image generation model is trained based on a sample virtual image that has target attributes;

[0190] The second acquisition unit 902 is configured to acquire a target virtual image corresponding to the original image of the target object based on the target image generation model. The target virtual image retains the object ontology features of the target object and has target attributes.

[0191] In one possible implementation, see Figure 10 , the device further comprises:

[0192] The third acquisition unit 903 is used to acquire a sample original image and a sample virtual image; based on the sample original image, a first image generation model is trained; based on the sample virtual image, a second image generation model is trained;

[0193] The fusion unit 904 is configured to fuse the first image generation model and the second image generation model to obtain a target image generation model.

[0194] In one possible implementation, the fusion unit 904 is used to determine at least one fusion method, any one of which is used to indicate a method of fusing the first image generation model and the second image generation model; based on the at least one fusion method, at least one candidate image generation model is obtained, and any one of the candidate image generation models is obtained by fusing the first image generation model and the second image generation model based on any one of the fusion methods; in response to the image generation function of the third image generation model satisfying the reference condition, the third image generation model is used as the target image generation model, and the third image generation model is a candidate image generation model that satisfies the selection condition among the at least one candidate image generation model.

[0195] In one possible implementation, the first image generation model and the second image generation model both include a reference number layer network, and any fusion method includes a method for determining the target network parameters corresponding to the reference number layer networks respectively; the fusion unit 904 is also used to determine the target network parameters corresponding to the reference number layer networks respectively based on the method for determining the target network parameters corresponding to the reference number layer networks respectively included in any fusion method; based on the target network parameters corresponding to the reference number layer networks respectively, the parameters of the reference number layer networks in the specified image generation model are adjusted, and the image generation model obtained after the adjustment is used as any candidate image generation model; wherein, the method for determining the target network parameters corresponding to any layer network included in any fusion method is used to indicate the relationship between the target network parameters corresponding to any layer network and at least one of the first network parameter and the second network parameter; the first network parameter is a parameter of any layer network in the first image generation model, and the second network parameter is a parameter of any layer network in the second image generation model.

[0196] In one possible implementation, see Figure 10 , the device further comprises:

[0197] A fourth acquiring unit 905 is configured to acquire a supplementary image in response to the image generation function of the third image generation model not satisfying the reference condition, where the supplementary image is used to supplement the sample virtual image;

[0198] The third acquisition unit 903 is further configured to train the second image generation model based on the supplementary image and the sample virtual image to obtain an updated second image generation model;

[0199] The fusion unit 904 is further configured to fuse the first image generation model and the updated second image generation model to obtain a fourth image generation model; in response to the image generation function of the fourth image generation model satisfying a reference condition, the fourth image generation model is used as a target image generation model.

[0200] In one possible implementation, the fourth acquisition unit 905 is configured to acquire enhanced image features, where the enhanced image features are used to enhance the image generation function of the second image generation model; and the second image generation model is called to process the enhanced image features to obtain a supplementary image.

[0201] In one possible implementation, the fourth acquisition unit 905 is used to obtain an image-driven model, which is trained based on a sample video with target attributes; calling the image-driven model to drive the sample virtual image to obtain an enhanced video corresponding to the sample virtual image; and extracting a video frame from the enhanced video corresponding to the sample virtual image as a supplementary image.

[0202] In one possible implementation, the second acquisition unit 902 is used to obtain original image features corresponding to the original image of the target object; obtain target image features based on the original image features; call the target image generation model to process the target image features to obtain a target virtual image corresponding to the original image of the target object.

[0203] In a possible implementation, the second acquisition unit 902 is further configured to convert the original image features using an image feature conversion method corresponding to the target image generation model, and use the image features obtained after the conversion as the target image features.

[0204] In one possible implementation, the second acquisition unit 902 is used to call the target image generation model, process the candidate image features, and obtain a candidate virtual image corresponding to the candidate image features; based on the candidate original image corresponding to the candidate image features and the candidate virtual image corresponding to the candidate image features, obtain the target image translation model, where the candidate original image corresponding to the candidate image features is an image identified by the candidate image features that retains the object's ontological features; call the target image translation model, process the original image of the target object, and obtain a target virtual image corresponding to the original image of the target object.

[0205] In one possible implementation, the second acquisition unit 902 is further used to call the initial image translation model, process the candidate original image corresponding to the candidate image feature, and obtain a candidate predicted image corresponding to the candidate image feature; determine a loss function based on the difference between the candidate predicted image corresponding to the candidate image feature and the candidate virtual image corresponding to the candidate image feature; and use the loss function to train the initial image translation model to obtain a target image translation model.

[0206] In one possible implementation, see Figure 10 , the device further comprises:

[0207] The acquisition unit 906 is configured to acquire an original image of the interactive object in response to a virtual image display instruction for a target attribute;

[0208] The display unit 907 is used to call the target image prediction model, process the original image of the interactive object, and obtain a virtual image corresponding to the original image of the interactive object. The target image prediction model is trained based on the original image of the target object and the target virtual image corresponding to the original image of the target object; and display the virtual image corresponding to the original image of the interactive object.

[0209] In an embodiment of the present application, a target image generation model for acquiring a virtual image is obtained by fusing a first image generation model and a second image generation model, wherein the first image generation model can focus on the object ontology features, and the second image generation model can focus on the target attributes. Therefore, the target image generation model can focus on both the object ontology features and the target attributes at the same time, which is beneficial to ensuring the quality of the acquired virtual image. In addition, compared to sample images that retain the object ontology features and have target attributes, the sample original images based on which the first image generation model is trained and the sample virtual images based on which the second image generation model is trained are both easier to obtain, which is beneficial to shortening the acquisition time of the target image generation model, thereby improving the efficiency of acquiring virtual images. In other words, the method for acquiring virtual images provided in an embodiment of the present application can improve the efficiency of acquiring virtual images while ensuring the quality of the acquired virtual images.

[0210] It should be noted that the apparatus provided in the above embodiments is merely illustrated by the division of the above functional modules when implementing its functions. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0211] In an exemplary embodiment, a computer device is also provided. The computer device includes a processor and a memory, wherein the memory stores at least one computer program. The at least one computer program is loaded and executed by one or more processors to enable the computer device to implement any of the aforementioned methods for acquiring a virtual image. Exemplarily, the computer device may be a server or a terminal. The following describes the structures of the server and terminal, respectively.

[0212] Figure 11 This is a schematic diagram of the structure of a server provided in an embodiment of the present application. The server may have relatively large differences due to different configurations or performances, and may include one or more processors (Central Processing Units, CPU) 1101 and one or more memories 1102, wherein the one or more memories 1102 store at least one computer program, and the at least one computer program is loaded and executed by the one or more processors 1101 to enable the server to implement the method for obtaining a virtual image provided by each of the above-mentioned method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and an input and output interface for input and output. The server may also include other components for implementing device functions, which will not be described in detail here.

[0213] Figure 12 This is a schematic diagram of the structure of a terminal provided in an embodiment of the present application. The terminal may be a smartphone, tablet computer, laptop computer, or desktop computer. The terminal may also be referred to as user equipment, portable terminal, laptop terminal, desktop terminal, or other names.

[0214] Typically, the terminal includes: a processor 1201 and a memory 1202 .

[0215] The processor 1201 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1201 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1201 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1201 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1201 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0216] The memory 1202 may include one or more computer-readable storage media, which may be non-transitory. The memory 1202 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1202 is used to store at least one computer program, which is executed by the processor 1201 to enable the terminal to implement the method for obtaining a virtual image provided in the method embodiment of the present application.

[0217] In some embodiments, the terminal may optionally include a peripheral device interface 1203 and at least one peripheral device. The processor 1201, memory 1202, and peripheral device interface 1203 may be connected via a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1203 via a bus, signal lines, or circuit boards. Specifically, the peripheral device may include at least one of a radio frequency circuit 1204, a display screen 1205, a camera assembly 1206, an audio circuit 1207, and a power supply 1209.

[0218] The peripheral device interface 1203 can be used to connect at least one I / O (Input / Output)-related peripheral device to the processor 1201 and the memory 1202. The radio frequency circuit 1204 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1204 communicates with a communication network and other communication devices via electromagnetic signals. The display screen 1205 is used to display a UI (User Interface). The UI may include images, text, icons, videos, and any combination thereof. The camera assembly 1206 is used to capture images or videos.

[0219] Audio circuit 1207 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting them into electrical signals that are then input into processor 1201 for processing or into RF circuit 1204 for voice communication. The speaker is used to convert electrical signals from processor 1201 or RF circuit 1204 into sound waves. Power supply 1209 is used to power various components in the terminal. Power supply 1209 can be AC, DC, disposable batteries, or rechargeable batteries.

[0220] In some embodiments, the terminal further includes one or more sensors 1210 , including but not limited to: an acceleration sensor 1211 , a gyroscope sensor 1212 , a pressure sensor 1213 , an optical sensor 1215 , and a proximity sensor 1216 .

[0221] The acceleration sensor 1211 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established by the terminal. The gyroscope sensor 1212 can detect the body direction and rotation angle of the terminal. The gyroscope sensor 1212 can cooperate with the acceleration sensor 1211 to collect the user's 3D actions on the terminal. The pressure sensor 1213 can be set on the side frame of the terminal and / or the lower layer of the display screen 1205. When the pressure sensor 1213 is set on the side frame of the terminal, it can detect the user's holding signal of the terminal, and the processor 1201 performs left and right hand recognition or shortcut operations based on the holding signal collected by the pressure sensor 1213. When the pressure sensor 1213 is set on the lower layer of the display screen 1205, the processor 1201 controls the operable controls on the UI interface based on the user's pressure operation on the display screen 1205.

[0222] The optical sensor 1215 is used to collect ambient light intensity. The proximity sensor 1216, also known as a distance sensor, is typically provided on the front panel of the terminal. The proximity sensor 1216 is used to collect the distance between the user and the front of the terminal.

[0223] Those skilled in the art will understand that Figure 12 The structure shown in the figure does not constitute a limitation on the terminal, and may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.

[0224] In some embodiments, the computer programs involved in the embodiments of the present application may be deployed and executed on a single computer device, or on multiple computer devices located in a single location, or on multiple computer devices distributed across multiple locations and interconnected via a communication network. Multiple computer devices distributed across multiple locations and interconnected via a communication network may constitute a blockchain system. In other words, the aforementioned servers and terminals may serve as node devices in the blockchain system.

[0225] In an exemplary embodiment, a computer-readable storage medium is further provided, in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor of a computer device to enable the computer to implement any of the above methods for obtaining a virtual image.

[0226] In one possible implementation, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, and the like.

[0227] In an exemplary embodiment, a computer program product or computer program is also provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform any of the aforementioned methods for acquiring a virtual image.

[0228] It should be understood that the term "plurality" used herein refers to two or more. "And / or" describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates an "or" relationship between the associated objects.

[0229] The above description is merely an exemplary embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A method for obtaining a virtual image, characterized in that: The method comprises: Acquire a sample original image retaining object ontology features of the sample object and a sample virtual image having target attributes; Based on the sample original image, a first image generation model is trained; based on the sample virtual image, a second image generation model is trained; the first image generation model and the second image generation model both include a reference number layer network; Determine at least one fusion method; any fusion method includes a method for determining target network parameters corresponding to the reference number of layer networks, and the method for determining the target network parameters corresponding to any layer network included in any fusion method is used to indicate the relationship between the target network parameters corresponding to any layer network and at least one of a first network parameter and a second network parameter, the first network parameter being a parameter of any layer network in the first image generation model, and the second network parameter being a parameter of any layer network in the second image generation model; Based on the at least one fusion method, at least one candidate image generation model is obtained; wherein, the method for obtaining any candidate image generation model includes: determining the target network parameters corresponding to the reference number of layer networks respectively based on the method for determining the target network parameters respectively corresponding to the reference number of layer networks included in any fusion method; adjusting the parameters of the reference number of layer networks in the specified image generation model based on the target network parameters respectively corresponding to the reference number of layer networks, and using the image generation model obtained after the adjustment as any candidate image generation model, wherein the specified image generation model is any model including the reference number of layer networks; In response to an image generation function of a third image generation model satisfying a reference condition, using the third image generation model as a target image generation model, the third image generation model being a candidate image generation model satisfying a selection condition among the at least one candidate image generation model; Based on the target image generation model, a target virtual image corresponding to the original image of the target object is obtained, where the target virtual image retains the object ontology features of the target object and has the target attributes.

2. The method according to claim 1, characterized in that The method further comprises: In response to the image generation function of the third image generation model not satisfying the reference condition, acquiring a supplementary image, wherein the supplementary image is used to supplement the sample virtual image; Training the second image generation model based on the supplementary image and the sample virtual image to obtain an updated second image generation model; fusing the first image generation model and the updated second image generation model to obtain a fourth image generation model; In response to the image generation function of the fourth image generation model satisfying the reference condition, the fourth image generation model is used as the target image generation model.

3. The method according to claim 2, characterized in that The obtaining of the supplementary image comprises: Acquire enhanced image features, where the enhanced image features are used to enhance the image generation function of the second image generation model; The second image generation model is called to process the enhanced image features to obtain the supplementary image.

4. The method according to claim 2, characterized in that The obtaining of the supplementary image comprises: Acquire an image-driven model, where the image-driven model is trained based on a sample video having the target attribute; Calling the image driving model to drive the sample virtual image to obtain an enhanced video corresponding to the sample virtual image; A video frame is extracted from the enhanced video corresponding to the sample virtual image as the supplementary image.

5. The method according to any one of claims 1 to 4, characterized in that: The step of obtaining a target virtual image corresponding to the original image of the target object based on the target image generation model includes: Obtaining original image features corresponding to the original image of the target object; Based on the original image features, obtaining target image features; The target image generation model is called to process the target image features to obtain a target virtual image corresponding to the original image of the target object.

6. The method according to claim 5, characterized in that The acquiring target image features based on the original image features includes: The original image features are converted using an image feature conversion method corresponding to the target image generation model, and the image features obtained after the conversion are used as the target image features.

7. The method according to any one of claims 1 to 4, characterized in that: The step of obtaining a target virtual image corresponding to the original image of the target object based on the target image generation model includes: calling the target image generation model to process candidate image features to obtain candidate virtual images corresponding to the candidate image features; the candidate image features are randomly generated image features, or image features obtained by extracting features from an image that retains the object's ontological features; acquiring a target image translation model based on a candidate original image corresponding to the candidate image feature and a candidate virtual image corresponding to the candidate image feature, wherein the candidate original image corresponding to the candidate image feature is an image identified by the candidate image feature and retaining the object ontology feature; The target image translation model is called to process the original image of the target object to obtain a target virtual image corresponding to the original image of the target object.

8. The method according to claim 7, characterized in that The obtaining of a target image translation model based on the candidate original image corresponding to the candidate image feature and the candidate virtual image corresponding to the candidate image feature includes: Calling the initial image translation model to process the candidate original image corresponding to the candidate image feature to obtain a candidate predicted image corresponding to the candidate image feature; determining a loss function based on a difference between a candidate predicted image corresponding to the candidate image feature and a candidate virtual image corresponding to the candidate image feature; The initial image translation model is trained using the loss function to obtain the target image translation model.

9. The method according to any one of claims 1 to 4, characterized in that: After acquiring the target virtual image corresponding to the original image of the target object based on the target image generation model, the method further includes: In response to a virtual image display instruction for the target attribute, capturing an original image of the interactive object; calling a target image prediction model to process the original image of the interactive object to obtain a virtual image corresponding to the original image of the interactive object, wherein the target image prediction model is trained based on the original image of the target object and a target virtual image corresponding to the original image of the target object; Display a virtual image corresponding to the original image of the interactive object.

10. A device for obtaining a virtual image, characterized in that: The device comprises: A third acquisition unit is configured to acquire a sample original image retaining object ontology features of the sample object and a sample virtual image having target attributes; based on the sample original image, a first image generation model is trained; based on the sample virtual image, a second image generation model is trained; the first image generation model and the second image generation model both include a reference scalar layer network; a fusion unit configured to determine at least one fusion method; obtain at least one candidate image generation model based on the at least one fusion method; and in response to an image generation function of a third image generation model satisfying a reference condition, use the third image generation model as a target image generation model, the third image generation model being a candidate image generation model that satisfies the selection condition among the at least one candidate image generation model; a second acquiring unit, configured to acquire, based on the target image generation model, a target virtual image corresponding to the original image of the target object, wherein the target virtual image retains the object ontology features of the target object and has the target attributes; Among them, any fusion method includes a method for determining the target network parameters corresponding to the reference number layer networks respectively, and the method for determining the target network parameters corresponding to any layer network included in any fusion method is used to indicate the relationship between the target network parameters corresponding to any layer network and at least one of the first network parameter and the second network parameter, the first network parameter is the parameter of any layer network in the first image generation model, and the second network parameter is the parameter of any layer network in the second image generation model; The fusion unit is used to obtain any candidate image generation model in the following manner: Based on the method for determining the target network parameters corresponding to the reference number-layer network included in any of the fusion methods, the target network parameters corresponding to the reference number-layer network are determined; based on the target network parameters corresponding to the reference number-layer network, the parameters of the reference number-layer network in the specified image generation model are adjusted, and the image generation model obtained after the adjustment is used as any of the candidate image generation models, and the specified image generation model is any model that includes the reference number-layer network.

11. The device according to claim 10, characterized in that The device further comprises: a fourth acquiring unit, configured to acquire a supplementary image in response to the image generation function of the third image generation model not satisfying the reference condition, wherein the supplementary image is used to supplement the sample virtual image; The third acquisition unit is further configured to train the second image generation model based on the supplementary image and the sample virtual image to obtain an updated second image generation model; The fusion unit is further used to fuse the first image generation model and the updated second image generation model to obtain a fourth image generation model; in response to the image generation function of the fourth image generation model satisfying the reference condition, the fourth image generation model is used as the target image generation model.

12. The device according to claim 11, characterized in that The fourth acquisition unit is used to acquire enhanced image features, where the enhanced image features are used to enhance the image generation function of the second image generation model; the second image generation model is called to process the enhanced image features to obtain the supplementary image.

13. The device according to claim 11, characterized in that The fourth acquisition unit is configured to acquire an image-driven model, where the image-driven model is trained based on a sample video having the target attribute; Calling the image driving model to drive the sample virtual image to obtain an enhanced video corresponding to the sample virtual image; A video frame is extracted from the enhanced video corresponding to the sample virtual image as the supplementary image.

14. The device according to any one of claims 10 to 13, characterized in that: The second acquisition unit is configured to acquire original image features corresponding to the original image of the target object; and acquire target image features based on the original image features; The target image generation model is called to process the target image features to obtain a target virtual image corresponding to the original image of the target object.

15. The device according to claim 14, characterized in that The second acquisition unit is configured to convert the original image features using an image feature conversion method corresponding to the target image generation model, and use the image features obtained after the conversion as the target image features.

16. The device according to any one of claims 10 to 13, characterized in that: The second acquisition unit is configured to call the target image generation model and process the candidate image features to obtain candidate virtual images corresponding to the candidate image features; the candidate image features are randomly generated image features or image features obtained by extracting features from an image that retains the object's ontological features; and obtain the target image translation model based on the candidate original images corresponding to the candidate image features and the candidate virtual images corresponding to the candidate image features, wherein the candidate original images corresponding to the candidate image features are images that retain the object's ontological features identified by the candidate image features. The target image translation model is called to process the original image of the target object to obtain a target virtual image corresponding to the original image of the target object.

17. The device according to claim 16, characterized in that The second acquisition unit is used to call the initial image translation model, process the candidate original image corresponding to the candidate image feature, and obtain the candidate predicted image corresponding to the candidate image feature; determine the loss function based on the difference between the candidate predicted image corresponding to the candidate image feature and the candidate virtual image corresponding to the candidate image feature; and use the loss function to train the initial image translation model to obtain the target image translation model.

18. The device according to any one of claims 10 to 13, characterized in that: The device further comprises: an acquisition unit, configured to acquire an original image of the interactive object in response to a virtual image display instruction for the target attribute; A display unit is used to call a target image prediction model, process the original image of the interactive object, and obtain a virtual image corresponding to the original image of the interactive object, wherein the target image prediction model is trained based on the original image of the target object and the target virtual image corresponding to the original image of the target object; and display the virtual image corresponding to the original image of the interactive object.

19. A computer device, characterized in that: The computer device includes a processor and a memory, wherein at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor, so that the computer device implements the method for obtaining a virtual image according to any one of claims 1 to 9.

20. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one computer program, and the at least one computer program is loaded and executed by a processor to enable a computer to implement the method for acquiring a virtual image according to any one of claims 1 to 9.

21. A computer program product, characterized in that The computer program product includes computer instructions, which are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device implements the method for obtaining a virtual image as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Image fusion method, model training method and related device

    CN109919888A

Cited By

  • Method and apparatus for obtaining virtual image, computer device, computer-readable storage medium, and computer program product

    EP4235491B1