User head portrait generation method and device based on deep learning, and medium
Generating personalized user avatars through deep learning solves the problems of homogeneity of user avatar system image and complex operations, and provides a convenient user interaction experience.
Patent Information
- Application Number
- CN202510558077.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-01
AI Technical Summary
The existing user avatar system cannot adapt to the user's diverse aesthetic preferences and cultural background, resulting in homogeneity of image, and the operation of users uploading avatars independently is complicated, especially unfriendly to elderly users.
By collecting user avatars and generating avatar data sets, using deep learning models to standardize images and text, combining LoRA technology to fine-tune the generative model, and automatically generate personalized avatars based on user description information, lowering the threshold for user operation.
It realizes personalized avatar generation, improves the recognition of digital identity, and simplifies user operation processes, especially the user experience of elderly users.
Smart Images

Figure CN120411290A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of avatar processing technology, and in particular to a method, device, and medium for generating a user avatar based on deep learning. Background Art
[0002] Current user avatar systems generally employ a model that combines user-uploaded avatars with a platform-default library. This traditional approach met basic needs in the early days of the internet, but its limitations have become increasingly apparent as the need to express digital identities has evolved. The default avatar libraries of most social platforms, forum systems, and enterprise software remain largely unupdated and typically contain no more than 50 standardized patterns, such as geometric shapes, animal silhouettes, or minimalist portraits. While universal, these designs fail to accommodate users' diverse aesthetic preferences and cultural backgrounds. More critically, the mechanical repetition of default avatars frequently leads users to encounter image homogeneity. A single gray silhouette icon can represent millions of different individuals, significantly reducing the ability to distinguish digital identities.
[0003] Currently, most social platforms and forum systems support users uploading their own profile pictures. However, these uploads often involve complex operations such as format and size restrictions, cropping, and compliance review. These complex operations create unnecessary operational barriers for users, especially elderly users. Summary of the Invention
[0004] To solve the above problems, this application proposes a user avatar generation method based on deep learning, including: Collect user portraits and generate a portrait data set for storing the user portraits and their corresponding annotated texts; Creating a data processing class for processing the user portrait, traversing the data set through the data processing class to convert the annotated text into a corresponding text annotation tensor, and preprocessing the user portrait; Loading a basic generative model, using the instance of the data processing class as a training data set for the basic generative model, and fine-tuning the basic generative model based on preset LoRA training parameters and the training data set to obtain a trained user avatar generation model; User description information input by the user is obtained, a prompt word is generated according to the user description information, and the user avatar generation model is called according to the prompt word, so as to output a target avatar corresponding to the user through the user avatar generation model.
[0005] In one implementation of the present application, generating an avatar dataset for storing the user avatar and its corresponding annotated text specifically includes: For each user avatar, determine the image description information corresponding to the user avatar; Concatenate the image description information with a preset trigger word to obtain the annotation text corresponding to the user avatar, and generate an annotation file corresponding to the annotation text; wherein, the file naming prefix of the annotation file is consistent with the avatar naming prefix of the user avatar corresponding to the annotation file; Generate a corresponding avatar dataset for the user avatar and the annotation file corresponding to the user avatar.
[0006] In an implementation manner of the present application, traversing the dataset through the data processing class specifically includes: Traverse the dataset through the data processing class to arrange the user avatars and their corresponding annotation files in the dataset according to the avatar naming prefix and the file naming prefix respectively, to obtain corresponding user avatar sequences and annotation file sequences; According to the order of the user avatar sequence, construct a path index array corresponding to the image paths based on the image paths corresponding to each user avatar in the user avatar sequence; Read the annotation files in sequence according to the order of the annotation file sequence to obtain the corresponding annotation text, and construct a text description array composed of the annotation text.
[0007] In an implementation manner of the present application, converting the annotation text into a corresponding text annotation tensor and preprocessing the user avatar specifically includes: Segment the annotation text in the text description array through a tokenizer, and convert the several text segments obtained after segmentation into corresponding text annotation tensors; Load the path index array to obtain the user avatar, and convert the user avatar into an RGB image; Convert the RGB image into an image tensor, and perform preprocessing operations on the image tensor to convert the user avatar into a tensor representation that can be input into a basic generative model.
[0008] In an implementation manner of the present application, traversing the dataset through the data processing class specifically includes: Through the indexing method in the data processing class, traverse the dataset according to the index value passed into the indexing method, so as to obtain the user avatar and the text annotation tensor whose corresponding array subscripts correspond to the index value from the path index array and the text description array according to the index value.
[0009] In an implementation manner of the present application, generating a prompt word according to the user description information specifically includes: Concatenate the user description information after the trigger word to obtain a corresponding prompt word; wherein, the user description information is used to represent the user's basic information, and the user's basic information includes the user's age, gender, occupation, and hobbies.
[0010] In an implementation manner of the present application, after the target avatar corresponding to the user is output by the user avatar generation model, the method further includes: Store the target avatar in a cloud storage platform, generate an access path corresponding to the target avatar, and construct a mapping relationship between the access path and its corresponding user.
[0011] In an implementation manner of the present application, the method further includes: In the case where the user has an avatar access requirement, obtain the mapping relationship corresponding to the user through the user identifier corresponding to the user; Obtain the access path according to the mapping relationship, and obtain the target avatar corresponding to the user according to the access path.
[0012] An embodiment of the present application provides a user avatar generation device based on deep learning, and the device includes: At least one processor; And a memory communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a method for generating a user avatar based on deep learning as described in any one of the above.
[0013] An embodiment of the present application provides a non-volatile computer storage medium storing computer-executable instructions, and the computer-executable instructions are set as: A method for generating a user avatar based on deep learning as described in any one of the above.
[0014] The method for generating a user avatar based on deep learning proposed by the present application can bring the following beneficial effects: By collecting user avatars and constructing a dataset, combining the standardized processing of images and texts by a data processing class, and fine-tuning a basic generative model using the LoRA technique, it is possible to automatically generate highly personalized target avatars according to the description information input by the user, fundamentally solving the problem of user image homogenization caused by the lagging update and single style of the default avatar library in traditional solutions, and endowing digital identities with unique visual recognition. At the same time, users do not need to go through complex avatar uploading, format adjustment, and compliance review processes. They only need to input natural language descriptions to obtain customized avatars, significantly reducing the operation threshold and providing a convenient interaction experience especially for elderly users. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings: Figure 1 is a schematic flowchart of a method for generating a user avatar based on deep learning provided by an embodiment of the present application; Figure 2 is a schematic structural diagram of a device for generating a user avatar based on deep learning provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0017] The following will detail the technical solutions provided by each embodiment of the present application in conjunction with the drawings.
[0018] As Figure 1 shown, a method for generating a user avatar based on deep learning provided by an embodiment of the present application includes: S101: Collect user avatars and generate an avatar dataset for storing user avatars and their corresponding annotation texts.
[0019] To reduce the complexity of user operations, the embodiments of the present application randomly generate user avatars according to the actual avatar needs of users. Compared with the method of users uploading avatars independently, this method breaks through the upload limitations brought by the image size and format, and users do not need to perform any operations on the avatar pictures, reducing the upload difficulty and being more suitable for elderly users. The automatic generation of user avatars relies on a deep learning model to achieve, and the training of the deep learning model requires a large number of labeled image data sets. Therefore, it is necessary to collect registered user avatars through various channels and perform image preprocessing on the user avatars to ensure that each user avatar is 512×512 pixels. Each of these user avatars carries specific information. To distinguish user avatars, it is necessary to generate annotation texts for describing the information carried by the avatars for each user avatar, and summarize the user avatars and their corresponding annotation texts to form a corresponding avatar data set.
[0020] In one embodiment, for each user avatar, determine its corresponding image description information. The image description information is a specific description of the image. For example, a 30-year-old chef who likes anime. After determining the image description information corresponding to each user avatar, splice the image description information with a preset trigger word. The trigger word refers to "a user avatar", which will be used as the trigger word for the subsequent fine-tuned model. The trigger word can make the trained model associated with the generation of user avatars, so as to iteratively train the generative model in the direction of generating user avatars. When splicing, the trigger word comes first and is used as a prompt word for calling the deep learning model, followed by the image description information, which contains the basic information of the user, such as the user's age, gender, occupation, hobbies, and style. After splicing the trigger word and the image description information, the annotation text corresponding to the user avatar can be obtained. For each annotation text, a corresponding annotation file is generated. The annotation file exists in txt format, and each annotation text corresponds to an annotation file. After obtaining the user avatar and its corresponding annotation file, they can be summarized into a data set, and the user avatars and annotation files in the data set are in one-to-one correspondence.
[0021] S102: Create a data processing class for processing user avatars. Through the data processing class, traverse the data set to convert the annotation text into the corresponding text annotation tensor and preprocess the user avatars.
[0022] After generating a dataset with data annotations, it is necessary to process it into a data type that can be understood by a deep learning model. This process needs to be achieved by inheriting the data processing class of torch.utils.data.Dataset. torch.utils.data.Dataset is an abstract class defined in the PyTorch framework, which is used to represent a dataset. By inheriting this class and implementing its methods, raw data in any format (such as images, text, audio) can be converted into a standardized data format that can be directly used by the PyTorch model. Therefore, it is necessary to create a data processing class avatar_dataset for processing user avatars. By implementing the methods of this data processing class, traverse the dataset, so as to convert the annotation text into the corresponding text annotation tensor, and preprocess the user avatars to form the corresponding tensor representation. Both the user avatars and the annotation text are processed into a tensor format that the model can recognize and process, which can ensure that its input form is compatible with the model structure.
[0023] In one embodiment, the dataset may contain hundreds or thousands of avatar files and annotation files. If not sorted according to rules, the storage order of the files may have nothing to do with the actual corresponding relationship, resulting in data loading chaos during subsequent model training. Therefore, it is necessary to traverse the dataset through the data processing class to match user avatars in different formats, so as to arrange the user avatars in the dataset according to their corresponding avatar naming prefixes to form a corresponding user avatar sequence. Similarly, for the annotation text, it also needs to be arranged according to its corresponding file naming prefix to obtain the corresponding annotation file sequence. For example, for user avatar files: user_001.jpg, user_002.jpg, the avatar naming prefixes are user_001 and user_002. Arranged in the order of the avatar naming prefixes, the obtained user avatar sequence is {user_001.jpg, user_002.jpg}. The order here can be in ascending order or descending order of the naming prefix, and this application does not limit this. After obtaining the user avatar sequence and the annotation file sequence, determine the image path corresponding to each user avatar in turn according to the order of the user avatar sequence, and form the corresponding path index array avatar_path array according to the image path. In addition, it is also necessary to read the annotation files in turn according to the order of the annotation file sequence to obtain the corresponding annotation text, and construct the text description array label array composed of the annotation text.
[0024] In one embodiment, during the process of traversing the dataset, each element traversed needs to be converted into a data format that the model can understand. Therefore, it is necessary to tokenize the annotated text in the text description array through the tokenizer CLIPTokenizer, and convert the several text segments obtained after tokenization into corresponding text annotation tensors. At the same time, load the path index array to obtain the corresponding user avatar according to the path index, and convert the user avatar into a three-channel RGB image. Then, it is also necessary to convert the RGB image into an image tensor torch.Tensor and perform preprocessing operations on the image tensor. Here, the preprocessing operations include flipping, cropping, scaling, normalization, and other operations. After completing the above preprocessing of the user avatar, the user avatar is converted into a tensor representation that can be input into the basic generative model.
[0025] In one embodiment, when traversing the dataset, the data processing class needs to implement an indexing method, that is, the __getitem__ method. The __getitem__ method is used to define the indexing operation of an object. When using this method, samples in the dataset can be obtained according to the passed index value. In the above process, the annotated text and the user avatars are sorted by name, and the positions of each user avatar and its corresponding annotated text in their respective arrays are corresponding. For example, the annotated text of user_001.jpg and user_001.txt corresponds and is both in the first position in the array. Therefore, by implementing the __getitem__ method, the dataset can be traversed according to the passed index value, and the user avatar and the annotated text corresponding to the array subscript corresponding to the index value can be obtained from the path index array and the text description array. For example, when the index value is 2, the obtained user avatar and the annotated text are the text annotation tensors corresponding to user_002.jpg and user_002.txt.
[0026] S103: Load the basic generative model, use the instance of the data processing class as the training dataset of the basic generative model, and fine-tune the basic generative model based on the preset LoRA training parameters and the training dataset to obtain the trained user avatar generation model.
[0027] After completing the construction of the avatar dataset, if you need to further implement the function of generating user avatars, you need to use a diffusion model for training or fine-tuning. The diffusers library provides an implementation environment for the diffusion model, including the model architecture, training process, and inference interface. After installing the diffusers library, load the basic generative model stabilityai / stable-diffusion-2-1-base under stable-diffusion, and create an instance of the data processing class for the train-dataset with some data processing. In this way, the instance of the data processing class can be used as the training dataset for the basic generative model and participate in the subsequent model training process. If you want to obtain a user avatar generation model that can directly generate user avatars, you need to perform parameter fine-tuning on the basis of the basic generative model. Therefore, you need to set the corresponding LoRA training parameters. For example, set height and width to 512 and the number of training steps to 500. In this way, by using the LoRA training parameters and the training dataset to fine-tune the basic generative model, you can obtain a trained user avatar generation model.
[0028] After generating the user avatar generation model, install the fastapi library and load the trained model, so as to convert the user avatar generation model into an avatar generation service that can be accessed through the network. When there is a need for user registration, after simply initiating an HTTP request to call this service, you can generate a target avatar that matches the user information.
[0029] S104: Obtain the user description information input by the user, generate a prompt according to the user description information, and call the user avatar generation model according to the prompt to output the target avatar corresponding to the user through the user avatar generation model.
[0030] The above process realizes the training of the user avatar generation model. During the actual application process of the model, after the user completes the registration process in the system, the system can correspondingly obtain the user description information input by the user, such as personal information like age, gender, occupation, hobbies, style, etc. After saving the user description information to the database, a unique user identifier for the user is generated. At the same time, the user data is delivered to the message queue, enabling the server to obtain the user description information input by the user by listening to the messages in the message queue. After obtaining the user description information, the server concatenates the user description information after the trigger word to obtain the corresponding prompt word. The generated prompt word is used as the input to call the user avatar generation model, and the model processes and analyzes the prompt word according to the mapping relationship learned internally to generate the target avatar corresponding to the user. The target avatar will be directly displayed on the web page. The user only needs to input the user description information to adaptively generate the user avatar without setting the format and size, reducing the complexity of user registration images, and the generated target avatar is relevant to the user's own information and is more easily accepted by the user.
[0031] After the target avatar is generated, it will be stored in the cloud storage platform. The server will correspondingly generate an access path and construct the mapping relationship between the access path and the user corresponding to the target avatar. In this way, when the user has a need to access the avatar, the corresponding mapping relationship of the user can be obtained through the user identifier, and then the access path can be obtained according to the mapping relationship, and thus the target avatar corresponding to the user can be obtained according to the access path.
[0032] The above is the method embodiment proposed by this application. Based on the same idea, some embodiments of this application also provide the devices and non-volatile computer storage media corresponding to the above method.
[0033] Figure 2 It is a schematic structural diagram of a user avatar generation device based on deep learning provided by an embodiment of this application. As Figure 2 shown, it includes: At least one processor; and, A memory communicatively connected to at least one processor; wherein, The memory stores instructions executable by at least one processor, and the instructions are executed by at least one processor so that at least one processor can perform a user avatar generation method based on deep learning as described in any one of the above.
[0034] An embodiment of this application provides a non-volatile computer storage medium storing computer-executable instructions, and the computer-executable instructions are set as: A user avatar generation method based on deep learning as described in any one of the above.
[0035] Each embodiment in this application is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.
[0036] The devices and media provided in the embodiments of this application correspond one-to-one with the methods. Therefore, the devices and media also have beneficial technical effects similar to those of their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be elaborated here.
[0037] Those skilled in the art should understand that the embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0038] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0039] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0040] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0041] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0042] Memory may include non-permanent storage in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0043] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can be implemented using any method or technology to store information. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change RAM (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0044] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0045] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A method for generating user avatars based on deep learning, characterized in that, The method includes: Collecting a user's head portrait and generating a head portrait dataset for storing the user's head portrait and its corresponding annotation text; Creating a data processing class for processing the user's head portrait. Through the data processing class, traverse the dataset to convert the annotation text into a corresponding text annotation tensor and preprocess the user's head portrait; Loading a basic generative model, using the instance of the data processing class as the training dataset of the basic generative model, and fine-tuning the basic generative model based on preset LoRA training parameters and the training dataset to obtain a trained user head portrait generation model; Obtaining user description information input by the user, generating a prompt word according to the user description information, and calling the user head portrait generation model according to the prompt word to output the target head portrait corresponding to the user through the user head portrait generation model.
2. The method for generating a user avatar based on deep learning according to claim 1, wherein, Generating a head portrait dataset for storing the user's head portrait and its corresponding annotation text specifically includes: For each user's head portrait, determining the image description information corresponding to the user's head portrait; Concatenating the image description information with a preset trigger word to obtain the annotation text corresponding to the user's head portrait and generating an annotation file corresponding to the annotation text; wherein, the file naming prefix of the annotation file is the same as the head portrait naming prefix of the user's head portrait corresponding to the annotation file; Generating a corresponding head portrait dataset for the user's head portrait and the annotation file corresponding to the user's head portrait.
3. The method for generating a user avatar based on deep learning according to claim 2, wherein, Traversing the dataset through the data processing class specifically includes: Traversing the dataset through the data processing class to arrange the user's head portraits and their corresponding annotation files in the dataset according to the head portrait naming prefix and the file naming prefix respectively to obtain corresponding user head portrait sequences and annotation file sequences; According to the order of the user head portrait sequence, construct a path index array corresponding to the image paths according to the image paths corresponding to each user head portrait in the user head portrait sequence; Read the annotation files in sequence according to the order of the annotation file sequence to obtain the corresponding annotation text and construct a text description array composed of the annotation text.
4. A method for generating a user avatar based on deep learning according to claim 3, characterized in that, Converting the annotation text into a corresponding text annotation tensor and preprocessing the user's head portrait specifically includes: Segmenting the annotation text in the text description array through a tokenizer and converting several text segments obtained after segmentation into corresponding text annotation tensors; Loading the path index array to obtain the user's head portrait and converting the user's head portrait into an RGB image; Converting the RGB image into an image tensor and performing preprocessing operations on the image tensor to convert the user's head portrait into a tensor representation that can be input into the basic generative model.
5. A method for generating a user avatar based on deep learning according to claim 4, characterized in that, Traversing the dataset through the data processing class specifically includes: Through the indexing method in the data processing class, traverse the data set according to the index value passed into the indexing method, so as to obtain the user avatar and text annotation tensor corresponding to the array subscript in the path index array and the text description array according to the index value.
6. The method for generating a user avatar based on deep learning according to claim 2, wherein, Generating a prompt according to the user description information specifically includes: Concatenating the user description information after the trigger word to obtain a corresponding prompt word; wherein, the user description information is used to represent the user's basic information, and the user's basic information includes the user's age, gender, occupation, and hobbies.
7. A method for generating a user avatar based on deep learning according to claim 1, characterized in that, After the target avatar corresponding to the user is output by the user avatar generation model, the method further includes: Storing the target avatar in a cloud storage platform, generating an access path corresponding to the target avatar, and constructing a mapping relationship between the access path and the corresponding user.
8. A method for generating a user avatar based on deep learning according to claim 7, characterized in that, The method further includes: In the case where the user has an avatar access requirement, obtain the mapping relationship corresponding to the user through the user identifier corresponding to the user; Obtain the access path according to the mapping relationship, and obtain the target avatar corresponding to the user according to the access path.
9. A user avatar generation device based on deep learning, characterized in that, The device includes: At least one processor; And a memory communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a method for generating a user avatar based on deep learning according to any one of claims 1-8.
10. A non-volatile computer storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions are set to: A method for generating a user avatar based on deep learning according to any one of claims 1-8.