Expression generation method and device, electronic equipment and storage medium
By using noisy images and preset expression generation conditions, combined with image segmentation and facial feature extraction methods, the efficiency and accuracy problems when generating personalized expressions in the prior art are solved, and efficient and accurate expression generation is achieved.
Patent Information
- Application Number
- CN202311781842.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-21
- Publication Date
- 2025-06-24
AI Technical Summary
When generating personalized expressions, the prior art requires individual training of character Lora sub-models, resulting in rapid growth in training time and storage space. The model is prone to conflicts between expressions and character IDs when used in combination, making it difficult to ensure the accuracy and efficiency of the generated results.
By obtaining the noise image to be processed and preset expression generation conditions, including character image and image description information, image segmentation, facial features are extracted, and noise-denoising is performed based on these features and description information to generate a designated expression image of the target character.
It realizes efficient and accurate generation of personalized expressions, avoiding the time and storage cost of training the character Lora sub-model alone, and ensuring the accuracy and efficiency of the generated results.
Smart Images

Figure CN120198536A_ABST
Abstract
Description
Background Art
[0002] With the rise of Artificial Intelligence Generated Content (AIGC) technology, a personalized expression generation technology has emerged in social products.
[0003] Specifically, the commonly used implementation method of this technology is as follows: Using multiple different images of a person to train a person Lora sub-model for this person, which can be used to achieve the consistency between the person ID information of the generated image and the person ID information of the input image; at the same time, training an expression Lora sub-model, which can be used to achieve various common human expressions; finally, combining the person Lora sub-model and the expression Lora sub-model and integrating them into the diffusion model (Stablediffuison) to achieve personalized expression generation.
[0004] Since this method is based on the Lora sub-model, for each person, a separate person Lora sub-model needs to be trained for them. When the number of people is tens of millions or more, the time and storage space required for training the model will increase exponentially; moreover, when the person Lora sub-model and the expression Lora sub-model are used in combination, there will be conflicts between the two sub-models, and it is impossible to ensure that the generated results simultaneously meet the specified expression and the specified person.
[0005] In summary, how to efficiently and accurately generate personalized expressions is an urgent problem to be solved. Summary of the Invention
[0006] The embodiments of the present application provide an expression generation method, device, electronic device and storage medium, which are used to efficiently and accurately generate personalized expressions and improve the efficiency and accuracy of expression generation.
[0007] An expression generation method provided by the embodiments of the present application includes:
[0008] Obtain a noise image to be processed and preset expression generation conditions for the noise image, where the expression generation conditions include: at least one person image of the same target person, and image description information including expression description words containing a specified expression;
[0009] Perform image segmentation on the at least one person image respectively to obtain corresponding segmentation images including the target person;
[0010] Extract the face features of the target person based on the obtained at least one segmentation image;
[0011] Based on the face features and the image description information, perform denoising processing on the noisy image to obtain a target image corresponding to the target person and conforming to the specified expression.
[0012] An expression generation device provided by an embodiment of the present application includes:
[0013] An acquisition unit, configured to acquire a noisy image to be processed and an expression generation condition preset for the noisy image, where the expression generation condition includes: at least one person image of the same target person and image description information including expression description words containing a specified expression;
[0014] A segmentation unit, configured to perform image segmentation on the at least one person image respectively to obtain corresponding segmentation images including the target person;
[0015] An extraction unit, configured to extract the face features of the target person based on the obtained at least one segmentation image;
[0016] A generation unit, configured to perform denoising processing on the noisy image based on the face features and the image description information to obtain a target image corresponding to the target person and conforming to the specified expression.
[0017] Optionally, the noisy image is a Gaussian noise image subject to a standard normal distribution.
[0018] Optionally, the extraction unit is specifically configured to:
[0019] Extract the face vectors of the target person included in each of the at least one segmentation image respectively;
[0020] Average the corresponding elements in the extracted face vectors to obtain the face features of the target person.
[0021] Optionally, if multiple target images are obtained, after obtaining the target image corresponding to the target person and conforming to the specified expression, the generation unit is further configured to, for each target image, perform the following operations respectively:
[0022] Determine the expression accuracy of the specified expression included in a target image based on a preset reference expression corresponding to the expression description words;
[0023] Determine the face similarity between the face included in the one target image and a preset reference face corresponding to the face features based on the face features;
[0024] Screen out target images corresponding to the expression accuracy and the face similarity that meet the preset threshold conditions from the multiple target images as target expression data.
[0025] Optionally, the generating unit is specifically configured to:
[0026] From the multiple target images, filter out the target images whose corresponding expression accuracy is greater than a preset accuracy threshold and whose corresponding face similarity is greater than a preset similarity threshold, and use them as target expression data.
[0027] Optionally, the device is implemented using a trained image generation model, and the trained image generation model is obtained by performing iterative training on the image generation model to be trained based on a first training sample set and a second training sample set;
[0028] Wherein, each first training sample in the first training sample set includes a person expression sample image and corresponding actually added noise, and each second training sample in the second training sample set includes a general sample image and corresponding actually added noise; the person expression sample images in the first training sample set are images of different sample persons with various expressions; the general sample images in the second training sample set are images with different contents of various types.
[0029] Optionally, the device further includes a model training unit, and the model training unit is configured to perform the following process during each iterative training:
[0030] Select a first training sample from the first training sample set and a second training sample from the second training sample set;
[0031] Input the first training sample into the image generation model to be trained to obtain the predicted added noise corresponding to the first training sample; and input the second training sample into the image generation model to be trained to obtain the predicted added noise corresponding to the second training sample;
[0032] Based on the differences between the predicted added noises corresponding to the first training sample and the second training sample and the corresponding actually added noises, construct a target loss function;
[0033] Based on the target loss function, adjust the parameters of the image generation model to be trained.
[0034] Optionally, the model training unit is specifically configured to:
[0035] Extract the first text description information corresponding to the person expression sample image in the first training sample, and add the expression vocabulary corresponding to the person expression sample image to the first text description information;
[0036] Perform image segmentation on the associated image of the human expression sample image to obtain a sample segmentation image corresponding to the associated image and containing the human part; the associated image and the human expression sample image are images of the same person with different expressions;
[0037] Input the human expression sample image, the first text description information, and the sample segmentation image into the image generation model to be trained;
[0038] Use the image generation model to be trained, combine the first text description information and the sample segmentation image, and perform denoising processing on the human expression sample image to obtain the predicted added noise corresponding to the first training sample.
[0039] Optionally, the model training unit is specifically configured to:
[0040] Extract the second text description information corresponding to the general sample image in the second training sample;
[0041] Use a preset segmentation image as the segmentation image corresponding to the general sample image, and each pixel value in the preset segmentation image is the same preset value;
[0042] Input the general sample image, the second text description information, and the preset segmentation image into the image generation model to be trained;
[0043] Use the image generation model to be trained, combine the second text description information and the preset segmentation image, and perform denoising processing on the general sample image to obtain the predicted added noise corresponding to the second training sample.
[0044] Optionally, the model training unit is specifically configured to:
[0045] Obtain the expression data difference between the predicted added noise corresponding to the first training sample and the corresponding actual added noise; and
[0046] Obtain the general data difference between the predicted added noise corresponding to the second training sample and the corresponding actual added noise;
[0047] Perform weighted summation on the expression data difference and the general data difference to obtain the target loss function.
[0048] Optionally, the obtaining unit is further configured to, before performing image segmentation on the at least one human image, for each human image, perform the following operations respectively:
[0049] Perform image detection on a human image;
[0050] If it is determined that the one person image does not contain a face, a prompt message indicating that the one person image is unqualified is fed back.
[0051] An electronic device provided by an embodiment of the present application includes a processor and a memory. Among them, the memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of any one of the above-mentioned expression generation methods.
[0052] An embodiment of the present application provides a computer-readable storage medium, which includes a computer program. When the computer program runs on an electronic device, the computer program is used to make the electronic device execute the steps of any one of the above-mentioned expression generation methods.
[0053] An embodiment of the present application provides a computer program product. The computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the electronic device executes the steps of any one of the above-mentioned expression generation methods.
[0054] The beneficial effects of the present application are as follows:
[0055] An embodiment of the present application provides an expression generation method, device, electronic device, and storage medium. Since the expression generation method in the present application mainly performs denoising processing on a to-be-processed noisy image based on preset expression generation conditions, and the expression generation conditions include two major parts. The first part is a person image, which is used to provide person ID information. The second part is image description information, which is an expression description vocabulary containing a specified expression and is used to provide expression information. Therefore, for different expressions of different people, only different expression generation conditions need to be set, and the to-be-processed noisy images can be kept consistent. Therefore, for any person, a complete set of expressions can be generated without separately training a person Lora sub-model.
[0056] In addition, in the present application, when generating a target image of a specified expression corresponding to a target person, the person ID information finally provided is the face feature extracted based on the segmented image, and the face feature can be extracted from multiple different images of the same person, which can effectively reduce the influence of environmental noise on the generation result.
[0057] In summary, the present application can generate personalized expressions more efficiently and accurately.
[0058] Other features and advantages of the present application will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of the present application. The objectives and other advantages of the present application may be realized and attained by the structure particularly pointed out in the written description, claims, as well as the drawings. Description of the Drawings
[0059] The drawings described herein are for further understanding of the present application, and form a part of the present application. The schematic embodiments and descriptions thereof of the present application are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings:
[0060] Figure 1 is an optional schematic diagram of an application scenario in an embodiment of the present application;
[0061] Figure 2 is a flowchart of an implementation of an expression generation method in an embodiment of the present application;
[0062] Figure 3 is a schematic diagram of a Gaussian noise image in an embodiment of the present application;
[0063] Figure 4A is a schematic diagram of a human image in an embodiment of the present application;
[0064] Figure 4B is another schematic diagram of a human image in an embodiment of the present application;
[0065] Figure 5A is a schematic diagram of a segmented image in an embodiment of the present application;
[0066] Figure 5B is another schematic diagram of a segmented image in an embodiment of the present application;
[0067] Figure 6A is a schematic diagram of an image of various types of different contents listed in the present application;
[0068] Figure 6B is another schematic diagram of an image of various types of different contents listed in the present application;
[0069] Figure 7 is a schematic diagram of a training process of an image generation model in an embodiment of the present application;
[0070] Figure 8A is a schematic diagram of a noise prediction process of a first training sample in an embodiment of the present application;
[0071] Figure 8B is a schematic diagram of a noise prediction process of a second training sample in an embodiment of the present application;
[0072] Figure 9 A logical schematic diagram for generating a target image using an image generation model in an embodiment of this application;
[0073] Figure 10 A schematic diagram of an application scenario related to a social product in an embodiment of this application;
[0074] Figure 11 A schematic diagram of a prompt message in an embodiment of this application;
[0075] Figure 12 A schematic diagram of a set of target images in an embodiment of this application;
[0076] Figure 13 A logical interaction schematic diagram between a terminal device and a server in an embodiment of this application;
[0077] Figure 14 A schematic diagram of the composition structure of an expression generation device in an embodiment of this application;
[0078] Figure 15 A schematic diagram of a hardware composition structure of an electronic device applying an embodiment of this application;
[0079] Figure 16 A schematic diagram of a hardware composition structure of another electronic device applying an embodiment of this application. Detailed implementation manners
[0080] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions of this application in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of the technical solutions of this application. Based on the embodiments recorded in this application document, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the technical solutions of this application.
[0081] The following introduces some concepts involved in the embodiments of this application.
[0082] Person identity (ID) information: refers to information that can uniquely identify a person, specifically some information characterizing the identity characteristics of a person, such as some characteristics in the appearance of a person, such as arched eyebrows, double eyelids, and other information.
[0083] Noisy image: Specifically, it can be an image obtained by randomly adding noise, where noise specifically refers to various factors in the image that hinder people from receiving its information. In the embodiments of the present application, by combining the expression generation conditions, the noisy image to be processed is denoised, and an expression image meeting the expected conditions can be generated. For any target person, the same noisy image or a similar type of noisy image can be used as the input for the subsequent model, and the noisy image is denoised, only the preset expression generation conditions need to be distinguished.
[0084] Image description information: It refers to the description information generated by extracting features from the image. This description information specifically describes the content contained in the image, which can be in the form of text, and the present application does not limit the language form of this text.
[0085] Expression accuracy: Specifically, it refers to the accuracy of the expression of the person in the image generated by the image generation model compared with the preset reference expression corresponding to the expression description vocabulary of the expression specified by the initial conditions. Specifically, the closer the expression of the person in the generated image is to the corresponding preset reference expression, the higher the corresponding expression accuracy; otherwise, it is lower.
[0086] Cosine similarity: It measures the similarity between two vectors by measuring the cosine value of the angle between them. The cosine value of a 0-degree angle is 1, and the cosine value of any other angle is not greater than 1; and its minimum value is -1. Thus, the cosine value of the angle between two vectors determines whether the two vectors generally point in the same direction. When the two vectors have the same direction, the value of the cosine similarity is 1; when the included angle between the two vectors is 90°, the value of the cosine similarity is 0; when the two vectors point in exactly opposite directions, the value of the cosine similarity is -1. This result is independent of the length of the vectors and is only related to the direction of the vectors. Cosine similarity is usually used in the positive space, so the value given is between 0 and 1.
[0087] The embodiments of the present application relate to artificial intelligence (AI) and machine learning technologies, and are designed based on computer vision technology (CV) and machine learning (ML) in artificial intelligence.
[0088] AI uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems that can perceive the environment, acquire knowledge, and use knowledge to achieve the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce an intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making.
[0089] AI technology covers a wide range of fields, including both hardware-level and software-level technologies. Specifically, the basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0090] Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes for tasks such as object recognition and measurement in machine vision, and further performing graphic processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems that can obtain information from images or multi-dimensional data. The large model technology has brought important changes to the development of computer vision technology. Pre-trained models in the field of vision such as Swin-transformer, Vision Transformer (ViT), Vision-Mixture of Experts (V-MOE), and Masked AutoEncoder (MAE) can be quickly and widely applied to downstream specific tasks after fine-tuning. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, Optical Character Recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0091] The image generation model in the embodiments of this application is trained using machine learning or deep learning techniques. Machine learning is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning generally include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. The pre-trained model is the latest development result of deep learning, integrating the above technologies.
[0092] After training the image generation model based on the above technologies, the image generation model can be applied to generate personalized expressions related to an object (such as a user, referring to a person using a terminal device).
[0093] In addition, it should be noted that the image generation model in the embodiments of this application can be trained online or offline, and no specific limitation is made here. In this article, offline training is taken as an example for illustration.
[0094] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common ones include smart home, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, digital twins, virtual humans, robots, AIGC, conversational interaction, intelligent healthcare, intelligent customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0095] The following briefly introduces the design concept of the embodiments of this application:
[0096] With the rapid development of information technology, instant messaging applications have become an important tool for people's social interactions.
[0097] Emoji is a way of expressing emotions using images. People use current popular stars, quotes, animations, and movie and TV screenshots as materials, and match them with a series of corresponding texts to express specific emotions. Emoji is a way of communication that became popular on the Internet after the active development of social applications, and basically everyone can use emojis. However, the types of existing emojis are relatively single.
[0098] Therefore, in social products, a technology for generating personalized expressions using AIGC has emerged. Besides the implementation methods listed in the background art, another common implementation method for this technology is as follows: Extract ten million face data from a large-scale open-source dataset for training the base model Stable Diffusion, so that the person ID information in the images generated by the base model is consistent with the input person ID information. By adding expression vocabulary to the input text information, personalized expression generation is achieved.
[0099] In the above method, although the training and storage costs of the model are reduced, on the one hand, since it adds ten million face data for training the base model, it will cause the model to forget the prior information originally obtained from the large-scale pre-trained dataset, and thus overfit to the face dataset, ultimately resulting in a weakened editable ability or even only being able to generate face results. On the other hand, since this method does not perform any processing on the generated person expressions and completely relies on the expression generation ability of the original base model, this leads to a very weak expression editing ability of this method.
[0100] In summary, neither the method of training the person Lora sub-model nor the above method can generate personalized expressions efficiently and accurately.
[0101] In view of this, the embodiments of the present application propose an expression generation method, device, electronic device, and storage medium. Since the expression generation method in the present application mainly performs denoising processing on the noise image to be processed based on preset expression generation conditions, and this expression generation condition includes two main parts. The first part is the person image, which is used to provide person ID information, and the second part is the image description information, which contains expression description vocabulary with specified expressions and is used to provide expression information. Therefore, for different expressions of different people, only different expression generation conditions need to be set, and the noise image to be processed can remain the same. Thus, for any person, a complete set of expressions can be generated without separately training the person Lora sub-model.
[0102] In addition, in the present application, when generating the target image of the specified expression corresponding to the target person, the person ID information finally provided is the face feature extracted based on the segmented image, and this face feature can be extracted from multiple different images of the same person, which can effectively reduce the impact of environmental noise on the generation result.
[0103] In summary, the present application can generate personalized expressions more efficiently and accurately.
[0104] The solution provided by the embodiments of the present application relates to the application of artificial intelligence technology in the field of AIGC, and is specifically illustrated through the following embodiments:
[0105] The preferred embodiments of the present application will be described below in conjunction with the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application. And without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.
[0106] As Figure 1 shown, it is a schematic diagram of the application scenario of the embodiment of the present application. The application scenario diagram includes two terminal devices 110 and a server 120.
[0107] In the embodiment of the present application, the terminal device 110 includes, but is not limited to, devices such as mobile phones, tablet computers, laptop computers, desktop computers, e-book readers, intelligent voice interaction devices, intelligent home appliances, vehicle-mounted terminals, etc.; a client related to expression generation can be installed on the terminal device, and the client can be software (such as a browser, instant messaging software, etc.), or a web page, a small program, etc. The server 120 is a background server corresponding to the software, web page, small program, etc., or a server dedicated to expression generation. The present application does not make specific limitations. The server 120 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (Content Delivery Network, CDN), and big data and artificial intelligence platforms.
[0108] It should be noted that the expression generation method in each embodiment of the present application can be executed by an electronic device, and the electronic device can be the terminal device 110 or the server 120, that is, the method can be executed independently by the terminal device 110 or the server 120, or jointly executed by the terminal device 110 and the server 120.
[0109] For example, when jointly executed by the terminal device 110 and the server 120, a client related to expression generation can be installed on the terminal device 110. The object can upload at least one personal image corresponding to itself based on this client, and describe the image description information corresponding to the expression image to be generated. The image description information should include the expression description vocabulary of the specified expression. The server 120 obtains the above-mentioned personal image and image description information through the terminal device 110, and uses these as the expression generation conditions for the noise image to be processed. Furthermore, the server 120 performs image segmentation on at least one personal image to obtain the corresponding segmented image containing the target person. Then, the server 120 extracts the facial features of the target person based on the obtained at least one segmented image. Subsequently, the server 120 performs denoising processing on the noise image based on the facial features and the image description information to obtain the target image corresponding to the target person that conforms to the specified expression. Finally, the server 120 can feedback the generated target image to the terminal device 110, and the terminal device 110 presents it to the object through the client. The object can download and save these target images, etc., which can be applied to subsequent scenarios such as chatting.
[0110] In an alternative embodiment, the terminal device 110 and the server 120 can communicate through a communication network.
[0111] In an alternative embodiment, the communication network is a wired network or a wireless network.
[0112] It should be noted that Figure 1 The above is only an example. In fact, the number of terminal devices and servers is not limited and is not specifically defined in the embodiments of the present application.
[0113] In the embodiments of the present application, when the number of servers is multiple, the multiple servers can form a blockchain, and the server is a node on the blockchain; for example, in the expression generation method disclosed in the embodiments of the present application, the data involved can be saved on the blockchain, such as noise images, expression generation conditions, segmented images, facial features, target images, target expression data, etc.
[0114] In addition, the embodiments of the present application can be applied to various scenarios, including but not limited to scenarios such as cloud technology, artificial intelligence, intelligent transportation, and assisted driving.
[0115] Next, in combination with the above-described application scenarios, the expression generation method provided by the exemplary embodiments of the present application will be described with reference to the accompanying drawings. It should be noted that the above application scenarios are only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited in this regard.
[0116] Refer to Figure 2As shown in the figure, it is a flowchart of the implementation of an expression generation method provided by an embodiment of the present application. Taking the server as the execution entity as an example, the specific implementation process of this method is as follows in S21 to S24:
[0117] S21: Obtain the noise image to be processed, and the preset expression generation conditions for the noise image. The expression generation conditions include: at least one person image of the same target person, and the image description information including the expression description vocabulary containing the specified expression.
[0118] In the embodiment of the present application, the expression generation conditions are divided into two major parts. The first part is the person image, which is used to provide the person ID information. The second part is the image description information, which contains the expression description vocabulary of the specified expression and is used to provide the expression information. In addition, the specified expression included in the image description information is in the form of a vocabulary, that is, the image description information is in text form, and can also be called the pre-set text description information, that is, the "expression text".
[0119] Specifically, the specified expression refers to the expression expected to be generated this time, and can specifically be various human expressions, such as sad, happy, afraid, surprised, calm, angry, disgusted, etc.
[0120] Based on the above expression generation conditions, the noise image to be processed can be denoised to ensure that the generated image after denoising is an image corresponding to the target person that conforms to the specified expression.
[0121] It should be noted that in the embodiment of the present application, the noise image to be processed can be an image obtained by randomly adding noise. In the embodiment of the present application, by combining the expression generation conditions to denoise the noise image to be processed, an expression image that meets the expected conditions can be generated. For any target person, the same noise image or a similar type of noise image can be used as the input of the subsequent model to denoise the noise image, and only the preset expression generation conditions need to be distinguished.
[0122] In the embodiment of the present application, the type of noise added to the noise image to be processed should be the same as the type of noise added to the sample image in the training process of the image generation model listed below. There are basically the following four common types of image noise: Gaussian noise, Poisson noise, multiplicative noise, and salt-and-pepper noise.
[0123] Optionally, the noise image is a Gaussian noise image that follows a standard normal distribution.
[0124] As Figure 3 shown, it is a schematic diagram of a Gaussian noise image in the embodiment of the present application. The probability density function of the noise in this image follows a standard normal distribution, that is, the noise in this image is Gaussian noise.
[0125] In the embodiments of the present application, in addition to the Gaussian noise images listed above, the noise image can also be other similar ones, such as Poisson noise images, multiplicative noise images, salt-and-pepper noise images, etc., which are not specifically limited herein.
[0126] In addition, the object can upload at least one person image of the same target person. For example, when the object is the target person, the target person can upload one or more of their own images.
[0127] It should be noted that during the generation process of an expression image, if multiple person images are obtained, it is necessary to ensure that the multiple person images contain the same person, that is, the target person to be generated.
[0128] Such as Figure 4A shown, it is a schematic diagram of a person image in the embodiments of the present application. Among them, Figure 4A the listed person image is a girl with long black hair wearing a blue shirt.
[0129] Such as Figure 4B shown, it is a schematic diagram of another person image in the embodiments of the present application. Among them, Figure 4B the listed multiple person images contain the same target person, specifically several expression images of a short-haired boy, such as Figure 4B the listed angry, disgusted, afraid, happy, calm, sad, and surprised.
[0130] It should be noted that the person images in the expression generation conditions of the present application are mainly used to provide person ID information. Therefore, whether the person in the person image has an expression, or which / what expressions it has, are not specifically limited herein. And the image description information in the expression generation conditions must include the specified expression.
[0131] S22: Perform image segmentation on at least one person image respectively to obtain corresponding segmentation images containing the target person.
[0132] This step is the step of preprocessing the person image. For each person image, image segmentation needs to be performed on the person image. The specific image segmentation process is as follows:
[0133] For each person image, first perform panoramic segmentation on the person image to obtain each entity contained in the person image; then, determine the category of each obtained entity to distinguish the person (referring to the target person) and the background (referring to other entities except the target person) in the person image; finally, retain the person and set the pixel values corresponding to the background to the same preset value, and the segmentation image corresponding to the person image containing the target person can be obtained.
[0134] Among them, when performing panoramic segmentation on a person image, the mask2former (Mask-attention MaskTransformer) model can be used. Mask2former is a new architecture that can solve any image segmentation task (panoramic, instance, or semantic). Of course, other models can also be used, such as mask-rcnn (Mask-Region-based Convolutional Neural Networks), Fully Convolutional Networks (FCN), Recurrent Neural Networks (RNN), Long Short-Term Memory (LSTM) networks, etc. This is not specifically limited in this article.
[0135] Among them, the preset value can be flexibly set according to actual needs. For example, it can be set to 0 (indicating black), or for another example, it can be set to 1 or 255 (indicating white), etc. This is not specifically limited in this article.
[0136] Refer to Figure 5A As shown, it is a schematic diagram of a segmented image in an embodiment of this application. Specifically, Figure 5A The segmented image shown refers to Figure 4A the result obtained by performing image segmentation on the person image shown respectively. This segmented image removes the background information in the original person image and only retains the target person, that is, a girl with long black hair wearing a blue shirt.
[0137] Refer to Figure 5B As shown, it is another schematic diagram of a segmented image in an embodiment of this application. Specifically, Figure 5B Each of the segmented images shown refers to Figure 4B the result obtained by performing image segmentation on each person image shown respectively. These segmented images all remove the background information in the original person image and only retain the target person, that is, a short-haired boy.
[0138] It should be noted that the above-listed process of image segmentation is only a simple example. In addition, other image segmentation methods are also applicable to the embodiments of this application and will not be elaborated one by one here.
[0139] S23: Extract the face features of the target person based on the obtained at least one segmented image.
[0140] In the embodiments of this application, after obtaining a segmented image that only contains the target person, the face features corresponding to the target person can be extracted based on the segmented image. These face features are mainly used in the subsequent denoising process to remove relevant noises and restore the ID information of the target person.
[0141] Optionally, if only one person image is included in the preset expression generation condition, then in S22, a single segmented image is obtained. In this case, directly extract the face vector of the target person in the segmented image as the face feature of the final target person.
[0142] Assume that the face vector extracted based on the segmented image is an n-dimensional vector, and the expression generation condition includes 1 person image. The face vector extracted from the segmented image of this person image is denoted as X1 = [X11, X12,..., X1n].
[0143] Then, the face feature A of the final target person = [A1, A2,..., An] = [X11, X12,..., X1n].
[0144] Optionally, if the preset expression generation condition includes multiple person images, then in S22, multiple segmented images are obtained, with each person image corresponding to one segmented image. In this case, respectively extract the face vectors of the target person in each segmented image, and then average the corresponding elements in the extracted face vectors as the face feature of the final target person.
[0145] Assume that the face vector extracted based on the segmented image is an n-dimensional vector, and the expression generation condition includes 7 person images. The face vectors extracted from the segmented images of these 7 person images are respectively denoted as X1 = [X11, X12,..., X1n], X2 = [X21, X22,..., X2n], X3 = [X31, X32,..., X3n], X4 = [X41, X42,..., X4n], X5 = [X51, X52,..., X5n], X6 = [X61, X62,..., X6n], X7 = [X71, X72,..., X7n].
[0146] Then, the face feature A of the final target person = [A1, A2,..., An]. Wherein:
[0147] A1 = (X11 + X21 + X31 + X41 + X51 + X61 + X71) / 7;
[0148] A2 = (X12 + X22 + X32 + X42 + X52 + X62 + X72) / 7;
[0149] ...;
[0150] An = (X1n + X2n + X3n + X4n + X5n + X6n + X7n) / 7.
[0151] In the above embodiments, by averaging the feature vectors of multiple different images of the same person, the influence of environmental noise on the generation result can be significantly reduced.
[0152] In addition, it should be noted that the above method of determining the facial features of the target person by averaging vectors is only a simple example. In addition, other methods can also be used, such as vector weighted summation, vector cross multiplication, etc. This article does not make specific limitations.
[0153] S24: Based on the facial features and image description information, perform denoising processing on the noisy image to obtain a target image corresponding to the target person with a specified expression.
[0154] In the embodiments of the present application, the facial features and image description information can be used as the information sources of the noisy image. By combining the facial features to denoise the noisy image, information such as the appearance of the target person can be gradually restored. By combining the image description information to denoise the noisy image, the expression of the target person can be gradually restored to conform to the preset specified expression.
[0155] Optionally, the above expression generation method can be implemented using a pre-trained image generation model.
[0156] Taking Stablediffusion as the base model as an example, based on the Stable Diffusion base model, directly train a base model with personalized expression generation. Without training, only by inputting at least one photo of a person and an expression vocabulary, a personalized expression picture can be generated.
[0157] To solve the existing problems, the present application proposes a personalized expression generation scheme based on prior preservation, cross-reference, and expression control, which is used to train a brand-new base model, and the base model trained based on this is a base model that retains prior information, decouples expression information and person ID information, and can independently control expressions. The object only needs the user to provide at least one photo, and a complete set of expressions can be generated for the object without training.
[0158] The training process of the image generation model will be described first below:
[0159] An optional implementation manner is that the pre-trained image generation model is obtained by performing iterative training on the image generation model to be trained based on the first training sample set and the second training sample set.
[0160] Among them, each first training sample in the first training sample set includes a sample image of a person's expression and the corresponding actually added noise, and each second training sample in the second training sample set includes a general sample image and the corresponding actually added noise; the sample images of different people with various expressions in the first training sample set; the general sample images in the second training sample set are images of different contents of various types.
[0161] Therefore, before performing iterative training on the image generation model to be trained, it is first necessary to prepare data and construct a training sample set.
[0162] As described above, the training sample set in this application is divided into two categories: the first training sample set is data for the model to learn expression control, and the second training sample set is data for the model to maintain prior information.
[0163] When constructing the first training sample set, a large amount of human expression data needs to be collected, specifically including but not limited to various human natural expressions such as sadness, happiness, fear, surprise, calmness, anger, disgust, etc. As Figure 4B shown, it is a schematic diagram of various different types of expressions of the same person listed in the embodiments of this application. Specifically, multiple Figure 4B such human expression images can be collected. For example, a total of 36,000 images are collected. These images are various different types of expressions of different people, and each image has a corresponding expression vocabulary.
[0164] For example, Figure 4B for the first row of the first human expression image, the corresponding expression vocabulary is anger, the second human expression image is disgust, the third human expression image is fear, and the fourth human expression image is happiness; for the first row of the second human expression image, the corresponding expression vocabulary is calmness, the second human expression image is sadness, and the third human expression image is surprise; and so on.
[0165] In the embodiments of this application, the above-mentioned images related to the collected expression data are the human expression sample images.
[0166] When constructing the second training sample set, it is necessary to collect a large amount of original data for Stablediffusion pre-training. These are general data, specifically including various types of images with different contents.
[0167] As Figure 6A 、 Figure 6B shown, it is a schematic diagram of two various types of images with different contents listed in this application. As Figure 6A listed in, the images include various types such as human figures, animal figures, architectural drawings, etc. In addition to humans, the contents included in the images can also be animals, buildings, etc. Another example is Figure 6B shown, Figure 6B the images listed in include various types such as human figures, landscape pictures, animal figures, ancient style pictures, etc. In addition to humans, the contents included in the images can also be animals, cartoon characters, buildings, and so on.
[0168] In addition to the aboveFigure 6A , Figure 6B In addition to the listed content, it can be plants, or transportation facilities such as vehicles and cruise ships, buildings such as hospitals and schools, or game characters, game equipment, game maps, etc., which will not be elaborated one by one here.
[0169] Specifically, multiple images such as Figure 6A , Figure 6B etc. can be collected, for example, a total of 600 million images are collected.
[0170] In the embodiment of the present application, the above-mentioned collected general data-related images are general sample images.
[0171] For the above two types of image data, the following processing can be pre-performed in the data preparation stage before model training: First, extract text descriptions; Second, perform image segmentation. It can also be processed as above after selecting samples in the model training stage. This is not specifically limited in this article.
[0172] In the embodiment of the present application, after preparing the above two types of sample images, noise can be randomly added to these sample images, that is, the actual noise addition in this article. Using this actual noise addition as a sample label, it is compared with the predicted added noise predicted by the subsequent image generation model to adjust the parameters of the model. Taking the example of performing the above processing after selecting samples in the model training stage, an optional model training method is as follows:
[0173] Refer to Figure 7 shown, which is a schematic diagram of the training process of an image generation model in the embodiment of the present application. Taking the server as the execution entity as an example, in each iterative training, the steps S71 to S75 shown in Figure 7 are executed:
[0174] S71: Select a first training sample from the first training sample set and a second training sample from the second training sample set.
[0175] In the embodiment of the present application, each time of training, assuming that the batch size is N, then N first training samples are taken from the first training sample set, and at the same time N second training samples are taken from the second training sample set and jointly added to the current iterative training process.
[0176] Among them, the batch size represents the number of samples used for training in a single iteration of the training process. For example, if there are 1000 data in the training sample set and the batch size is set to 100, then first, the first 100 training samples in the training sample set, that is, the 1st to 100th training samples, are used to train the model. After the training is completed, the model parameters (such as weights) are updated, and then the 101st to 200th training samples are used for training until the 1000 training samples in the training sample set are used up for the tenth time and then the process stops.
[0177] In the embodiments of the present application, by setting the batch size, the use of memory can be effectively reduced, and the training speed can be improved.
[0178] It should be noted that the above-listed method of selecting training samples is only a simple example, and other methods are also applicable to the embodiments of the present application, and no specific limitations are made here.
[0179] After selecting N first training samples and N second training samples, each training sample can be used to predict noise based on the following method:
[0180] S72: Input the first training sample into the image generation model to be trained, and obtain the predicted added noise corresponding to the first training sample.
[0181] In the embodiments of the present application, the above-listed processes of extracting text descriptions and performing image segmentation can be executed in this step. Similar to the processing process in the model application process, in the model application process as shown in Figure 2 , the text description of the image is preset. In addition, the human part can be extracted through image segmentation. Furthermore, by combining the text description of the image and the face features related to the human as control conditions, the noise image can be denoised, and this denoising process is the process of predicting and adding noise. In this step, that is, by combining the text description information of the sample image of the human expression category and the segmented image as control conditions, the original sample image of the human expression is denoised.
[0182] In the model training stage, the text description of the sample image of the human expression is not preset. For each sample image of the human expression, relevant descriptions need to be additionally extracted; similarly, for each sample image of the human expression, image segmentation also needs to be additionally performed.
[0183] An optional implementation method is that step S72 can be implemented according to the sub-steps S721 to S724 as shown in Figure 7 :
[0184] S721: Extract the first text description information corresponding to the sample image of the human expression in the first training sample, and add the expression vocabulary corresponding to the sample image of the human expression to the first text description information.
[0185] In the embodiments of the present application, the second-generation Bootstrapping Language-Image Pretraining (Blip-2) model can be used to extract the image caption of an image. For the facial expression sample image in the first training sample, its corresponding image caption is the first text description information, which is the same type of information as the above-listed image description information.
[0186] For example, the image caption extracted from Figure 4A the listed facial expression sample image is "A woman with long black hair and a blue shirt", and the image caption extracted from Figure 4B the first facial expression sample image in the first row is "An angry boy with short hair"; and so on.
[0187] It should be noted that in the embodiments of the present application, the first text description information extracted from the facial expression sample image may contain expression-related words, such as "angry" in the above-listed "An angry boy with short hair", or may not contain expression-related words, such as the above-listed "A woman with long black hair and a blue shirt". Whether it contains or not, the expression word corresponding to the facial expression sample image needs to be further added to the first text description information corresponding to the image.
[0188] For example, for the above "A woman with long black hair and a blue shirt", after adding the expression word "smile", it becomes: "A woman with long black hair and a blue shirt, Smile".
[0189] Another example is the above "An angry boy with short hair", after adding the expression word "anger", it becomes: "An angry boy with short hair, Anger"; or updated to "A boy with short hair, Anger"; and so on.
[0190] It should be noted that, in addition to the Blip-2 model listed above, other models can also be used to extract the text description of the image, and this application does not make specific limitations.
[0191] S722: Perform image segmentation on the associated image of the person's expression sample image to obtain a sample segmentation image including the person part corresponding to the associated image; the associated image and the person's expression sample image are images of the same person with different expressions.
[0192] Similar to the image segmentation process in S22 above, for the sample images of person expressions, image segmentation can also be performed in the manner listed above.
[0193] It should be noted that this application proposes a learning method for decoupling expressions and person IDs designed based on cross-reference. Specifically, this learning method divides the control conditions (also known as expression generation conditions) into two parts, namely the person images (specifically segmentation images) listed above for controlling person IDs, and the description information for controlling person expressions. This method allows the model to decouple expressions and person IDs from different expressions of the same person ID.
[0194] In the model training stage, for the first training sample, the above two control conditions are the sample segmentation image corresponding to the person's expression sample image, and the first description information.
[0195] As Figure 8A shown, it is a schematic diagram of the noise prediction process of a first training sample in an embodiment of this application. For the expression data, an Anger image of a person is input, and it is required that the model output the same Anger image. However, in this application, the input control conditions are two. One input is an image of another expression of the same person ID, such as Figure 8A in which is a Disgust image of a person, and this image is the segmentation image corresponding to the original image, with the background information removed, making the model focus more on the person; the other input is the text description information of the Anger image. Because it is expression data, so Figure 8A the text description in it adds the expression word "Anger".
[0196] This application hopes that the model can learn the person ID information from the segmentation image of the input Disgust image and learn the expression information from the text description. For example, Figure 8A in which, it is hoped that the model can learn the expression semantic information from "Anger".
[0197] Therefore, in step S722, instead of performing image segmentation on the facial expression sample image in the first training sample, image segmentation is performed on the associated image of the facial expression sample image, that is, for the image of the facial expression sample image containing different expressions of the same person, image segmentation is performed to obtain a sample segmentation image of the part containing the person corresponding to the associated image.
[0198] In the embodiment of the present application, the mask2former model can be used to perform image segmentation on the associated image, and the specific process is as follows:
[0199] First, perform panoramic segmentation on the associated image to obtain each entity contained in the associated image; furthermore, determine the categories of the obtained entities to distinguish the people and the background in the associated image; finally, retain the people and set the pixel values corresponding to the background to the same preset value (such as set to 0), and the sample segmentation image corresponding to the associated image can be obtained.
[0200] For example, the Figure 4A corresponding segmentation image of the shown image is as Figure 5A , the Figure 4B corresponding segmentation image of the shown image is as 5B, and the segmentation processes and effects of other images are similar, which will not be repeated here.
[0201] It should be noted that in addition to the mask2former model listed above, other models can also be used to perform image segmentation. For details, please refer to the above embodiments, and the repeated parts will not be elaborated.
[0202] S723: Input the facial expression sample image, the first text description information, and the sample segmentation image into the image generation model to be trained.
[0203] S724: Use the image generation model to be trained, combine the first text description information and the sample segmentation image, and perform denoising processing on the facial expression sample image to obtain the predicted added noise corresponding to the first training sample.
[0204] In the embodiment of the present application, the image generation model can be a generation model such as the Stable Diffusion base model, Imagen, Dalle, etc. There is no specific limitation in this article. Here, the Stable Diffusion base model is taken as an example for illustration:
[0205] After preparing the above data, the Stable Diffusion base model can be trained using the above data. Specifically, input the facial expression sample image, the first text description information, and the sample segmentation image into the image generation model to be trained. The model will predict the noise added to the facial expression sample image and perform denoising to restore the original facial expression sample image.
[0206] S73: Input the second training sample into the image generation model to be trained, and obtain the predicted added noise corresponding to the second training sample.
[0207] In the embodiments of the present application, the above-mentioned processes of extracting text descriptions and performing image segmentation can be executed in this step. Similar to the processing process during model application, in the model application process as shown in Figure 2 , the text description of the image is preset. In addition, the human part can be extracted through image segmentation. Furthermore, by combining the text description of the image and the face features related to the human as control conditions, the noise image can be denoised, and this denoising process is the process of predicting and adding noise. In this step, that is, by combining the text description information of the general sample image and the segmented image as control conditions, the original general sample image is denoised.
[0208] In the model training stage, the text description of the general sample image is not preset. For each general sample image, relevant descriptions need to be extracted additionally; similarly, for each general sample image, image segmentation also needs to be performed additionally.
[0209] It should be noted that since the general sample image is an image containing various types and contents, and may not contain people. Therefore, the image segmentation of the general sample image in the present application can be understood as a process of using the preset segmented image as the segmentation result of the sample image.
[0210] An optional implementation method is that step S73 can be implemented according to the sub-steps S731 to S734 as shown in Figure 7 :
[0211] S731: Extract the second text description information corresponding to the general sample image in the second training sample.
[0212] This step is similar to the above S72, and the image caption of the image can be extracted using the second-generation Blip-2 model. For the general sample image in the second training sample, its corresponding text description is the first text description information, which is the same type of information as the above-listed image description information.
[0213] For example, the image caption extracted from the first row, third person's expression sample image in Figure 6A is "An Owl in flight"; and so on.
[0214] It should be noted that in the embodiments of the present application, different from the person expression sample image, there is no need to add expression vocabulary to the second text description information corresponding to the general sample image.
[0215] S732: Use the preset segmented image as the segmented image corresponding to the general sample image, where each pixel value in the preset segmented image is the same preset value.
[0216] In the model training stage, for the second training sample, the above two control conditions are the preset segmented image corresponding to the general sample image and the second description information.
[0217] Among them, the preset value can be flexibly set according to actual needs. For example, it can be set to 0 (indicating black), or it can be set to 1 or 255 (indicating white), etc. This is not specifically limited in this article.
[0218] S733: Input the general sample image, the second text description information, and the preset segmented image into the image generation model to be trained.
[0219] S734: Use the image generation model to be trained, combine the second text description information and the preset segmented image, and perform denoising processing on the general sample image to obtain the predicted added noise corresponding to the second training sample.
[0220] Taking the preset value as 0 as an example, in the embodiments of the present application, for general data, the processing method is the same as that of expression data. The difference is that the segmented image used to control the person ID for general data is an image with all pixel values being 0, and there are no expression words in the text description information.
[0221] In the embodiments of the present application, after preparing the above data, the StableDiffusion base model can be trained using the above data. Specifically, input the general sample image, the second text description information, and the preset segmented image into the image generation model to be trained. The model will predict the noise added to the general sample image and perform denoising to restore the initial general sample image.
[0222] Such as Figure 8B shown, which is a schematic diagram of the noise prediction process of a second training sample in the embodiments of the present application. Among them, for general data, a general sample image is input. For example Figure 8B the owl image in, and it is required that the model outputs the same owl image. Specifically, there are two control conditions input in the present application. One is the preset segmented image with all pixel values being 0, and the other is the text description information of this owl image (i.e., the second text description information. Specifically, the text description information can be Chinese, English, etc. This is not specifically limited in this article. It should be noted that the second text description information in the present application does not contain expression words, such as Figure 8BIn the example, the text description information is “An Owl inflight” if it is expressed in English, and is “An Owl inflight” if it is expressed in Chinese. After the general sample image, the preset segmented image, and the second text description information are input into the image generation model to be trained, the model can add noise to the prediction of the general sample image in combination with the above two control conditions.
[0223] It should be noted that in the embodiment of the present application, the execution order between S72 and S73 is not specifically limited. S72 can be executed first and then S73, or S73 can be executed first and then S72, or S72 and S73 can be executed at the same time. This article does not make any specific limitations.
[0224] After obtaining the predicted added noises corresponding to the first training sample and the second training sample respectively through the above steps, a loss function can be constructed, and parameters of the image generation model can be adjusted based on the loss function.
[0225] S74: constructing a target loss function based on the difference between the predicted added noise corresponding to the first training sample and the second training sample and the corresponding actual added noise.
[0226] Specifically, after obtaining the predicted added noise corresponding to the first training sample and the second training sample respectively through the above steps, the predicted added noise of the first training sample can be compared with the actual added noise to determine the loss of expression data; similarly, the predicted added noise of the second training sample can also be compared with the actual added noise to determine the loss of common data.
[0227] Therefore, the target loss function in this application combines the loss of expression data and the loss of general data. Among them, through the above two control conditions, the model can learn the ability to decouple expression and character ID, ensuring that the trained model can obtain character ID information from the image contained in the control condition (i.e., the expression generation condition mentioned above), and obtain the character's expression information from the text description information contained in the control condition. In addition, through the loss of general data, the model can learn the ability to maintain a priori, so that the model can maintain the prior information obtained from pre-training and avoid the effect of overfitting.
[0228] Specifically, when calculating the target loss function, an optional implementation is to follow the following method: Figure 7 Step S74 is implemented by using sub-steps S741 to S743 shown in:
[0229] S741: Obtain the expression data difference between the predicted added noise corresponding to the first training sample and the corresponding actual added noise.
[0230] S742: Obtain the general data difference between the predicted added noise corresponding to the second training sample and the corresponding actual added noise.
[0231] S743: Perform weighted summation on the expression data difference and the general data difference to obtain the target loss function.
[0232] Specifically, in the embodiments of this application, the learning objective of the model is a prior-preserving target loss function. After adjusting the parameters of the model based on this loss function, the model can maintain the prior information obtained from pre-training and avoid overfitting. An optional calculation formula for this target loss function is as follows:
[0233]
[0234] where, w t and w t` are hyperparameters set manually. The former controls the weight of the loss corresponding to the first training sample (i.e., the weight of the expression data difference), and the latter controls the weight of the loss corresponding to the second training sample (i.e., the weight of the general data difference).
[0235] where, t and t` are time steps, and the value range can be 0 to 1000. That is, t in the above formula can be any value from 0 to 1000, indicating the number of times of adding noise (such as Gaussian noise) to the original facial expression sample image of a person. For example, if t = 100, then X t means the facial expression sample image of a person obtained after adding Gaussian noise 100 times to the original facial expression sample image of a person. Similarly, t` in the above formula can also be any value from 0 to 1000, indicating the number of times of adding noise (such as Gaussian noise) to the original general sample image. For example, if t` = 200, then X t`,pr means the general sample image obtained after adding Gaussian noise 200 times to the original general sample image. Specifically, the values of t and t` can be the same or different.
[0236] where, C and C pr are the control conditions corresponding to the facial expression sample image of a person and the general sample image respectively. Among them, the control condition refers to the two conditions of the control person ID and the control facial expression of a person listed above.
[0237] Therefore, in the above formula, represents the expression data difference, where, ∈ θ (X t , t, C) represents the predicted added noise determined by the model for the first training sample, and ∈ is the actual added noise corresponding to the first training sample; similarly, Indicates the general data difference, where ∈ θ (X t`,pr , t`, C pr ) indicates that the prediction determined by the model for the second training sample adds noise, ∈ pr is the actual added noise corresponding to the second training sample.
[0238] In the above embodiments, the role of the general data is to ensure that during the training process of the model, the powerful prior information learned by the model from the pre-training of the large-scale raw data is still retained, and to prevent the model from overfitting to a small amount of expression data.
[0239] It should be noted that in the embodiments of this application, the execution order between S741 and S742 is not specifically limited. S741 can be executed first and then S742, or S742 can be executed first and then S741, or S741 and S742 can be executed simultaneously. This is not specifically limited in this article.
[0240] S75: Adjust the parameters of the image generation model to be trained based on the target loss function.
[0241] During the model training process, the model parameters can be adjusted by minimizing the above target loss function. For example, methods such as gradient descent are used to minimize the above-listed target loss function to update the parameters of the image generation model to be trained. Specifically, it refers to the parameters related to each network layer, including but not limited to: bias, weight, trainable parameter matrix, convolution kernel size of the convolution layer, etc.
[0242] Among them, before the first iteration of training, the image generation model to be trained can be randomly initialized, or some image generation models that have been trained currently, etc. This is not specifically limited in this article.
[0243] It should be noted that when training the model with the above two types of data, this application designs a prior-preserving target loss function to achieve the effect of allowing the model to retain the prior information obtained from the pre-training and preventing overfitting. Moreover, this application also designs a decoupled learning method for expressions and person IDs based on cross-reference, enabling the model to learn to obtain person ID information from the input image and obtain person's expression information from the input text.
[0244] After training the image generation model based on the above method, the object only needs to provide at least one photo, and the base model can generate personalized expressions without retraining the sub-model for each object. The base model can make various expression changes, and can ensure that the person ID of the input picture is the same as the person ID of the output expression picture, and has good generation ability for several common human expressions.
[0245] In addition, the model retains the prior information learned during the original pre-training process, has strong editability, and avoids overfitting.
[0246] In summary, after the trained image generation model is obtained through the above method, the trained image generation model can be used to generate various personalized expressions of various characters.
[0247] Specifically, in order to further reduce the impact of environmental noise on the model, the present application proposes a method of averaging the feature vectors of multiple different images of the same person as the input of the model.
[0248] See also Figure 9 As shown, it is a logical schematic diagram of generating a target image using an image generation model in an embodiment of the present application.
[0249] There are some differences from the above-mentioned model training stage. In the model application stage, the input of the trained image generation model is no longer an image of a human expression or a general type, but a Gaussian noise that obeys a standard normal distribution; in the input control conditions, the image part can be one image or multiple images. In the embodiment of the present application, it is necessary to segment the input image to obtain a segmented image that ultimately contains only the human part, and then take the average of the vectors extracted from each segmented image as the final control condition. This method can greatly reduce the impact of environmental noise on the generation results.
[0250] Specifically, Figure 9 As shown, in addition to the Gaussian noise image to be processed, the input control conditions include multiple images of the same person and preset image description information. The image description information is similar to the first text description information listed above, and contains the description words corresponding to the specified expression. Specifically, the text description information can be in Chinese, English, etc., which is not specifically limited in this article. Figure 9 In the example, the image description information is expressed in English as "A boy with short hair, angry", and in Chinese as "a boy with short hair, angry". The specified expression is angry, and the corresponding English expression description word is "Anger", and the corresponding Chinese expression description word is "angry".
[0251] After the Gaussian noise image, multiple person images, and image description information are input into the trained image generation model, the model can combine the above two control conditions to predict the noise added to the Gaussian noise image, and perform denoising, and finally obtain a target image with a specified expression corresponding to the person (referring to the target person) in the multiple person images, such as Figure 9 In the example above, the model outputs an angry boy.
[0252] It should be noted that Figure 8A 、 Figure 8B or Figure 9 The description information in both English and Chinese forms is listed. In fact, only the description information in one language needs to be input. The above drawings are just simple examples. Besides the English and Chinese listed above, of course, it can also be in other language forms, which are not specifically limited in this article.
[0253] Specifically, the above-listed image generation model can be applied to the social product scenario. For example, integrating the above-trained image generation model into a social product with an expression generation function or a client specifically for generating expression images, and then deploying a complete inference pipeline. The object can generate the desired personalized expression based on this social product or client, and then use the generated personalized expression in scenarios such as chatting in the social product. The following combines Figure 10 to briefly describe this inference pipeline:
[0254] Refer to Figure 10 shown in the figure, which is a schematic diagram of an application scenario related to a social product in an embodiment of the present application.
[0255] First, for the model input part, the input image is processed. As Figure 10 shown in the figure, the object can upload one or more photos in a social product with an expression generation function or a client specifically for generating expression images, and the model (referring to the above-listed trained image generation model) will preprocess the photos uploaded by the object.
[0256] In the embodiment of the present application, for the input image, certain preprocessing is required. This preprocessing is mainly the detection of the input image. The input image needs to include a human face to be considered qualified.
[0257] Specifically, the image input by the object may not contain a person, and in this case, it will definitely not contain a human face; in addition, even if the image input by the object contains a person, it may not contain a human face. In either of the above two cases, in the embodiment of the present application, it belongs to an unqualified situation. Here, the image input by the object can be called a person image (it may not contain a person, but in principle, this image should contain a person to be possibly qualified).
[0258] Therefore, before performing image segmentation on at least one person image respectively, these person images can be pre-detected. In principle, a person image needs to include a human face to be considered qualified. An optional implementation method is to perform the following operations on each person image respectively:
[0259] First, perform image detection on the person image; if it is determined that the person image does not contain a face, feedback a prompt message indicating that the person image is unqualified to prompt the object to re-upload a qualified person image; if it is determined that the person image contains a face, it means that the person image is qualified, and then the person image can be segmented to remove the background part in the image.
[0260] As Figure 11 shown, it is a schematic diagram of a prompt message in an embodiment of the present application. Among them, Figure 11 it means that although the object uploads an image containing a person, the image shows the back of the person and cannot provide information related to the face of the person. Therefore, the person image is unqualified, and a prompt message as Figure 11 shown, "It is detected that the image you uploaded does not contain a face. Please re-upload a suitable image" can be feedback.
[0261] In the above implementation manner, through the pre-processing step, feedback reminder of unqualified input is given to the object, so that the object can immediately update the input image, thereby effectively reducing the entire expression generation process and improving the expression generation efficiency.
[0262] For each person image that passes the detection, the person image can be segmented to remove the background part in the image to obtain a segmented image containing only the person part; furthermore, a face vector is extracted for each segmented image; if the object uploads multiple person images, the face vectors extracted from the multiple person images also need to be averaged to be used as the final control condition.
[0263] Secondly, in the model inference part, the person image and the image description information are jointly input into the image generation model, and inference is performed in combination with the previously processed image input and 7 common human expression words. The image generation model will automatically generate one or a group of personalized expressions, that is, the target image of the specified expression corresponding to the target person in this article.
[0264] It should be noted that if there are 7 common expressions during model training, the 7 common expression words should also be used during model inference and application. If other expressions are also included during model training, other expression words may also be included during model inference and application, which is not specifically limited here. This article only takes the above-listed 7 common expressions as an example.
[0265] In addition, different from the above model training stage, in this stage, the model can output one or more target images of the specified expression corresponding to the target person, and it is not limited to only output one.
[0266] Refer to Figure 12 shown, it is a schematic diagram of a group of target images in an embodiment of the present application. Among them, Figure 12Listed are examples of target images corresponding to 6 common expressions, such as Figure 12 lists 6 types: anger, disgust, fear, happiness, sadness, and surprise. Each expression corresponds to two target images. In Figure 12 the upper and lower two images in the same column correspond to the same expression.
[0267] In the embodiments of the present application, for any object, the generated expressions can all achieve various personalized edits, and the editing ability of the model is very powerful.
[0268] Finally, post-process the output result of the model.
[0269] In the embodiments of the present application, if the model outputs multiple target images, post-processing can also be performed on these multiple target images. Generally speaking, this post-processing refers to screening the above-generated expression images (i.e., target images), selecting high-quality results and removing inferior results. Finally, save the generated expression data, and the object can use it in various social products.
[0270] An optional screening method is as follows:
[0271] If multiple target images are obtained, after obtaining the target images corresponding to the specified expression of the target person, for each target image, based on the preset reference expression corresponding to the expression description vocabulary, determine the expression accuracy of the specified expression included in the target image; and, based on the facial features, determine the facial similarity between the face included in the target image and the preset reference face corresponding to the facial features.
[0272] Furthermore, after obtaining the two parameters of the expression accuracy and facial similarity corresponding to each target image, the target images corresponding to the expression accuracy and facial similarity that meet the preset threshold conditions can be screened out from the multiple target images as the target expression data.
[0273] Specifically, the above screening process mainly relies on two parameters: expression accuracy and facial similarity, and combines these two parameters to screen out the expressions that meet the requirements and display them to the object.
[0274] Among them, there are many ways to evaluate the expression accuracy. For example, an expression classification model can be used to classify the expression of the target image to determine the expression accuracy.
[0275] Taking the expression classification model as the DeepFace model as an example, assuming that the specified expression vocabulary input is Anger, the target image generated by the above image generation model can be input into the DeepFace classification model, and based on the DeepFace classification model, judge the probability that the expression of the person in the target image is Anger, and this probability is the expression accuracy corresponding to the target image.
[0276] Specifically, it can be understood that based on the DeepFace classification model, the specified expression included in the target image is compared with a preset reference expression corresponding to the expression description vocabulary, and then a probability that the specified expression included in the target image conforms to the preset reference expression is determined, and this probability is the expression accuracy.
[0277] It should be noted that the above takes the DeepFace model as an example of the expression classification model. In addition, other expression classification models can also be used, or other evaluation methods can be used. Any evaluation method for determining the expression accuracy based on the preset reference expression corresponding to the expression description vocabulary is applicable to the embodiments of the present application, and no specific limitation is made here.
[0278] Among them, there are also many evaluation methods for face similarity. For example, the evaluation value can be determined by calculating the similarity between features.
[0279] An optional implementation manner is as follows: First, use the Multi-task Cascaded Convolutional Networks (MTCNN) to extract the face region of the generated target image, and then use the Residual Network (ResNet) to extract vectors from the extracted region to determine the face vector corresponding to the generated target image.
[0280] Among them, the MTCNN model uses a multi-cascaded structure to predict the face and the corresponding feature coordinate positions from coarse to fine, and can be applied to the detection of complex face scenes under various natural conditions, and can achieve face detection.
[0281] Finally, calculate the cosine similarity of the face vectors of the input person image and the generated target image. Assume that the face vector of the input person image is denoted as A, and the face vector of the generated target image is denoted as B. Assume that A and B are two n-dimensional vectors, A is [A1, A2,..., An], and B is [B1, B2,..., Bn]. Then the above face similarity can be expressed as the cosine similarity between vector A and B. Specifically, the calculation formula of the cosine similarity is as follows:
[0282]
[0283] Since in the embodiments of the present application, the object can input one or more person images. Among them, when the object only inputs one person image, vector A can be regarded as determined by using the MTCNN model to extract the face region of the person image and then using the ResNet model to extract vectors from the extracted region; it can also be regarded as determined by extracting face features from the segmented image of the person image.
[0284] Similarly, when multiple portrait images are input by the object, vector A can be regarded as the average value of the face vectors determined after using the MTCNN model to extract the face regions for each portrait image and then using the ResNet model to extract vectors from the extracted regions; it can also be regarded as the average value of the face vectors determined after extracting face vectors from the segmented images of the portrait images.
[0285] It should be noted that the above are just simple examples of the MTCNN, ResNet models, etc. Other feature extraction models can also be used. In addition, in addition to the cosine similarity calculation method, other similarity calculation methods can also be used, such as calculating the Euclidean distance, Minkowski distance, etc. Any evaluation method for determining face similarity based on face features is applicable to the embodiments of the present application and will not be specifically limited here.
[0286] In the embodiments of the present application, the preset threshold condition means comparing the above-mentioned expression accuracy and face similarity with the preset threshold, and filtering the images according to the comparison results. Generally, the target images that meet the preset threshold conditions can be retained, while the target images that do not meet the preset threshold conditions can be filtered out.
[0287] Specifically, when filtering the target images based on the preset threshold condition, an optional implementation method is as follows:
[0288] From multiple target images, filter out the target images whose corresponding expression accuracy is greater than the preset accuracy threshold and whose corresponding face similarity is greater than the preset similarity threshold as the target expression data.
[0289] In the embodiments of the present application, the values of the preset similarity threshold and the preset accuracy threshold can be the same or different.
[0290] Taking both of these thresholds as 0.5 as an example, it means filtering out the target images whose corresponding expression accuracy and face similarity are both greater than 0.5 from multiple target images as the target expression data. That is:
[0291] If the expression accuracy and face similarity corresponding to a target image are both greater than 0.5, it means that this generation result is considered qualified and can be retained. If at least one of the expression accuracy and face similarity corresponding to a target image is not greater than 0.5, it means that this generation result is considered unqualified and can be filtered out.
[0292] In the above implementation method, by combining the expression accuracy and face similarity to filter the target images, the quality of the finally generated target expression data can be ensured, and the finally generated expression data can be made more in line with the actual expectations of the object, thereby effectively improving the usage stickiness of the object.
[0293] Refer to Figure 13 As shown, it is a schematic diagram of the interaction logic between a terminal device and a server in an embodiment of the present application. Assume that a social product with an expression generation function or a client specifically for generating expression images is installed on the terminal device. An object can generate a desired personalized expression based on the social product or the client, and then use the generated personalized expression in scenarios such as chatting in the social product.
[0294] As Figure 13 shown, an object can upload one or more images for controlling a character ID and input image description information corresponding to the image that includes a specified expression. For example, Figure 13 as shown, an object uploads a portrait image and inputs the image description information "a boy with short hair, laughing", where the specified expression is "laughing". A trained image generation model is deployed on the server side. After the server obtains the above control conditions from the terminal device side, it can denoise a Gaussian noise image that follows a standard normal distribution based on this to generate a target image of the laughing expression corresponding to the target character. As Figure 13 shown, after the server generates the target image, it can feedback it to the terminal device, and the terminal device presents it to the object.
[0295] In addition, pre-processing and post-processing steps may also be included in the above process. For example, in the pre-processing step, the server detects the image uploaded by the object. If it is determined to be unqualified, relevant prompts can be feedback to the terminal device, and the terminal device presents them to the object, as shown in the above Figure 11 figure; Another example is the post-processing step. The server can also generate target images of the laughing expression corresponding to multiple target characters, and then calculate the expression accuracy and face similarity of each of the multiple target images through the server. Based on these two parameters, the target images are screened, and finally the screened target expression data is feedback to the terminal device, and the terminal device presents it to the object; and so on.
[0296] In summary, for any object, the present application can generate a complete set of expressions without training, and the generated expressions can achieve various personalized edits. In addition, the object can generate its own personalized expressions and can use them in any social product at any time.
[0297] It should be noted that the above-listed social product scenarios are only simple examples. In addition, of course, it can also be applied to other scenarios, which are not specifically limited in this article.
[0298] Based on the same inventive concept, an embodiment of the present application also provides an expression generation device. As Figure 14 shown, it is a schematic diagram of the structure of the expression generation device 1400, which may include:
[0299] An acquisition unit 1401, configured to acquire a noisy image to be processed and an expression generation condition preset for the noisy image, where the expression generation condition includes: at least one person image of the same target person and image description information including expression description words containing a specified expression;
[0300] A segmentation unit 1402, configured to perform image segmentation on at least one person image respectively to obtain corresponding segmented images including the target person;
[0301] An extraction unit 1403, configured to extract the face features of the target person based on the obtained at least one segmented image;
[0302] A generation unit 1404, configured to perform denoising processing on the noisy image based on the face features and the image description information to obtain a target image corresponding to the target person and conforming to the specified expression.
[0303] Optionally, the noisy image is a Gaussian noise image subject to a standard normal distribution.
[0304] Optionally, the extraction unit 1403 is specifically configured to:
[0305] Extract the face vectors of the target person included in each of the at least one segmented image respectively;
[0306] Average the corresponding elements in the extracted face vectors to obtain the face features of the target person.
[0307] Optionally, if multiple target images are obtained, the generation unit 1404 is further configured to, after obtaining the target image corresponding to the target person and conforming to the specified expression, for each target image, perform the following operations respectively:
[0308] Determine the expression accuracy of the specified expression included in a target image based on a preset reference expression corresponding to the expression description words;
[0309] Determine the face similarity between the face included in a target image and a preset reference face corresponding to the face features based on the face features;
[0310] Screen out the target images whose corresponding expression accuracy and face similarity meet the preset threshold conditions from the multiple target images as target expression data.
[0311] Optionally, the generation unit 1404 is specifically configured to:
[0312] Screen out the target images whose corresponding expression accuracy is greater than a preset accuracy threshold and whose corresponding face similarity is greater than a preset similarity threshold from the multiple target images as target expression data.
[0313] Optionally, the device is implemented using a trained image generation model, and the trained image generation model is obtained by performing iterative training on the image generation model to be trained based on a first training sample set and a second training sample set;
[0314] Wherein, each first training sample in the first training sample set includes a facial expression sample image of a person and the corresponding actually added noise, and each second training sample in the second training sample set includes a general sample image and the corresponding actually added noise; the facial expression sample images in the first training sample set are images of different sample persons with various expressions; the general sample images in the second training sample set are images of different contents of various types.
[0315] Optionally, the device further includes a model training unit 1405, and the model training unit 1405 is configured to perform the following processes during each iterative training:
[0316] Select a first training sample from the first training sample set and a second training sample from the second training sample set;
[0317] Input the first training sample into the image generation model to be trained to obtain the predicted added noise corresponding to the first training sample; and input the second training sample into the image generation model to be trained to obtain the predicted added noise corresponding to the second training sample;
[0318] Construct an objective loss function based on the differences between the predicted added noise corresponding to the first training sample and the second training sample and the corresponding actually added noise;
[0319] Adjust the parameters of the image generation model to be trained based on the objective loss function.
[0320] Optionally, the model training unit 1405 is specifically configured to:
[0321] Extract the first text description information corresponding to the facial expression sample image in the first training sample, and add the facial expression vocabulary corresponding to the facial expression sample image to the first text description information;
[0322] Perform image segmentation on the associated image of the facial expression sample image of the person to obtain a sample segmentation image including the person part corresponding to the associated image; the associated image and the facial expression sample image of the person are images of the same person with different expressions;
[0323] Input the facial expression sample image of the person, the first text description information, and the sample segmentation image into the image generation model to be trained;
[0324] Use the image generation model to be trained, in combination with the first text description information and the sample segmentation image, to perform denoising processing on the facial expression sample image of the person to obtain the predicted added noise corresponding to the first training sample.
[0325] Optionally, the model training unit 1405 is specifically configured to:
[0326] Extract the second text description information corresponding to the general sample image in the second training sample;
[0327] Use the preset segmented image as the segmented image corresponding to the general sample image, and each pixel value in the preset segmented image is the same preset value;
[0328] Input the general sample image, the second text description information, and the preset segmented image into the image generation model to be trained;
[0329] Use the image generation model to be trained, combine the second text description information and the preset segmented image, and perform denoising processing on the general sample image to obtain the predicted added noise corresponding to the second training sample.
[0330] Optionally, the model training unit 1405 is specifically configured to:
[0331] Obtain the expression data difference between the predicted added noise corresponding to the first training sample and the corresponding actual added noise; and
[0332] Obtain the general data difference between the predicted added noise corresponding to the second training sample and the corresponding actual added noise;
[0333] Perform weighted summation on the expression data difference and the general data difference to obtain the target loss function.
[0334] Optionally, the obtaining unit 1401 is further configured to, before performing image segmentation on at least one person image respectively, for each person image, perform the following operations respectively:
[0335] Perform image detection on a person image;
[0336] If it is determined that a person image does not contain a face, feedback a prompt message indicating that the person image is unqualified.
[0337] Since the expression generation method in this application mainly performs denoising processing on the noise image to be processed based on the preset expression generation conditions, and the expression generation conditions include two major parts. The first part is the person image, which is used to provide person ID information, and the second part is the image description information, which contains the expression description vocabulary of the specified expression and is used to provide expression information. Therefore, for different expressions of different people, only different expression generation conditions need to be set, and the noise image to be processed can be kept consistent. Therefore, for any person, a complete set of expressions can be generated without separately training a person Lora sub-model.
[0338] In addition, in the present application, when generating a target image of a specified expression corresponding to a target person, the person ID information finally provided is based on the face features extracted from the segmented image, and the face features can be extracted from multiple different images of the same person, which can effectively reduce the impact of environmental noise on the generation result.
[0339] In summary, the present application can generate personalized expressions more efficiently and accurately.
[0340] For the convenience of description, the above parts are divided into various modules (or units) according to their functions and described separately. Of course, when implementing the present application, the functions of the various modules (or units) can be implemented in the same or multiple software or hardware.
[0341] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the functions of the module or unit.
[0342] After introducing the expression generation method and device of the exemplary embodiment of the present application, next, an electronic device according to another exemplary embodiment of the present application will be introduced.
[0343] Those skilled in the art can understand that various aspects of the present application can be implemented as a system, method, or program product. Therefore, various aspects of the present application can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.
[0344] Based on the same inventive concept as the above method embodiment, an electronic device is also provided in the embodiments of the present application. In one embodiment, the electronic device can be a server, such as Figure 1 the server 120 shown. In this embodiment, the structure of the electronic device can be as Figure 15 shown, including a memory 1501, a communication module 1503, and one or more processors 1502.
[0345] The memory 1501 is used to store the computer program executed by the processor 1502. The memory 1501 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system and programs required to run the instant messaging function, etc.; the data storage area can store various instant messaging information and operation instruction sets, etc.
[0346] The memory 1501 can be a volatile memory, such as a random-access memory (RAM); the memory 1501 can also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or the memory 1501 is any other medium that can be used to carry or store a desired computer program in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1501 can be a combination of the above memories.
[0347] The processor 1502 can include one or more central processing units (CPUs) or be a digital processing unit, etc. The processor 1502 is used to implement the above-mentioned expression generation method when calling the computer program stored in the memory 1501.
[0348] The communication module 1503 is used to communicate with the terminal device and other servers.
[0349] In the embodiments of the present application, the specific connection medium between the above-mentioned memory 1501, communication module 1503, and processor 1502 is not limited. In the embodiments of the present application Figure 15 it is described that the memory 1501 and the processor 1502 are connected through a bus 1504, and the bus 1504 is described by a thick line in Figure 15 The connection manners between other components are only for illustrative purposes and are not to be construed as limiting. The bus 1504 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of description, Figure 15 it is only described by a thick line in
[0350] The memory 1501 stores a computer storage medium, and the computer storage medium stores computer-executable instructions for implementing the expression generation method of the embodiments of the present application. The processor 1502 is used to execute the above-mentioned expression generation method, as Figure 2 shown.
[0351] In another embodiment, the electronic device can also be other electronic devices, such as Figure 1 the terminal device 110 shown. In this embodiment, the structure of the electronic device can be as Figure 16As shown, it includes components such as a communication component 1610, a memory 1620, a display unit 1630, a camera 1640, a sensor 1650, an audio circuit 1660, a Bluetooth module 1670, and a processor 1680.
[0352] The communication component 1610 is used to communicate with a server. In some embodiments, it may include a Wireless Fidelity (WiFi) module. The WiFi module belongs to short - range wireless transmission technology, and through the WiFi module, the electronic device can help users send and receive information.
[0353] The memory 1620 can be used to store software programs and data. The processor 1680 executes various functions and data processing of the terminal device 110 by running the software programs or data stored in the memory 1620. The memory 1620 may include high - speed random - access memory, and may also include non - volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non - volatile solid - state storage devices. The memory 1620 stores an operating system that enables the terminal device 110 to run. In this application, the memory 1620 can store the operating system and various application programs, and can also store a computer program for executing the expression generation method of the embodiments of this application.
[0354] The display unit 1630 can also be used to display information input by the user or information provided to the user, as well as the graphical user interface (GUI) of various menus of the terminal device 110. Specifically, the display unit 1630 may include a display screen 1632 disposed on the front of the terminal device 110. Among them, the display screen 1632 can be configured in the form of a liquid crystal display, a light - emitting diode, etc. The display unit 1630 can be used to display the operation interface of the client in the embodiments of this application.
[0355] The display unit 1630 can also be used to receive input digital or character information, and generate signal inputs related to the user settings and function controls of the terminal device 110. Specifically, the display unit 1630 may include a touch screen 1631 disposed on the front of the terminal device 110, which can collect touch operations of the user on or near it, such as clicking buttons, dragging scroll boxes, etc.
[0356] Among them, the touch screen 1631 can cover the display screen 1632, or the touch screen 1631 and the display screen 1632 can be integrated to implement the input and output functions of the terminal device 110. After integration, it can be simply called a touch display screen. In this application, the display unit 1630 can display application programs and corresponding operation steps.
[0357] The camera 1640 can be used to capture static images, and users can publish the images captured by the camera 1640 through an application. There can be one or more cameras 1640. An object generates an optical image through a lens and projects it onto a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the processor 1680 to be converted into a digital image signal.
[0358] The terminal device may further include at least one sensor 1650, such as an acceleration sensor 1651, a distance sensor 1652, a fingerprint sensor 1653, and a temperature sensor 1654. The terminal device may also be configured with other sensors such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, a light sensor, and a motion sensor.
[0359] The audio circuit 1660, the speaker 1661, and the microphone 1662 can provide an audio interface between the user and the terminal device 110. The audio circuit 1660 can transmit the electrical signal converted from the received audio data to the speaker 1661, and the speaker 1661 converts it into a sound signal for output. The terminal device 110 may also be configured with volume buttons for adjusting the volume of the sound signal. On the other hand, the microphone 1662 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1660 and converted into audio data, and then the audio data is output to the communication component 1610 to be sent to, for example, another terminal device 110, or the audio data is output to the memory 1620 for further processing.
[0360] The Bluetooth module 1670 is used to interact with other Bluetooth devices having Bluetooth modules through the Bluetooth protocol. For example, the terminal device can establish a Bluetooth connection with a wearable electronic device (such as a smart watch) that also has a Bluetooth module through the Bluetooth module 1670 to perform data interaction.
[0361] The processor 1680 is the control center of the terminal device, connecting various parts of the entire terminal through various interfaces and circuits. By running or executing software programs stored in the memory 1620 and calling data stored in the memory 1620, it performs various functions of the terminal device and processes data. In some embodiments, the processor 1680 may include one or more processing units; the processor 1680 may also integrate an application processor and a baseband processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the baseband processor mainly processes wireless communication. It can be understood that the above baseband processor may not be integrated into the processor 1680. In this application, the processor 1680 can run the operating system, application programs, user interface display, and touch response, as well as the expression generation method of the embodiments of this application. In addition, the processor 1680 is coupled to the display unit 1630.
[0362] In some possible implementation manners, various aspects of the expression generation method provided in this application can also be implemented in the form of a program product, which includes a computer program. When the program product runs on an electronic device, the computer program is used to cause the electronic device to execute the steps in the expression generation method according to various exemplary embodiments of this application described above in this specification. For example, the electronic device can execute the steps as shown in Figure 2 shown.
[0363] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0364] The program product of the embodiments of this application can adopt a portable compact disc read-only memory (CD-ROM) and include a computer program, and can run on an electronic device. However, the program product of this application is not limited to this. In this document, the readable storage medium can be any tangible medium that contains or stores a program that can be used by or in combination with a command execution system, apparatus, or device.
[0365] A readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium can also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with a command execution system, apparatus, or device.
[0366] The computer program contained on a readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.
[0367] The computer program for performing the operations of the present application can be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The computer program can be executed entirely on the user's electronic device, partially on the user's electronic device, executed as a stand-alone software package, partially on the user's electronic device and partially on a remote electronic device, or entirely on a remote electronic device or server. In the case of a remote electronic device, the remote electronic device can be connected to the user's electronic device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external electronic device (e.g., by connecting through the Internet using an Internet service provider).
[0368] It should be noted that although several units or subunits of the apparatus are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more of the above-mentioned units can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.
[0369] In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the operations shown must be performed to achieve the desired result. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution.
[0370] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable computer programs.
[0371] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0372] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that realizes the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0373] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0374] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present application.
[0375] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.
Claims
1. A method for generating an expression, characterized in that, The method includes: obtaining a noisy image to be processed and an expression generation condition preset for the noisy image, where the expression generation condition includes: at least one person image of the same target person, and image description information including expression description words containing a specified expression; performing image segmentation on the at least one person image respectively to obtain corresponding segmented images including the target person; extracting a face feature of the target person based on the obtained at least one segmented image; performing denoising processing on the noisy image based on the face feature and the image description information to obtain a target image corresponding to the target person and conforming to the specified expression.
2. The method according to claim 1, wherein The noisy image is a Gaussian noise image subject to a standard normal distribution.
3. The method according to claim 1, wherein The extracting the face feature of the target person based on the obtained at least one segmented image includes: extracting face vectors of the target person included in the at least one segmented image respectively; averaging corresponding elements in the extracted face vectors to obtain the face feature of the target person.
4. The method according to claim 1, characterized in that, If multiple target images are obtained, after obtaining the target image corresponding to the target person and conforming to the specified expression, the method further includes: for each target image, performing the following operations respectively: determining an expression accuracy of the specified expression included in one target image based on a preset reference expression corresponding to the expression description words; determining a face similarity between the face included in one target image and a preset reference face corresponding to the face feature based on the face feature; screening out target images corresponding to the expression accuracy and the face similarity that meet preset threshold conditions from the multiple target images as target expression data.
5. The method according to claim 4, wherein The screening out target images corresponding to the expression accuracy and the face similarity that meet preset threshold conditions from the multiple target images as target expression data includes: screening out target images corresponding to the expression accuracy greater than a preset accuracy threshold and the face similarity greater than a preset similarity threshold from the multiple target images as target expression data.
6. The method according to any one of claims 1 to 5, characterized in that, The method is implemented by using a trained image generation model, and the trained image generation model is obtained by performing cyclic iterative training on an image generation model to be trained based on a first training sample set and a second training sample set; wherein each first training sample in the first training sample set includes a person expression sample image and corresponding actually added noise, and each second training sample in the second training sample set includes a general sample image and corresponding actually added noise; the person expression sample images in the first training sample set are images of different sample persons with various expressions; the general sample images in the second training sample set are images with different contents of various types.
7. The method according to claim 6, wherein Each iterative training performs the following process: selecting a first training sample from the first training sample set and a second training sample from the second training sample set; inputting the first training sample into the image generation model to be trained to obtain predicted added noise corresponding to the first training sample; Further, input the second training sample into the image generation model to be trained, and obtain the predicted added noise corresponding to the second training sample; Based on the differences between the predicted added noises corresponding to the first training sample and the second training sample and the corresponding actual added noises, construct an objective loss function; Based on the objective loss function, adjust the parameters of the image generation model to be trained.
8. The method according to claim 7, wherein The step of inputting the first training sample into the image generation model to be trained and obtaining the predicted added noise corresponding to the first training sample includes: Extract the first text description information corresponding to the character expression sample image in the first training sample, and add the expression vocabulary corresponding to the character expression sample image to the first text description information; Perform image segmentation on the associated image of the character expression sample image to obtain a sample segmentation image corresponding to the associated image and including the character part; the associated image and the character expression sample image are images of the same person with different expressions; Input the character expression sample image, the first text description information, and the sample segmentation image into the image generation model to be trained; Use the image generation model to be trained, combine the first text description information and the sample segmentation image, and perform denoising processing on the character expression sample image to obtain the predicted added noise corresponding to the first training sample.
9. The method according to claim 7, wherein The step of inputting the second training sample into the image generation model to be trained and obtaining the predicted added noise corresponding to the second training sample includes: Extract the second text description information corresponding to the general sample image in the second training sample; Use the preset segmentation image as the segmentation image corresponding to the general sample image, and each pixel value in the preset segmentation image is the same preset value; Input the general sample image, the second text description information, and the preset segmentation image into the image generation model to be trained; Use the image generation model to be trained, combine the second text description information and the preset segmentation image, and perform denoising processing on the general sample image to obtain the predicted added noise corresponding to the second training sample.
10. The method according to claim 7, characterized in that, The step of constructing an objective loss function based on the differences between the predicted added noises corresponding to the first training sample and the second training sample and the corresponding actual added noises includes: Obtain the expression data difference between the predicted added noise corresponding to the first training sample and the corresponding actual added noise; and Obtain the general data difference between the predicted added noise corresponding to the second training sample and the corresponding actual added noise; Perform weighted summation on the expression data difference and the general data difference to obtain the objective loss function.
11. The method according to any one of claims 1 to 5, characterized in that, Before performing image segmentation on each of the at least one person image, the method further includes: For each person image, perform the following operations respectively: Perform image detection on a person image; If it is determined that the person image does not include a face, feedback a prompt message indicating that the person image is unqualified.
12. An expression generation device, characterized in that, including: An acquisition unit, configured to acquire a noisy image to be processed and an expression generation condition preset for the noisy image, where the expression generation condition includes: at least one person image of the same target person and image description information including expression description words containing a specified expression; A segmentation unit, configured to perform image segmentation on the at least one person image respectively to obtain corresponding segmentation images including the target person; An extraction unit, configured to extract the face features of the target person based on the obtained at least one segmentation image; A generation unit, configured to perform denoising processing on the noisy image based on the face features and the image description information to obtain a target image corresponding to the target person and conforming to the specified expression.
13. An electronic device, characterized in that, It includes a processor and a memory. Among them, the memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, It includes a computer program. When the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute the steps of any one of claims 1 to 11.
15. A computer program product, characterized in that, It includes a computer program, and the computer program is stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the electronic device executes the steps of any one of claims 1 to 11.