Image processing method and related equipment

By fusing features from the first image and the 3D face model image, and using an encoder and a generative model to generate images that meet user needs, the problem of uncontrollable details such as facial expressions in existing technologies is solved, and high-quality image generation is achieved.

CN120976030APending Publication Date: 2025-11-18HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410623164.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-17
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing technologies cannot effectively control details such as facial expressions in images, thus failing to meet users' customized needs for facial details.

Method used

By acquiring a first image and a 3D face model image with target attributes, feature fusion is performed using an encoder and a generative model to generate an image that meets the user's needs, including control over facial expressions and lighting.

Benefits of technology

It achieves precise control over facial details, improves the quality and realism of generated images, and reduces the risk of image errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976030A_ABST
    Figure CN120976030A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method and related equipment, relates to the field of artificial intelligence, and is used for controlling details of a human face in an image when a portrait customization task is executed so as to meet the requirement of a user for customization of the details of the human face. In the method, a first device obtains a first image and a second image, the first image comprises a first face, the second image is an image of a 3D face model with target attributes, and the target attributes comprise shadow, face expression or face orientation; the first device obtains an encoding result through an encoder according to the first image and the second image, and the encoding result is fused with features of the first image and the second image; the first device obtains a third image through a generation model according to the coding result, the third image comprises a second face with the target attribute, and the second face and the first face correspond to the same person.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the field of artificial intelligence, and in particular, to an image processing method and related equipment. BACKGROUND

[0002] At present, portrait customization tasks need to be performed in more and more scenarios, for example, in a passport photo customization software, the hair and clothes of a person in an image input by a user can be customized. The portrait customization task refers to analyzing and processing a portrait by computer vision and artificial intelligence technology, and realizing personalized customization of the portrait. According to different user needs, the portrait customization task involves customization of clothes of a person, body posture of the person, orientation of a face, and position of a person in a group photo.

[0003] In order to realize portrait customization, a text-to-image model in which a control condition such as a skeleton point is additionally input can be used. After obtaining a portrait image that needs to be customized, the running of the text-to-image model is controlled by the control condition of the skeleton point. Specifically, a skeleton point constraint can be additionally input as a control condition of the text-to-image model, and the text-to-image model is controlled to generate a portrait image that meets the skeleton point constraint, so as to finally control the posture of a portrait in the generated portrait image. However, details such as expressions of a person cannot be accurately controlled by the skeleton point. It can be seen that the scheme cannot be applied to a scenario in which a user requests to customize details such as expressions of a face in an image.

[0004] Therefore, when performing a portrait customization task, how to control details of a face in an image is a technical problem to be solved. SUMMARY

[0005] The present application provides an image processing method and related equipment for controlling details of a face in an image when performing a portrait customization task, so as to meet the needs of a user for customizing details of a face.

[0006] In a first aspect, the present application provides an image processing method, which is executed by a first device, or executed by a part of components (such as a processor, a chip or a chip system, etc.) in the first device, or can also be implemented by a logic module or software which can realize all or part of the functions of the first device. In the first aspect and its possible implementation manners, the image processing method is taken as an example to be executed by the first device. The first device acquires a first image and a second image. The first image contains a first face, and the second image is an image of a 3D face model with a target attribute, which includes light and shadow, expression of the face, or orientation of the face. The first device obtains an encoding result through an encoder according to the first image and the second image. The encoding result fuses the features of the first image and the second image. The first device obtains a third image through a generation model according to the encoding result. The third image contains a second face with the target attribute. The second face corresponds to the same person as the first face.

[0007] In the above first aspect, the first device realizes the fusion of the features of the first image and the second image, and further processes the encoding result obtained by the encoder through the generation model, so as to improve the quality of the generated image. Finally, the third image obtained by the present application fuses the features of the first image and the second image. Specifically, by fusing the features of the first image, the third image can correspond to the same person as the face in the first image. By fusing the features of the second image, the expression, light and shadow, etc. of the face in the third image can be controlled. It can be seen that the present application is based on the face in the first image, and controls the expression, light and shadow, etc. of the face through the second image to generate the third image. Therefore, the present application can control the details of the face, such as expression, light and shadow, etc. For example, when a user needs to customize the expression and light and shadow, etc. of a face in an image, the image can be taken as the first image, the expression and light and shadow, etc. of the face specified by the user can be taken as the target attribute, and an image of a 3D face model with the target attribute can be acquired as the second image. Then, the face in the image is controlled to have the target attribute through the present application, so as to meet the needs of the user to customize the details of the face.

[0008] In addition, compared with fusing the first image and the second image through a mapping rendering, etc., the present application fuses the features of the first image and the second image through the encoder. For example, the encoder in the present application can be an encoder in a neural network. The features of the image can be better fused through the neural network technology. The quality of the generated image can be further improved through the processing of the encoding result of the encoder by the generation model. Through the cooperation of the encoder and the generation model, the delicacy of the image can be improved, and the risk of image errors (Bug) can be reduced.

[0009] Optionally, the encoder structure includes convolutional layers and transformer layers. The first convolutional layer in the convolutional layer generates a feature map based on the second image. This feature map is then input to the transformer layer, while the features of the first image are input to the transformer layer via a cross-attention mechanism. The transformer layer processes the feature map and the features of the first image to obtain the encoded result.

[0010] Optionally, the second image in this solution is a 3D face model image without textures. Optionally, the first device can create a "texture-free" 3D face model using a 3D Morphable Face Model (3DMM).

[0011] Optionally, the first device can generate a corresponding original 3D face model image based on the face in the first image, and then adjust the parameters of the original 3D face model image according to user needs or system presets. For example, if the user instructs that the expression in the first image be adjusted to a smile, the parameters of the original 3D face model image can be adjusted so that the expression in the generated second image is a smile. For example, some image software has a system preset photo template with specific lighting and shadows on the face in the template, so the parameters of the original 3D face model image can be adjusted so that the lighting and shadows in the generated second image match the lighting and shadows of the template.

[0012] In one possible implementation of the first aspect, the first device acquires a first result, which is obtained based on the second image and the noisy image, the first result being fused with features of the second image and the noisy image; the first device obtains an encoding result by the encoder based on the first image and the first result, the encoding result being fused with features of the first result and the first image.

[0013] In this implementation, the first result incorporates features from both the second image and the noisy image. By fusing the features of the first image with those of the noisy image, the second image can be enhanced. When error information appears in the second image, this feature fusion can cover up the error information in the second image.

[0014] Optionally, the first device can fuse the second image and the noisy image using a fusion module. Optionally, the fusion module can first align the sizes of the second image and the noisy image in the spatial dimension, and then perform channel-dimensional connection.

[0015] The above implementation can solve the problem of different resolutions between the 3D face model image and the noisy image. Generally speaking, the resolution of the 3D face model image is 8 times that of the noisy image. This solution ensures that the number of elements before and after fusion does not change through spatial alignment and channel dimension connection, as well as the connection of pixels in the same spatial position in the channel dimension, thereby ensuring lossless information and realizing lossless fusion of the 3D face model image and the noisy image.

[0016] In one possible implementation of the first aspect, the noisy image includes: a completely noisy image or a fourth image, which is obtained by adding noise to a fifth image, which is an image obtained in the adjacent previous round through the generative model.

[0017] In this implementation, noise is added to the fifth image obtained from the adjacent previous generation model to obtain the fourth image. Then, based on the fourth image and the second image, a first result is obtained through a fusion module. That is, the output result from the previous round is used in this round of image processing, and this result is fused with the second image. It is evident that the above implementation ensures that subsequent iterations can better utilize the generation results of preceding iterations, improving the quality of the generated image and increasing the similarity between the generated image and the face IDs in the obtained original image. For example, it improves the similarity between the face IDs in the third image and the first image.

[0018] Optionally, in the first round, the first device can obtain a fusion result of the fully noisy image and the second image through a fusion module based on the fully noisy image and the second image. Then, based on the second image and the noisy image of the previous round's output result, i.e., the fourth image, the fusion module obtains a fusion result of the fourth image and the second image. Furthermore, during the iteration process, the degree of noise addition decreases with each iteration when adding noise to the output result of the previous round's generation model.

[0019] Optionally, the first device can first extract the descriptive token of the first image using a retrained pre-trained model, such as a retrained Contrastive Language-Image Pre-Training (CLIP) model, and then input the descriptive token into the encoder through Cross-Attention.

[0020] Optionally, the process of processing the output of the fusion module and the features of the first image through the encoder can be executed multiple times. Each execution will generate the features of the first face's identity document (ID). Multiple executions will yield features of multiple IDs. The features of multiple IDs can be summed to obtain the total features of the first face.

[0021] In one possible implementation of the first aspect, the first image includes: a face image and a face warp image, the face image being obtained by extracting a face from a sixth image, the sixth image being an acquired image containing the first face, and the face warp image being obtained by distorting the sixth image, the face warp image having the target attribute.

[0022] In this implementation, the first device acquires a sixth image containing a face, such as a source image input by the user. The sixth image is then preprocessed in different ways to obtain a face image and a face warp image. The face warp image is obtained by distorting the sixth image, giving it the target attributes. Therefore, the face warp image possesses features not found in the face image, which are adapted to the target attributes and can be better integrated with the second image. For example, when the target attribute is a face facing 30 degrees, the face warp image will have the features of a face facing 30 degrees. The face image and face warp image are obtained by processing the sixth image. However, the image distortion process may result in poor realism and lack of refinement, and texture bugs may appear in occluded areas. The face image, on the other hand, is obtained by directly extracting the face from the sixth image, thus retaining the features of the first face in the image and being less prone to bugs. Therefore, it can serve as a supplement to the warp image. In this scheme, both the face image and the face warp are fused with the second image through an encoder. The features of the face image and the face warp image complement each other, providing richer and more diverse face-related features, reducing the risk of bugs, and thus making the final fused image more detailed, more realistic, and of higher quality.

[0023] In one possible implementation of the first aspect, the encoding result includes: a first encoding result and a second encoding result, wherein the first device obtains the first encoding result by the encoder based on the face image and the second image, and the first encoding result incorporates features of the face image and the second image; the first device obtains the second encoding result by the encoder based on the face warp image and the second image, and the second encoding result incorporates features of the face image and the second image.

[0024] In this embodiment, the processing of the face image and the face warp image by the first device is performed separately. Specifically, the first device obtains a first encoding result using an encoder based on the face image and the second image, and obtains a second encoding result using an encoder based on the face warp image and the second image. The encoder's processing of the face image and the encoder's processing of the face warp image are performed independently. Figure 3As shown, the face image and the face warp are input into the encoder separately. The reason for processing them independently is that fusing the face image and the second image through the encoder results in a final encoded image with high similarity to the source image's ID. This also avoids issues such as poor realism in the face warp image affecting the final generated image. Furthermore, fusing the second image and the face warp image through the encoder allows for the extraction of edited features from the first face's ID, enabling the modification of features such as the face's orientation or expression.

[0025] In one possible implementation of the first aspect, the first device obtains a first output result through the generative model based on the first encoding result; the first device obtains a second output result through the generative model based on the second encoding result; the first device sums the first output result and the second output result to obtain the third image.

[0026] In this implementation, the generative model processes the encoded results obtained from the face image and the encoded results obtained from the face warp image, respectively. Optionally, the generative model includes: a generative adversarial model, a variational autoencoder, a streaming model, or a diffusion model. Among them, obtaining the third image through a diffusion model is a typical implementation of this scheme.

[0027] Optionally, the diffusion model in this scheme can be the Stable Diffusion model Unet, that is, the process of obtaining the third image based on the encoding result can be performed in the latent space of the Variational Auto-Encoder (VAE), and finally the third image needs to be obtained through the VAE decoder.

[0028] Optionally, the feature map resolution in the diffusion model will first decrease and then increase. As mentioned earlier, the ID injection model outputs feature maps at multiple resolutions. Before each resolution reduction, the feature map in the diffusion model is added to the corresponding resolution feature map generated by the ID injection model, and then subsequent processing is performed.

[0029] Optionally, the first device obtains the third image through the generation model based on the encoding result and the fourth image, and the third image incorporates features from the encoding result and the fourth image.

[0030] Alternatively, the first device can also input the prompt into the diffusion model via Cross-Attention.

[0031] In one possible implementation of the first aspect, the first face includes: a third face and a fourth face; the first image includes: a seventh image containing the third face and an eighth image containing the fourth face; the second image includes: a ninth image for controlling the attributes of the seventh image and a tenth image for controlling the attributes of the eighth image; the encoding result includes: a third encoding result and a fourth encoding result; and the first device...

[0032] In the above implementation, when a user inputs multiple face images, or obtains multiple face images from the cloud, and receives an instruction to generate a group photo of these multiple faces, this solution will combine a single face with its corresponding 3D face model image to obtain an encoding result. Thus, the face in the final image corresponds to the 3D face model. The 3D face model image is used to control the attributes of the face, i.e., control conditions. Therefore, this solution can ensure that the face in the final image corresponds to the control conditions, avoiding mutual interference between different IDs.

[0033] In one possible implementation of the first aspect, the attributes of the face in the first image are different from the target attributes, and the first face corresponds to the same object as the face in the second image.

[0034] This implementation does not specify whether the above conditions are met during model inference or model training. It is important to note that when the above implementation is applied to model training, for example, the first image and the second image are a set of training data for the model. The attributes of the face in the first image are different from the target attributes, and the first face corresponds to the same object as the face in the second image. It can be understood that the first image controls the face ID of the final generated photo, while the second image controls the attributes of the final generated face. This method decouples the face ID from the attribute control conditions, ensuring that the expression, angle, and position of the final generated content are consistent with the control conditions while the face ID of the source image is preserved.

[0035] Secondly, this application provides an image processing apparatus, which includes an acquisition module and a processing module for performing all or part of the operations described in the first aspect. The communication device can be a network device such as an access network element, or a component within the network device used to perform the relevant operations, such as a line card or interface board, or a chip system used to perform the relevant operations. The chip system may include one or more chips. When the communication device is a chip system, the acquisition module and processing module can be, for example, the interface circuit of the chip, and the processing module can be, for example, the processing circuit of the chip.

[0036] For example, when executing the method of the first aspect, an acquisition module is used to acquire a first image and a second image, the first image containing a first face, and the second image being an image of a 3D face model with target attributes, the target attributes including: lighting, facial expression, or facial orientation; a processing module is used to obtain an encoding result by an encoder based on the first image and the second image, the encoding result incorporating features of the first image and the second image; and to obtain a third image by a generation model based on the encoding result, the third image containing a second face with the target attributes, the second face corresponding to the first face being the same person.

[0037] In one possible implementation of the second aspect, the acquisition module is further configured to: acquire a first result, the first result being obtained based on the second image and the noisy image, the first result being fused with features of the second image and the noisy image; the processing module is specifically configured to: obtain an encoding result by the encoder based on the first image and the first result, the encoding result being fused with features of the first result and the first image.

[0038] In one possible implementation of the second aspect, the noisy image includes: a completely noisy image or a fourth image, which is obtained by adding noise to a fifth image, which is an image obtained in the adjacent previous round through the generation model.

[0039] In one possible implementation of the second aspect, the first image includes: a face image and a face warp image, the face image being obtained by extracting a face from a sixth image, the sixth image being an acquired image containing the first face, and the face warp image being obtained by distorting the sixth image, the face warp image having the target attribute.

[0040] In one possible implementation of the second aspect, the encoding result includes: a first encoding result and a second encoding result. The processing module is specifically used to: obtain a first encoding result by the encoder based on the face image and the second image, wherein the first encoding result incorporates features of the face image and the second image; and obtain a second encoding result by the encoder based on the face warp image and the second image, wherein the second encoding result incorporates features of the face image and the second image.

[0041] In one possible implementation of the second aspect, the processing module is specifically used to: obtain a first output result through the generation model based on the first encoding result; and obtain a second output result through the generation model based on the second encoding result.

[0042] The first output and the second output are summed to obtain the third image.

[0043] In one possible implementation of the second aspect, the first face includes a third face and a fourth face; the first image includes a seventh image containing the third face and an eighth image containing the fourth face; the second image includes a ninth image for controlling the attributes of the seventh image and a tenth image for controlling the attributes of the eighth image; the encoding result includes a third encoding result and a fourth encoding result; the acquisition module is further configured to receive a first instruction for instructing the generation of a combined photo containing the third face and the fourth face; the processing module is specifically configured to: obtain the third encoding result by the encoder based on the seventh image and the ninth image, the third encoding result incorporating features of the seventh image and the ninth image; and obtain the fourth encoding result by the encoder based on the eighth image and the tenth image, the fourth encoding result incorporating features of the eighth image and the tenth image.

[0044] In one possible implementation of the second aspect, the attributes of the face in the first image are different from the target attributes, and the first face corresponds to the same object as the face in the second image.

[0045] Thirdly, this application provides an image processing apparatus, which includes: a processor, a memory, an input / output device, and a bus; the memory stores computer instructions; when the processor executes the computer instructions in the memory, the memory stores computer instructions; when the processor executes the computer instructions in the memory, it is used to implement any of the embodiments of the first aspect.

[0046] Fourthly, this application provides an image processing apparatus, which includes: a processor, a memory, an input / output device, and a bus; the memory stores computer instructions; when the processor executes the computer instructions in the memory, the memory stores computer instructions; when the processor executes the computer instructions in the memory, it is used to implement any of the embodiments of the second aspect.

[0047] Fifthly, embodiments of this application provide a chip system including a processor and input / output ports. The processor is used to implement the processing functions involved in the method described in the first aspect above, and the input / output ports are used to implement the transmission and reception functions involved in the method described in the first aspect above.

[0048] In one possible design, the chip system also includes a memory for storing program instructions and data for implementing the functions involved in the method described in the first aspect above.

[0049] This chip system can consist of chips or include chips and other discrete components.

[0050] Sixthly, embodiments of this application provide a computer-readable storage medium. The computer-readable storage medium stores computer instructions; when the computer instructions are executed on a computer, the computer causes the computer to perform the method as described in any of the possible implementations of the first aspect.

[0051] The technical effects of the second to sixth aspects or any of the possible implementations can be found in the first aspect or the related possible implementations of the first aspect, and will not be repeated here. Attached Figure Description

[0052] Figures 1A to 1C This application provides an illustration of an application scenario.

[0053] Figure 2 A flowchart illustrating an image processing method provided in this application;

[0054] Figure 3 A schematic diagram of an image processing procedure provided in this application;

[0055] Figure 4 A schematic diagram illustrating a process for generating a noisy image as provided in this application;

[0056] Figure 5 A schematic diagram illustrating the process of image processing using a diffusion model provided in this application;

[0057] Figure 6 A schematic diagram illustrating the process of training a model as provided in this application;

[0058] Figure 7 A schematic diagram illustrating a scenario for generating a group photo using a related technology provided in this application;

[0059] Figure 8 A schematic diagram illustrating a scenario where this solution generates a group photo for the purposes of this application;

[0060] Figure 9 A flowchart illustrating an embodiment of this application;

[0061] Figure 10 A schematic diagram illustrating the result of face editing provided in this application;

[0062] Figure 11 A comparison diagram of the relevant technologies provided for this application and the solution applied to a group photo scenario;

[0063] Figure 12 A schematic diagram of the structure of an image processing device provided in this application;

[0064] Figure 13 A schematic diagram of the structure of yet another image processing apparatus provided in this application;

[0065] Figure 14 This is a schematic diagram of the structure of a chip provided in this application. Detailed Implementation

[0066] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, some terms involved in the embodiments of this application are explained below.

[0067] This application relates to the application of neural networks. In order to better understand the solutions of this application, the relevant terms and concepts of neural networks that may be involved in this application will be introduced below.

[0068] (1) Transformer Network

[0069] The Transformer model was first used in natural language processing, employing attention layers to effectively capture long-range connections between words. With further research into the Transformer network structure, researchers in computer vision have improved the performance of various tasks such as detection, classification, and generation by designing the Vision Transformer. By segmenting images into blocks and mapping them to sequential features, the attention mechanism within the Transformer structure is used to extract long-range information from the image, mitigating the problem of the convolutional neural network's small receptive field.

[0070] The Transformer network structure includes:

[0071] ① Multi-head Self-attention: Multi-head attention expands the model's representational power by weighting and merging individual attentions.

[0072] ② Multi-Layer perceptron (MLP): The MLP in Transformer is a neural network containing two fully connected layers.

[0073] ③ Layer Normalization (LN): The introduction of layer normalization can speed up network training and ensure network convergence.

[0074] ④ Positional embedding: To distinguish the positions of different vectors, positional embedding is introduced as an additional piece of information in the Transformer result. Common positional embedding methods include: sin / cos function positional embedding that does not require learning and learnable positional embedding.

[0075] In addition, the Transformer network has a Cross-Attention operation, which can be used to inject image or text information into the neural network.

[0076] (2) Diffusion model

[0077] The diffusion model is an unsupervised generative model, comprising a diffusion process and a reverse process (generation process). The forward process is the diffusion process, and the reverse process is the generation process. In the forward process, a pure noise sample x0 can be obtained by continuously adding T steps of Gaussian noise to the initial sample x0. T The reverse process involves continuously denoising the noise to generate a clean image.

[0078] Classifier-free guidance: Without adding additional input to the diffusion model, users cannot effectively control the results. Therefore, control conditions are often needed to limit the generation of results.

[0079] Stable Diffusion: A latent space-based diffusion generation model that extends the diffusion process from the pixel level to the latent space, saving significant time and computational resources. Stable Diffusion mainly consists of three modules: ① Variational Auto-Encoder (VAE): Responsible for compressing the three-channel source image into the latent space, and then the decoder uses the vectors in the latent space to restore the three-channel red, green, blue (RGB) image. ② Backbone Network UNet: Responsible for predicting the noise added to the latent space vectors. The main algorithm of diffusion models is the process of adding and removing noise to the image or latent space tensor, while Stable Diffusion mainly completes the diffusion generation of the image by adding and removing noise in the latent space. ③ CLIP Text Encoder: Used to first decompose the given text description into tokens, then encode them into text vectors, and inject information into UNet for information fusion through a cross-modal attention mechanism.

[0080] This solution can be applied to the field of face image processing. The following describes some terminology used in face image processing:

[0081] (1) Face ID: Each person has a unique face ID, which can be identified and distinguished based on the features of the face.

[0082] (2) Controllable and editable face ID: refers to changing the expression, angle, lighting and shadow of a face without changing the face ID.

[0083] (3) Generation: refers to image generation, which is to obtain a complete image based on the input through a neural network model and certain processing, such as ID photos, artistic photos, etc.

[0084] (4) Downsampling: An operation that reduces the size of an image.

[0085] To facilitate the description of the plan, some terms are used. To prevent misunderstandings, these terms are explained below.

[0086] (1) The terms “first,” “second,” etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate; this is merely a way of distinguishing objects with the same attributes in the description of embodiments of this application. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of units is not necessarily limited to those units, but may include other units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0087] (2) The terms “substantially,” “about,” and similar terms used herein are used as approximations rather than as terms of degree, and are intended to take into account the inherent biases of measurements or calculations known to those skilled in the art. Furthermore, the use of “may” in describing embodiments of the invention means “one or more possible embodiments.” Additionally, the term “exemplary” is intended to refer to an instance or illustration.

[0088] The following examples illustrate the application fields and application scenarios of the embodiments of this application.

[0089] The method provided in this application can be applied to the field of image processing. For example, this solution can be applied to the field of virtual photography, an art form that uses computer graphics technology to create virtual scenes. It involves reconstructing, combining, and redesigning elements of the real world to create new virtual worlds. For example, this solution can be applied to image customization and image generation in virtual photography. For example, this solution can also be applied to the field of virtual humans, which are simulated human bodies created through computer graphics and virtual reality technology, typically referring to characters in virtual worlds, such as characters in online games, special effects characters in movies, animation, and virtual reality. For example, this solution can be used for modeling virtual humans and editing their actions and expressions. For example, this solution can also be applied to terminal devices such as mobile phones and computers to provide photo editing functions. Furthermore, this solution can also be applied to other image editing, image customization, or image generation fields.

[0090] For example, when this solution is executed by a cloud server and applied in the field of virtual photography, the user first inputs or the cloud server retrieves an image containing a human figure, for example... Figure 1A At this point, the user or cloud server will automatically select customized content, such as... Figure 1B As shown, the customization options include: whether to enable a smile, whether to maintain the hairstyle, whether to maintain the pose, and a prompt describing the desired content, which is: monk. After selecting the customization, the cloud server will edit the aforementioned photo based on the customization to generate a new photo, for example, based on the prompt... Figure 1A Editing makes Figure 1A Object A in the story wears various types of monk's clothing and obtains... Figure 1C If the user clicks "Submit Now" at the end, the cloud server will return the result to the user. Figure 1C .

[0091] The above is an example of an application scenario for this solution. A typical application scenario is to customize faces in images, such as facial expressions, lighting, or facial orientation. In addition, this solution can also be used to customize other features of images, including: clothing, body orientation, animal expressions, fur color, or movement.

[0092] The application fields and scenarios of the image processing method provided in the embodiments of this application have been described above. The execution process of the image processing method provided in the embodiments of this application will be described in detail below. Please refer to... Figure 2 , Figure 2 This is a schematic flowchart illustrating an image processing method provided in an embodiment of this application. Figure 2As shown, the image processing method provided in this application embodiment includes the following steps S201-S203.

[0093] S201, The first device acquires the first image and the second image;

[0094] It should be noted that the first image contains a first human face, and the second image is an image of a 3D face model with target attributes, including lighting, facial expression, or facial orientation. For ease of description, the 3D face model image mentioned below refers to the second image.

[0095] It should be noted that the first device includes: terminal equipment or server. Terminal equipment includes: mobile phones (or "cellular" phones), mobile phones, computers, and data cards, for example, portable, pocket-sized, handheld, computer-embedded, or vehicle-mounted mobile devices or wearable devices that exchange voice and / or data with the wireless access network. Examples include personal communication service phones, cordless phones, tablets, computers with wireless transceiver capabilities, etc. Terminal equipment can also be referred to as a system, subscriber unit, subscriber station, mobile station (MS), remote station, access point (AP), remote terminal, access terminal, user terminal, user agent, subscriber station (SS), customer premises equipment (CPE), terminal, user equipment (UE), and mobile terminal. Terminal (MT), etc. Servers include: cloud servers, computing servers, media servers, supercomputing servers, virtual servers, and physical servers.

[0096] It should be noted that, as Figure 3 As shown, Figure 3 The second image is planar, but the face model in the second image has a three-dimensional structure, capable of representing the features of a face in a three-dimensional way. For example, this three-dimensional face model has depth and can represent the features of a frontal or side view of a face. Optionally, the second image can be a two-dimensional projection of the 3D face model.

[0097] It should be noted that the first face in this solution can be the face of a real person or the face of a virtual character. For example, it can be the face of a cartoon character, the face of a virtual character in an online game, the face of a special effects character in a movie, or a virtual character created in the fields of virtual reality.

[0098] It should be noted that the above-mentioned light and shadow can be the light and shadow of the entire second image, or the light and shadow of the face in the second image. The so-called light and shadow includes: brightness, contrast, color level or exposure.

[0099] It should be noted that the facial expressions mentioned above include the direction or state of the eye muscles, facial muscles, or mouth muscles. Generally, when the eye muscles, facial muscles, or mouth muscles change, the emotion conveyed by the face also changes, resulting in emotions such as smiling and happiness. Based on the different emotions conveyed, facial expressions can be categorized, for example, into types such as crying or smiling.

[0100] Optionally, the orientation of the face can be the angle between the face and the plane containing the entire body, or the angle between the face and a direction perpendicular to the second image or a direction parallel to the second image. For example, the face can be oriented towards a direction of 30°, where 30° is the angle between the entire face and a direction parallel to the second image.

[0101] Optionally, the second image in this scheme is a 3D face model image without textures, such as... Figure 3 As shown, this 3D face model image does not undergo texture mapping or other processing; instead, it is a 3D shape model of the face. Optionally, the first device can create a "texture-free" 3D face model using a 3DMM model.

[0102] It should be noted that "having target attributes" does not mean that all images with attributes such as expression, lighting, or face orientation are second images. For example, in this solution, the second image can have specific attributes such as expression, lighting, or face orientation. For instance, by setting specific values ​​to the parameters of the 3D face model, the expression, angle, and lighting of the 3D face model can be edited to give the final 3D face model image specific attributes.

[0103] It should be noted that this solution does not limit the first device to simultaneously acquiring the first image and the second image; the first device may also acquire the first image and the second image at different times. To help understand how the first device acquires the first image and the second image, several examples are given: The first device can first receive the image input by the user, preprocess the image to obtain the first image; the first device can match the corresponding second image from the database according to user requirements or system presets, or generate the corresponding second image; the first device can simultaneously acquire the image containing a face and the second image from memory or the cloud, and preprocess the image containing the face to obtain the first image, or acquire the image containing a face and the second image separately, and preprocess the image containing the face to obtain the first image. For ease of description, the image that needs to be edited, whether input by the user or acquired by the device from the cloud, is referred to as the source image.

[0104] Optionally, the first device can generate a corresponding original 3D face model image based on the face in the first image, and then adjust the parameters of the original 3D face model image according to user needs or system presets. For example, if the user instructs that the expression in the first image be adjusted to a smile, the parameters of the original 3D face model image can be adjusted so that the expression in the generated second image is a smile. For example, some image software has a system preset photo template with specific lighting and shadows on the face in the template, so the parameters of the original 3D face model image can be adjusted so that the lighting and shadows in the generated second image match the lighting and shadows of the template.

[0105] In one possible implementation, the first image includes: a face image and a face warp image. The face image is obtained by extracting a face from a sixth image, which is an image containing the first face. The face warp image is obtained by distorting the sixth image and has the target attribute.

[0106] In this implementation, the first device acquires a sixth image containing a face, such as a source image input by the user. The sixth image is then preprocessed in different ways to obtain a face image and a face warp image. The face warp image is obtained by distorting the sixth image, giving it the target attributes. Therefore, the face warp image possesses features not found in the face image, which are adapted to the target attributes and can be better integrated with the second image. For example, when the target attribute is a face facing 30 degrees, the face warp image will have the features of a face facing 30 degrees. The face image and face warp image are obtained by processing the sixth image. However, the image distortion process may result in poor realism and lack of refinement, and texture bugs may appear in occluded areas. The face image, on the other hand, is obtained by directly extracting the face from the sixth image, thus retaining the features of the first face in the image and being less prone to bugs. Therefore, it can serve as a supplement to the warp image. In this scheme, both the face image and the face warp are fused with the second image through an encoder. The features of the face image and the face warp image complement each other, providing richer and more diverse face-related features, reducing the risk of bugs, and thus making the final fused image more detailed, more realistic, and of higher quality.

[0107] S202, the first device obtains the encoding result through an encoder based on the first image and the second image;

[0108] It should be noted that the encoding result incorporates features from both the first and second images.

[0109] In one possible implementation, the first device acquires a first result, which is obtained based on the second image and the noisy image, and the first result incorporates features of the second image and the noisy image. The processing module is specifically used to: the first device obtains an encoding result through the encoder based on the first image and the first result, and the encoding result incorporates features of the first result and the first image.

[0110] In this implementation, the first result incorporates features from both the second image and the noisy image. By fusing the features of the first image with those of the noisy image, the second image can be enhanced. When error information appears in the second image, this feature fusion can cover up the error information in the second image.

[0111] Optionally, the first device can fuse the second image and the noisy image using a fusion module. Optionally, the fusion module can first align the sizes of the second image and the noisy image in the spatial dimension, and then perform channel-level concatenation. For example, let the shape of the 3D face model image be [H, W, C], where H is the height, W is the width, and C is the number of channels. During spatial alignment, the first device connects adjacent 8*8 pixels in the 3D face model image in space through the channel dimension, i.e., 8*8*C → 1*1*8. 2 C. Obviously, the number of elements remains unchanged before and after alignment, thus ensuring information integrity. During channel-dimensional concatenation, the first device spatially aligns the 3D face model image with the noisy image, then concatenates pixels at the same spatial location along the channel dimension, i.e., 1*1*8. 2 C and 1*1*C′→1*1*(8 2 (C+C′). At this point, since the noise in the noisy image does not affect the element values ​​in the 3D face model image, the information remains lossless.

[0112] The above implementation can solve the problem of different resolutions between the 3D face model image and the noisy image. Generally speaking, the resolution of the 3D face model image is 8 times that of the noisy image. This solution ensures that the number of elements before and after fusion does not change through spatial alignment and channel dimension connection, as well as the connection of pixels in the same spatial position in the channel dimension, thereby ensuring lossless information and realizing lossless fusion of the 3D face model image and the noisy image.

[0113] Optionally, the fusion module and encoder can also be located in the same device, such as the aforementioned first device. In this case, the fusion module and encoder are connected in series, and the first result output by the fusion module can be input into the encoder for further processing. The encoder and fusion module together form an ID injection model, which is used to process the acquired first and second images. Alternatively, the fusion module and encoder can be located in different devices. For example, the encoder may be located in the first device, while the fusion module may be located in another device, such as the second device. In this case, the fusion module in the second device processes the second image to obtain a first result, which the second device can then send to the first device. The first device then further processes the first result through the encoder.

[0114] Optionally, the encoder can be located on the cloud side, while the fusion module can be located on the cloud side or the edge side.

[0115] Optionally, the encoder can be a program running in a processor, which should have strong computing power, such as a graphics processing unit (GPU). The fusion module can also be a program running in a processor, with relatively low requirements for the processor's computing power, such as a central processing unit (CPU) or GPU.

[0116] In one possible implementation, the noisy image includes: a completely noisy image or a fourth image, which is obtained by adding noise to a fifth image, which is an image obtained in the previous round through the generative model.

[0117] It should be noted that a completely noisy image is an image that is entirely filled with noise. This application provides an iterative scheme, which iteratively uses the fifth image output by the previous generation model and adds noise to the fifth image to obtain a fourth image. For example, in this application, step S202 is executed multiple times. The following steps S203 and their possible implementations involve obtaining the fifth image output by the previous generation model before executing step S202 in this round, adding noise to the fifth image to obtain a fourth image, and then fusing the fourth image and the second image through a fusion module to obtain a first result.

[0118] In this implementation, noise is added to the fifth image obtained from the adjacent previous generation model to obtain the fourth image. Then, based on the fourth image and the second image, a first result is obtained through a fusion module. That is, the output result from the previous round is used in this round of image processing, and this result is fused with the second image. It is evident that the above implementation ensures that subsequent iterations can better utilize the generation results of preceding iterations, improving the quality of the generated image and increasing the similarity between the generated image and the face IDs in the obtained original image. For example, it improves the similarity between the face IDs in the third image and the first image.

[0119] Optionally, in the first round, the first device can obtain a fusion result of the fully noisy image and the second image through a fusion module based on the fully noisy image and the second image. Then, based on the second image and the noisy image of the previous round's output result, i.e., the fourth image, the fusion module obtains a fusion result of the fourth image and the second image. Furthermore, during the iteration process, the degree of noise addition decreases with each iteration when adding noise to the output result of the previous round's generation model.

[0120] by Figure 4For example, in each round of generating a noisy image, the image output by the previous generation model is first acquired. This image is noise-free and can therefore be called the noiseless output. Then, noise is added to the noiseless output. Specifically, when adding noise to the first round of output image, the intensity of the added noise is high; when adding noise to the second round of output image, the intensity of the added noise is medium; and when adding noise to the third round of output image, the intensity of the added noise is low.

[0121] Optionally, the output of the fusion module is a fused image of the second image and the noisy image, i.e., the first result. This first result is input into the encoder, which then fuses the first result with the first image. Figure 3 As shown, the output of the fusion module is input to the encoder and further processed by the encoder.

[0122] The encoder is described below:

[0123] The encoder in this application is a neural network encoder, and the model in the encoder is a neural network model. Optionally, the network structure of the encoder can be consistent with the encoder of Stable Diffusion's UNet, but the parameters in this network structure need to be retrained. The specific training process is described below. For example, the structure of the encoder in this scheme includes convolutional layers and transformer layers. The first convolutional layer in the convolutional layer is input to the first result output by the fusion module. The convolutional layer can extract the features of the first result to obtain a feature map. The feature map is input to the transformer layer, while the features of the first image are input to the transformer layer through Cross-Attention. The transformer layer processes the feature map and the features of the first image to obtain the encoding result.

[0124] Understandably, in this scheme, the encoder also fuses the features of the first image and the first result during the process of fusing them. This fusion is performed because the ID of the first face in the first image is the same as that in the source image. By fusing with the 3D face model image, attributes such as facial expression, lighting, and face orientation in the first image can be edited, thus meeting the user's customization needs when required. Since the 3D face model image can control the attributes of the final generated image, it can be referred to as the control condition.

[0125] Optionally, the first device can first extract the description token of the first image using a retrained pre-trained model, such as a retrained CLIP model, and then input the description token into the encoder through Cross-Attention. Here, "token" is usually used to represent the smallest unit in text or sequence data.

[0126] Optionally, the process of processing the output of the fusion module and the features of the first image by the encoder can be executed multiple times. Each execution will generate the features of the ID of the first face. Multiple executions will result in features of multiple IDs. The features of multiple IDs can be summed to obtain the total features of the first face.

[0127] Optionally, to facilitate integration with subsequent generative models, the encoder can generate a feature of spatial resolution ID in each execution process. For example, this feature can be in the form of a feature map. For instance, when the generative model is a Diffusion model, the spatial resolution of the input to the Diffusion model is [H, W], where H is the height and W is the width. Therefore, in the encoder and decoder of the Diffusion model, there exists [H, W]... Feature maps at four resolutions. The encoder of the ID injection model also generates [H,W] maps. ID feature maps at four different resolutions. (Example) Figure 5 As shown, Figure 5 The four parallelograms output by the ID injection model represent feature maps at four different resolutions. Figure 5 In the diffusion model, the four polygons used for feature map addition represent the feature maps of the corresponding resolution in the diffusion model. Figure 5 It can be seen that the ID injection model where the encoder is located outputs feature maps of four resolutions and adds them together. Element-wise summation can be performed during the feature map addition process.

[0128] As mentioned above, in this scheme, the first image includes a face image or a face warp image. When the encoder processes the first image and the 3D face model image, it can process the face image and the 3D face model image, or the face warp image and the 3D face model image, respectively, as detailed below:

[0129] In one possible implementation, the encoding result includes: a first encoding result and a second encoding result. The first device obtains the first encoding result by the encoder based on the face image and the second image, and the first encoding result incorporates features of the face image and the second image. The first device obtains the second encoding result by the encoder based on the face warp image and the second image, and the second encoding result incorporates features of the face image and the second image.

[0130] In this embodiment, the processing of the face image and the face warp image by the first device is performed separately. Specifically, the first device obtains a first encoding result using an encoder based on the face image and the second image, and obtains a second encoding result using an encoder based on the face warp image and the second image. The encoder's processing of the face image and the encoder's processing of the face warp image are performed independently. Figure 3 As shown, the face image and the face warp are input into the encoder separately. The reason for processing them independently is that fusing the face image and the second image through the encoder results in a final encoded image with high similarity to the source image's ID. This also avoids issues such as poor realism in the face warp image affecting the final generated image. Furthermore, fusing the second image and the face warp image through the encoder allows extraction of edited features of the first face's ID, i.e., features resulting from changes in the first face's orientation and expression.

[0131] In one possible implementation, the first device obtains a first output result through the generative model based on the first encoding result; the first device obtains a second output result through the generative model based on the second encoding result; the first device sums the first output result and the second output result to obtain the third image.

[0132] Understandably, in the above implementation, the generative model obtains corresponding output results from the encoding results obtained based on the face image and the encoding results obtained based on the face warp image, respectively. In this scheme, firstly, the face image and the face warp image are processed independently in parallel, and then the output results of the two are summed to obtain the final third image. Based on the foregoing analysis, it is known that the face image and the face warp image can complement each other, ensuring a high similarity between the ID of the final third image and the source image, while also allowing for better editing of facial attributes such as expressions.

[0133] S203. The first device obtains the third image by generating a model based on the encoding result.

[0134] It should be noted that the third image contains a second face with the target attribute, and the second face corresponds to the same person as the first face.

[0135] Optionally, the generative model includes: a generative adversarial model, a variational autoencoder, a streaming model, or a diffusion model. Obtaining the third image through a diffusion model is a typical implementation of this scheme. Optionally, the diffusion model in this scheme can be Stable Diffusion's Unet, meaning the process of obtaining the third image based on the encoding result can be performed in the latent space of the VAE, and the final image needs to be obtained through the VAE's decoder.

[0136] Optionally, the feature map resolution in the diffusion model will first decrease and then increase. As mentioned earlier, the ID injection model outputs feature maps at multiple resolutions. Before each resolution reduction, the feature map in the diffusion model is added to the corresponding resolution feature map generated by the ID injection model, and then subsequent processing is performed, such as feature extraction through convolutional layers and resolution reduction. Finally, the third image generated by the diffusion model is a noise-free image.

[0137] Optionally, the first device obtains the third image through the generation model based on the encoding result and the fourth image, and the third image incorporates features from the encoding result and the fourth image.

[0138] Optionally, the generative model can be located on the cloud side. For example, the generative model can be a program running on a processor on the cloud side, which should have strong computing power, such as a graphics processing unit (GPU).

[0139] Optionally, a typical application scenario of this solution is that the generative model and the encoder can also be located in the same device, such as the aforementioned first device. In this case, the generative model and the encoder are connected in series, and the output of the encoder can be input into the generative model for further processing. However, the generative model and the encoder can also be located in different devices. For example, the encoder is located in the first device, while the generative model is located in another device, such as the third device. In this case, the encoder in the first device performs image processing to obtain the encoding result, and the first device can send the first result to the third device. Then, the third device further processes the encoding result through the generative model.

[0140] It is understandable that during this round of image processing, the generative model can also input the noisy image of the previous round's output, so as to iterate the output results of the previous round, thereby ensuring the similarity of IDs between rounds and improving the similarity between the final generated image and the source image ID.

[0141] Optionally, the first device can also input the prompt into the diffusion model via Cross-Attention, and then utilize the powerful text-to-image generation capability of the diffusion model to edit the background, clothing, pose, etc. of the person in the generated image based on the prompt. For example... Figure 5 As shown, the prompt word could be: "A pretty girl." This prompt word can be used to control the image processing, ensuring that the final generated image matches the prompt word.

[0142] In this scheme, "the third image contains a second face with the target attribute, and this second face corresponds to the same person as the first face." That is, the face in the third image corresponds to the same person as the face in the first image, and the attributes of the face in the third image are the same as those in the second image. In other words, this scheme generates an image with the same face ID as the source image and edits aspects such as expression, lighting, and face orientation based on the source image. This editing capability is related to the model training process. In this scheme, model training generally refers to training a combination of the ID injection model and the generation model, but this scheme does not limit this. The specific training process of the model is as follows:

[0143] In one possible implementation, the attributes of the face in the first image are different from the target attributes, and the first face corresponds to the same object as the face in the second image.

[0144] It is understood that the above implementation method is not limited to satisfying the above conditions during model inference or model training. It should be specifically noted that this solution provides an optional model training method. In this method, the first image and the second image are a set of training data for the model. The attributes of the face in the first image are different from the target attribute, and the first face corresponds to the same object as the face in the second image. It is understood that the first image is used to control the face ID of the final generated photo, while the second image is used to control the attributes of the final generated face. Figure 6 As shown in the right figure, the face ID is obtained from image A. Then, based on image B, which has the same ID but different attributes, control conditions such as expression, angle, and face orientation are generated. Based on these two conditions, image processing is performed to finally generate image B. This method decouples the face ID from the attribute control conditions, ensuring that the expression, angle, and position of the final generated content are consistent with the control conditions, while the face ID of the source image is preserved.

[0145] Understandably, this decoupling is necessary because the first image contains not only the face ID but also facial attributes, such as expression. If the model extracts the face ID and control conditions for facial attributes from the first image, the final image generated by the model might overlap with the first image, making it impossible to edit the facial attributes in the first image. Therefore, decoupling training is required to allow the model to learn how to extract the face ID and control conditions from different images separately.

[0146] The image processing method provided in this application can be applied to scenarios involving group photos of multiple people, as detailed below:

[0147] In one possible implementation, the first face includes a third face and a fourth face; the first image includes a seventh image containing the third face and an eighth image containing the fourth face; the second image includes a ninth image for controlling the attributes of the seventh image and a tenth image for controlling the attributes of the eighth image; the encoding result includes a third encoding result and a fourth encoding result; the first device receives a first instruction to instruct the generation of a combined photo containing the third face and the fourth face; the first device obtains the third encoding result by the encoder based on the seventh image and the ninth image, the third encoding result incorporating features of the seventh image and the ninth image; the first device obtains the fourth encoding result by the encoder based on the eighth image and the tenth image, the fourth encoding result incorporating features of the eighth image and the tenth image.

[0148] Understandably, when a user inputs multiple face images, or retrieves multiple face images from the cloud, and receives an instruction to generate a group photo of these multiple faces, this solution will input each individual face and its corresponding 3D face model image into the ID injection model in a combined manner. The 3D face model image is used to control the attributes of the face, i.e., control conditions. The output of the ID injection model is then input into the generation model, thereby ensuring that the face in the final image corresponds to the control conditions and avoiding mutual interference between different IDs.

[0149] like Figure 7 As shown, in existing image processing technologies, the device directly inputs images 1, 2, and 3, along with control conditions 1, 2, and 3, into the diffusion model. This results in images containing faces not corresponding to the control conditions. Taking the example that image 1 could be an image with face ID A, image 2 could be an image with face ID B, and image 3 could be an image with face ID C, the final model obtains photos of the average faces of A, B, and C, and generates multiple photos of this average face based on the control conditions. It is evident that in scenarios involving group photos, this existing technology causes interference between different IDs, making it impossible to independently correspond to the control conditions.

[0150] by Figure 8For example, consider the following scenario: The user inputs Image 1, Image 2, and Image 3. Image 1 could be an image with face ID A, Image 2 could be an image with face ID B, and Image 3 could be an image with face ID C. The user or system template indicates that Image 1 corresponds to control condition 1, Image 2 to control condition 2, and Image 3 to control condition 3. For example, control condition 1 could be a smile, control condition 2 could be crying, and control condition 3 could be laughing. Then, obtain the 3D face model images for the smiling, crying, and laughing faces respectively. Input Image 1 with control condition 1, Image 2 with control condition 2, and Image 3 with control condition 3 into the ID injection model.

[0151] To facilitate understanding, the above image processing methods will be explained below with specific examples.

[0152] This embodiment is an example of an image processing method. This embodiment is an example of the above image processing method being applied to the following scenario: a user uploads an image, and the device performs customized editing on the face in the image. This embodiment is an example of a diffusion model as the generative model.

[0153] Please see Figure 9 , Figure 9 This is a flowchart illustrating an embodiment provided in this application. Figure 9 The specific process includes:

[0154] S901, The first device receives image A input by the user;

[0155] It should be noted that the image A includes a human face, which has attribute a. For example, attribute a includes a smiling expression and the orientation of the face being 30° off from the vertical direction of the photograph.

[0156] S902, The first device determines the control parameters corresponding to image A;

[0157] It should be noted that this control parameter is used to control the attributes of the 3D face model image.

[0158] Optionally, the first device can determine control parameters based on user instructions. For example, if a user clicks a control on the screen of the first device to instruct the expression of a face in an image to be edited to a smile, the first device can determine that the control parameter is used to control the expression of the 3D face model image to be a smile.

[0159] Optionally, the first device can also determine the control parameters based on a preset template. For example, some software with fixed photo templates may preset the orientation of the face to be parallel to the vertical direction of the photo. In this case, the first device can determine that the control parameter is used to control the orientation of the 3D face model to be parallel to the vertical direction of the photo.

[0160] S903, The first device creates a corresponding 3D face model based on image A and the control parameters;

[0161] Optionally, the first device can first create an original 3D face model image based on image A, and then adjust the parameters in the original 3D face model image to the control parameters in step S902, thereby adjusting the expression, lighting and other attributes in the original 3D face model image.

[0162] S904. The first device preprocesses image A to obtain a face image and a face warp image;

[0163] The preprocessing includes: extracting faces from image A to obtain face images; processing image A according to the attributes of a 3D face model; and rendering the face textures from image A onto the 3D face model to obtain a textured image. Figure 3 The 3D face model image is used to convert image A into a face warp image that matches the shape attributes such as angle and expression of the 3D face model image.

[0164] S905, The first device fuses the 3D face model image and the noisy image through the fusion module to obtain image B;

[0165] As mentioned earlier, the first device can fuse the 3D face model image and the noisy image through spatial alignment and channel connection. When the process of steps S905-S912 is executed for the first time, the noisy image is a completely noisy image. When the process of steps S905-S912 is executed for the second time, the noisy image is the image G obtained in step S911 when the process of steps S905-S912 is executed for the first time. When the process of steps S905-S912 is executed for the third time, the noisy image is the image G obtained in step S911 when the process of steps S905-S912 is executed for the second time, and so on.

[0166] S906. The first device fuses the face image with image B through an encoder to obtain image C;

[0167] Optionally, the first device can first extract the features of the face image through the clip model, then input the features into the encoder through cross-attention, and then fuse the features with image B through the encoder to obtain image C.

[0168] S907. The first device fuses the face warp image with image B using an encoder to obtain image D;

[0169] Optionally, the first device can first extract features of the face warp image through the clip model, then input the features into the encoder through cross-attention, and then fuse the features with image B through the encoder to obtain image D.

[0170] S908, The first device inputs the prompt into the diffusion model through cross-attention;

[0171] Optionally, the above prompt can be a prompt input by the user, such as: a girl with yellow hair, or the above prompt can be a prompt obtained based on an image input by the user.

[0172] S909. The first device processes image C using a diffusion model to obtain image E;

[0173] It should be noted that the diffusion model can be used to add and remove noise from image C, preventing erroneous information in image C from affecting the final generated image.

[0174] S910. The first device processes image D using a diffusion model to obtain image F;

[0175] Refer to the description of step S909.

[0176] S911. The first device performs a weighted summation on images E and F to obtain image G;

[0177] Optionally, when performing weighted summation here, the weights of images E and F can be set by the user or be the system default weights.

[0178] In step S911, the image G generated by the diffusion model does not contain noise. After generating image G, the first device can add noise to image G to obtain a noisy image. Then, it returns to step S905 to input the noisy image into the fusion module for the next round of fusion of the noisy image and the 3D face model image. It then returns to steps S909 and S910 to input the noisy image into the diffusion model. The diffusion model processes the noisy image and image C or image D so that the diffusion model can iterate the output results of the previous round.

[0179] S912, The first device returns the image G to the user.

[0180] For example, when the first device is a mobile phone, the image G can be displayed to the user on the mobile phone screen.

[0181] The above embodiments represent a significant improvement over related portrait customization solutions, as detailed below:

[0182] First, most of the related technologies require customized training for the specific face ID obtained. For example, in text-based image models, there exists a learnable text S. * The S * To describe the target object. Because customized training is required for each ID, the number of training samples is relatively small, which can lead to the model overfitting to the given training set. This will have the following consequences: S * Originally, it should only describe the face IDs in the image, but through this training method, S * It will also record common backgrounds, clothing, etc., in the images in the training set, thus leading to S-based... * During generation, only fixed backgrounds and clothing can be generated and cannot be modified. It is evident that in related technologies, customized training for specific face IDs leads to model overfitting, causing the model to lose generative diversity. In this invention, the difference in face IDs has little impact on the execution of the encoder and diffusion model. This solution does not require retraining of the encoder and diffusion model for specific IDs, saving significant training time and computational resources, and avoiding problems such as overfitting and loss of generative diversity.

[0183] Furthermore, in some related technologies, the device cannot precisely control details such as facial expressions in the image, resulting in low ID similarity. This invention, however, creates an ID injection model, introducing a fusion module and an encoder, and also incorporates 3D face technology when editing faces. This solution controls facial attributes through a 3D face model image, and through the mutual supplementation of the face image and face warp image, as well as iteration of the previous output image, this solution has significant advantages in controllability and ID preservation capabilities. Figure 10 As shown, the face ID in the image input by the user is A. This solution can generate multiple images with the same face ID A, but with edited expressions and facial orientation based on the user-input image. For example, the user-input face may not be smiling or having its eyes closed, and the face may be centered. Based on this solution, the expression can be edited to a smile, a laugh, or closed eyes, etc. The face orientation can also be adjusted to the left or right, and the clothing can be adjusted by setting prompts.

[0184] Furthermore, in related technologies, the device directly inputs the acquired images containing faces into the generation model, causing interference between the ID information corresponding to different faces when multiple images containing faces are input. For example... Figure 11As shown, in related technologies, when a user uploads an image with face ID A and an image with face ID B, the faces in the generated group photo are the average of A and B's faces. However, this application's embodiment creates an ID injection model, eliminating the need to directly input images containing faces into the generation model. The ID injection model can run independently on multiple images containing faces, generating features corresponding to each face ID separately, and then injecting them all into the generation model to generate a group photo containing these multiple face IDs. Figure 11 As shown in the diagram, in this solution, when a user uploads an image with face ID A and an image with face ID B, a composite photo of the faces with ID A and ID B will be generated.

[0185] The method flow provided in this application has been described above. The apparatus provided in this application will now be described based on the aforementioned method flow.

[0186] See Figure 12 The present application provides a schematic diagram of the structure of an image processing apparatus, comprising:

[0187] The acquisition module 1201 is used to acquire a first image and a second image. The first image contains a first face, and the second image is an image of a 3D face model with target attributes, including: lighting, facial expression, or facial orientation.

[0188] The processing module 1202 is configured to obtain an encoding result by an encoder based on the first image and the second image, the encoding result being fused with features of the first image and the second image; and to obtain a third image by a generation model based on the encoding result, the third image containing a second face with the target attribute, the second face corresponding to the same person as the first face.

[0189] In one possible implementation, the acquisition module 1201 is further configured to: acquire a first result, the first result being obtained based on the second image and the noisy image, the first result being fused with features of the second image and the noisy image; the processing module 1202 is specifically configured to: obtain the encoding result through the encoder based on the first image and the first result, the encoding result being fused with features of the first result and the first image.

[0190] In one possible implementation, the noisy image includes: a completely noisy image or a fourth image, which is obtained by adding noise to a fifth image, which is an image obtained in the previous round through the generative model.

[0191] In one possible implementation, the first image includes: a face image and a face warp image. The face image is obtained by extracting a face from a sixth image, which is an image containing the first face. The face warp image is obtained by distorting the sixth image and has the target attribute.

[0192] In one possible implementation, the encoding result includes: a first encoding result and a second encoding result. The processing module 1202 is specifically used to: obtain a first encoding result by the encoder based on the face image and the second image, wherein the first encoding result incorporates features of the face image and the second image; and obtain a second encoding result by the encoder based on the face warp image and the second image, wherein the second encoding result incorporates features of the face image and the second image.

[0193] In one possible implementation, the processing module 1202 is specifically used to: obtain a first output result through the generation model based on the first encoding result; and obtain a second output result through the generation model based on the second encoding result.

[0194] The first output and the second output are summed to obtain the third image.

[0195] In one possible implementation, the first face includes a third face and a fourth face; the first image includes a seventh image containing the third face and an eighth image containing the fourth face; the second image includes a ninth image for controlling the attributes of the seventh image and a tenth image for controlling the attributes of the eighth image; the encoding result includes a third encoding result and a fourth encoding result; the acquisition module 1201 is further configured to receive a first instruction, which instructs the generation of a combined photo containing the third face and the fourth face; the processing module 1202 is specifically configured to: obtain the third encoding result by the encoder based on the seventh image and the ninth image, the third encoding result incorporating features of the seventh image and the ninth image; and obtain the fourth encoding result by the encoder based on the eighth image and the tenth image, the fourth encoding result incorporating features of the eighth image and the tenth image.

[0196] In one possible implementation, the attributes of the face in the first image are different from the target attributes, and the first face corresponds to the same object as the face in the second image.

[0197] Please see Figure 13 The following is a schematic diagram of another image processing apparatus provided in this application.

[0198] The image processing apparatus may include a processor 1301 and a memory 1302. The processor 1301 and the memory 1302 are interconnected via a circuit. The memory 1302 stores program instructions and data.

[0199] The aforementioned are stored in memory 1302 Figure 2 or Figure 9 The steps in the code include the corresponding program instructions and data.

[0200] Processor 1301 is used to perform the aforementioned Figure 2 or Figure 9 The method steps performed by the image processing apparatus shown in any of the embodiments.

[0201] Optionally, the image processing apparatus may also include a transceiver 1303 for receiving or sending data.

[0202] This application also provides a computer-readable storage medium storing a program that, when run on a computer, causes the computer to perform the aforementioned actions. Figure 2 or Figure 9 The steps in the method described in the illustrated embodiment.

[0203] Alternatively, the aforementioned Figure 13 The image processing device shown is a chip.

[0204] This application also provides an image processing apparatus, which may also be referred to as a digital processing chip or a chip. The chip includes a processing unit and a communication interface. The processing unit obtains program instructions through the communication interface, and the program instructions are executed by the processing unit. The processing unit is used to perform the aforementioned... Figure 2 or Figure 9 The method steps performed by the image processing apparatus shown in any of the embodiments.

[0205] This application also provides a digital processing chip. This digital processing chip integrates circuitry for implementing the processor 1301 described above, or the functions of processor 1301, and one or more interfaces. When the digital processing chip integrates a memory, it can complete the method steps of any one or more of the foregoing embodiments. When the digital processing chip does not integrate a memory, it can be connected to an external memory via a communication interface. The digital processing chip implements the actions performed by the image processing device in the foregoing embodiments according to the program code stored in the external memory.

[0206] This application also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned actions. Figure 2 or Figure 9 The method steps described in the illustrated embodiment.

[0207] The image processing apparatus or image processing device provided in this application embodiment can be a chip, which includes a processing unit and a communication unit. The processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, pins, or circuits. The processing unit can execute computer execution instructions stored in the storage unit to cause the chip in the server to perform the above-mentioned operations. Figure 2 or Figure 9 The method described in the illustrated embodiment. Optionally, the storage unit is a storage unit within the chip, such as a register, cache, etc. The storage unit can also be a storage unit located outside the chip within the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, random access memory (RAM), etc.

[0208] Specifically, the aforementioned processing unit or processor can be a central processing unit (CPU), a neural-network processing unit (NPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0209] For example, please refer to Figure 14 , Figure 14 This is a schematic diagram of a chip provided in an embodiment of this application. The chip can be represented as a neural network processor (NPU) 1400. The NPU 1400 is mounted as a coprocessor on the host CPU, and tasks are assigned by the host CPU. The core part of the NPU is the arithmetic circuit 1403, which is controlled by a controller 1404 to retrieve matrix data from the memory and perform multiplication operations.

[0210] In some implementations, the arithmetic circuit 1403 internally includes multiple process engines (PEs). In some implementations, the arithmetic circuit 1403 is a two-dimensional pulsating array. The arithmetic circuit 1403 can also be a one-dimensional pulsating array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1403 is a general-purpose matrix processor.

[0211] For example, suppose we have an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from the weight memory 1402 and caches it in each PE of the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from the input memory 1401 and performs matrix operations with matrix B. The partial result or the final result of the obtained matrix is ​​stored in the accumulator 1408.

[0212] Unified memory 1406 is used to store input and output data. Weight data is directly transferred to weight memory 1402 via direct memory access controller (DMAC) 1405. Input data is also transferred to unified memory 1406 via DMAC.

[0213] The bus interface unit (BIU) 1410 is used for interaction between the AXI bus and the DMAC and instruction fetch buffer (IFB) 1409. It is used for the instruction fetch buffer 1409 to fetch instructions from external memory, and also for the memory access controller 1405 to fetch the original data of the input matrix A or the weight matrix B from external memory.

[0214] The DMAC is mainly used to move input data from external memory DDR to unified memory 1406, or to weight data to weight memory 1402, or to input data to input memory 1401.

[0215] The vector computation unit 1407 includes multiple arithmetic processing units that further process the output of the computation circuit as needed, such as vector multiplication, vector addition, exponential operations, logarithmic operations, size comparisons, etc. It is mainly used for computation in non-convolutional / fully connected layers of neural networks, such as batch normalization, pixel-level summation, and upsampling of feature planes.

[0216] In some implementations, the vector computation unit 1407 can store the processed output vector in the unified memory 1406. For example, the vector computation unit 1407 can apply linear and / or nonlinear functions to the output of the computation circuit 1403, such as performing linear interpolation on feature planes extracted by convolutional layers, or accumulating a vector of values ​​to generate activation values. In some implementations, the vector computation unit 1407 generates normalized values, pixel-level summed values, or both. In some implementations, the processed output vector can be used as activation input to the computation circuit 1403, for example, for use in subsequent layers of the neural network.

[0217] The instruction fetch buffer 1409 connected to the controller 1404 is used to store the instructions used by the controller 1404;

[0218] Unified memory 1406, input memory 1401, weighted memory 1402, and instruction fetch memory 1409 are all on-chip memories. External memory is proprietary to this NPU hardware architecture.

[0219] The processor mentioned above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits used to control the above. Figure 2 or Figure 9 The procedure of the method.

[0220] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

[0221] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, the specific working process of the systems, devices and units described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0222] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0223] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0224] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0225] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0226] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0227] Finally, it should be noted that the above are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of this application.

Claims

1. An image processing method, characterized in that, The method includes: Acquire a first image and a second image, wherein the first image contains a first face and the second image is an image of a 3D face model with target attributes, the target attributes including: lighting, facial expression or facial orientation; Based on the first image and the second image, an encoding result is obtained by an encoder, and the encoding result incorporates features from the first image and the second image. Based on the encoding result, a third image is obtained through a generation model. The third image contains a second face with the target attribute, and the second face corresponds to the same person as the first face.

2. The image processing method according to claim 1, characterized in that, The method further includes: A first result is obtained based on the second image and the noisy image, and the first result incorporates features from the second image and the noisy image. The step of obtaining the encoding result by an encoder based on the first image and the second image includes: Based on the first image and the first result, the encoder obtains the encoding result, which incorporates features of the first result and the first image.

3. The image processing method according to claim 2, characterized in that, The noisy image includes: a completely noisy image or a fourth image, wherein the fourth image is obtained by adding noise to the fifth image, and the fifth image is an image obtained in the previous round through the generation model.

4. The image processing method according to any one of claims 1 to 3, characterized in that, The first image includes a face image and a face warp image. The face image is obtained by extracting the face from a sixth image, which is an image containing the first face. The face warp image is obtained by distorting the sixth image and has the target attribute.

5. The image processing method according to claim 4, characterized in that, The encoding result includes: a first encoding result and a second encoding result. The step of obtaining the encoding result through an encoder based on the first image and the second image includes: Based on the face image and the second image, a first encoding result is obtained through the encoder, wherein the first encoding result incorporates features from the face image and the second image; Based on the face warp image and the second image, a second encoding result is obtained by the encoder, and the second encoding result incorporates features from the face image and the second image.

6. The image processing method according to claim 5, characterized in that, The step of obtaining the third image through a generation model based on the encoding result includes: Based on the first encoding result, a first output result is obtained through the generation model; Based on the second encoding result, a second output result is obtained through the generation model; The first output result and the second output result are summed to obtain the third image.

7. The image processing method according to any one of claims 1 to 6, characterized in that, The first face includes a third face and a fourth face; the first image includes a seventh image containing the third face and an eighth image containing the fourth face; the second image includes a ninth image for controlling the attributes of the seventh image and a tenth image for controlling the attributes of the eighth image; the encoding result includes a third encoding result and a fourth encoding result; the method further includes: Receive a first instruction, the first instruction being used to instruct the generation of a group photo including the third face and the fourth face; The step of obtaining the encoding result by an encoder based on the first image and the second image includes: Based on the seventh image and the ninth image, the encoder obtains the third encoding result, which incorporates features from the seventh image and the ninth image. Based on the eighth image and the tenth image, the encoder obtains the fourth encoding result, which incorporates features from the eighth image and the tenth image.

8. The image processing method according to any one of claims 1 to 7, characterized in that, The attributes of the face in the first image are different from the target attributes, and the first face corresponds to the same object as the face in the second image.

9. An image processing apparatus, characterized in that, include: The acquisition module is used to acquire a first image and a second image. The first image contains a first face, and the second image is an image of a 3D face model with target attributes, including: lighting, facial expression, or facial orientation. The processing module is configured to obtain an encoding result by an encoder based on the first image and the second image, the encoding result being a fusion of features from the first image and the second image; and to obtain a third image by a generation model based on the encoding result, the third image containing a second face having the target attribute, the second face corresponding to the first face being the same person.

10. The image processing apparatus according to claim 9, characterized in that, The acquisition module is further configured to: acquire a first result, wherein the first result is obtained based on the second image and the noisy image, and the first result incorporates features of the second image and the noisy image; The processing module is specifically used to: obtain the encoding result through the encoder based on the first image and the first result, wherein the encoding result incorporates features of the first result and the first image.

11. The image processing apparatus according to claim 10, characterized in that, The noisy image includes: a completely noisy image or a fourth image, wherein the fourth image is obtained by adding noise to the fifth image, and the fifth image is an image obtained in the previous round through the generation model.

12. The image processing apparatus according to any one of claims 9 to 11, characterized in that, The first image includes a face image and a face warp image. The face image is obtained by extracting the face from a sixth image, which is an image containing the first face. The face warp image is obtained by distorting the sixth image and has the target attribute.

13. The image processing apparatus according to claim 12, characterized in that, The encoding result includes: a first encoding result and a second encoding result, and the processing module is specifically used for: Based on the face image and the second image, a first encoding result is obtained through the encoder, wherein the first encoding result incorporates features from the face image and the second image; Based on the face warp image and the second image, a second encoding result is obtained by the encoder, and the second encoding result incorporates features from the face image and the second image.

14. The image processing apparatus according to claim 12, characterized in that, The processing module is specifically used for: Based on the first encoding result, a first output result is obtained through the generation model; Based on the second encoding result, a second output result is obtained through the generation model; The first output result and the second output result are summed to obtain the third image.

15. The image processing apparatus according to any one of claims 9 to 14, characterized in that, The first face includes a third face and a fourth face; the first image includes a seventh image containing the third face and an eighth image containing the fourth face; the second image includes a ninth image for controlling the attributes of the seventh image and a tenth image for controlling the attributes of the eighth image; the encoding result includes a third encoding result and a fourth encoding result; the acquisition module is further configured to receive a first instruction, the first instruction being used to instruct the generation of a group photo containing the third face and the fourth face; The processing module is specifically used to: obtain the third encoding result by the encoder based on the seventh image and the ninth image, wherein the third encoding result incorporates features of the seventh image and the ninth image; and obtain the fourth encoding result by the encoder based on the eighth image and the tenth image, wherein the fourth encoding result incorporates features of the eighth image and the tenth image.

16. The image processing apparatus according to any one of claims 9 to 15, characterized in that, The attributes of the face in the first image are different from the target attributes, and the first face corresponds to the same object as the face in the second image.

17. A communication device, characterized in that, include: Communication interface and processor; The communication interface and the processor perform the method as described in any one of claims 1 to 8.

18. A computer-readable storage medium, characterized in that, The medium stores instructions that, when executed by a processor, implement the method of any one of claims 1 to 8.

19. A computer program product, characterized in that, Includes instructions that, when executed on a processor, perform the method as described in any one of claims 1 to 8.

20. A chip, characterized in that, It includes at least one processing unit and an interface circuit, the interface circuit being used to provide program instructions or data to the at least one processing unit, the at least one processing unit being used to execute the program instructions to implement the method of any one of claims 1 to 8.