Image editing method and device and electronic equipment

By constructing an image library and training a model based on it, and using real image data as a reference, the problems of uncontrollability and unrealistic image editing in existing technologies are solved, thereby improving the realism and controllability of image editing.

CN120976008APending Publication Date: 2025-11-18HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410612853.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-16
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing image editing methods based on stable diffusion models employ a free-associative content generation strategy, resulting in uncontrollable and unrealistic images after processing.

Method used

By constructing an image library and training a model based on it, and using real image data as a reference, image editing is performed to ensure that the edited images are consistent with images in the same category of the image library. Sub-models are used to refine and expand the images, thereby improving the realism and controllability of image processing.

Benefits of technology

It improves the realism and controllability of edited images, enhancing the effectiveness of image editing and creation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976008A_ABST
    Figure CN120976008A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image editing method, and the method comprises the steps: responding to an operation for triggering the editing of a first image, and determining a target category which is a category to which the first image belongs; determining whether an image of the target category exists in a first image library or not, wherein the first image library comprises one or more groups of images in one-to-one correspondence with the one or more categories; when it is determined that the image of the target category exists in the first image library, the first image is edited through a first model, and the first model is obtained through training based on the first image library. Through the method, when image editing is performed on the to-be-processed image, the real image data of the same category as the to-be-processed image can be taken as a reference, and the authenticity and controllability of the edited image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image editing, and more specifically, to an image editing method, apparatus, and electronic device. Background Technology

[0002] Based on the stable diffusion model, intelligent image expansion, object removal, and facial detail retouching can be achieved. However, the current image editing and creation based on the stable diffusion model essentially adopts a free-associative content generation strategy. The expanded content is generated purely by model guessing, and the processed images often have problems such as being uncontrollable and unrealistic. Summary of the Invention

[0003] This application provides a method, apparatus, and electronic device for image editing. When editing an image, this method, apparatus, and electronic device can use real image data of the same category as the image to be processed as a reference. Compared with existing image editing methods (which use a free-associative content generation strategy and do not refer to image library data), this method can improve the realism and controllability of the edited image and play an important role in the fields of image editing and content creation.

[0004] In a first aspect, an image editing method is provided, the method comprising: in response to an operation that triggers editing of a first image, determining a target category, the target category being the category to which the first image belongs; determining whether an image of the target category exists in a first image library, the first image library including one or more sets of images, the one or more sets of images corresponding one-to-one with one or more categories; when it is determined that an image of the target category exists in the first image library, editing the first image using a first model, the first model being trained based on the first image library.

[0005] In some embodiments, the first image library includes a face image library and / or a scene image library. The face image library includes one or more sets of face images, each set of face images corresponding to one or more face categories. When the face image library includes multiple sets of face images, any two sets of face images in the multiple sets of face images correspond to different face categories. The scene image library includes one or more sets of scene images, each set of scene images corresponding to one or more scene categories. When the scene image library includes multiple sets of scene images, any two sets of scene images in the multiple sets of scene images correspond to different scene categories.

[0006] When the first image library includes multiple sets of images, any two sets of images in the multiple sets of images correspond to different categories.

[0007] It is understood that the image editing described in this application embodiment can be editing the existing content in the image (e.g., refining the face in the image), or it can be expanding the image. The expanded content can be content that is not in the original image. Users can achieve image creation through the image editing method provided in this application.

[0008] In some embodiments, one or more face categories correspond one-to-one with one or more users, that is, one user corresponds to one face category (i.e., one user corresponds to one model), for example: Zhang San corresponds to one face category, and Li Si corresponds to another face category.

[0009] In some other embodiments, a user's pre-makeup image may correspond to one face category, and the user's post-makeup image may correspond to another face category. For example, Zhang San's pre-makeup image may correspond to one face category, and this face category may correspond to a model used to edit Zhang San's pre-makeup image. Zhang San's post-makeup image may correspond to another face category, and this face category may correspond to a model used to edit Zhang San's post-makeup image.

[0010] In this embodiment, when editing an image, real image data of the same category as the image to be processed can be used as a reference. Compared with existing image editing methods (which use a free-associative content generation strategy and do not refer to image library data), this method can improve the authenticity and controllability of the edited image and play an important role in the fields of image editing and content creation.

[0011] In addition, before processing the image to be processed, it is first determined whether the category to which the image to be processed belongs exists in the image library used to train the image processing model. If it exists, it means that the category to which the image to be processed belongs has completed data construction and feature model pre-training, and can support the editing process. Only then will the image to be processed be further processed through the image processing model, which can further improve the realism and controllability of the image after editing.

[0012] In conjunction with the first aspect, in one possible implementation, the first model is trained based on images of the target category.

[0013] In conjunction with the first aspect, in one possible implementation, the first model includes one or more sub-models, one of which is trained based on images of the target category, and the one or more sub-models correspond one-to-one with the one or more categories.

[0014] In conjunction with the first aspect, in one possible implementation, editing the first image through the first model includes: editing the first image through a target sub-model, which is trained based on one or more images of the target category existing in the first image library, and the target sub-model is integrated into the first model.

[0015] In this embodiment, the model is trained based on images under each category in the image library to obtain a sub-model specifically for processing images of the corresponding category. The final inference model includes multiple sub-models, which correspond one-to-one with multiple categories. Taking the face category (e.g., the faces of the same user are one category) as an example, this means that multiple images of the same face are processed by a dedicated sub-model corresponding to that face. In this way, when an image needs to be processed, it will be processed by the sub-model corresponding to the category of the image, which can improve the realism and controllability of image processing.

[0016] In conjunction with the first aspect, in one possible implementation, the target category includes a first target category and a second target category. Editing the first image using a first model includes: editing the first image using a first target sub-model and a second target sub-model. The first target sub-model is trained based on one or more images of the first target category existing in the first image library, and the second target sub-model is trained based on one or more images of the second target category existing in the first image library. Both the first target sub-model and the second target sub-model are integrated into the first model.

[0017] In some embodiments, both the first target category and the second target category are face categories, such as the faces of different users, or the faces of the same user before and after makeup.

[0018] In some embodiments, both the first target category and the second target category are scene categories, such as different scenes.

[0019] In some embodiments, the first target category and the second target category can be a mixture of face category and scene category.

[0020] It should be noted that the number of target categories corresponding to the first image is not limited; it can be one or more, for example, it can include the faces of multiple users, and / or include multiple scenarios.

[0021] In this embodiment of the application, when there are multiple target categories in the image to be processed, multiple sub-models corresponding one-to-one with the multiple target categories can be mixed for image editing processing, which makes the application scope of the solution of this application wider and the realism and controllability of image processing better.

[0022] In conjunction with the first aspect, in one possible implementation, editing the first image through the first model includes: when the first image contains a first object and the first image library includes the category to which the first object belongs, generating a first mask image corresponding to the first image based on the area to be processed in the first image, in which the area to be processed is masked; inputting the first mask image into the first model so that the first model outputs the first image after image editing.

[0023] In one implementation, the first object includes a first face, the area to be processed in the first image includes a first area to be processed, the first area to be processed includes a partial or complete area of ​​the first face, and / or the first area to be processed includes an area to be expanded of the first face, and the first image after image editing is the first image after the first face is refined and / or the first image after the first face is expanded.

[0024] In this embodiment of the application, it is possible to use real personal face image data from the image library as reference content to edit face images belonging to the same category. For example, using the real face image of the first user as a reference, the face image of the first user to be processed can be processed, which can improve the authenticity and controllability of the edited face image.

[0025] In another implementation, the first object includes a first scene, the area to be processed of the first image includes a second area to be processed, the second area to be processed includes the area to be expanded of the first scene, and the first image after image editing is the first image after the first scene is expanded.

[0026] In some embodiments, when the area to be processed of the first image includes only the first area to be processed, the first image after image editing is the first image after the first face is refined and / or the first image after the first face is expanded.

[0027] In some other embodiments, when the area to be processed of the first image only includes the second area to be processed, the first image after image editing is the first image after the first scene is expanded;

[0028] In some other embodiments, when the area to be processed of the first image includes the first area to be processed and the second area to be processed, the first image after image editing is the first face refinement and / or the first face expansion, and the first image after the first scene expansion.

[0029] In this embodiment, the scene of the image to be processed can be expanded using real scene image data from the image library as reference content. For example, the local first scene in the image to be processed can be expanded using the first scene image as reference, so that the expanded scenes are all real scenes, which can improve the authenticity and controllability of the expanded scene image.

[0030] In conjunction with the first aspect, in one possible implementation, before generating a first mask image corresponding to the first image based on the region to be processed of the first image, the method further includes: determining a defect region of the first image by performing face detection on the first image; determining the defect region of the first image as the region to be processed of the first image; and / or determining the region to be processed of the first image in response to a first user interaction operation.

[0031] In this embodiment of the application, before processing the image to be processed, the area to be processed of the image can be determined by automatic detection, or the area to be processed of the image can be determined based on the user's interactive operation (e.g., drawing circles, smearing, zooming in, zooming out, etc.), so that the image to be processed can be precisely edited. For example, if the area to be processed determined by the user includes the mouth of the first user, then the first model uses the sub-model corresponding to the first user to refine the mouth of the first user.

[0032] In conjunction with the first aspect, in one possible implementation, the method further includes: clustering images containing faces in the image library to obtain a second image library; performing face detection on the images in the second image library to filter out images in the second image library whose face resolution is less than a first resolution threshold or whose face occlusion area ratio is greater than a first ratio; and performing face segmentation on the images in the second image library after the filtering process to generate the face image library.

[0033] In this embodiment, a face image library can be constructed by performing face clustering on real face images in the image library. Furthermore, low-quality images in the face image library will be filtered out, which makes the image editing capability of the model trained by the face image library stronger, that is, it can further improve the realism and controllability of the edited image.

[0034] In conjunction with the first aspect, in one possible implementation, the method further includes: obtaining a third image library by clustering scene images in the image library; and generating the scene image library by filtering out images in the third image library whose resolution is less than a second resolution threshold.

[0035] In some embodiments, after filtering out images in the third image library whose resolution is less than a second resolution threshold, the scene image library is generated by performing data augmentation on the images in the third image library after the filtering process.

[0036] In this embodiment, scene image library can be constructed by performing scene clustering on real scene images in the image library. It can also filter out low-quality images in the scene image library and perform data augmentation on the images in the scene image library. This makes the image editing capability of the model trained by the scene image library stronger, that is, it can further improve the realism and controllability of the edited image.

[0037] In conjunction with the first aspect, in one possible implementation, the method further includes: sending to a server the features of one or more images in the i-th category of the first image library and the features of one or more masked images obtained by masking one or more images in the i-th category, where i is a natural number, i = 1, 2, 3, ...; receiving a first model sent by the server, wherein the first model is trained based on the features of one or more images in the i-th category and the features of one or more masked images, the i-th category being the target category, or the i-th sub-model included in the first model is trained based on the features of one or more images in the i-th category and the features of one or more masked images.

[0038] In some embodiments, after the image of the i-th category in the first image library is updated, the features of one or more images of the updated i-th category and the features of one or more masked images obtained by masking the one or more images of the updated i-th category can also be sent to the server; receiving the updated first model sent by the server can also be receiving the updated i-th sub-model sent by the server; or receiving the update data related to the i-th sub-model sent by the server, wherein the update data related to the i-th sub-model refers to the update data relative to the data of the i-th sub-model before the update.

[0039] In this embodiment, the model is trained based on images under each category in the image library to obtain a sub-model specifically for processing the image to be processed in the corresponding category. The final inference model includes multiple sub-models, which correspond one-to-one with multiple categories. When an image to be processed needs to be processed, the image to be processed will be processed by the sub-model corresponding to the category of the image to be processed, which can improve the realism and controllability of image processing.

[0040] In conjunction with the first aspect, in one possible implementation, the method further includes: using one or more masked images obtained by masking one or more images under the i-th category in the first image library and one or more images under the i-th category as training samples to train the model, thereby obtaining the i-th sub-model. The i-th sub-model is used to process the images to be processed belonging to the i-th category, where i is a natural number, i = 1, 2, 3, ..., and the i-th sub-model is the first model, the i-th category is the target category, or the first model includes the i-th sub-model.

[0041] In this embodiment, the model is trained based on images under each category in the image library to obtain a sub-model specifically for processing the image to be processed in the corresponding category. The final inference model includes multiple sub-models, which correspond one-to-one with multiple categories. When an image to be processed needs to be processed, the image to be processed will be processed by the sub-model corresponding to the category of the image to be processed, which can improve the realism and controllability of image processing.

[0042] In conjunction with the first aspect, in one possible implementation, the one or more masked images are obtained by randomly masking one or more images under the i-th category.

[0043] In this embodiment, the one or more masked images are obtained by randomly masking one or more images under the i-th category, which results in better performance of the trained model.

[0044] Secondly, an image editing method is provided, comprising: training a model using one or more masked images obtained by masking one or more images of the i-th category in a first image library and one or more images of the i-th category as training samples to obtain an i-th sub-model, the i-th sub-model being used to process images belonging to the i-th category to be processed, wherein i is a natural number, i = 1, 2, 3, ..., wherein the first image library includes one or more sets of images, the one or more sets of images corresponding one-to-one with one or more categories.

[0045] In some embodiments, the first image library includes a face image library and / or a scene image library, wherein the face image library includes one or more sets of face images, each set of face images corresponding to one or more face categories, and the scene image library includes one or more sets of scene images, each set of scene images corresponding to one or more scene categories.

[0046] In this embodiment, the model is trained based on images under each category in the image library to obtain a sub-model specifically for processing the image to be processed in the corresponding category. The final inference model includes multiple sub-models, which correspond one-to-one with multiple categories. When an image to be processed needs to be processed, the image to be processed will be processed by the sub-model corresponding to the category of the image to be processed, which can improve the realism and controllability of image processing.

[0047] In conjunction with the second aspect, in one possible implementation, the method further includes: in response to triggering an operation to edit the first image, determining a target category, which is the category to which the first image belongs; determining whether an image of the target category exists in the first image library; and when it is determined that an image of the target category exists in the first image library, editing the first image through a sub-model corresponding to the target category.

[0048] In this embodiment of the application, when creating image content, real personal image data from the image library can be used as reference content. Compared with existing image editing methods (which use a free-associative content generation strategy and do not refer to image library data), this method can improve the authenticity and controllability of the edited image and play an important role in the fields of image editing and content creation.

[0049] In addition, before processing the image to be processed, it is first determined whether the category to which the image to be processed belongs exists in the image library used to train the image processing model. If it exists, it means that the category to which the image to be processed belongs has completed data construction and feature model pre-training, and can support the editing process. Only then will the image to be processed be further processed through the image processing model, which can further improve the realism and controllability of the image after editing.

[0050] Thirdly, an image editing method is provided, comprising: receiving features of one or more images in an i-th category from a first image library sent by a first device, and features of one or more masked images obtained by masking one or more images in the i-th category, wherein the first image library includes one or more sets of images, each set of images corresponding to one or more categories, i being a natural number, i = 1, 2, 3, ...; training a model based on the features of one or more images in the i-th category and the features of the one or more masked images to obtain an i-th sub-model, the i-th sub-model being used to process images belonging to the i-th category; and sending a first model to the first device, wherein the first model is the i-th sub-model, the i-th category is the target category, or the first model includes the i-th sub-model.

[0051] In some embodiments, after the image of the i-th category in the first image library is updated, the server can receive the features of one or more images of the updated i-th category and the features of one or more masked images obtained by masking the one or more images of the updated i-th category; perform model training based on the features of the one or more images of the updated i-th category and the features of the one or more masked images obtained by masking the one or more images of the updated i-th category to obtain an updated i-th sub-model; the server can send the updated first model to the first device, or send the updated i-th sub-model included in the first model to the first device; or send update data related to the i-th sub-model to the first device, wherein the update data related to the i-th sub-model refers to the update data relative to the data of the i-th sub-model before the update.

[0052] In some embodiments, the first image library includes a face image library and / or a scene image library, wherein the face image library includes one or more sets of face images, each set of face images corresponding to one or more face categories, and the scene image library includes one or more sets of scene images, each set of scene images corresponding to one or more scene categories.

[0053] In this embodiment, the model is trained based on images under each category in the image library to obtain a sub-model specifically for processing the image to be processed in the corresponding category. The final inference model includes multiple sub-models, which correspond one-to-one with multiple categories. When an image to be processed needs to be processed, the image to be processed will be processed by the sub-model corresponding to the category of the image to be processed, which can improve the realism and controllability of image processing.

[0054] In conjunction with the third aspect, in one possible implementation, the method further includes: in response to triggering an operation to edit the first image, determining a target category, which is the category to which the first image belongs; determining whether an image of the target category exists in the first image library; and when it is determined that an image of the target category exists in the first image library, editing the first image through the sub-model corresponding to the target category.

[0055] In this embodiment of the application, when creating image content, real personal image data from the image library can be used as reference content. Compared with existing image editing methods (which use a free-associative content generation strategy and do not refer to image library data), this method can improve the authenticity and controllability of the edited image and play an important role in the fields of image editing and content creation.

[0056] In addition, before processing the image to be processed, it is first determined whether the category to which the image to be processed belongs exists in the image library used to train the image processing model. If it exists, it means that the category to which the image to be processed belongs has completed data construction and feature model pre-training, and can support the editing process. Only then will the image to be processed be further processed through the image processing model, which can further improve the realism and controllability of the image after editing.

[0057] Fourthly, an image editing apparatus is provided, comprising: a determining module, configured to determine a target category in response to an operation triggering the editing of a first image, the target category being the category to which the first image belongs; the determining module is further configured to determine whether an image of the target category exists in a first image library, the first image library including one or more sets of images, the one or more sets of images corresponding one-to-one with one or more categories; and an inference module, configured to edit the first image using a first model when it is determined that an image of the target category exists in the first image library, the first model being trained based on the first image library.

[0058] In some embodiments, the first image library includes a face image library and / or a scene image library. The face image library includes one or more sets of face images, each set of face images corresponding to one or more face categories. When the face image library includes multiple sets of face images, any two sets of face images in the multiple sets of face images correspond to different face categories. The scene image library includes one or more sets of scene images, each set of scene images corresponding to one or more scene categories. When the scene image library includes multiple sets of scene images, any two sets of scene images in the multiple sets of scene images correspond to different scene categories.

[0059] When the first image library includes multiple sets of images, any two sets of images in the multiple sets of images correspond to different categories.

[0060] It is understood that the image editing described in this application embodiment can be editing the existing content in the image (e.g., refining the face in the image), or it can be expanding the image. The expanded content can be content that is not in the original image. Users can create images using the image editing device provided in this application.

[0061] In some embodiments, one or more face categories correspond one-to-one with one or more users, that is, one user corresponds to one face category (i.e., one user corresponds to one model), for example: Zhang San corresponds to one face category, and Li Si corresponds to another face category.

[0062] In some other embodiments, a user's pre-makeup image may correspond to one face category, and the user's post-makeup image may correspond to another face category. For example, Zhang San's pre-makeup image may correspond to one face category, and this face category may correspond to a model used to edit Zhang San's pre-makeup image. Zhang San's post-makeup image may correspond to another face category, and this face category may correspond to a model used to edit Zhang San's post-makeup image.

[0063] In this embodiment of the application, when creating image content, real personal image data from the image library can be used as reference content. Compared with existing image editing devices (which use a free-associative content generation strategy and do not refer to image library data), this can improve the authenticity and controllability of the edited image and play an important role in the fields of image editing and content creation.

[0064] In addition, before processing the image to be processed, it is first determined whether the category to which the image to be processed belongs exists in the image library used to train the image processing model. If it exists, it means that the category to which the image to be processed belongs has completed data construction and feature model pre-training, and can support the editing process. Only then will the image to be processed be further processed through the image processing model, which can further improve the realism and controllability of the image after editing.

[0065] In conjunction with the fourth aspect, in one possible implementation, the first model is trained based on images of the target category.

[0066] In conjunction with the fourth aspect, in one possible implementation, the first model includes one or more sub-models, one of which is trained based on images of the target category, and the one or more sub-models correspond one-to-one with the one or more categories.

[0067] In conjunction with the fourth aspect, in one possible implementation, the inference module is specifically used to: edit the first image through a target sub-model, which is trained based on one or more images of the target category existing in the first image library, and the target sub-model is integrated into the first model.

[0068] In this embodiment, the model is trained based on images under each category in the image library to obtain a sub-model specifically for processing images of the corresponding category. The final inference model includes multiple sub-models, which correspond one-to-one with multiple categories. Taking the face category (e.g., the faces of the same user are one category) as an example, this means that multiple images of the same face are processed by a dedicated sub-model corresponding to that face. In this way, when an image needs to be processed, it will be processed by the sub-model corresponding to the category of the image, which can improve the realism and controllability of image processing.

[0069] In conjunction with the fourth aspect, in one possible implementation, the target category includes a first target category and a second target category. The inference module is specifically used to: edit the first image through a first target sub-model and a second target sub-model. The first target sub-model is trained based on one or more images of the first target category existing in the first image library, and the second target sub-model is trained based on one or more images of the second target category existing in the first image library. Both the first target sub-model and the second target sub-model are integrated into the first model.

[0070] In some embodiments, both the first target category and the second target category are face categories, such as the faces of different users, or the faces of the same user before and after makeup.

[0071] In some embodiments, both the first target category and the second target category are scene categories, such as different scenes.

[0072] In some embodiments, the first target category and the second target category can be a mixture of face category and scene category.

[0073] It should be noted that the number of target categories corresponding to the first image is not limited; it can be one or more, for example, it can include the faces of multiple users, and / or include multiple scenarios.

[0074] In this embodiment of the application, when there are multiple target categories in the image to be processed, multiple sub-models corresponding one-to-one with the multiple target categories can be mixed for image editing processing, which makes the application scope of the solution of this application wider and the realism and controllability of image processing better.

[0075] In conjunction with the fourth aspect, in one possible implementation, the inference module is specifically used to: when the first image contains a first object and the first image library includes the category to which the first object belongs, generate a first mask image corresponding to the first image based on the area to be processed in the first image, in which the area to be processed is masked; input the first mask image into the first model so that the first model outputs the first image after image editing.

[0076] In one implementation, the first object includes a first face, the area to be processed in the first image includes a first area to be processed, the first area to be processed includes a partial or complete area of ​​the first face, and / or the first area to be processed includes an area to be expanded of the first face, and the first image after image editing is the first image after the first face is refined and / or the first image after the first face is expanded.

[0077] In this embodiment of the application, it is possible to use real personal face image data from the image library as reference content to edit face images belonging to the same category. For example, using the real face image of the first user as a reference, the face image of the first user to be processed can be processed, which can improve the authenticity and controllability of the edited face image.

[0078] In another implementation, the first object includes a first scene, the area to be processed of the first image includes a second area to be processed, the second area to be processed includes the area to be expanded of the first scene, and the first image after image editing is the first image after the first scene is expanded.

[0079] In some embodiments, when the area to be processed of the first image includes only the first area to be processed, the first image after image editing is the first image after the first face is refined and / or the first image after the first face is expanded.

[0080] In some other embodiments, when the area to be processed of the first image only includes the second area to be processed, the first image after image editing is the first image after the first scene is expanded;

[0081] In some other embodiments, when the area to be processed of the first image includes the first area to be processed and the second area to be processed, the first image after image editing is the first face refinement and / or the first face expansion, and the first image after the first scene expansion.

[0082] In this embodiment, the scene of the image to be processed can be expanded using real scene image data from the image library as reference content. For example, the local first scene in the image to be processed can be expanded using the first scene image as reference, so that the expanded scenes are all real scenes, which can improve the authenticity and controllability of the expanded scene image.

[0083] In conjunction with the fourth aspect, in one possible implementation, the determining module is further configured to: determine the defective region of the first image by performing face detection on the first image; determine the defective region of the first image as the region to be processed of the first image; and / or determine the region to be processed of the first image in response to a first user interaction operation.

[0084] In this embodiment of the application, before processing the image to be processed, the area to be processed of the image can be determined by automatic detection, or the area to be processed of the image can be determined based on the user's interactive operation (e.g., drawing circles, smearing, zooming in, zooming out, etc.), so that the image to be processed can be precisely edited. For example, if the area to be processed determined by the user includes the mouth of the first user, then the first model uses the sub-model corresponding to the first user to refine the mouth of the first user.

[0085] In conjunction with the fourth aspect, in one possible implementation, the device further includes: an image library construction module, used to construct the first image library by clustering images in the image library; specifically, the image library construction module is used to: obtain a second image library by clustering images containing faces in the image library; filter out images in the second image library whose face resolution is less than a first resolution threshold or whose face occlusion area ratio is greater than a first ratio by performing face detection on the images in the second image library; and generate the face image library by performing face segmentation on the images in the second image library after the filtering process.

[0086] In this embodiment, a face image library can be constructed by performing face clustering on real face images in the image library. Furthermore, low-quality images in the face image library will be filtered out, which makes the image editing capability of the model trained by the face image library stronger, that is, it can further improve the realism and controllability of the edited image.

[0087] In conjunction with the fourth aspect, in one possible implementation, the image library construction module is also specifically used to: obtain a third image library by clustering scene images in the image library; and generate the scene image library by filtering out images in the third image library whose resolution is less than a second resolution threshold.

[0088] In some embodiments, the inference module is specifically used to: filter out images in the third image library whose image resolution is less than a second resolution threshold; and generate the scene image library by performing data augmentation on the images in the third image library after the filtering process.

[0089] In this embodiment, scene image library can be constructed by performing scene clustering on real scene images in the image library. It can also filter out low-quality images in the scene image library and perform data augmentation on the images in the scene image library. This makes the image editing capability of the model trained by the scene image library stronger, that is, it can further improve the realism and controllability of the edited image.

[0090] In conjunction with the fourth aspect, in one possible implementation, the device further includes: a transceiver module, configured to send to the server the features of one or more images in the i-th category of the first image library and the features of one or more masked images obtained by masking one or more images in the i-th category, where i is a natural number, i = 1, 2, 3, ...; the transceiver module is further configured to receive a first model sent by the server, wherein the first model is trained based on the features of one or more images in the i-th category and the features of one or more masked images, the i-th category being the target category, or the i-th sub-model included in the first model is trained based on the features of one or more images in the i-th category and the features of one or more masked images.

[0091] In this embodiment, the model is trained based on images under each category in the image library to obtain a sub-model specifically for processing the image to be processed in the corresponding category. The final inference model includes multiple sub-models, which correspond one-to-one with multiple categories. When an image to be processed needs to be processed, the image to be processed will be processed by the sub-model corresponding to the category of the image to be processed, which can improve the realism and controllability of image processing.

[0092] In conjunction with the fourth aspect, in one possible implementation, the apparatus further includes: a training module, configured to train a model using one or more masked images obtained by masking one or more images of the i-th category in the first image library and one or more images of the i-th category as training samples to obtain an i-th sub-model, wherein the i-th sub-model is used to process images to be processed belonging to the i-th category, where i is a natural number, i = 1, 2, 3, ..., wherein the i-th sub-model is the first model, the i-th category is the target category, or the first model includes the i-th sub-model.

[0093] In this embodiment, the model is trained based on images under each category in the image library to obtain a sub-model specifically for processing the image to be processed in the corresponding category. The final inference model includes multiple sub-models, which correspond one-to-one with multiple categories. When an image to be processed needs to be processed, the image to be processed will be processed by the sub-model corresponding to the category of the image to be processed, which can improve the realism and controllability of image processing.

[0094] In conjunction with the fourth aspect, in one possible implementation, the one or more masked images are obtained by randomly masking one or more images under the i-th category.

[0095] In this embodiment, the one or more masked images are obtained by randomly masking one or more images under the i-th category, which results in better performance of the trained model.

[0096] Fifthly, an image editing apparatus is provided, comprising: a training module, configured to train a model using one or more masked images obtained by masking one or more images of the i-th category in a first image library and one or more images of the i-th category as training samples to obtain an i-th sub-model, the i-th sub-model being used to process images to be processed belonging to the i-th category, wherein i is a natural number, i = 1, 2, 3, ..., wherein the first image library includes one or more sets of images, the one or more sets of images corresponding one-to-one with one or more categories.

[0097] In some embodiments, the first image library includes a face image library and a scene image library, wherein the face image library includes one or more sets of face images, each set of face images corresponding to one or more face categories, and the scene image library includes one or more sets of scene images, each set of scene images corresponding to one or more scene categories.

[0098] In this embodiment, the model is trained based on images under each category in the image library to obtain a sub-model specifically for processing the image to be processed in the corresponding category. The final inference model includes multiple sub-models, which correspond one-to-one with multiple categories. When an image to be processed needs to be processed, the image to be processed will be processed by the sub-model corresponding to the category of the image to be processed, which can improve the realism and controllability of image processing.

[0099] In conjunction with the fifth aspect, in one possible implementation, the apparatus further includes: a determining module, configured to determine a target category in response to an operation that triggers editing of the first image, the target category being the category to which the first image belongs; the determining module is further configured to determine whether an image of the target category exists in the first image library; and an inference module, configured to edit the first image by means of a sub-model corresponding to the target category when it is determined that an image of the target category exists in the first image library.

[0100] In this embodiment of the application, when creating image content, real personal image data from the image library can be used as reference content. Compared with existing image editing devices (which use a free-associative content generation strategy and do not refer to image library data), this can improve the authenticity and controllability of the edited image and play an important role in the fields of image editing and content creation.

[0101] In addition, before processing the image to be processed, it is first determined whether the category to which the image to be processed belongs exists in the image library used to train the image processing model. If it exists, it means that the category to which the image to be processed belongs has completed data construction and feature model pre-training, and can support the editing process. Only then will the image to be processed be further processed through the image processing model, which can further improve the realism and controllability of the image after editing.

[0102] A sixth aspect provides an image editing apparatus, comprising: a transceiver module for receiving features of one or more images of an i-th category from a first image library sent by a first device, and features of one or more masked images obtained by masking the one or more images of the i-th category, wherein the first image library includes one or more sets of images, each set of images corresponding to one or more categories, i being a natural number, i = 1, 2, 3, ...; a training module for training a model based on the features of the one or more images of the i-th category and the features of the one or more masked images to obtain an i-th sub-model, wherein the i-th sub-model is used to process images to be processed belonging to the i-th category; the transceiver module is further configured to send a first model to the first device, wherein the first model is the i-th sub-model, the i-th category is the target category, or the first model includes the i-th sub-model.

[0103] In some embodiments, the first image library includes a face image library and / or a scene image library, wherein the face image library includes one or more sets of face images, each set of face images corresponding to one or more face categories, and the scene image library includes one or more sets of scene images, each set of scene images corresponding to one or more scene categories.

[0104] In this embodiment, the model is trained based on images under each category in the image library to obtain a sub-model specifically for processing the image to be processed in the corresponding category. The final inference model includes multiple sub-models, which correspond one-to-one with multiple categories. When an image to be processed needs to be processed, the image to be processed will be processed by the sub-model corresponding to the category of the image to be processed, which can improve the realism and controllability of image processing.

[0105] In a seventh aspect, an electronic device is provided, comprising a memory and a processor, wherein the memory is used to store computer program code, and the processor is used to execute the computer program code stored in the memory to implement the method in the first aspect or any possible implementation of the first aspect, or to implement the method in the second aspect or any possible implementation of the second aspect, or to implement the method in the third aspect or any possible implementation of the third aspect.

[0106] Eighthly, a computer-readable storage medium is provided, which stores a computer program or instructions that, when executed, implement the method of the first aspect or any possible implementation thereof, or implement the method of the second aspect or any possible implementation thereof, or implement the method of the third aspect or any possible implementation thereof.

[0107] Ninthly, a chip is provided, wherein instructions are stored therein, which, when executed on a device, cause the chip to perform the method of the first aspect or any possible implementation thereof, or to perform the method of the second aspect or any possible implementation thereof, or to perform the method of the third aspect or any possible implementation thereof.

[0108] In a tenth aspect, a computer program product is provided, which stores a computer program or instructions that, when executed, implement the method in the first aspect or any possible implementation of the first aspect, or implement the method in the second aspect or any possible implementation of the second aspect, or implement the method in the third aspect or any possible implementation of the third aspect. Attached Figure Description

[0109] Figure 1 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application;

[0110] Figure 2 This is a software structure block diagram of an electronic device provided in an embodiment of this application;

[0111] Figure 3 This is a schematic flowchart of an intelligent map expansion method;

[0112] Figure 4 This is a schematic flowchart illustrating an image editing method provided in an embodiment of this application;

[0113] Figure 5 This is a schematic flowchart illustrating a method for constructing a face image library according to an embodiment of this application;

[0114] Figure 6 This is a schematic flowchart illustrating a method for constructing a scene image library according to an embodiment of this application;

[0115] Figure 7 This is a schematic flowchart illustrating a method for training a database model provided in an embodiment of this application;

[0116] Figure 8 These are schematic diagrams of several masking images provided in the embodiments of this application;

[0117] Figure 9 This is a schematic flowchart illustrating an image editing method provided in an embodiment of this application;

[0118] Figure 10 This is a schematic diagram of an image editing process provided in an embodiment of this application;

[0119] Figure 11 This is a schematic diagram of the end-to-cloud interaction corresponding to an image editing process provided in an embodiment of this application;

[0120] Figure 12 This is a schematic diagram of the end-to-cloud interaction corresponding to another image editing process provided in the embodiments of this application;

[0121] Figure 13 This is a schematic diagram of the functional modules of an image editing device provided in an embodiment of this application. Detailed Implementation

[0122] The technical solutions of this application will now be described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments.

[0123] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; "and / or" in this text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "plural" or "multiple" refers to two or more than two.

[0124] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this embodiment, unless otherwise stated, "a plurality of" means two or more.

[0125] The terminology used in the following embodiments is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to also include expressions such as “one or more,” unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of this application, “at least one” and “one or more” refer to one, two, or more than two. The term “and / or” is used to describe the relationship between related objects, indicating that three relationships may exist; for example, A and / or B can indicate: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character “ / ” generally indicates that the preceding and following related objects are in an “or” relationship.

[0126] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "one embodiment," "some embodiments," "another embodiment," "other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0127] The method provided in this application can be applied to electronic devices with time display or time recognition functions, such as mobile phones, tablets, wearable devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), smart home devices, and other electronic devices. This application does not impose any restrictions on the specific type of electronic device.

[0128] For example, Figure 1A schematic diagram of the structure of electronic device 100 is shown. Electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, antenna 1, antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0129] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0130] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.

[0131] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of fetching and executing instructions.

[0132] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0133] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0134] USB port 130 is a USB standard compliant interface, specifically a Mini USB port, Micro USB port, USB Type-C port, etc. USB port 130 can be used to connect a charger to charge electronic device 100, and can also be used for data transfer between electronic device 100 and peripheral devices. It can also be used to connect headphones for audio playback. This interface can also be used to connect other electronic devices, such as AR devices.

[0135] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.

[0136] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device via the power management module 141.

[0137] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, external memory, display screen 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.

[0138] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.

[0139] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with tuning switches.

[0140] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.

[0141] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through an audio device (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 194. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 150 or other functional modules.

[0142] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0143] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling electronic device 100 to communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or satellite-based augmentation systems (SBAS).

[0144] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0145] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 100 may include one or N displays 194, where N is a positive integer greater than 1.

[0146] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.

[0147] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.

[0148] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0149] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.

[0150] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, electronic device 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0151] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.

[0152] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.

[0153] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.

[0154] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.

[0155] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.

[0156] The speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or make hands-free calls through the speaker 170A.

[0157] The receiver 170B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the electronic device 100 answers a telephone call or voice message, the receiver 170B can be brought close to the ear to listen to the voice.

[0158] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. Electronic device 100 may have at least one microphone 170C. In some embodiments, electronic device 100 may have two microphones 170C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, electronic device 100 may also have three, four, or more microphones 170C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.

[0159] The 170D headphone jack is used to connect wired headphones. The 170D headphone jack can be a USB 130 interface or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, a CTIA (Cellular Telecommunications Industry Association of the USA) standard interface.

[0160] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. Electronic device 100 can receive button input and generate key signal inputs related to user settings and function control of electronic device 100.

[0161] Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. For example, different vibration feedback effects can correspond to touch operations performed on different applications (such as taking photos, playing audio, etc.). Motor 191 can also correspond to different vibration feedback effects for touch operations performed on different areas of the display screen 194. Different application scenarios (such as time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also be customized.

[0162] Indicator 192 can be an indicator light, used to indicate charging status, power changes, or to indicate messages, missed calls, notifications, etc.

[0163] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to make contact with and separate from the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 simultaneously. The multiple cards can be of the same or different types. The SIM card interface 195 is also compatible with different types of SIM cards. The SIM card interface 195 is also compatible with external memory cards. The electronic device 100 interacts with the network through the SIM card to realize functions such as calls and data communication. In some embodiments, the electronic device 100 uses an embedded SIM (eSIM) card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100.

[0164] It should be understood that the phone cards in the embodiments of this application include, but are not limited to, SIM cards, eSIM cards, universal subscriber identity modules (USIM), universal integrated circuit cards (UICC), etc.

[0165] The software system of electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment uses the layered architecture Android system as an example to exemplify the software structure of electronic device 100.

[0166] Figure 2This is a software structure block diagram of an electronic device 100 according to an embodiment of this application. The layered architecture divides the software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer. The application layer may include a series of application packages.

[0167] like Figure 2 As shown, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and SMS.

[0168] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.

[0169] like Figure 2 As shown, the application framework layer may include a window manager, content provider, view system, phone manager, resource manager, notification manager, etc.

[0170] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.

[0171] Content providers store and retrieve data, making that data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.

[0172] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.

[0173] The phone manager is used to provide communication functions for electronic device 100. For example, it manages call status (including connection and disconnection).

[0174] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.

[0175] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.

[0176] The Android Runtime consists of core libraries and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.

[0177] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.

[0178] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0179] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.

[0180] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.

[0181] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.

[0182] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0183] A 2D graphics engine is a graphics engine for 2D drawing.

[0184] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.

[0185] It should be understood that the technical solutions in the embodiments of this application can be used in systems such as Android, iOS, and HarmonyOS.

[0186] The technical solutions of this application embodiment can be applied to image content editing and creation scenarios. For example, they can be applied to scenarios such as image aesthetic composition, intelligent image enlargement, image background editing, portrait retouching, and object removal in images.

[0187] Among them, electronic devices can be televisions, desktop computers, laptops, or portable electronic devices such as mobile phones, foldable screens, tablets, cameras, camcorders, and video recorders. They can also be smart home devices such as refrigerators, washing machines, robot vacuums, and any other electronic devices with image processing capabilities. Furthermore, they can be electronic devices in 5G networks or in future evolved public land mobile networks (PLMNs).

[0188] For example, Figure 3 A schematic flowchart of an intelligent map expansion method 300 is shown. Figure 3 As shown, the method 300 includes:

[0189] S301: When intelligent image augmentation is required on the original image, input the original image and the descriptive terms of the augmented content into the stable diffusion (SD) model.

[0190] In some embodiments, it is also possible to omit the descriptive terms for the expanded content.

[0191] S302: The stable diffusion model intelligently expands the original image by adding and removing noise based on the descriptive terms of the expanded content and its understanding of the original image.

[0192] In the process of intelligently expanding the original image, the central content of the image remains unchanged. That is, the original image is used as the central content, and the expansion is carried out based on the edges of the original image.

[0193] S303: The output of the stable diffusion model is the expanded original image.

[0194] This method, based on a stable diffusion model, can achieve AI-based image enlargement, object removal, and facial detail retouching. However, the editing and creation of images based on the stable diffusion model essentially adopts a free-associative content generation strategy. The expanded content is generated purely by model guessing. As a result, the processed images often have problems such as being uncontrollable and unrealistic. For example, the expanded content (e.g., clothing) may not be present in the image's subject. Another example is that the expanded image may present some illogical scenes (e.g., the person in the image has messy hands and feet, multiple hands and feet, or distorted face and hands).

[0195] In conclusion, the edited images obtained using current image editing methods are prone to appearing unrealistic and unreasonable, which seriously affects the user experience.

[0196] In view of this, embodiments of this application provide an image editing method, apparatus, and electronic device. In this method, when creating image content editing, real personal image data from a stock photo library can be used as reference content. Furthermore, the model used to process the image to be processed is a model that matches the identifier of the image to be processed. Compared with existing image editing methods (which use a free-associative content generation strategy and do not refer to stock photo library data), this method can improve the authenticity and controllability of the edited image and play an important role in the fields of image editing and content creation.

[0197] For example, Figure 4 A schematic flowchart of an image editing method 400 provided in an embodiment of this application is shown. Figure 4 As shown, the method 400 includes:

[0198] S401: In response to the first image uploaded by the user, determine the identifier (ID) corresponding to the first image, and denot it as the first ID.

[0199] In some embodiments, the identifier corresponding to the first image may include an identifier of a face image contained in the first image, and may also include an identifier of a scene image contained in the first image.

[0200] When a user needs to process the first image, the user uploads the first image.

[0201] S402: Determine whether an image with the ID "first ID" exists in the first image library. If an image with the ID "first ID" exists in the first image library, then proceed to S403.

[0202] Specifically, determining whether an image with ID 1 exists in the first image library can be done as follows: when the first ID is a face ID, determine whether a face image with ID 1 exists in the first image library; when the first ID is a scene ID, determine whether a scene image with ID 1 exists in the first image library.

[0203] In some embodiments, if an image with the ID of the first ID does not exist in the first image library, the processing flow of the first image is terminated.

[0204] In one example, if the first image contains a face and the ID of the face is A, then it is determined whether a face image with ID A exists in the first image library.

[0205] In another example, if the first image contains a scene and the ID of the contained scene is B, then it is determined whether a scene image with ID B exists in the first image library.

[0206] In another example, if the first image contains both a face and a scene, and the ID of the face is A and the ID of the scene is B, then it is determined whether there is a face image with ID A and a scene image with ID B in the first image library.

[0207] It should be understood that the number of face images contained in the first image can be one or more. When the number of face images contained in the first image is multiple, the ID of the first image is determined. Specifically, this can be done by determining the IDs of the multiple face images in the first image and determining the scene ID of the first image.

[0208] The first image library may include a face image library and / or a scene image library. The face image library includes one or more sets of images, each set of images corresponding to one or more face IDs. Any two sets of images in the set of images correspond to different face IDs.

[0209] In one example, the face image library includes three sets of images. The first set of images has an ID of A, meaning that each image in the first set contains a face image with ID A. The second set of images has an ID of B, meaning that each image in the second set contains a face image with ID B. The third set of images has an ID of C, meaning that each image in the third set contains a face image with ID C.

[0210] In one example, the scene image library includes 3 sets of images. The first set of images has an ID of A, meaning that each image in the first set contains a scene image with ID A. The second set of images has an ID of B, meaning that each image in the second set contains a scene image with ID B. The third set of images has an ID of C, meaning that each image in the third set contains a scene image with ID C.

[0211] In some embodiments, the first image contains multiple face images, which correspond to multiple face IDs. In this case, the first image can be clustered into multiple groups of face images simultaneously.

[0212] In one example, the first image contains a face image with ID A and a face image with ID B. In this case, the first image can be clustered into both the first group of images and the second group of images mentioned above.

[0213] In some embodiments, one user may correspond to one face ID, or one user may correspond to multiple face IDs. For example, a first user may correspond to one face ID before makeup and another face ID after makeup.

[0214] In some embodiments, the first image library may be obtained by clustering multiple images in the image library (album) of an electronic device.

[0215] Since electronic devices contain a wealth of personal images from users, using these images as a reference for editing and creating images can improve the authenticity and reliability of the resulting output images.

[0216] S403: Edit the first image using a first model, wherein the first model is trained based on a first image library.

[0217] In some embodiments, the first model is trained based on images clustered under a first ID in the first image library, meaning that different face IDs correspond to different models.

[0218] In some other embodiments, the first model may integrate one or more sub-models, which correspond one-to-one with one or more ID clusters in the first image library. In one example, the model is trained based on the images under the first ID cluster in the first image library to obtain the first sub-model, which is integrated into the first model and corresponds to the first ID cluster.

[0219] Editing the first image using the first model can be further understood as: editing the first image using the sub-model in the first model that corresponds to the ID of the first image.

[0220] The editing operations performed on the first image may include any one or more of the following: aesthetic composition, intelligent image expansion, background editing, portrait retouching, object removal, and background editing. In addition, other image processing operations may also be included, which are not limited in this application.

[0221] In one implementation, if the first image contains both a face image and a scene image, and the face image has ID A and the scene image has ID B, then when the first image library contains a face image with ID A but not a scene image with ID B, the face image with ID A is refined and / or expanded when the first image is edited (e.g., if the first image lacks a left face image, the left face image is expanded). The specific image editing operation can be determined based on human interaction.

[0222] In one implementation, if the first image contains both a face image and a scene image, and the ID of the face image is A and the ID of the scene image is B, then when the first image library does not contain a face image with ID A but contains a scene image with ID B, the scene image with ID B is edited when the first image is edited. This process can be, for example, scene expansion, object removal, background editing, etc., and the specific editing operation can be determined based on the human interaction.

[0223] In one implementation, if the first image contains both a face image and a scene image, and the face image has ID A and the scene image has ID B, then when the first image library contains both a face image with ID A and a scene image with ID B, then when editing the first image, the face image with ID A is refined and / or expanded, and the scene image with ID B is edited. The specific editing operation can be determined based on human interaction.

[0224] S404: Output the first image after editing.

[0225] In some embodiments, the edited first image is saved to the image library of the electronic device and also serves as the basis for training the first model. This allows for continuous optimization of the first model, which in turn gradually improves the realism and reliability of the images processed by the first model, giving users a "better and better" experience.

[0226] In this embodiment of the application, when creating image content, real personal image data from the image library can be used as reference content. Compared with existing image editing methods (which use a free-associative content generation strategy and do not refer to image library data), this method can improve the authenticity and controllability of the edited image and play an important role in the fields of image editing and content creation.

[0227] In addition, before processing the image to be processed, it is first determined whether the face ID or scene ID corresponding to the image to be processed exists in the image library used to train the image processing model. If it exists, it means that the face ID has completed the face ID data construction and feature model pre-training, which can support the fine retouching process. Only then will the image to be processed be further processed through the image processing model, which can further improve the realism and controllability of the image after editing and creation.

[0228] For example, Figure 5 A schematic flowchart of a method 500 for constructing a face image library according to an embodiment of this application is shown. Figure 5 As shown, the method 500 includes:

[0229] S501: A second image library is obtained by clustering images containing human faces in the image library.

[0230] The clustering of images containing faces in the image library is based on the ID of the face contained in the image. That is, multiple images containing the same face ID are clustered into the same group.

[0231] The second image library includes one or more sets of images, each set of images corresponding to one or more face IDs, and any two sets of images in the set of images correspond to different face IDs.

[0232] S502: Detect face regions in the images in the second image library.

[0233] In some embodiments, face regions are detected in images from a second image library based on a face recognition algorithm, wherein the face recognition algorithm may be, for example, the dlib algorithm.

[0234] It can be understood that this step essentially involves determining the location region of the face in each image of the second image library, that is, locating the face in each image of the second image library to facilitate subsequent operations such as quality screening and segmentation of the face images.

[0235] S503: Based on the face region of each image in the second image library, filter out images in the second image library whose face resolution is less than the first resolution threshold or whose face region occlusion area ratio is greater than the first ratio.

[0236] Taking image 1 in the second image library as an example, when the resolution of the face region in image 1 is less than the first resolution threshold, image 1 is filtered out; when the ratio of the occluded area of ​​the face region in image 1 to the area of ​​the face region is greater than the first ratio, image 1 is filtered out.

[0237] In some embodiments, the first resolution threshold can be 224×224, or any resolution greater than 224×224.

[0238] In some embodiments, the first ratio can be 1 / 8, or any ratio greater than 1 / 8, such as 1 / 6, 1 / 5, 1 / 4, 1 / 3, 1 / 2, etc.

[0239] In some embodiments, images with obvious blurriness and low quality can be filtered out. Specifically, image quality metrics can be used to measure this. These image quality metrics may include, for example, any one or more of the following: mean opinion score (MOS), mean squared error (MSE), and peak signal to noise rate (PSNR), and may also include other image quality metrics.

[0240] S504: Perform face segmentation and / or facial feature segmentation on the images in the second image library after the filtering process.

[0241] Among them, face segmentation of an image refers to cutting out the face from the face region of the image to obtain the face image corresponding to the face ID; similarly, facial feature segmentation of an image refers to cutting out the facial feature images (e.g., ear image, mouth image, eye image, nose image, eyebrow image) from the face region of the image to obtain the facial feature image corresponding to the face ID.

[0242] It can be understood that the purpose of performing face segmentation and / or facial feature segmentation on images in the second image library is to obtain a reference object at the smallest unit during subsequent image retouching. For example, when retouching the mouth part of a face image with ID A, the corresponding reference object can be the mouth image under the cluster with ID A in the face image library.

[0243] In some embodiments, a secondary screening step may be performed before S504: screening out groups of images with fewer than N images in one or more groups of images included in the second image library after screening, where N is a positive integer greater than or equal to 2. That is, the number of images in each group of images for face segmentation and facial feature segmentation is greater than or equal to N, which makes the trained model more reliable.

[0244] S505: Generate a facial image library.

[0245] The face image database includes one or more sets of images, each set of images corresponding to one or more face IDs. When the face image database includes multiple sets of images, any two sets of images in the multiple sets of images correspond to different face IDs.

[0246] Taking the first group of images in the set of one or more sets of images as an example, the first group of images includes one or more face images with ID A; it also includes one or more facial images corresponding to the one or more face images with ID A; and it also includes multiple facial feature images corresponding to the one or more face images with ID A.

[0247] In one example, the images included in the face image library can be seen in Table 1 below.

[0248] Table 1

[0249]

[0250]

[0251] In this embodiment, face ID clustering is performed on face images in the image library. This enables the model to be trained based on the image library after face clustering. In this way, when the trained model is used to process the face image to be processed, it can first be determined whether the face ID corresponding to the face image to be processed exists in the image library used to train the model. If it exists, it means that the face ID has completed face ID data construction and feature model pre-training, and can support the fine-tuning process. Only then will the model be used to process the image to be processed, which can significantly improve the realism and controllability of the image processed by the model.

[0252] For example, Figure 6 A schematic flowchart of a method 600 for constructing a scene image library according to an embodiment of this application is shown. Figure 6 As shown, the method 600 includes:

[0253] S601: A third image library is obtained by clustering scene images in the image library.

[0254] The clustering of scene images in the image library is based on the ID of the scene contained in the image. That is, multiple images containing the same scene ID are clustered into the same group.

[0255] The third image library includes one or more sets of images, each set of images corresponding to one or more scene IDs, and any two sets of images in the set of images correspond to different scene IDs.

[0256] S602: Filter out images in the third image library whose resolution is less than the second resolution threshold.

[0257] In some embodiments, the second resolution threshold can be 224×224, or any resolution greater than 224×224.

[0258] In some embodiments, images with obvious blurriness and low quality can be filtered out. Specifically, image quality metrics can be used to measure this. These image quality metrics may include, for example, any one or more of the following: mean opinion score (MOS), mean squared error (MSE), and peak signal to noise rate (PSNR), and may also include other image quality metrics.

[0259] S603: Perform data augmentation on the images in the third image library after the filtering process.

[0260] The data enhancement performed on the image includes any one or more of random cropping, affine transformation, rotation, translation, and lighting changes, and may also include other data enhancement processes, which are not limited in this application.

[0261] In some embodiments, S603 is an optional step.

[0262] In some embodiments, a secondary screening step may be performed before S603: screening out groups of images in one or more groups of images included in the third image library after screening process that contain fewer than N images, where N is a positive integer greater than or equal to 2, that is, the number of images in each group of images that is finally data augmented is greater than or equal to N.

[0263] S604: Generate a scene image library.

[0264] The scene image library includes one or more sets of images, each set of images corresponding to one or more scene IDs, and any two sets of images in the set of images correspond to different scene IDs.

[0265] Taking the first group of images in one or more groups of images as an example, the first group of images includes one or more scene images with ID D.

[0266] In one example, the images included in the scene image library can be seen in Table 2 below.

[0267] Table 2

[0268]

[0269] In this embodiment, face ID clustering is performed on scene images in the image library. This enables the model to be trained based on the image library after scene clustering. In this way, when the trained model is used to process the scene image to be processed, it can first be determined whether the scene ID corresponding to the scene image to be processed exists in the image library used to train the model. If it exists, it means that the scene ID has completed scene ID data construction and feature model pre-training and can support scene processing. Only then will the model be used to process the image to be processed, which can significantly improve the realism and controllability of the image processed by the model.

[0270] For example, Figure 7 A schematic flowchart of a method 700 for training a database model according to an embodiment of this application is shown. Figure 7 As shown, the method 700 includes:

[0271] S701: Using one or more images under the first ID cluster, and one or more mask images obtained by masking one or more images under the first ID cluster as training samples, the model is trained to obtain the first sub-model, which is used to edit the image to be processed corresponding to the first ID.

[0272] In one implementation, model training is performed on the edge. Specifically, the model can be trained by taking one or more masked images obtained by masking one or more images under the first ID cluster as input and one or more images under the first ID cluster as output, thereby obtaining the first sub-model.

[0273] In another implementation, model training is performed by the server. Specifically, the client sends the features of one or more masked images obtained by masking one or more images under the first ID cluster, and the features of one or more images under the first ID cluster to the server. The server uses the features of the one or more masked images obtained by masking one or more images under the first ID cluster as input and the features of one or more images under the first ID cluster as output to train the model, thereby obtaining the first sub-model, and then sends the first sub-model to the client.

[0274] In some embodiments, one or more masked images are obtained by randomly masking one or more images under the first ID cluster.

[0275] The method 700 further includes: using one or more images under the second ID cluster, and one or more masked images obtained by masking one or more images under the second ID cluster as training samples, to train the model and obtain a second sub-model, which is used to process the image to be processed corresponding to the second ID.

[0276] Similarly, when the image library includes multiple ID clusters, the above model training method can be used to obtain multiple models that correspond one-to-one with the multiple ID clusters.

[0277] In some embodiments, the multiple models that correspond one-to-one with multiple ID clusters can be independent models; the multiple models that correspond one-to-one with multiple ID clusters can also be integrated into an inference model (e.g., the first model). That is, the inference model integrates multiple sub-models, which correspond one-to-one with multiple IDs (which can be face IDs or scene IDs), and each sub-model is used to process the image to be processed under the corresponding ID cluster.

[0278] In one example, the image library contains face images of user A and user B. The images in the library are categorized into two groups: a cluster of user A's face images (where the ID of this cluster is user A) and a cluster of user B's face images (where the ID of this cluster is user B). The model is trained using user A's face image and a masked image obtained by masking user A's face image as training samples. Similarly, the model is trained using user B's face image and a masked image obtained by masking user B's face image as training samples. The models for user A and user B can be independent. They can also be integrated into an inference model. When a user uploads an image to be processed from user A, the model for user A will be used to edit that image. Likewise, when a user uploads an image to be processed from user B, the model for user B will be used to edit that image.

[0279] In other words, when the face image database includes m face ID clusters (i.e., face images of m different users), each ID cluster in the m face ID clusters and the corresponding mask image of the ID cluster will be used as training samples for model training. The resulting model can be integrated into the inference model as a sub-model. Similarly, when the scene image database includes n scene ID clusters, each ID cluster in the n scene ID clusters and the corresponding mask image of the ID cluster will be used as training samples for model training. The resulting model can be integrated into the inference model as a sub-model.

[0280] In one specific embodiment, after obtaining the face image database, each group of images in the face image database (all images in each group correspond to the same ID, and different groups of images correspond to different IDs) is used as a model training unit (i.e., each face ID cluster is used as a model training unit). Taking the first group of images in the face image database as an example, one or more masked images obtained by randomly masking one or more images in the first group are used as input, and the first group of images is used as output for model training, to obtain a model corresponding to the ID of the first group of images. This model is used to process the images to be processed that have the same ID as the images in the first group of images. In this way, images under the first face ID cluster in the image database can be used as reference images for the processing of images to be processed that belong to the same first face ID cluster, so as to guide the processing of images to be processed that belong to the same first face ID cluster as the reference images through the feature information of the reference images.

[0281] In another specific embodiment, after obtaining the scene image library, each group of images in the scene image library (all images in each group correspond to the same ID, and different groups of images correspond to different IDs) is used as a model training unit (i.e., each scene ID cluster is used as a model training unit). Taking the first group of images in the scene image library as an example, one or more masked images obtained by randomly masking one or more images in the first group are used as input, and the first group of images is used as output for model training, to obtain a model corresponding to the ID of the first group of images. This model is used to process the images to be processed that have the same ID as the images in the first group of images. In this way, images under the first scene ID cluster in the image library can be used as reference images for the processing of images to be processed that belong to the same first scene ID cluster, so as to guide the processing of images to be processed that belong to the same first scene ID cluster as the reference images through the feature information of the reference images.

[0282] In some embodiments, during model training, the image input to the model can actually be the extracted features of the image.

[0283] In a more specific implementation, the model architecture of the above model is based on the stablediffusioninpainting generative model. The image data within each training unit is randomly sampled, encoded by variational autoencoders (VAEs), and noise is added in the feature dimension. At the same time, the text is encoded by a contrastive language-image pre-training encoder. The text features, image features, and mask features after random masking are fused together and trained into the stablediffusioninpainting UNet. Considering the training cost, this application chooses to use LoRa fine-tuning to fine-tune only the cross attention layer in the UNet, thereby supporting multiple LoRa modules on one model architecture and supporting model fine-tuning of multiple faces and multiple scenes in the image library in a low-cost manner.

[0284] Stable diffusion inpainting is a very practical function. Its principle is to locally repaint the image, which can not only repair image flaws but also produce more stunning effects by modifying local image content. UNet consists of two parts: a feature extraction part and an upsampling part. Because its network structure resembles a U-shape, it is called the UNet network. In image processing, a mask is typically used to control the region or process of image processing. Specifically, a mask in image processing can be considered a kind of "occlusion" or "template" used to occlude, protect, or otherwise manipulate specific parts of an image. This occlusion can be selectively applied to the entire image or only to a local area. In this way, masks help us precisely control which pixels are modified and how they are modified; LoRa models extract and refine features for specific styles. They are usually small in size and can be stacked to achieve a fusion effect. The full name of LoRa models is: low-rank adaptation of large language models. It can be understood as a plugin in stable-diffusion. It is a model that only requires a small amount of data to train. When generating images, LoRa models are used in combination with large models to adjust the output image results; Crossattention refers to the cross-attention layer between the encoder and decoder. In this layer, the decoder adjusts the attention of the encoder's output to obtain encoder information related to the current decoding position.

[0285] In this embodiment, after clustering all images in the image library based on their IDs, a model is trained based on the images in each ID cluster to obtain a sub-model for processing the image to be processed with that ID. This results in multiple sub-models, each corresponding to a specific ID. Taking face IDs as an example, this ensures that multiple images to be processed corresponding to the same face are processed by a dedicated sub-model. Thus, when an image needs processing, it will be processed by the sub-model corresponding to its ID, improving the realism and controllability of image processing.

[0286] To more clearly understand the definition of the masking image provided in the embodiments of this application, exemplarily, Figure 8 The illustration shows several masking images provided in the embodiments of this application.

[0287] Figure 8 (a) and Figure 8 (b) shows a schematic diagram of the masking image used in the two inference processes (i.e., the process of editing the image to be processed by the model) provided in the embodiments of this application.

[0288] like Figure 8 As shown in (a), image B is a masked image obtained by masking the mouth region 810 of image A, where the mouth region 810 of image A is the region to be processed.

[0289] like Figure 8 As shown in (b), image D is a masked image obtained by masking the extended region 820 of image C. The extended region 820 is the region obtained by expanding outward based on the edge of image C, and the extended region 820 is the region to be processed.

[0290] In some implementations, a corresponding mask image can be generated in response to user interaction, for example:

[0291] When the first control box completely overlaps with the image C, the user can determine the expansion area 820 by moving the boundary line of the control box outward, thereby generating the mask image D; the user can also determine the expansion area 820 by scaling the image C within the first control box, thereby generating the mask image D.

[0292] Figure 8 (c) and Figure 8 (d) in the figure shows a schematic diagram of the masking images used in the training process of the two models provided in the embodiments of this application.

[0293] like Figure 8As shown in (c), image F is a masked image obtained by masking the mouth region 830 of image E. The masked image F serves as a training sample in the model training process, and image E serves as the output target of the model.

[0294] like Figure 8 As shown in (d), image G is a masked image obtained by masking region 840 of image E, where the masked image G serves as a training sample in the model training process, and image E serves as the output target of the model.

[0295] It can be understood that during the model training process, the edge regions of image E are masked, the masked image is used as the training sample, and the image before masking is used as the output target, so that the trained model has the ability to edit images.

[0296] For example, Figure 9 A schematic flowchart of an image editing method 900 provided in an embodiment of this application is shown. Figure 9 As shown, the method 900 includes:

[0297] S901: After the user uploads the first image to be processed, determine whether the first image contains a face. If it does, execute S902; if it does not, execute S908.

[0298] In some embodiments, face detection is performed on the first image to determine whether the first image contains a face.

[0299] S902: Determine whether there are any flaws in the faces contained in the first image. If there are, execute S903; if not, execute S908.

[0300] In some embodiments, determining whether a face in the first image has defects can be done using a binary classification neural network structure based on a deep convolutional neural network (e.g., VGG19). Defects on the face may include, for example, closed eyes, hair covering the face, or partial missing parts of the face.

[0301] S903: Determine the ID of the face contained in the first image, and denot it as the first face ID.

[0302] S904: Determine whether the face corresponding to the first face ID exists in the face image database. If it exists, execute S905.

[0303] It can be understood that if a face corresponding to a first face ID exists in the face image database, it means that the images clustered by the first face ID have completed face ID data construction and feature model pre-training. That is, it means that the first model integrates a sub-model obtained by training the model with the images clustered by the first face ID as input. The first image can be edited through the first model. Specifically, it is edited through the sub-model corresponding to the first face ID. The editing operation can include face refinement and face expansion, for example. The specific editing operation can be determined based on the user's interaction behavior. For example, the user can indicate the area to be refined by drawing a circle on the first image, and the user can also indicate face expansion by drawing a circle on the area to be expanded.

[0304] S905: Determine the defective areas in the first image and generate the first mask image.

[0305] In some embodiments, the defective regions of the first image are determined by the defective region detection and discrimination module.

[0306] In one implementation, the pixels of the defective area of ​​the first image are set to 255 (i.e., the color is white), and the pixels of the non-defective area of ​​the first image are set to 0 (i.e., the color is black), which enables the localization of the defective area of ​​the first image, thus obtaining the first localization image of the first image; further, based on the first image and the first localization image of the first image, a first mask image is generated.

[0307] The first image may consist of a defective region and a non-defective region.

[0308] It can be understood that the first positioning image presents a black and white image composed of black areas (corresponding to non-defective areas) and white areas (corresponding to defective areas).

[0309] In some embodiments, the defective areas of the first image can also be determined based on the user's interactive operation on the first image. For example, the area circled by the user on the first image can be determined as the defective area of ​​the first image.

[0310] S906: Input the first mask image into the first model.

[0311] In some embodiments, features extracted from the first mask image may be input into the first model.

[0312] In another implementation, the first image and the first positioning image can be input into the first model, and the first model can generate the first masking image based on the first image and the first positioning image of the first image.

[0313] In some embodiments, features extracted from the first image and features extracted from the first localization image may be input into the first model.

[0314] The input to the training process of the first model may include images from a face image database; that is, the first model may be trained based on images from a face image database, and when the image to be processed is processed by the first model, the output of the first model is a face retouched image and / or a face augmentation image.

[0315] S907: Output the first image after portrait retouching and / or portrait augmentation.

[0316] In some embodiments, a first mask image and a first text feature are input into a first model to obtain a first image after portrait retouching, wherein the first text feature is a text description feature corresponding to the first image.

[0317] Essentially, this involves inputting the first mask image and the first text features into the sub-model corresponding to the first face ID in the first model to obtain the first image after portrait retouching and / or portrait augmentation.

[0318] The first text feature refers to the descriptive features of the first image, such as descriptive text like "a photo of..." or "spring outing photo".

[0319] S908: Determine the scene ID corresponding to the first image, and denot it as the first scene ID.

[0320] S909: Determine whether the scene corresponding to the first scene ID exists in the scene image library. If it exists, execute S910.

[0321] It can be understood that if the scene image library contains a scene corresponding to the first scene ID, it means that the images clustered by the first scene ID have completed scene ID data construction and feature model pre-training. That is, it means that the first model integrates a sub-model obtained by training the model with the images clustered by the first scene ID as input. The first image can be edited by the first model, specifically by the sub-model corresponding to the first scene ID. The editing operation can include scene expansion, object elimination, background editing, etc. The specific editing operation can be determined based on the user's interaction behavior. For example, the user can indicate the part to be eliminated by drawing a circle on the first image, and the user can also indicate scene expansion by drawing a circle on the position to be expanded.

[0322] S910: In response to a first operation by the user based on the first image, determine the area to be processed in the first image and generate a second mask image.

[0323] In some embodiments, the defective regions of the first image (i.e., the regions of the first image to be processed) are determined by the defective region detection and discrimination module.

[0324] In one implementation, the pixels of the defective area of ​​the first image are set to 255 (i.e., the color is white), and the pixels of the non-defective area of ​​the first image are set to 0 (i.e., the color is black), which can realize the localization of the defective area of ​​the first image, that is, obtain the second localization image of the first image; further, based on the first image and the second localization image of the first image, a second mask image is generated.

[0325] The first image may consist of a defective region and a non-defective region.

[0326] In some embodiments, the user's first operation based on the first image can be: when dissatisfied with the original composition of the first image and wishing to expand the first image, the user performs an operation of using two fingers to shrink the first image and drag its position. At this time, a blank area appears on the interface of the electronic device due to the shrinking of the first image, which is the area to be expanded. Correspondingly, the area to be processed in the first image is the blank area. The pixels of the blank area are set to 255 (i.e., the color is white), and the pixels of the area occupied by the shrunken first image are set to 0 (i.e., the color is black), thus obtaining a second positioning image corresponding to the first image; or, the pixels of the blank area are set to 255 (i.e., the color is white), thus obtaining a second masking image corresponding to the first image.

[0327] It should be understood that in the aforementioned blank area to be expanded, the image to be expanded may include scene images and face images. For example, in the original first image, only the left face of the person exists. During expansion, while expanding the scene, the right face of the person will also be expanded. In the process, the models corresponding to the expansion of the scene and the expansion of the right face are different. The expansion of the scene is completed based on the sub-model corresponding to the first scene ID, and the expansion of the right face is completed based on the sub-model corresponding to the first face ID.

[0328] In some other embodiments, the user's first operation based on the first image may be: when the user wants to achieve local elimination of the first image, the user performs interactive operations such as local smearing or drawing circles on the first image; correspondingly, the area to be processed in the first image is the area of ​​the local smearing or the area of ​​the circle, the pixels of the area of ​​the local smearing or the area of ​​the circle are set to 255 (i.e., the color is white), and the pixels of the area outside the area of ​​the local smearing and / or the area outside the area of ​​the circle are set to 0 (i.e., the color is black), thus obtaining a second positioning image corresponding to the first image; or, the pixels of the blank area are set to 255 (i.e., the color is white), thus obtaining a second mask image corresponding to the first image.

[0329] It can be understood that the second positioning image presents a black and white image composed of black and white areas.

[0330] S911: Input the second mask image into the first model.

[0331] In some embodiments, features extracted from the second mask image may be input into the first model.

[0332] In another implementation, the first image and the second positioning image can be input into the first model, and the first model can generate the second masking image based on the first image and the second positioning image of the first image.

[0333] In some embodiments, features extracted from the first image and features extracted from the second localization image may be input into the first model.

[0334] The input to the training process of the first model includes images from the scene image library; that is, the first model can be trained based on images from the scene image library, and when the image to be processed is processed by the first model, the output of the first model is the scene-edited image.

[0335] S912: Outputs the first image after scene editing.

[0336] In some embodiments, a second mask image and a second text description are input into a first model to obtain a first image after scene editing.

[0337] Essentially, this involves inputting the second mask image into the sub-model corresponding to the first scene ID in the first model to obtain the first image after scene editing.

[0338] The second text description refers to the description of the first image, such as "a photo of…" or "spring outing photo".

[0339] In some embodiments, after S907 is completed, S908 to S912 may be further executed.

[0340] In this embodiment of the application, when editing and creating image content, real personal image data from the image library can be used as reference content, which can improve the authenticity and controllability of the edited image and play an important role in the fields of image editing and content creation.

[0341] In addition, before processing the image to be processed, it is first determined whether the face ID or scene ID corresponding to the image to be processed exists in the image library used to train the image processing model. If it exists, it means that the face ID has completed the face ID data construction and feature model pre-training. That is, the image processing model integrates a sub-model for processing the image to be processed corresponding to the face ID, which can support the editing process of the image to be processed. Only then will the image to be processed be further processed by the sub-model of the image to be processed corresponding to the face ID in the image processing model, which can further improve the realism and controllability of the image after editing.

[0342] To better understand the image editing method provided in this application, the following, by way of example, Figure 10 A schematic diagram of an image editing process provided in an embodiment of this application is shown.

[0343] Figure 10 Image (a) shows a schematic diagram of an image library constructed according to an embodiment of this application, as shown in Figure (a). Figure 10 As shown in (a), the image library includes a face image library and a scene image library. The face image library includes multiple images of a first user, such as image 1 and image 2, both of which contain the face image of the first user. It also includes multiple images of a second user, such as image 3 and image 4, both of which contain the face image of the second user. The scene image library includes multiple images of a first scene, such as image 2 and image 5, both of which contain images of the first scene (the first scene shown in the figure includes a tower). It also includes multiple images of a second scene, such as image 3 and image 6, both of which contain images of the second scene (the second scene shown in the figure includes a house).

[0344] Furthermore, with Figure 10 (a) shows multiple images of the first user in the image database (i.e., the first face cluster) as model input for model training, resulting in the first sub-model corresponding to the first user; Figure 10 (a) shows multiple images of the second user in the image database (i.e., the second face cluster) as model input for model training, resulting in a second sub-model corresponding to the second user;Figure 10 As shown in (a), multiple images of the first scene in the image library (i.e., the first scene cluster) are used as model inputs for model training to obtain a third sub-model corresponding to the first scene; Figure 10 Multiple images of the second scene in the image library shown in (a) are used as model inputs for model training to obtain the fourth sub-model corresponding to the second scene.

[0345] Integrate the first, second, third, and fourth sub-models mentioned above into... Figure 10 In the first model shown in (b) of the diagram.

[0346] exist Figure 10 Based on the embodiment shown in (a), Figure 10 (b) shows a schematic diagram of an image editing process provided in an embodiment of this application.

[0347] like Figure 10 As shown in (b), the image to be processed 7 and the corresponding localization image are input into the first model. The localization image corresponding to the image to be processed 7 is used to indicate the area to be processed; specifically, the white area represents the area to be processed. Figure 10 In the embodiment shown in (b), in the positioning image corresponding to the image to be processed 7, the white area 1001 corresponds to the right eye in the image to be processed (viewed from the perspective of the person facing it), indicating that the right eye of the portrait in the image to be processed needs to be refined. The white areas 1002 and 1003 are the expanded areas relative to the image to be processed 7, indicating that the image to be processed 7 needs to be expanded (portrait expansion and / or scene expansion). The expanded areas are areas 1002 and 1003.

[0348] Since the face in image 7 is that of the first user, the first sub-model in the first model (a sub-model trained with multiple face images of the first user as input) will be used to refine the face and expand the portrait in image 7. That is, refine the eyes of the portrait in image 7 and expand the left half of the portrait. At the same time, since the scene in image 7 is the first scene, the second sub-model in the first model (a sub-model trained with multiple images of the first scene as input) will be used to expand the scene in image 7.

[0349] The image after image editing using the first model is shown below. Figure 10As shown in the output image 8 (b), the image to be processed 7 has undergone right eye refinement (the right eye in the image to be processed 7 is in a closed state), portrait expansion, and scene expansion. The effect of the right eye refinement, the expanded portrait, and the expanded scene can all be found in the multiple images corresponding to the first user and the multiple images corresponding to the first scene in the face image library. That is, the image after image editing by the solution of this application is more realistic and controllable.

[0350] In another example, such as Figure 10 As shown in (c), the image to be processed 9 and the corresponding localization image are input into the first model. The localization image corresponding to the image to be processed 9 is used to indicate the area to be processed; specifically, the white area represents the area to be processed. Figure 10 In the embodiment shown in (c), in the positioning image corresponding to the image to be processed 9, the white area 1004 is the area expanded relative to the image to be processed 9, indicating that the image to be processed 9 needs to be expanded (portrait expansion and / or scene expansion), and the expanded area is area 1004.

[0351] Since the face in the image to be processed 9 is the face of the second user, the second sub-model in the first model (a sub-model trained with multiple face images of the second user as input) will be used to augment the face in the image to be processed 9. At the same time, since the scene in the image to be processed 9 is the second scene, the fourth sub-model in the first model (a sub-model trained with multiple images of the second scene as input) will be used to augment the scene in the image to be processed 9.

[0352] The image after image editing using the first model is shown below. Figure 10 As shown in (c) of the output image 10, the image to be processed 9 has been augmented with human portrait and scene. The augmented human portrait and scene can be found in the multiple images corresponding to the second user and the multiple images corresponding to the second scene in the face image library. That is, the image after image editing by the solution of this application is more realistic and controllable.

[0353] The generation of the positioning image corresponding to the image to be processed can be automatically generated, such as automatically performing face recognition and scene recognition, detecting defects based on the face recognition and scene recognition results, and identifying defective areas as areas to be processed (marked as white areas). Alternatively, it can be automatically determined based on the size ratio of the portrait in the image, and the expanded area is identified as the area to be processed (marked as white areas). It can also be generated based on user interaction operations, such as responding to the user's operation of drawing circles on the defective areas of the image to be processed, identifying the circled areas as areas to be processed (marked as white areas, and the processing task is retouching), responding to the user's operation of smearing on the image to be processed, identifying the smeared areas as areas to be processed (marked as white areas, and the processing task is object removal), and responding to the user's operation of shrinking the entire image to be processed within a fixed frame, identifying the blank areas within the fixed frame as areas to be processed (marked as white areas, and the processing task is scene expansion).

[0354] The expansion area can be determined by using the positioning image corresponding to the image to be processed. Figure 10 (b) and Figure 11 In the embodiments shown in (c), it is possible to determine, based on the image to be processed and the corresponding localized image, that the current expansion region needs to expand the local human figure and the large background area. A specific judgment mechanism could be, for example, human skeleton point detection. When it is found that some human skeleton points are missing and there is a large mask covering an area much larger than the human body, it is determined that the current expansion region needs to expand both the local human figure and the large background area. Correspondingly, when processing the image to be processed using the model, multiple sub-models can be mixed to process the image.

[0355] In this embodiment, when there is a large area of ​​human limb completion scene in the image to be processed, it can be considered that the current image needs to be expanded in both human figure and scene. Multiple sub-models can be mixed for image editing, which makes the application scope of this solution wider and the image editing effect better.

[0356] For example, Figure 11 This illustration shows a schematic diagram of the edge-cloud interaction corresponding to an image editing process provided in an embodiment of this application. For example... Figures 4 to 9 As shown, the interaction process includes the following steps:

[0357] S1101: The first device generates and maintains the first image library.

[0358] The process of generating and maintaining the first image library (i.e., updating the first image library), as well as a detailed explanation of the first image library, have already been provided. Figures 4 to 9The embodiments shown are described in detail, and for the sake of brevity, they will not be repeated here.

[0359] S1102: The first device encodes one or more images under the first ID cluster in the first image library, and one or more masked images obtained by masking one or more images under the first ID cluster, to obtain first encoded information.

[0360] The explanations regarding one or more images under the first ID cluster, and one or more masked images obtained by randomly masking one or more images under the first ID cluster, have already been provided. Figure 9 The embodiments shown are described in detail, and for the sake of brevity, they will not be repeated here.

[0361] In some embodiments, feature information of one or more images under the first ID cluster and feature information of one or more masked images obtained by masking one or more images under the first ID cluster are encoded to obtain first encoded information.

[0362] S1103: The first device sends the first encoded information to the server.

[0363] In some embodiments, the server may be in the cloud.

[0364] S1104: The server trains the model using the first encoded information as training samples to obtain the first sub-model, which is used to process the image to be processed with the ID of the first ID.

[0365] In some embodiments, after receiving the first encoded information, the server decodes the first encoded information to obtain the features of one or more images under the first ID cluster, and the features of one or more masked images obtained by masking one or more images under the first ID cluster; then, using the features of one or more masked images obtained by randomly masking one or more images under the first ID cluster as input, and using one or more images under the first ID cluster as output targets, the server trains the model to obtain the first sub-model.

[0366] In some embodiments, decoding the first encoded information yields feature information of one or more images under the first ID cluster, and feature information of one or more masked images obtained by masking one or more images under the first ID cluster.

[0367] In some embodiments, the first sub-model is integrated into the first model, that is, the image features under the first ID cluster are solidified into the UNet network of stable diffusion inpainting (considered as the first model).

[0368] Similarly, S1102 to S1104 can be executed synchronously for each ID cluster in all ID clusters in the first image library, or S1102 to S1104 can be executed sequentially for each ID cluster in all ID clusters in the first image library to obtain one or more sub-models corresponding one-to-one with one or more ID clusters in the first image library.

[0369] In some embodiments, one or more sub-models are integrated into the first model, that is, the image features under each ID cluster in one or more ID clusters in the first image library are solidified into the UNet network of stable diffusion inpainting (considered as the first model).

[0370] S1105: In response to the operation of triggering the editing of the image to be processed, the first device encodes the image to be processed and the positioning image corresponding to the image to be processed to obtain the second encoded information.

[0371] The process of generating the localization image corresponding to the image to be processed is described in [reference needed]. Figure 10 and Figure 12 For the sake of brevity, the descriptions in the illustrated embodiments will not be repeated here.

[0372] In some embodiments, the feature information of the image to be processed and the feature information of the corresponding positioning image are encoded to obtain second encoded information.

[0373] In some embodiments, the second encoded information also includes text features of the image to be processed.

[0374] Optionally, in response to triggering an operation to edit the image to be processed, the first device generates a mask image of the first image based on the image to be processed and the positioning image corresponding to the image to be processed, and encodes the mask image of the first image to obtain second encoded information.

[0375] S1106: The first device sends the second encoded information to the server.

[0376] S1107: After receiving the second encoded information, the server inputs the second encoded information into the sub-model corresponding to the ID of the image to be processed in the first model for image editing, and obtains the edited image.

[0377] In some embodiments, after receiving the second encoded information, the server decodes the second encoded information to obtain the features of the image to be processed and the features of the positioning image corresponding to the image to be processed; then, the features of the image to be processed and the features of the positioning image corresponding to the image to be processed are input into the sub-model corresponding to the ID of the image to be processed for image editing to obtain the edited image.

[0378] In some embodiments, after receiving the second encoded information, the server decodes the second encoded information to obtain the features of the mask image corresponding to the image to be processed; then, the features of the mask image corresponding to the image to be processed are input into the sub-model corresponding to the ID of the image to be processed for image editing to obtain the edited image.

[0379] S1108: The server sends the edited image to the first device, which then displays the edited image to the user.

[0380] In some embodiments, the edited image obtained through S1107 is an image in an encoded state. The server can directly send the image in the encoded state to the first device, which will then receive and decode it; alternatively, the server can decode the image in the encoded state and send it to the first device.

[0381] For example, Figure 12 This illustration shows a schematic diagram of the edge-cloud interaction corresponding to another image editing process provided in an embodiment of this application. For example... Figure 11 As shown, the interaction process includes the following steps:

[0382] S1201 to S1204 and Figures 4 to 9 S1101 to S1104 in the illustrated embodiment are the same, and will not be described again here for the sake of simplicity.

[0383] S1205: After generating one or more sub-models, the server sends the one or more sub-models to the first device.

[0384] Among them, the description of one or more sub-models has already been... Figure 9 The embodiments shown are explained in detail, and for the sake of brevity, they will not be repeated here.

[0385] S1206: In response to the operation of triggering the editing of the image to be processed, the first device inputs the image to be processed and the positioning image corresponding to the image to be processed into the sub-model corresponding to the ID of the image to be processed for image editing, and obtains the edited image.

[0386] The process for generating the mask image corresponding to the image to be processed is described in [link to documentation]. Figure 10 and Figure 11 For the sake of brevity, the descriptions in the illustrated embodiments will not be repeated here.

[0387] In some embodiments, the first device determines a mask image corresponding to the image to be processed based on the image to be processed and the positioning image corresponding to the image to be processed, and inputs the mask image corresponding to the image to be processed into a sub-model corresponding to the ID of the image to be processed for image editing to obtain an edited image.

[0388] In some embodiments, the feature information of the image to be processed and the feature information of the corresponding location image are input into a sub-model corresponding to the ID of the image to be processed.

[0389] In some embodiments, the text features of the image to be processed can also be input into a sub-model corresponding to the ID of the image to be processed.

[0390] In some embodiments, after acquiring the edited image, the first device displays the edited image to the user.

[0391] Note: The above Figure 12 The illustrated embodiments and Figure 13 The embodiments shown all involve image editing via edge-cloud interaction. However, this does not limit the implementation of this application. When the chip capabilities on the edge are sufficient, the training process of the first model can also be deployed on the edge. That is, the generation of the first image library, the training of the first model, and the editing process of the image to be processed through the first model are all implemented on the edge.

[0392] For example, Figure 13 This diagram illustrates the functional modules of an image editing apparatus 1300 provided in an embodiment of this application. For example... Figures 4 to 9 As shown, the device 1300 includes:

[0393] Image library building module 1310 is used to generate an image library, which includes a face image library and / or a scene image library.

[0394] Specifically, the image library construction module 1310 includes a face image library construction module 1311 and a scene image library construction module 1312.

[0395] The face image library construction module 1311 is used to construct a face image library by clustering images containing faces in the library.

[0396] The face image database includes one or more sets of images, each set of images corresponding to one or more face IDs, and any two sets of images in the set of images correspond to different face IDs.

[0397] In some embodiments, the face image library construction module 1311 is specifically used to perform the following steps:

[0398] (1) A second image library is obtained by clustering images containing human faces in the image library.

[0399] The second image library includes one or more sets of images, each set of images corresponding to one or more face IDs, and any two sets of images in the set of images correspond to different face IDs.

[0400] (2) Detect face regions in the images in the second image library.

[0401] (3) Based on the face region of each image in the second image library, filter out images in the second image library whose face resolution is less than the first resolution threshold or whose face region occlusion area ratio is greater than the first ratio.

[0402] Taking image 1 in the second image library as an example, when the resolution of the face region in image 1 is less than the first resolution threshold, image 1 is filtered out; when the ratio of the occluded area of ​​the face region in image 1 to the area of ​​the face region is greater than the first ratio, image 1 is filtered out.

[0403] In some embodiments, the first resolution threshold can be 224×224, or any resolution greater than 224×224.

[0404] In some embodiments, the first ratio can be 1 / 8, or any ratio greater than 1 / 8, such as 1 / 6, 1 / 5, 1 / 4, 1 / 3, 1 / 2, etc.

[0405] (4) Perform face segmentation and / or facial feature segmentation on the images in the second image library after the filtering process.

[0406] Among them, face segmentation of an image refers to cutting out the face from the face region of the image to obtain the face image corresponding to the face ID; similarly, facial feature segmentation of an image refers to cutting out the facial feature images (e.g., ear image, mouth image, eye image, nose image, eyebrow image) from the face region of the image to obtain the facial feature image corresponding to the face ID.

[0407] In some embodiments, the face image library construction module 1311 can also be used to: before performing face segmentation and facial feature segmentation on the images in the second image library after the filtering process, filter out groups of one or more groups of images included in the second image library after the filtering process that contain fewer than N images, where N is a positive integer greater than or equal to 2, that is, the number of images in each group of images that are finally used for face segmentation and facial feature segmentation is greater than or equal to N.

[0408] (5) Generate a face image library.

[0409] The face image database includes one or more sets of images, each set of images corresponding to one or more face IDs. When the face image database includes multiple sets of images, any two sets of images in the multiple sets of images correspond to different face IDs.

[0410] Taking the first group of images in the set of one or more sets of images as an example, the first group of images includes one or more face images with ID A; it also includes one or more facial images corresponding to the one or more face images with ID A; and it also includes multiple facial feature images corresponding to the one or more face images with ID A.

[0411] The scene image library construction module 1312 is used to construct a scene image library by clustering scene images in the library.

[0412] The scene image library includes one or more sets of images, each set of images corresponding to one or more scene IDs. When the scene image library includes multiple sets of images, any two sets of images in the multiple sets of images correspond to different scene IDs.

[0413] In some embodiments, the scene image library construction module 1312 is specifically used to perform the following steps:

[0414] (1) A third image library is obtained by clustering the scene images in the image library.

[0415] The clustering of scene images in the image library is based on the ID of the scene contained in the image. That is, multiple images containing the same scene ID are clustered into the same group.

[0416] The third image library includes one or more sets of images, each set of images corresponding to one or more scene IDs, and any two sets of images in the set of images correspond to different scene IDs.

[0417] (2) Filter out images in the third image library whose resolution is less than the second resolution threshold.

[0418] In some embodiments, the second resolution threshold can be 224×224, or any resolution greater than 224×224.

[0419] (3) Perform data augmentation on the images in the third image library after the filtering process.

[0420] The data enhancement performed on the image includes any one or more of random cropping, affine transformation, rotation, translation, and lighting changes, and may also include other data enhancement processes, which are not limited in this application.

[0421] In some embodiments, the scene image library construction module 1312 is further configured to: before performing data augmentation on the images in the third image library after the filtering process, filter out groups of one or more groups of images included in the third image library after the filtering process that contain fewer than N images, where N is a positive integer greater than or equal to 2, that is, the number of images in each group of images that is finally subjected to data augmentation is greater than or equal to N.

[0422] (4) Generate a scene image library.

[0423] The scene image library includes one or more sets of images, each set of images corresponding to one or more scene IDs, and any two sets of images in the set of images correspond to different scene IDs.

[0424] Taking the first group of images in one or more groups of images as an example, the first group of images includes one or more scene images with ID D.

[0425] The model training module 1320 is used to train the model using one or more images under the first ID cluster and one or more masked images obtained by masking one or more images under the first ID cluster as training samples to obtain a first sub-model. The first sub-model is used to process the image to be processed with the ID of the first ID.

[0426] The model training module 1320 is also used to train the model using one or more images under the second ID cluster and one or more masked images obtained by masking one or more images under the second ID cluster as training samples, thereby obtaining a second sub-model, which is used to process the image to be processed corresponding to the second ID.

[0427] Similarly, when the image library includes multiple ID clusters, the model training module 1320 can train the images under each ID cluster to obtain multiple models that correspond one-to-one with the multiple ID clusters.

[0428] In some embodiments, the multiple models that correspond one-to-one with multiple ID clusters can be integrated into a single inference model (e.g., the first model). That is, the inference model integrates multiple sub-models, which correspond one-to-one with multiple IDs (which can be face IDs or scene IDs). Each sub-model is used to process the image to be processed under the corresponding ID cluster.

[0429] The task reasoning module 1330 is used to edit the image to be processed using a model corresponding to the ID of the image to be processed when it receives the image to be processed uploaded by the user, so as to obtain the edited image.

[0430] In some embodiments, the task reasoning module 1330 includes a determining module and a reasoning module, specifically:

[0431] The determination module is used to determine the corresponding ID of the image to be processed in response to the image uploaded by the user, which is denoted as the first ID.

[0432] In some embodiments, the ID corresponding to the first image may include the ID of the face image contained in the first image, and may also include the ID of the scene image contained in the first image.

[0433] It should be understood that the number of face images contained in the image to be processed can be one or more. When the number of face images contained in the image to be processed is multiple, the determining module determines the ID corresponding to the image to be processed. Specifically, the determining module can determine the IDs of multiple face images in the image to be processed and determine the scene ID of the image to be processed.

[0434] This determining module is also used to determine whether an image with the ID of the first ID exists in the image library. If an image with the ID of the first ID exists in the image library, the inference module uses the trained first model to edit the image to be processed.

[0435] In some embodiments, if an image with the ID of the first ID does not exist in the image library, the processing flow of the image to be processed ends.

[0436] The editing operations performed on the image to be processed may include any one or more of the following: aesthetic composition, intelligent image expansion, background editing, portrait retouching, object removal, and background editing. In addition, other image editing operations may also be included, which are not limited in this application.

[0437] In some embodiments, the edited image is saved to the library of the electronic device, or it can be used by the model training module 1320 as a training sample for training the first model. In this way, the first model can be continuously optimized, and the realism and reliability of the processed image are gradually improved, giving users a "the more you use it, the better it gets" experience.

[0438] The specific process of editing the image to be processed by the task reasoning module 1330 is as follows: ​ The embodiments shown have been described in detail, and for the sake of brevity, they will not be repeated here.

[0439] In this embodiment of the application, when creating image content, real personal image data from the image library can be used as reference content. Compared with existing image editing methods (which use a free-associative content generation strategy and do not refer to image library data), this method can improve the authenticity and controllability of the edited image and play an important role in the fields of image editing and content creation.

[0440] In addition, before editing the image to be processed, it will first determine whether the face ID or scene ID corresponding to the image to be processed exists in the image library used to train the model. If it exists, it means that the face ID has completed the face ID data construction and feature model pre-training, and can support the fine-tuning process. Only then will the image to be processed be further edited through the image processing model, which can further improve the realism and controllability of the edited image.

[0441] One or more modules or units described herein can be implemented in software, hardware, or a combination of both. When any of the above modules or units are implemented in software, the software exists as computer program instructions and is stored in memory. A processor can be used to execute the program instructions and implement the above method flow. The processor can include, but is not limited to, at least one of the following: a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a microcontroller unit (MCU), or an artificial intelligence processor, etc., and various computing devices that run software. Each computing device may include one or more cores for executing software instructions to perform calculations or processing. The processor can be built into a SoC (System-on-a-Chip) or an application-specific integrated circuit (ASIC), or it can be a separate semiconductor chip. In addition to the cores within the processor for executing software instructions to perform calculations or processing, it may further include necessary hardware accelerators, such as field-programmable gate arrays (FPGAs), PLDs (programmable logic devices), or logic circuits that implement dedicated logic operations.

[0442] When the modules or units described herein are implemented in hardware, the hardware may be any one or any combination of a CPU, microprocessor, DSP, MCU, artificial intelligence processor, ASIC, SoC, FPGA, PLD, application-specific digital circuit, hardware accelerator, or non-integrated discrete device, which may run the necessary software or perform the above method flow independently of software.

[0443] When the modules or units described herein are implemented using software, they can be implemented, in whole or in part, in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0444] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0445] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0446] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0447] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0448] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0449] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0450] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for image editing, characterized in that, The method includes: In response to an operation that triggers editing of the first image, a target category is determined, wherein the target category is the category to which the first image belongs; Determine whether an image of the target category exists in a first image library, wherein the first image library includes one or more sets of images, and the one or more sets of images correspond one-to-one with one or more categories; When it is determined that an image of the target category exists in the first image library, the first image is edited using a first model, which is trained based on the first image library.

2. The method according to claim 1, characterized in that, The first model is trained based on images of the target category.

3. The method according to claim 1, characterized in that, The first model includes one or more sub-models, one of which is trained based on images of the target category, and the one or more sub-models correspond one-to-one with the one or more categories.

4. The method according to claim 3, characterized in that, The step of editing the first image using the first model includes: The first image is edited using a target sub-model, which is trained based on one or more images of the target category existing in the first image library, and the target sub-model is integrated into the first model.

5. The method according to claim 3 or 4, characterized in that, The target categories include a first target category and a second target category. Editing the first image using the first model includes: The first image is edited using a first target sub-model and a second target sub-model. The first target sub-model is trained on one or more images of the first target category existing in the first image library, and the second target sub-model is trained on one or more images of the second target category existing in the first image library. Both the first target sub-model and the second target sub-model are integrated into the first model.

6. The method according to any one of claims 1 to 5, characterized in that, The first image library includes a face image library and / or a scene image library, wherein the face image library includes one or more sets of face images, and the one or more sets of face images correspond one-to-one with one or more face categories; the scene image library includes one or more sets of scene images, and the one or more sets of scene images correspond one-to-one with one or more scene categories.

7. The method according to claim 6, characterized in that, The one or more face categories correspond one-to-one with one or more users.

8. The method according to any one of claims 1 to 7, characterized in that, The step of editing the first image using the first model includes: When the first image contains a first object and the first image library includes the category to which the first object belongs, a first masking image corresponding to the first image is generated according to the area to be processed in the first image, and the area to be processed is masked in the first masking image. The first mask image is input into the first model so that the first model outputs the first image after image editing.

9. The method according to claim 8, characterized in that, The first object includes a first face, the area to be processed in the first image includes a first area to be processed, the first area to be processed includes a partial or complete area of ​​the first face, and / or the first area to be processed includes an area to be expanded of the first face, and the first image after image editing is the first image after the first face is refined and / or the first image after the first face is expanded.

10. The method according to claim 8 or 9, characterized in that, The first object includes a first scene, and the area to be processed in the first image includes a second area to be processed, which includes the area to be expanded in the first scene. The first image after image editing is the first image after the first scene is expanded.

11. The method according to any one of claims 8 to 10, characterized in that, Before generating a first mask image corresponding to the first image based on the region to be processed in the first image, the method further includes: By performing face detection on the first image, the defective areas of the first image are determined; The defective areas of the first image are identified as the areas to be processed in the first image, and / or In response to the user's first interactive operation, the area to be processed in the first image is determined.

12. The method according to claim 6, characterized in that, The method further includes: A second image library is obtained by clustering images containing human faces in the image library; By performing face detection on the images in the second image library, images in the second image library whose face resolution is less than the first resolution threshold or whose face occlusion area ratio is greater than the first ratio are filtered out. The face image library is generated by performing facial segmentation on the images in the second image library after the filtering process.

13. The method according to claim 6 or 12, characterized in that, The method further includes: A third image library is obtained by clustering scene images in the image library; The scene image library is generated by filtering out images in the third image library whose resolution is less than the second resolution threshold.

14. The method according to any one of claims 1 to 13, characterized in that, The method further includes: Send to the server the features of one or more images in the i-th category of the first image library and the features of one or more masked images obtained by masking one or more images in the i-th category, where i is a natural number, i = 1, 2, 3, ...; Receive the first model sent by the server, wherein, The first model is trained based on the features of one or more images in the i-th category and the features of one or more masked images, where the i-th category is the target category, or The first model includes an i-th sub-model that is trained based on the features of one or more images under the i-th category and the features of one or more masked images.

15. The method according to any one of claims 1 to 13, characterized in that, The method further includes: The model is trained using one or more masked images obtained by masking one or more images of the i-th category in the first image library and one or more images of the i-th category as training samples to obtain the i-th sub-model. The i-th sub-model is used to process the images to be processed belonging to the i-th category, where i is a natural number, i = 1, 2, 3, ... The i-th sub-model is the first model, and the i-th category is the target category, or The first model includes the i-th sub-model.

16. The method according to claim 14 or 15, characterized in that, The one or more masked images are obtained by randomly masking one or more images under the i-th category.

17. An image editing apparatus, characterized in that, The device includes: A determination module is configured to determine a target category in response to an operation that triggers editing of a first image, wherein the target category is the category to which the first image belongs; The determining module is further configured to determine whether an image of the target category exists in the first image library, wherein the first image library includes one or more sets of images, and the one or more sets of images correspond one-to-one with one or more categories; The inference module is used to edit the first image by means of a first model when it is determined that an image of the target category exists in the first image library. The first model is trained based on the first image library.

18. The apparatus according to claim 17, characterized in that, The first model is trained based on images of the target category.

19. The apparatus according to claim 17, characterized in that, The first model includes one or more sub-models, one of which is trained based on images of the target category, and the one or more sub-models correspond one-to-one with the one or more categories.

20. The apparatus according to any one of claims 17 to 19, characterized in that, The reasoning module is specifically used for: The first image is edited using a target sub-model, which is trained based on one or more images of the target category existing in the first image library, and the target sub-model is integrated into the first model.

21. The apparatus according to any one of claims 17 to 20, characterized in that, The target categories include a first target category and a second target category, and the reasoning module is specifically used for: The first image is edited using a first target sub-model and a second target sub-model. The first target sub-model is trained on one or more images of the first target category existing in the first image library, and the second target sub-model is trained on one or more images of the second target category existing in the first image library. Both the first target sub-model and the second target sub-model are integrated into the first model.

22. The apparatus according to any one of claims 17 to 21, characterized in that, The first image library includes a face image library and / or a scene image library, wherein the face image library includes one or more sets of face images, and the one or more sets of face images correspond one-to-one with one or more face categories; the scene image library includes one or more sets of scene images, and the one or more sets of scene images correspond one-to-one with one or more scene categories.

23. The apparatus according to claim 22, characterized in that, The one or more face categories correspond one-to-one with one or more users.

24. The apparatus according to any one of claims 17 to 23, characterized in that, The reasoning module is specifically used for: When the first image contains a first object and the first image library includes the category to which the first object belongs, a first masking image corresponding to the first image is generated according to the area to be processed in the first image, and the area to be processed is masked in the first masking image. The first mask image is input into the first model so that the first model outputs the first image after image editing.

25. The apparatus according to claim 24, characterized in that, The first object includes a first face, the area to be processed in the first image includes a first area to be processed, the first area to be processed includes a partial or complete area of ​​the first face, and / or the first area to be processed includes an area to be expanded of the first face, and the first image after image editing is the first image after the first face is refined and / or the first image after the first face is expanded.

26. The apparatus according to claim 24 or 25, characterized in that, The first object includes a first scene, the area to be processed of the first image includes a second area to be processed, the second area to be processed includes the area to be expanded of the first scene, and the first image after image editing is the first image after the first scene is expanded.

27. The apparatus according to any one of claims 24 to 26, characterized in that, The determining module is also used for: By performing face detection on the first image, the defective areas of the first image are determined; The defective areas of the first image are identified as the areas to be processed in the first image, and / or In response to the user's first interactive operation, the area to be processed in the first image is determined.

28. The apparatus according to claim 22, characterized in that, The device further includes: An image library construction module is used to construct the first image library by clustering the images in the image library; The image library construction module is specifically used for: A second image library is obtained by clustering images containing human faces in the image library; By performing face detection on the images in the second image library, images in the second image library whose face resolution is less than the first resolution threshold or whose face occlusion area ratio is greater than the first ratio are filtered out. The face image library is generated by performing facial segmentation on the images in the second image library after the filtering process.

29. The apparatus according to claim 22 or 28, characterized in that, The image library construction module is also specifically used for: A third image library is obtained by clustering scene images in the image library; The scene image library is generated by filtering out images in the third image library whose resolution is less than the second resolution threshold.

30. The apparatus according to any one of claims 17 to 29, characterized in that, The device further includes: The transceiver module is used to send to the server the features of one or more images in the i-th category of the first image library and the features of one or more masked images obtained by masking one or more images in the i-th category, where i is a natural number, i = 1, 2, 3, ...; The transceiver module is further configured to receive the first model sent by the server, wherein... The first model is trained based on the features of one or more images in the i-th category and the features of one or more masked images, where the i-th category is the target category, or The first model includes an i-th sub-model that is trained based on the features of one or more images under the i-th category and the features of one or more masked images.

31. The apparatus according to any one of claims 17 to 29, characterized in that, The device further includes: The training module is used to train a model using one or more masked images obtained by masking one or more images of the i-th category in the first image library and one or more images of the i-th category as training samples, to obtain the i-th sub-model. The i-th sub-model is used to process images belonging to the i-th category, where i is a natural number, i = 1, 2, 3, ... The i-th sub-model is the first model, and the i-th category is the target category, or The first model includes the i-th sub-model.

32. The apparatus according to claim 30 or 31, characterized in that, The one or more masked images are obtained by randomly masking one or more images under the i-th category.

33. An electronic device, characterized in that, include: One or more processors; One or more memory units; And one or more computer programs, wherein the one or more computer programs are stored in the one or more memories, the one or more computer programs including instructions that, when executed by the one or more processors, cause the electronic device to perform the method as described in any one of claims 1 to 16.

34. A computer-readable storage medium, characterized in that, The storage medium stores a program or instructions that, when executed, implement the method as described in any one of claims 1 to 16.

35. A chip, characterized in that, The chip stores instructions that, when executed, implement the method as described in any one of claims 1 to 16.

36. A computer program product, characterized in that, The computer program product stores a program or instructions that, when executed, implement the method as described in any one of claims 1 to 16.