Image generation method and apparatus, device, and storage medium
By using a target diffusion model to adjust the initial group photo image during the generation of multiple group photos, the problem of model confusion regarding human features is solved, resulting in more natural and higher-quality image generation.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2026-03-12
AI Technical Summary
Existing personalized image generation models are prone to confusing human features when generating group photos of multiple people, resulting in low similarity between the people in the generated images and the input images, and poor realism and image quality.
By acquiring portrait images and reference group photos, the initial group photo image is adjusted using a target diffusion model to generate the target group photo image. By combining object features and fine-tuning training of the diffusion model, the naturalness and clarity of the image are improved.
It improves the flexibility and diversity of generating group photos of multiple people, enhances the naturalness and realism of the images, and improves image quality.
Smart Images

Figure CN2025118749_12032026_PF_FP_ABST
Abstract
Description
Image generation method, device, apparatus and storage medium
[0001] Cross-reference to Related Applications
[0002] This application claims priority to Chinese Patent Application No. 202411253037.5, filed September 6, 2024, the disclosure of which is incorporated herein in its entirety by this reference as part of the present application. TECHNICAL FIELD
[0003] The present disclosure relates to an image generation method, device, apparatus and storage medium. BACKGROUND
[0004] With the development of image technology, personalized image generation schemes based on user portrait images to generate images with certain styles and effects have emerged to improve image interest.
[0005] At present, the above-mentioned personalized image generation scheme is mainly model-based image generation. However, this scheme is mainly applicable to single-person image generation. If it is extended to multi-person image generation, the model is prone to confusion in reasoning the characteristics of each person, resulting in low similarity between the person in the final generated image and the person in the input image, and causing the problems of poor authenticity and low image quality of the generated personalized image. SUMMARY
[0006] To solve the above technical problems, the embodiments of the present disclosure provide an image generation method, device, apparatus and storage medium.
[0007] In a first aspect, the embodiments of the present disclosure provide an image generation method, which comprises:
[0008] In response to a target interaction operation, at least one portrait image is obtained; wherein the portrait image contains one or more target objects to be processed;
[0009] A reference group photo image is determined; wherein the number of reference objects in the reference group photo image is the same as the number of target objects;
[0010] Based on the portrait image, the target objects are fused into the reference objects in the reference group photo image to generate an initial group photo image;
[0011] Based on a target diffusion model corresponding to the target object, at least one object feature of a fusion object corresponding to the target object in the initial group photo image is adjusted to generate a target group photo image; wherein the target diffusion model is obtained by fine-tuning training an initial diffusion model based on image samples of the target object.
[0012] In a second aspect, the embodiments of the present disclosure further provide an image generation apparatus, the apparatus comprising:
[0013] an image acquisition module configured to acquire at least one portrait image in response to a target interaction operation, wherein the portrait image contains one or more target objects to be processed;
[0014] a reference group photo image determination module configured to determine a reference group photo image, wherein the number of reference objects in the reference group photo image is the same as the number of the target objects;
[0015] an initial group photo image generation module configured to fuse the target objects to the reference objects in the reference group photo image based on the portrait image to generate an initial group photo image;
[0016] a target group photo image generation module configured to adjust at least one object feature of a fusion object corresponding to the target object in the initial group photo image based on a target diffusion model corresponding to the target object to generate a target group photo image, wherein the target diffusion model is obtained by fine-tuning training an initial diffusion model based on image samples of the target object.
[0017] In a third aspect, the embodiments of the present disclosure further provide an electronic device, the electronic device comprising:
[0018] a processor;
[0019] a memory configured to store executable instructions;
[0020] wherein the processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the image generation method described in any of the embodiments of the present disclosure.
[0021] In a fourth aspect, the embodiments of the present disclosure further provide a computer readable storage medium, the storage medium storing a computer program, when the computer program is executed by a processor, causing the processor to implement the image generation method described in any of the embodiments of the present disclosure.
[0022] In a fifth aspect, the embodiments of the present disclosure further provide a computer program product, the computer program product being configured to execute the image generation method described in any of the embodiments of the present disclosure.
[0023] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiments of the present disclosure are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal. BRIEF DESCRIPTION OF DRAWINGS
[0024] The above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent upon consideration of the following detailed description, taken in conjunction with the accompanying drawings. Throughout the drawings, like or similar reference numerals designate identical or similar elements throughout the several views. It should be understood that the drawings are schematic and elements and features are not necessarily to scale.
[0025] FIG. 1 is a flow diagram of an image generation method according to an embodiment of the present disclosure;
[0026] FIG. 2 is a detailed flow diagram of S130 in the image generation method shown in FIG. 1;
[0027] FIG. 3 is a detailed flow diagram of S140 in the image generation method shown in FIG. 1;
[0028] FIG. 4 is a processing process diagram of an image generation method according to an embodiment of the present disclosure, taking a person image 1 and a person image 2 as examples;
[0029] FIG. 5 is a structural diagram of an image generation apparatus according to an embodiment of the present disclosure; and
[0030] FIG. 6 is a structural diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0031] Embodiments of the present disclosure will be described more fully hereinafter with reference to the accompanying drawings, in which some, but not all, embodiments of the present disclosure are shown. Like or similar designations can be used throughout the various drawings and elements of the drawings can be used to indicate like or similar elements. It will be understood that the present disclosure can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and fully convey the scope of the present disclosure to those skilled in the art.
[0032] It should be understood that the various steps of the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the execution of the steps shown. The scope of the present disclosure is not limited in this respect.
[0033] As used herein, the term "includes" and its variants are open-ended, meaning that "includes but is not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment." The term "another embodiment" means "at least one additional embodiment." The term "some embodiments" means "at least some embodiments." Related terms have corresponding meanings.
[0034] It should be noted that the terms "first", "second", and the like in the present disclosure are used only to distinguish different devices, modules or units, and do not imply the order or sequence of the functions performed by the devices, modules or units.
[0035] It should be noted that the terms "one", "multiple" in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that "one or more" should be understood unless otherwise explicitly stated in the context.
[0036] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0037] In the related art, there are mainly three methods for generating a group photo image of multiple persons. The first method is a generative model-based image generation method, which generates personalized person images using a trained model. However, when the model is extended to the group photo generation scenario of multiple persons, the model is easy to confuse the features of multiple persons, resulting in a low similarity between the persons in the generated group photo image and the originally input persons, and the generated persons are not natural enough, and the image quality is low. The second method is to introduce a mask-based attention mechanism into the above model to shield the information of other persons in the process of generating a group photo image of multiple persons. However, the group photo image generated by this method is still not natural enough, and the image quality is low; in addition, it limits the content of the group photo image, resulting in low flexibility and lack of diversity of the generated group photo image. The third method is to fuse the features of persons in the input image into the group photo image. However, this method needs to perform multiple rounds of fusion of person features, and the processing process is relatively complex, and the persons in the generated group photo image have obvious fusion traces, and the naturalness of the persons is poor.
[0038] Based on the above, the embodiment of the present disclosure provides an image generation scheme to preliminarily fuse each target object in the input portrait image into a group photo image template (i.e. a reference group photo image) by one processing to obtain an initial group photo image; and then adjust at least one object feature in the local image corresponding to the corresponding target object in the initial group photo image by using a target diffusion model adapted to each target object to obtain a target group photo image with more natural and clear characters. In this way, the content of the group photo image is not limited to improve the flexibility and diversity of the generation of the group photo image, and the naturalness and image quality of the group photo image are improved.
[0039] The image generation method provided by the embodiment of the present disclosure can be applied to the scene of synthesizing a group photo image by using an input image containing multiple objects. The method can be executed by an image generation device, which can be realized by software and / or hardware, and the device can be integrated in an electronic device with image processing function. The electronic device can include but is not limited to a smart phone, a Tablet PC, a PDA, a notebook computer, a desktop computer, a mobile workstation or a server, etc.
[0040] FIG. 1 shows a flowchart of an image generation method provided by the embodiment of the present disclosure. As shown in FIG. 1, the image generation method can include the following steps:
[0041] S110, in response to a target interaction operation, obtaining at least one portrait image; the portrait image contains one or more target objects to be processed.
[0042] The target interaction operation is an interaction operation that triggers the input of the portrait image, which can be an image upload operation, an image selection operation, an image acquisition / shooting operation, etc. The portrait image is a portrait image to be processed, which contains one or more target objects. The target object is an object that the user focuses on and wants to migrate to the group photo image, which can be a part of the character (such as facial features, limbs, etc.) or some wearing articles of the character, etc. The character / portrait here can be a real person or a computer-synthesized virtual person.
[0043] Specifically, the electronic device provides a group photo image generation function, and provides an interaction interface for using the function. A user can perform a target interaction operation according to a prompt in the interaction interface. The electronic device determines an image determined after the target interaction operation as a portrait image to be processed. The portrait image can be an image of a single person or an image containing multiple persons. There can be one or multiple portrait images. For example, a user wants to make a group photo image of four persons, and then the user can input four single-person portrait images, two double-person portrait images, two single-person portrait images and one double-person portrait image, or one single-person portrait image and one three-person portrait image, and so on.
[0044] In some embodiments, S110 includes: in response to the second image selection operation, determining the selected image in the image library as the portrait image.
[0045] The selected image contains one or more target objects.
[0046] The second image selection operation is an interaction operation of selecting an image corresponding to a portrait image input function. The image library stores multiple images, which can be a photo album provided by a user side or a candidate image library provided by a platform side, and so on.
[0047] Specifically, the electronic device can determine the image selected by a user as a portrait image in response to a second image selection operation performed by the user. For example, a user triggers a function button of selecting an image from a photo album, and then performs an operation of selecting a photo from the photo album, and then the electronic device can determine each image selected by the user as a corresponding portrait image. For another example, in order to provide an experience function of generating a group photo image, the electronic device can provide some candidate portrait images for the user to select. At this time, the electronic device determines the candidate portrait image selected by the user as a portrait image.
[0048] In other embodiments, S110 includes: in response to an image acquisition operation, determining the acquired image as the portrait image.
[0049] The acquired image contains one or more target objects.
[0050] Specifically, a user can acquire an image in real time in addition to selecting an existing image. For example, a user wants to obtain a multi-person image with a certain artistic style, and the current shooting environment does not have the artistic style, and then the user can acquire an image of multiple persons in real time to obtain an acquired single-person image or a multi-person image. At this time, the electronic device can determine the acquired image as a portrait image.
[0051] S120, determining a reference group photo image; the number of reference objects in the reference group photo image is the same as the number of target objects.
[0052] The reference group photo image is an image providing a group photo image template function. The reference object is an object of the same type as the target object included in the reference group photo image.
[0053] Specifically, the electronic device can determine the reference group photo image to provide other image features (such as overall composition, color matching, light and shadow effect, etc.) before the target object when generating a subsequent group photo image. The reference group photo image can be an image uploaded / selected by the user, an image automatically generated by the electronic device according to the user's demand, an image automatically selected by the electronic device from the candidate group photo images provided by the background according to the user's demand, and the like.
[0054] In some embodiments, S120 includes: in response to an image uploading operation, determining the received group photo image as the reference group photo image.
[0055] Specifically, the user can upload the group photo image he wants to achieve by himself. At this time, the electronic device determines the image input by the user as the reference group photo image in response to the image uploading operation.
[0056] In other embodiments, S120 includes: in response to a first image selection operation, determining the selected candidate group photo image as the reference group photo image.
[0057] The first image selection operation is an interactive operation of selecting an image corresponding to the function of determining a group photo template. The candidate group photo image is a pre-set group photo image.
[0058] Specifically, the electronic device can provide a plurality of candidate group photo images. When the user performs the first image selection operation, the electronic device can determine the candidate group photo image selected by the user as the reference group photo image.
[0059] It should be noted that whether the user uploads the reference group photo image or the user selects the reference group photo image, it is necessary to judge whether the number of reference objects in the reference group photo image is consistent with the number of target objects. If consistent, continue to execute the subsequent process; if not consistent, prompt the user that the reference group photo image is incorrect, and / or the electronic device automatically generates a suitable reference group photo image.
[0060] In yet other embodiments, S120 includes: determining a target text containing a group photo description; constructing an image generation prompt word based on the target text and a second preset prompt word, and calling an image generation model based on the image generation prompt word to generate the reference group photo image.
[0061] The target text is a text used to describe a desired group photo image. The second preset prompt word is a pre-constructed model prompt word prompt, which contains some pre-designed text content, such as the process of generating a group photo image and the output format. The image generation prompt word is a prompt word used to trigger the model to generate a reference group photo image. The image generation model is a generative model that has the function of generating an image from a text, which can be pre-trained according to business needs.
[0062] Specifically, if the user does not determine the reference group photo image, the electronic device can generate the reference group photo image by itself. In this process, the electronic device first obtains the target text. Then, the target text is combined into the second preset prompt word, such as filling the text content in the target text into the text placeholder in the second preset prompt word, to generate an image generation prompt word. Then, the electronic device inputs the image generation prompt word into the image generation model, and generates the reference group photo image through the operation of the model.
[0063] In an example, the "determining the target text containing the group photo description" in the above embodiment includes: in response to a text input operation, determining the received text as the target text.
[0064] Specifically, the user can input the target text according to his own needs. In this case, the electronic device determines the received text as the target text in response to the text input operation.
[0065] In another example, the "determining the target text containing the group photo description" in the above embodiment includes: generating the target text based on object information of the target object.
[0066] Specifically, if the user inputs an inappropriate text or the user does not want to input the text or the image, the electronic device can automatically generate a suitable target text according to the object information of the target object in the portrait image input by the user, such as the number, gender, and style of the target object. For example, the electronic device can use the object information as input data to call a text generation model to output the target text.
[0067] In yet another example, the "determining the target text containing the group photo description" in the above embodiment includes: generating the target text based on the object information of the target object and the target group photo style.
[0068] The target group photo style is a style of the group photo image. The target group photo style can be determined according to behavior data of the user in the process of using the group photo image generation function in a historical time period. For example, the target group photo style can be determined by analyzing reference group photo images historically uploaded by the user, positive feedback or negative feedback of the user on the output final synthesized group photo image (i.e., the target group photo image), and the like. The target group photo style can also be a preferred style set by the user. The style here can refer to the overall composition of the image (spatial sense and hierarchical sense of element layout, etc.), color, tone, light, element type (such as buildings, natural scenery, home, etc.).
[0069] Specifically, in addition to considering the object information of the target object, the electronic device can also combine the target group photo style when automatically generating the target text, so as to further improve the degree of fit between the target text and the user demand. For example, the electronic device can use the object information and the target group photo style as input data, call a text generation model, and output the target text.
[0070] In still some embodiments, S120 includes: based on the object information of the target object, screening a reference group photo image from a plurality of candidate group photo images.
[0071] Specifically, the electronic device can provide a plurality of candidate group photo images. In a case where it is determined that the user does not determine the target text or the reference group photo image by himself / herself, the electronic device can automatically screen a candidate group photo image with the highest matching degree from the candidate group photo images as the reference group photo image according to the matching degree between the object information of the target object in the input image and the object information of the reference object in each candidate group photo image.
[0072] In still some embodiments, S120 includes: based on the object information of the target object and the target group photo style, screening a reference group photo image from a plurality of candidate group photo images.
[0073] Specifically, in addition to the object information of the target object, the electronic device can also consider the target group photo style adapted by the user when automatically selecting the candidate group photo image, so as to further improve the degree of fit between the screened reference group photo image and the user demand.
[0074] S130, based on the portrait image, fusing the target object to the reference object in the reference group photo image to generate an initial group photo image.
[0075] Specifically, the electronic device can extract the object features of the target object from the portrait image. Then, image fusion processing is performed on each object feature and the reference object adapted in the reference group photo image, to obtain a group photo image in which the target object is preliminarily fused (i.e., the initial group photo image). The image fusion process can complete the fusion processing of multiple target objects at a time, reducing the fusion times in the process of generating a group photo image of multiple persons.
[0076] In some embodiments, S130 comprises: generating a first object feature of each reference object based on the reference group photo image; extracting a second object feature of each target object from the portrait image; fusing the target object to the reference object in the reference group photo image based on the first object feature and the second object feature to generate an initial group photo image.
[0077] The first object feature is an object feature of the reference object, which can be an object mask image or an object class identifier (i.e., a first object class identifier) of the reference object. The object mask image is a binary image, in which the image region where the reference object is located has a value of 1, and other image regions have a value of 0. The second object feature is an object feature of the target object, which can be an object description feature describing the structure key points, texture, edge, details, etc. of the target object, or an object class identifier (i.e., a second object class identifier) of the target object, etc. The object class identifier is information representing the object class to which an object belongs. The object class is a pre-set class that can distinguish different objects, which can be determined by analyzing the information of multiple dimensions of the object. Taking a person as an example, the object class can be a makeup style class, such as a bare makeup class, a heavy makeup class, a business makeup class, a life makeup class, a cute makeup class, a mature makeup class, etc. The object class can also be a dressing class, such as a dressing class determined by a skirt, high heels, etc., and another dressing class determined by a suit, a tie, etc., and can also be a formal class, a casual class, an urban class, a classical / national / rustic class, etc. The object class can also be different classes according to whether the face structure is three-dimensional, whether the face has wrinkles, the length of the hair, the color of the hair, etc.
[0078] Specifically, the process of the electronic device preliminarily fusing the target objects is: extracting the first object feature of each reference object from the reference group photo image, and extracting the second object feature of each target object from the portrait image. Then, based on each first object feature and each second object feature, a corresponding relationship between the target objects and the reference objects is established. Then, according to the corresponding relationship, each target object is fused into the image region of the corresponding reference object in the reference group photo image to generate an initial group photo image. In this way, object fusion can be performed with pertinence, the probability of object feature confusion can be reduced to some extent, the accuracy of object fusion can be improved, and the portrait similarity of the group photo image can be further improved.
[0079] S140, based on the target diffusion model corresponding to the target object, adjusting at least one object feature of the fused object corresponding to the target object in the initial group photo image to generate a target group photo image.
[0080] The target diffusion model is obtained by fine-tuning training of an initial diffusion model based on image samples of the target object. The initial diffusion model is a diffusion model obtained by model training using image samples of the target object in advance. The object feature refers to an identifiable attribute or characteristic of the fused object in the image. For example, the object feature can be a structural feature, a shape feature or an edge feature, a pose feature, a texture feature, a depth information feature, a light feature or a shadow feature of the fused object. Taking the head of the target object as a person as an example, the object feature can be a distribution relationship feature between parts such as facial features and hair contained in the head region, a texture feature of the skin, a pose feature of the head, an edge feature between the face and the hair, a depth difference feature of each part of the head region, a light and shadow relationship feature of each part of the head region, and the like.
[0081] Specifically, due to the limitation of the image fusion technology, the reference object (i.e., the fused object) in the initial group photo image obtained above, which fuses the object feature of the target object, can have the problems of unnatural mask effect and unclear image quality. Therefore, the image adjustment process is added in the embodiments of the present disclosure to greatly reduce the above image problems.
[0082] Since the initial group photo image contains multiple fused objects, each of which corresponds to a different target object and has different object features. Therefore, in order to improve the effect of image repair, the initial diffusion model can be fine-tuned for each target object in the embodiments of the present disclosure to obtain a candidate diffusion model that is more suitable for the target object. Then, for each target object, the electronic device calls the candidate diffusion model corresponding to the target object (i.e., the target diffusion model), adjusts one or more object features in the image region where the fused object in the initial fused image is located, so that the fused object has a higher similarity with the target object in one or more object features, improves the similarity between the adjusted fused object and the target object corresponding thereto, and makes the fused object (such as a face object of a person) maintain reasonable coherence and better visual naturalness between one or more object features and other objects (such as a hair object, a neck object, or a shoulder object of a person) in the local image where the fused object is located, and improves the naturalness of the adjusted fused object. After the processing process, the electronic device can perform targeted image adjustment on each fused object in the initial fused image to obtain a target group photo image.
[0083] The image generation method provided by the above embodiments of the present disclosure can obtain a portrait image containing a plurality of target objects to be processed, and determine a reference group photo image containing the same number of reference objects; then, each target object is preliminarily fused into the reference group photo image respectively to obtain an initial group photo image containing a plurality of fused objects, and at least one object feature of the fused object corresponding to the corresponding target object in the initial group photo image is adjusted by using a target diffusion model obtained by fine-tuning training based on the image sample of the target object adapted to each target object, so as to greatly weaken the image problems of insufficient similarity and unnaturalness of the fused object in the initial group photo image, and obtain a final target group photo image. On the one hand, the input of the portrait image of the target object is not limited, and the personalized group photo image is generated, which improves the flexibility and diversity of image generation. On the other hand, the similarity between each object contained in the target group photo image and the corresponding target object is improved, and the naturalness of each object in the target group photo image is improved, thereby improving the authenticity and image quality of personalized image generation.
[0084] FIG. 2 is a detailed flowchart of S130 in the image generation method shown in FIG. 1 according to an embodiment of the present disclosure. As shown in FIG. 2, in the case where the first object feature includes an object mask image and a first object category label, and the second object feature includes an object description feature and a second object category label, S130 "fusing the target object into the reference object in the reference group photo image based on the portrait image to generate an initial group photo image" includes the following steps:
[0085] S210, performing object category identification and mask segmentation on each reference object in the reference group photo image to generate an object mask image and a first object category label of each reference object.
[0086] Specifically, the electronic device can perform mask segmentation processing on the reference object in the reference group photo image to obtain an object mask image corresponding to each reference object. In addition, the electronic device can use an object classification algorithm in / pretrained in a related technology to perform object category identification on each reference object in the reference group photo image to obtain a first object category label of each reference object. In this way, the object mask image and the first object category label of each reference object in the reference group photo image can be obtained.
[0087] S220, extracting an object description feature of each target object from the portrait image, and performing object category identification on the portrait image to determine a second object category label of each target object.
[0088] Specifically, the electronic device can perform feature extraction on the portrait image to obtain object description features of each target object. Moreover, the electronic device can perform object category recognition on each target object in the portrait image by using an object classification algorithm in / pre-trained in the related art, to obtain a second object category identifier of each target object. In this way, the object description features and the second object category identifier of each target object in the portrait image can be obtained.
[0089] In S230, a matching relationship between the target object and the reference object is determined based on the first object category identifier and the second object category identifier.
[0090] Specifically, when the user does not explicitly specify the matching relationship, the electronic device can determine the matching relationship between the target object and the reference object by using the matching relationship between the first object category identifier and the second object category identifier.
[0091] For example, when the portrait image contains two target objects of two different object categories (for example, object category 1 and object category 2), and the reference group photo image also contains two reference objects of object category 1 and object category 2, the matching relationship between the target object of object category 1 and the reference object of object category 1 can be established according to the complete matching relationship between the first object category identifier and the second object category identifier, and the matching relationship between the target object of object category 2 and the reference object of object category 2 can also be established.
[0092] For another example, when the portrait image contains four target objects of two object categories 1 and two object categories 2, and the reference group photo image also contains four reference objects of two object categories 1 and two object categories 2, the matching relationship between any target object of object category 1 and any reference object of object category 1, and the matching relationship between the remaining target object of object category 1 and the remaining reference object of object category 1 can be established on the premise that the first object category identifier and the second object category identifier match, combined with a random matching rule. Meanwhile, the matching relationship between any target object of object category 2 and any reference object of object category 2, and the matching relationship between the remaining target object of object category 2 and the remaining reference object of object category 2 can also be established.
[0093] For another example, when the first object category identifier corresponding to the portrait image and the second object category identifier corresponding to the reference group photo image cannot be well matched, the matching relationship between the target object and the reference object can be randomly established. For example, when the portrait image contains three target objects of object category 1, and the reference group photo image contains three reference objects of object category 2, the target object and the reference object can be randomly matched.
[0094] S240, if the reference group photo is determined by the user input operation, the matching relationship is determined in response to the pairing operation of the target object and the reference object.
[0095] Specifically, when the reference group photo is user-specified, the user can see the input image and the reference group photo. At this time, the electronic device can provide an interactive function of specifying the matching relationship between the target object and the reference object.
[0096] For example, the above-mentioned interactive function is implemented as an interactive operation of sequentially pairing the objects. Then, when the user triggers the interactive function and performs an interactive operation of sequentially crossing and selecting the target object in the portrait image and the reference object in the reference group photo, the electronic device can establish the matching relationship between each target object and each reference object according to the selection order of the user.
[0097] For another example, the above-mentioned interactive function is implemented as an interactive operation of a drop-down box or a selection item. Then, the electronic device can detect the selection operation of the user and establish the matching relationship between the target object and the reference object according to the selection result.
[0098] S250, based on the matching relationship, the object description feature is fused into the object mask image to generate each initial fusion image.
[0099] Specifically, according to the above-mentioned matching relationship, the object description feature is fused into the corresponding object mask image to generate the fusion result (i.e. the initial fusion image) of the corresponding target object. The object mask region in the initial fusion image is fused with the object description feature, and the pixel value of the other region is 0.
[0100] S260, the initial fusion image is fused with the reference group photo to generate an initial group photo.
[0101] Specifically, each initial fusion image is fused with the reference group photo to retain the pixel value of the object mask region of each initial fusion image and the pixel value of the region of the reference group photo other than the object mask region, to obtain the initial group photo.
[0102] It should be understood that the above-mentioned fusion process of generating the initial group photo can perform certain transition processing on the boundary region between each object mask region and its adjacent non-object mask region, to improve the smoothness of the initial group photo.
[0103] The above-mentioned embodiments of the present disclosure improve the matching accuracy of the target object and the reference object to a certain extent by introducing the object category identifier in the image fusion process, weaken the unnatural problem of the fusion result caused by the image fusion of different object categories, and further improve the naturalness and portrait similarity of the target group photo.
[0104] FIG. 3 is a detailed flowchart of S140 in the image generation method shown in FIG. 1 according to an embodiment of the present disclosure. As shown in FIG. 3, S140 "adjusting at least one object feature of the fusion object corresponding to the target object in the initial group photo image based on the target diffusion model corresponding to the target object, to generate a target group photo image" includes the following steps:
[0105] S310, extracting an initial local image in which the fusion object is located from the initial group photo image.
[0106] Specifically, the candidate diffusion model obtained by pre-training is an image processing model for a specific single target object, and the processing effect on other target objects can be poor. Therefore, the electronic device first extracts the image of the local area in which each fusion object is located from the initial group photo image to obtain each initial local image. In this way, one candidate diffusion model can be applied only to the initial local image corresponding thereto to adjust the object features of the adaptive fusion object, which can greatly improve the image adjustment effect. Moreover, by extracting the local image, the data calculation amount can be reduced to a certain extent and the image adjustment efficiency can be improved.
[0107] S320, selecting a target diffusion model adaptive to the target object from the plurality of candidate diffusion models based on the object identifier of the target object corresponding to the fusion object.
[0108] Specifically, according to the description of the foregoing embodiments, there is a corresponding relationship between the reference object, the fusion object, and the target object with the highest portrait similarity to the fusion object in the embodiments of the present disclosure, and there is a corresponding relationship between the object identifier of the target object and the candidate diffusion model. Therefore, the electronic device can determine the corresponding target object according to the fusion object in the initial group photo image, and then determine the candidate diffusion model suitable for the fusion object according to the object identifier of the target object as the target diffusion model.
[0109] In some embodiments, before S320, the method further includes: if the matching relationship between the target object and the reference object is random matching, determining the object similarity between the fusion object and each target object for any fusion object, and determining that the target object corresponding to the maximum object similarity corresponds to the fusion object.
[0110] Specifically, the matching relationship between the target object and the reference object constructed by the foregoing embodiments is used to fuse the target object into a suitable object mask image. However, based on various processing in the image fusion process, the portrait similarity between the obtained fused object and the target object is not necessarily optimal. Therefore, when the candidate diffusion model is directly screened according to the matching relationship, the model application may not match, resulting in a worse image restoration effect. Therefore, before screening the candidate model, a corresponding relationship between the fused object and the target object with the highest portrait similarity to the fused object can be constructed.
[0111] For the case of random matching between the target object and the reference object, the electronic device can calculate the object similarity between any fused object and each target object, screen the target object corresponding to the maximum object similarity, and then establish the corresponding relationship between the fused object and the target object. According to this process, the best matching target object for each fused object can be obtained, thereby improving the screening accuracy of the diffusion model, and further improving the portrait similarity and naturalness of the target group photo image.
[0112] In some other embodiments, before S320, the method further includes: if the matching relationship between the target object and the reference object is one-to-one correspondence matching, the target object corresponding to the reference object corresponds to the fused object.
[0113] Specifically, for the case of one-to-one correspondence matching between the target object and the reference object, it reflects the user's fusion demand, so the matching relationship can be directly determined as the corresponding relationship between the fused object and the target object, to further improve the fit degree of the target group photo image and the user's demand.
[0114] S330, based on the image adjustment prompt word and the initial local image corresponding to the target object, calling the target diffusion model to perform image adjustment to generate a target local image.
[0115] The image adjustment prompt word is a prompt word prompt for triggering the running of the target diffusion model with image adjustment function, which can be a pre-constructed prompt, or can be obtained by information filling based on the pre-constructed prompt.
[0116] Specifically, after determining the target diffusion model corresponding to a certain fused object, the electronic device can further determine the image adjustment prompt word suitable for it, and input the image adjustment prompt word and the initial local image corresponding to the fused object into the target diffusion model. After model operation, the object adjustment result of the initial local image, i.e., the target local image, is output.
[0117] In some embodiments, before S330, the method further comprises: constructing an image adjustment prompt word based on the first object category identifier of the reference object corresponding to the fusion object and the first preset prompt word.
[0118] The first preset prompt word is a prompt word prompt of a model applied to an image adjustment function, which contains some pre-designed text content, such as an object category identifier placeholder, an image adjustment process, an output format, and the like.
[0119] Specifically, since the final target group photo image mainly reflects the information of the reference object, the electronic device can limit the adjustment effect of the fusion object according to the first object category identifier of the reference object. Therefore, the electronic device can fill the first object category identifier of the reference object into the object category identifier placeholder in the first preset prompt word to obtain the image adjustment prompt word corresponding to the reference object. In this way, when the image adjustment model refers to the image adjustment prompt word for image adjustment, the object fusion effect can be further improved, for example, the adjustment result of the fusion object of the urban category is more intellectual, and the adjustment result of the fusion object of the leisure category is more sunny, thereby further improving the naturalness of the target group photo image.
[0120] S340, fuse the target local image into the initial group photo image to generate a target group photo image.
[0121] Specifically, the electronic device fuses each target local image back into the corresponding image area of the initial group photo image to obtain the target group photo image.
[0122] The above embodiments of the present disclosure can reduce the data processing amount, improve the image adjustment efficiency, and improve the portrait similarity between the corresponding fusion object and the target object in the target group photo image, and the naturalness of the final fused portrait by applying the target diffusion model adapted to the target object on the initial local image corresponding to each fusion object to perform targeted adjustment processing on at least one object feature of each fusion object.
[0123] In combination with the above embodiments, taking a portrait image 1 containing a single object category 1, a portrait image 2 containing a single object category 2, and a reference group photo image containing two people of one object category 1 and one object category 2 as examples, the generation process of the personalized target group photo image is described. Referring to FIG. 4, the implementation process of the image generation method is specifically as follows:
[0124] The user inputs the portrait image 1, the portrait image 2, and the reference group image. ① The reference group image is subjected to object class recognition and mask segmentation respectively, to obtain an object mask image 1 carrying an object class 1 identifier and an object mask image 2 carrying an object class 2 identifier. ② The matching relationship of the target objects in the object mask image 1 and the portrait image 1, and the matching relationship of the target objects in the object mask image 2 and the portrait image 2 are established. Then, the target objects in the portrait image 1 are fused into the object mask image 1 to obtain an initial fusion image 1, and the target objects in the portrait image 2 are fused into the object mask image 2 to obtain an initial fusion image 2. After that, the initial fusion image 1, the initial fusion image 2, and the reference group image are fused to obtain an initial group image. As can be seen from the initial group image, there is an obvious fusion edge in the head region (such as the lower jaw, hairline, etc.) of the two target objects, and the texture and details of the face region are relatively rough, and there is an obvious mask feeling. ③ The initial local images corresponding to the two persons are extracted from the initial group image. Then, the image adjustment prompt corresponding to the object class 1 identifier prompt and the image adjustment prompt corresponding to the object class 2 identifier prompt are obtained respectively by using the object class identifier. And according to the corresponding relationship between the fusion object and the target object, the target diffusion model 1 corresponding to the person of object class 1 and the target diffusion model 2 corresponding to the person of object class 2 are selected respectively. After that, the initial local image corresponding to the person of object class 1 is adjusted by using the image adjustment prompt corresponding to the object class 1 identifier prompt and the target diffusion model 1, to obtain a target local image 1. And the initial local image corresponding to the person of object class 2 is adjusted by using the image adjustment prompt corresponding to the object class 2 identifier prompt and the target diffusion model 2, to obtain a target local image 2. ④ The target local image 1, the target local image 2, and the initial group image are fused to obtain a target group image.
[0125] The following is an embodiment of an image generation apparatus provided by the embodiments of the present disclosure, which belongs to the same inventive concept as the image generation method of each of the above embodiments. Details not described in detail in the embodiment of the image generation apparatus can be referred to the embodiments of the image generation method described above.
[0126] FIG. 5 shows a structural schematic diagram of an image generation apparatus according to an embodiment of the present disclosure. As shown in FIG. 5, the image generation apparatus 500 can include:
[0127] The portrait image acquisition module 510 is configured to acquire at least one portrait image in response to a target interaction operation; wherein the portrait image contains one or more target objects to be processed.
[0128] The reference group photo image determination module 520 is configured to determine a reference group photo image; wherein the number of reference objects in the reference group photo image is the same as the number of target objects.
[0129] The initial group photo image generation module 530 is configured to fuse the target objects into the reference objects in the reference group photo image based on the portrait images to generate an initial group photo image.
[0130] The target group photo image generation module 540 is configured to adjust at least one object feature of a fusion object corresponding to a target object in the initial group photo image based on a target diffusion model corresponding to the target object to generate a target group photo image; wherein the target diffusion model is obtained by fine-tuning training of an initial diffusion model based on an image sample of the target object.
[0131] The image generation apparatus provided by the embodiments of the present disclosure can obtain a portrait image containing a plurality of target objects to be processed, and determine a reference group photo image containing the same number of reference objects; then, each target object is preliminarily fused into the reference group photo image to obtain an initial group photo image containing a plurality of fusion objects, and a target diffusion model adapted to each target object and obtained by fine-tuning training based on an image sample of the target object is used to adjust at least one object feature of a local fusion object corresponding to a corresponding target object in the initial group photo image, so as to greatly weaken the image problems of insufficient similarity and unnatural visual effect of the fusion objects in the initial group photo image, and obtain a final target group photo image; on the one hand, personalized group photo images can be generated by inputting portrait images of target objects without limitation, and the flexibility and diversity of image generation are improved; on the other hand, the similarity between each object contained in the target group photo image and the corresponding target object is improved, and the naturalness of each object in the target group photo image is improved, thereby improving the authenticity and image quality of personalized image generation.
[0132] In some embodiments, the initial group photo image generation module 530 comprises:
[0133] The first object feature generation submodule is configured to generate a first object feature of each reference object based on the reference group photo image;
[0134] The second object feature extraction submodule is configured to extract a second object feature of each target object from the portrait image;
[0135] The initial group photo image generation submodule is configured to fuse the target objects into the reference objects in the reference group photo image based on the first object feature and the second object feature to generate an initial group photo image.
[0136] In some embodiments, the first object feature comprises an object mask image and a first object category identifier; and the second object feature comprises an object description feature and a second object category identifier.
[0137] Correspondingly, the initial group photo image generation submodule is specifically configured to:
[0138] determine a matching relationship between the target object and the reference object based on the first object category identifier and the second object category identifier;
[0139] fuse the object description feature to the object mask image based on the matching relationship to generate an initial fusion image;
[0140] fuse the initial fusion image with the reference group photo image to generate an initial group photo image.
[0141] In some embodiments, the initial group photo image generation submodule is further configured to:
[0142] before fusing each object description feature to each object mask image based on the matching relationship to generate each initial fusion image, if the reference group photo image is determined by a user input operation, then in response to a pairing operation of the target object and the reference object, the matching relationship is determined.
[0143] In some embodiments, the target group photo image generation module 540 includes:
[0144] an initial local image extraction submodule configured to extract an initial local image in which the fused object is located from the initial group photo image;
[0145] a target diffusion model screening submodule configured to screen a target diffusion model that is suitable for the target object from a plurality of candidate diffusion models based on an object identifier of the target object corresponding to the fused object;
[0146] a target local image generation submodule configured to call the target diffusion model to perform adjustment based on the image adjustment prompt word and the initial local image corresponding to the target object to generate a target local image;
[0147] a target synthetic image generation submodule configured to fuse the target local image to the initial group photo image to generate a target group photo image.
[0148] In some embodiments, the target group photo image generation module 540 further includes an image repair prompt word construction submodule configured to:
[0149] before calling the target diffusion model to perform image adjustment based on the image adjustment prompt word and the initial local image corresponding to the target object to generate the target local image, the image adjustment prompt word is constructed based on the first object category identifier of the reference object corresponding to the fused object and the first preset prompt word.
[0150] In some embodiments, the target group photo image generation module 540 further includes an object corresponding submodule configured to:
[0151] Before screening a target diffusion model that fits the target object from a plurality of candidate diffusion models based on the object identification of the target object corresponding to the fusion object, if the matching relationship between the target object and the reference object is random matching, for any fusion object, determining the object similarity between the fusion object and each target object, and the target object corresponding to the maximum object similarity corresponds to the fusion object;
[0152] If the matching relationship between the target object and the reference object is one-to-one matching, the target object corresponding to the reference object corresponds to the fusion object.
[0153] In some embodiments, the reference group image determination module 520 is specifically configured to:
[0154] In response to an image uploading operation, the received group image is determined as the reference group image;
[0155] Alternatively, in response to a first image selection operation, the selected candidate group image is determined as the reference group image.
[0156] In other embodiments, the reference group image determination module 520 is specifically configured to:
[0157] Determine the target text containing the group description;
[0158] Based on the target text and the second preset prompt word, an image generation prompt word is constructed, and an image generation model is called based on the image generation prompt word to generate the reference group image.
[0159] In some embodiments, the reference group image determination module 520 is further specifically configured to determine the target text containing the group description by any of the following:
[0160] In response to a text input operation, the received text is determined as the target text;
[0161] Based on the object information of the target object, the target text is generated;
[0162] Based on the object information of the target object and the target group style, the target text is generated.
[0163] In yet other embodiments, the reference group image determination module 520 is specifically configured to:
[0164] Based on the object information of the target object, the reference group image is screened from a plurality of candidate group images;
[0165] Alternatively, based on the object information of the target object and the target group style, the reference group image is screened from a plurality of candidate group images.
[0166] In some embodiments, the portrait image acquisition module 510 is specifically configured to:
[0167] In response to the second image selection operation, the selected image in the image library is determined as a portrait image; wherein the selected image contains one or more target objects;
[0168] Alternatively, in response to the image acquisition operation, the acquired image is determined as a portrait image; wherein the acquired image contains one or more target objects.
[0169] The image generation apparatus provided by the embodiments of the present disclosure can perform the image generation method provided by any of the embodiments of the present disclosure, and has the corresponding function modules and beneficial effects of performing the method.
[0170] It is worth noting that in the above embodiments of the image generation apparatus, each module and sub-module included is only divided according to functional logic, but is not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional module / sub-module are only for the convenience of mutual differentiation, and do not serve to limit the protection scope of the present disclosure.
[0171] The embodiments of the present disclosure further provide an electronic device, which can include a processor and a memory. The memory can be configured to store executable instructions. The processor can be configured to read the executable instructions from the memory and execute the executable instructions to implement the image generation method in the above embodiments.
[0172] FIG. 6 shows a structural schematic diagram of an electronic device provided by an embodiment of the present disclosure.
[0173] As shown in FIG. 6, the electronic device 600 can include a processing apparatus 601 (such as a central processor, a graphics processor, etc.), which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 602 or loaded from a storage apparatus 608 to a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing apparatus 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output interface (I / O interface) 605 is also connected to the bus 604.
[0174] Generally, the following apparatuses can be connected to the I / O interface 605: an input apparatus 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output apparatus 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage apparatus 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication apparatus 609. The communication apparatus 609 can allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data.
[0175] It should be noted that the electronic device 600 illustrated in FIG. 6 is merely an example and should not bring any limitation to the function and use range of the embodiments of the present disclosure. That is, although FIG. 6 illustrates the electronic device 600 with various devices, it should be understood that all of the illustrated devices are not required to be implemented or provided. More or less devices can be alternatively implemented or provided.
[0176] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the methods illustrated in the flowcharts. In such embodiments, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the image generation method of any embodiment of the present disclosure are executed.
[0177] Embodiments of the present disclosure also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to implement the image generation method in any embodiment of the present disclosure.
[0178] It should be noted that the computer-readable medium described above in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium, for example, can be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination thereof. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal that propagates in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take any of a variety of forms, including, but not limited to, an electromagnetic signal, an optical signal, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium that is not a storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium can be transmitted using any suitable medium, including, but not limited to, wire, cable, RF (radio frequency), or the like, or any suitable combination thereof.
[0179] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP, and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.
[0180] The computer-readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device and be not assembled in the electronic device.
[0181] The computer-readable medium described above carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the image generation method described in any embodiment of the present disclosure.
[0182] Computer program code for carrying out operations of the present disclosure can be written in any one or more of a variety of programming languages or combinations of languages including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages such as C or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0183] The computer program instructions can also be loaded onto a computer or other programmable information processing apparatus to cause a series of operations to be performed on the computer or other programmable information processing apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable information processing apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0184] The functions described herein above can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Program-specific Integrated Circuits (ASICs), Program-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0185] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include a lined- up electrical connection, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0186] The above description is only the preferred embodiment of the present disclosure and the explanation of the principles of the applied technology. It should be understood by those skilled in the art that the disclosure range involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the technical features described above, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the technical solutions formed by replacing the above features with the technical features disclosed in the present disclosure (but not limited to) having similar functions.
[0187] In addition, although each operation is described in a particular order, this should not be understood as requiring the operations to be performed in the specific order shown or in a sequential order. In certain circumstances, multitasking and parallel processing can be advantageous. Similarly, although several implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be combined in a single embodiment. Conversely, various features described in the context of a single embodiment can also be separated and implemented in multiple embodiments.
[0188] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. An image generation method, comprising: obtaining at least one portrait image in response to a target interaction operation; wherein the portrait image contains one or more target objects to be processed; determining a reference group image; wherein the number of reference objects in the reference group image is the same as the number of target objects; fusing the target objects to the reference objects in the reference group image based on the portrait image to generate an initial group image; adjusting at least one object feature of a fused object corresponding to the target object in the initial group image based on a target diffusion model corresponding to the target object to generate a target group image; wherein the target diffusion model is obtained by fine-tuning training an initial diffusion model based on image samples of the target object.
2. The method of claim 1, wherein, The method further comprises: generating a first object feature of each reference object based on the reference group image; extracting a second object feature of each target object from the portrait image; fusing the target objects to the reference objects in the reference group image based on the first object feature and the second object feature to generate the initial group image.
3. The method of claim 2, wherein, The first object feature comprises an object mask image and a first object category identifier; and the second object feature comprises an object description feature and a second object category identifier. The method further comprises: determining a matching relationship between the target objects and the reference objects based on the first object category identifier and the second object category identifier; fusing the object description feature to the object mask image based on the matching relationship to generate an initial fusion image; fusing the initial fusion image with the reference group image to generate the initial group image.
4. The method of claim 3, wherein, The method further comprises: if the reference group image is determined by a user input operation, determining the matching relationship in response to a pairing operation of the target objects and the reference objects.
5. The method according to any one of claims 1 to 4, wherein, The method further comprises: extracting an initial local image in which the fused object is located from the initial group image; selecting the target diffusion model that is suitable for the target object from a plurality of candidate diffusion models based on an object identifier of the target object corresponding to the fused object; calling the target diffusion model to perform image adjustment based on an image adjustment prompt word and the initial local image corresponding to the target object to generate a target local image; fusing the target local image to the initial group image to generate the target group image.
6. The method of claim 5, wherein, Before the image adjustment prompt word is constructed based on the initial local image of the target object corresponding to the fusion object, and the target diffusion model is called to perform image adjustment on the initial local image of the target object corresponding to the fusion object to generate a target local image, the method further includes: The image adjustment prompt word is constructed based on the first object category identifier and the first preset prompt word of the reference object corresponding to the fusion object.
7. The method of claim 5 or 6, wherein, Before the target diffusion model that adapts to the target object is selected from a plurality of candidate diffusion models based on the object identifier of the target object corresponding to the fusion object, the method further includes: If the matching relationship between the target object and the reference object is random matching, for any fusion object, an object similarity between the fusion object and each target object is determined, and the target object corresponding to the maximum object similarity corresponds to the fusion object. If the matching relationship between the target object and the reference object is one-to-one correspondence matching, the target object corresponding to the reference object corresponds to the fusion object.
8. The method of any one of claims 1-7, wherein, The reference group photo image is determined, including: In response to an image uploading operation, a received group photo image is determined as the reference group photo image; Or, in response to a first image selection operation, a selected candidate group photo image is determined as the reference group photo image.
9. The method of any one of claims 1-8, wherein, The reference group photo image is determined, including: A target text containing a group photo description is determined; Based on the target text and a second preset prompt word, an image generation prompt word is constructed, and an image generation model is called based on the image generation prompt word to generate the reference group photo image.
10. The method of claim 9, wherein, The target text containing a group photo description is determined, including any one of the following: In response to a text input operation, a received text is determined as the target text; Based on the object information of the target object, the target text is generated; Based on the object information of the target object and a target group photo style, the target text is generated.
11. The method of any one of claims 1-10, wherein, The reference group photo image is determined, including: Based on the object information of the target object, the reference group photo image is selected from a plurality of candidate group photo images; Or, based on the object information of the target object and a target group photo style, the reference group photo image is selected from a plurality of candidate group photo images.
12. The method of any one of claims 1-11, wherein, In response to a target interaction operation, at least one portrait image is acquired, including: In response to a second image selection operation, a selected image in an image library is determined as the portrait image; wherein the selected image contains one or more target objects; Or, in response to an image acquisition operation, an acquired image is determined as the portrait image; wherein the acquired image contains one or more target objects.
13. An image generation apparatus, comprising: a portrait image acquisition module configured to acquire at least one portrait image in response to a target interaction operation; wherein the portrait image contains one or more target objects to be processed; a reference group photo image determination module configured to determine a reference group photo image; wherein the number of reference objects in the reference group photo image is the same as the number of target objects; An initial group photo image generation module is configured to fuse the target object to the reference object in the reference group photo image based on the portrait image to generate an initial group photo image; A target group photo image generation module is configured to adjust at least one object feature of a fused object corresponding to the target object in the initial group photo image based on a target diffusion model corresponding to the target object to generate a target group photo image, wherein the target diffusion model is obtained by fine-tuning training an initial diffusion model based on image samples of the target object.
14. An electronic device comprising: a processor; a memory configured to store executable instructions; wherein the processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the image generation method of any one of claims 1-12.
15. A computer readable storage medium storing a computer program, wherein, The computer program, when executed by the processor, causes the processor to implement the image generation method of any one of claims 1-12.
Citation Information
Patent Citations
Multi-person group photo synthesis method and device
CN117372239A
Special effect processing method and device, electronic equipment and storage medium
CN117876208A
Machine learning diffusion model with image encoder trained for synthetic image generation
US20240282016A1