Image generation method and device, electronic equipment and medium
By obtaining the image description information of multiple groups of objects to be photographed and the group image description information, the group image matching the description information is generated, and the problem of fixed posture and stiff effect in the prior art is solved, and the group image generation with high freedom and realism is achieved.
Patent Information
- Application Number
- CN202311753288.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-19
- Publication Date
- 2025-06-20
AI Technical Summary
The existing group photo image generation methods rely on image segmentation and synthesis technology, resulting in fixed portrait postures in group photo images, stiff effects and lack of beauty.
By acquiring multiple sets of images of different objects to be photographed and group image description information input by the user, a group image containing the objects to be photographed and matched with the description information is generated based on these information. The method includes appearance feature extraction, fine-tuning the image generation model, and fusion description information to generate a group image.
It is possible to create a photo image matching the photo image description information without a shooting device, which improves the freedom and playability of the generated image, making the generated photo image closer to the real shooting effect.
Smart Images

Figure CN120182399A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of electronic devices, and in particular, to a method, an apparatus, an electronic device, and a medium for generating an image. Background Art
[0002] Image generation refers to the process of generating images using artificial intelligence technology, which can simulate factors such as light, material, and color in the real world to generate artistic images of various styles, such as oil paintings, watercolor paintings, pencil sketches, etc. Image generation is widely used in fields such as computer games, film production, virtual reality, and visual effects. The existing group photo image generation methods mainly apply image segmentation and image synthesis technologies. Since the group photo image is composed of the portraits segmented from the original image, the poses of the portraits in the group photo image are fixed, the group photo effect is rigid, and there is a lack of beauty. Summary of the Invention
[0003] To overcome the problems in the related art, the present disclosure provides a method, an apparatus, an electronic device, and a medium for generating an image.
[0004] According to a first aspect of an embodiment of the present disclosure, there is provided a method for generating an image, the method including:
[0005] Obtaining multiple groups of images, each group of images including at least one first sample image containing an object to be included in the group photo, and the objects to be included in the group photo in the first sample images of different groups of images being different;
[0006] Obtaining group photo image description information, where the group photo image description information is used to describe the group photo generation effect of the objects to be included in the group photo in each group of images;
[0007] Generating a group photo image containing the objects to be included in the group photo in each group of images and matching the group photo image description information based on the multiple groups of images and the group photo image description information.
[0008] In some exemplary embodiments of the present disclosure, the generating a group photo image containing the objects to be included in the group photo in each group of images and matching the group photo image description information based on the multiple groups of images and the group photo image description information includes:
[0009] Respectively extracting appearance feature information of the objects to be included in the group photo in each group of images to obtain the appearance feature information of the objects to be included in the group photo in each group of images;
[0010] Obtaining a target group photo image generation model based on the appearance feature information of the objects to be included in the group photo in each group of images and an initial group photo image generation model;
[0011] Generating the group photo image based on the group photo image description information and the target group photo image generation model.
[0012] In some exemplary embodiments of the present disclosure, the extracting the appearance feature information of the object to be grouped in each group of images respectively, includes:
[0013] Fine-tuning the initial image generation model respectively based on each group of images to obtain the appearance feature models of the objects to be grouped in each group of images, and using the appearance feature models of the objects to be grouped in each group of images as the appearance feature information of the objects to be grouped in each group of images;
[0014] The obtaining the target group photo image generation model based on the appearance feature information of the objects to be grouped in each group of images and the initial group photo image generation model, includes:
[0015] Fusing the appearance feature models of the objects to be grouped in each group of images with the initial group photo image generation model to obtain the target group photo image generation model;
[0016] Generating the group photo image based on the group photo image description information and the target group photo image generation model.
[0017] In some exemplary embodiments of the present disclosure, the fine-tuning the initial image generation model respectively based on each group of images to obtain the appearance feature models of the objects to be grouped in each group of images, includes:
[0018] Performing the following operations for each group of images:
[0019] Cropping each of the first sample images according to the position of the object to be grouped included therein to obtain first sub-sample images;
[0020] Cutting at least some of the first sub-sample images according to the body parts of the object to be grouped to obtain a plurality of second sub-sample images;
[0021] Using the first sub-sample images and the plurality of second sub-sample images as first training samples to perform fine-tuning training on the initial image generation model to obtain the appearance feature model.
[0022] In some exemplary embodiments of the present disclosure, the group photo image description information includes the objects to be grouped included in each group of images, and the poses of the objects to be grouped included in each group of images and / or the positional relationships between the objects to be grouped included in each group of images.
[0023] In some exemplary embodiments of the present disclosure, the initial group photo image generation model is obtained by the following method:
[0024] Fine-tune and train the initial image generation model with the second training sample to obtain a group photo pose feature model. The second training sample includes a plurality of second sample images with annotations. The second sample images are group photo images, and the annotations are used to indicate the group photo objects in the second sample images, as well as the poses of each group photo object and / or the positional relationships between each group photo object.
[0025] Fuse the group photo pose feature model with the initial image generation model to obtain the initial group photo image generation model.
[0026] In some exemplary embodiments of the present disclosure, the group photo image description information includes scene information.
[0027] According to a second aspect of the embodiments of the present disclosure, there is provided an image generation device, the device includes:
[0028] An acquisition module, configured to acquire multiple groups of images, each group of images includes at least one first sample image containing objects to be group-photographed, and the objects to be group-photographed included in the first sample images of different groups of images are different;
[0029] The acquisition module is further configured to acquire group photo image description information, and the group photo image description information is used to describe the group photo generation effect of the objects to be group-photographed included in each group of images;
[0030] An execution module, configured to generate a group photo image containing the objects to be group-photographed included in each group of images and matching the group photo image description information based on the multiple groups of images and the group photo image description information.
[0031] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, the electronic device includes:
[0032] A processor;
[0033] A memory for storing executable instructions executable by the processor;
[0034] Wherein, the processor is configured to execute the executable instructions in the memory to implement the image generation method provided in the first aspect of the present disclosure.
[0035] According to a fourth aspect of the embodiments of the present disclosure, there is provided a non-transitory computer-readable storage medium, on which executable instructions are stored, and when the executable instructions are executed by a processor, the image generation method provided in the first aspect of the present disclosure is implemented.
[0036] Adopting the above method of the present disclosure has the following beneficial effects: In the process of generating a group photo image of the present disclosure, there is no need to use a shooting device. Based on multiple groups of images of different objects to be photographed and the description information of the group photo image input by the user, a group photo image matching the description information of the group photo image can be created, with high degrees of freedom and playability. The generated group photo image is closer to the real shooting effect. Even if the user and their relatives and friends are separated by a long distance, a group photo image of the user and their relatives and friends can be generated, improving the user experience.
[0037] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Brief Description of the Drawings
[0038] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.
[0039] Figure 1 is a flowchart of a method for generating an image shown according to an exemplary embodiment.
[0040] Figure 2 is a flowchart of a method for generating an image shown according to an exemplary embodiment.
[0041] Figure 3 is a flowchart of a method for generating an image shown according to an exemplary embodiment.
[0042] Figure 4 is a block diagram of a method for generating an image shown according to an exemplary embodiment.
[0043] Figure 5 is a block diagram of a method for generating an image shown according to an exemplary embodiment.
[0044] Figure 6 is a block diagram of a device for generating an image shown according to an exemplary embodiment.
[0045] Figure 7 is a block diagram of an electronic device shown according to an exemplary embodiment. Detailed Description of the Embodiments
[0046] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are only examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0047] Existing group photo image generation methods mainly apply image segmentation and image synthesis technologies. For example, when a user has a video call with another user, the video frames of the video stream containing the user can be first segmented to obtain a portrait of the user without the background, and then the portraits of the two users can be synthesized into a preset background image using an image synthesis algorithm to generate a new image, that is, a group photo image containing the portraits of multiple users having a video call. Since the group photo image is composed of the portraits segmented from the original images, the poses of the portraits in the group photo image are fixed, the group photo effect is rigid, and there is a lack of beauty.
[0048] To solve the above problems, the present disclosure provides an image generation method, which obtains multiple groups of images containing different objects to be grouped in a photo, obtains group photo image description information describing the group photo generation effect of the objects to be grouped in each group of images, and generates a group photo image containing the objects to be grouped in each group of images and matching the group photo image description information based on the multiple groups of images and the group photo image description information. The entire process of generating the group photo image does not require the use of a shooting device. Based on multiple groups of images of different objects to be grouped in a photo and the group photo image description information input by the user, a group photo image matching the group photo image description information can be created, with high freedom and playability. The generated group photo image is closer to the real shooting effect. Even if the user and their relatives and friends are separated by a distance, a group photo image of the user and their relatives and friends can be generated, improving the user experience.
[0049] An exemplary embodiment of the present disclosure provides an image generation method, which is applied to an electronic device, specifically, intelligent devices such as mobile phones, tablet computers, notebooks, smart watches, etc. As Figure 1 shown, the image generation method shown in the present disclosure includes:
[0050] S101. Obtain multiple groups of images, each group of images includes at least one first sample image containing an object to be grouped in a photo, and the objects to be grouped in the first sample images of different groups of images are different;
[0051] S102. Obtain group photo image description information, which is used to describe the group photo generation effect of the objects to be grouped in each group of images;
[0052] S103. Generate a group photo image containing the objects to be grouped in each group of images and matching the group photo image description information based on the multiple groups of images and the group photo image description information.
[0053] In step S101, the electronic device can obtain an image containing the objects to be group-photographed collected by the imaging device by maintaining a communication connection with the imaging device; the electronic device can also obtain the images of the objects to be group-photographed through a network disk or a communication application. For example, the images of the objects to be group-photographed can be searched for through a network disk, or the images of the objects to be group-photographed can be received through WeChat, and the images are downloaded and saved on the electronic device. Since the group-photograph image includes at least two objects to be group-photographed, in order to ensure that the appearance features of each object to be group-photographed can be accurately extracted, multiple groups of images need to be obtained. Each group of images includes at least one first sample image containing the object to be group-photographed, and the objects to be group-photographed included in the first sample images of different groups of images are different, that is, at least one single-person image of each object to be group-photographed needs to be obtained as the first sample image for subsequent generation of the group-photograph image. The first sample image in each group of images can be a close-up of the face of the object to be group-photographed, a half-body photo of the object to be group-photographed, or a full-body photo of the object to be group-photographed in various poses. The more appearance feature information of the object to be group-photographed is included in the first sample image, the more the finally generated group-photograph image conforms to the actual situation of the object to be group-photographed.
[0054] For example, the objects to be group-photographed are a child and an old person. Therefore, two groups of images of the child and the old person need to be obtained respectively. Three images of the old person from the album can be saved as a group of images as the first sample images. The three images are a facial image of the old person, a full-body photo sitting on the sofa, and a full-body photo standing in the living room. Four images of the child obtained through WeChat are saved as another group of images as the first sample images. The four images are a facial image of the child, a full-body photo lying on the lawn, a half-body photo standing on the lawn, and a full-body photo sitting on a deck chair.
[0055] In addition to the multiple groups of images being single-person photos of multiple objects to be group-photographed, the multiple groups of images can also be group photos of multiple objects to be group-photographed. The user selects multiple objects to be group-photographed on the electronic device in advance. When the electronic device obtains multiple groups of images including the selected multiple objects to be group-photographed, the multiple selected objects to be group-photographed included in each group of images can be automatically recognized. In order to improve the accuracy of image recognition, each group of images can be cut according to the positions of the objects to be group-photographed included in each group of images, so that one image is cut into multiple images each only containing one object to be group-photographed, in order to extract the appearance feature information of the objects to be group-photographed.
[0056] For example, if the user selects a child and an elderly person as the subjects to be included in a group photo, the electronic device can obtain at least one set of images containing the child and the elderly person from the photo album, and another set of images containing the child and the elderly person from WeChat. All of the images are group photos of the child and the elderly person. Since the user has already selected the subjects to be included in the group photo, even if there are other people in the images, the electronic device will not identify them. The electronic device will only identify the child and the elderly person, and cut the two sets of images to obtain individual images containing only the child or the elderly person. The electronic device can extract the appearance feature information of the child and the elderly person based on the cut individual photos.
[0057] In step S102, the group photo image description information can be in text form or voice form. For example, the electronic device can obtain the group photo image description information dictated by the user through a microphone device, and the electronic device can also obtain the group photo image description information input by the user through touch on the touch screen. The group photo image description information is used to describe the generation effect of the group photo of the subjects to be included in each set of images, that is, the group photo image description information determines the specific content of the finally generated group photo image. Therefore, the group photo image description information can include the subjects to be included in each set of images, and the poses of the subjects to be included in each set of images and / or the positional relationship between the subjects to be included in each set of images. In addition, the group photo image description information can also include scene information, weather information, etc. The more specific the content of the group photo image description information is, the more vivid and rich the generated group photo image will be. For example, the user inputs the group photo image description information in text form, and the specific content is: the subjects to be included in the group photo are three people, A, B, and C, and A, B, and C are sitting on the sofa; another example is that the user inputs the group photo image description information in voice form, and the specific content is: the subjects to be included in the group photo are three people, A, B, and C, A and B are sitting on the sofa, and C is standing beside the sofa, where A is sitting on the left of B and C is standing on the right of B.
[0058] In step S103, the electronic device can, for example, generate a group photo image based on a pre-stored neural network model. The basis of the neural network model can be a diffusion model, a generative adversarial network, a variational autoencoder, etc. During the training process of the neural network model, multiple groups of images of the objects to be grouped in the photo and the description information of the group photo image are obtained and marked as learning samples, which are then input into the neural network model for training. Among them, the images of each group containing the objects to be grouped in the photo can be directly input into the neural network model for training at one time, or the images of each group containing the objects to be grouped in the photo can be input into the neural network model for training in multiple batches to obtain better training results. After the training is completed, the neural network model can generate and output a group photo image that includes each object to be grouped in the photo and matches the description information of the group photo image based on the description information of the group photo image. The generated group photo image can be saved in the electronic device and displayed on the screen. The user can also send the group photo image to other users for sharing based on personal needs. It should be noted that since the process of generating a group photo image requires a large amount of computing power and computing resources, and has high requirements for the computing power of the electronic device, in order to ensure that the generation process of the group photo image does not affect the normal use of the electronic device, the neural network model used to generate the group photo image can be pre-stored in the cloud, and the electronic device can call the neural network model through applications and other means. Of course, if the computing power of the electronic device can meet the generation of the group photo image, the neural network model can be pre-stored in the electronic device, and the electronic device can omit the process of sending multiple groups of images and the description information of the group photo image to the cloud.
[0059] In the present disclosure, the generation process of the group photo image does not require the use of a photographing device. Based on multiple groups of images of different objects to be grouped in the photo and the description information of the group photo image input by the user, a group photo image that matches the description information of the group photo image can be created, with high degrees of freedom and playability. The generated group photo image is closer to the real shooting effect. Even if the user and their relatives and friends are separated by a long distance, a group photo image of the user and their relatives and friends can be generated, improving the user experience.
[0060] According to an exemplary embodiment, as Figure 2 shown, the method for generating an image in this embodiment includes:
[0061] S201. Obtain multiple groups of images, each group of images including at least one first sample image containing an object to be grouped in the photo, and the objects to be grouped in the photo included in the first sample images of different groups of images are different;
[0062] S202. Extract the appearance feature information of the objects to be grouped in the photo included in each group of images respectively to obtain the appearance feature information of the objects to be grouped in the photo included in each group of images;
[0063] S203. Based on the appearance feature information of the objects to be grouped in the photo included in each group of images and the initial group photo image generation model, obtain the target group photo image generation model;
[0064] S204. Obtain group photo image description information, where the group photo image description information includes the objects to be grouped in each group of images, as well as the poses of the objects to be grouped in each group of images and / or the positional relationships between the objects to be grouped in each group of images;
[0065] S205. Generate a group photo image based on the group photo image description information and the target group photo image generation model.
[0066] Among them, the implementation manners of steps S201 and S204 are the same as those of steps S101 and S102 in the above embodiments, and will not be elaborated here.
[0067] In step S202, the electronic device extracts the appearance features of each object to be grouped based on each group of images. Among them, the appearance features include the facial features of the objects to be grouped and the body features other than the face. The appearance feature information can reflect the looks, body types, etc. of the objects to be grouped. The electronic device can obtain the appearance feature information of the objects to be grouped included in each group of images by extracting the feature points of a group of images corresponding to each object to be grouped. For example, multiple first sample images in a group of images include the facial close-ups and full-body photos of the objects to be grouped. Based on the extraction of facial feature points, the feature information of the arrangements of facial features such as the eyes, mouth, and nose of the objects to be grouped can be obtained, and based on the extraction of body feature points, the feature information of the body types of the objects to be grouped can be obtained. In addition, the electronic device can also directly use each group of images to train a neural network model respectively, so that the neural network model can extract the appearance features of each object to be grouped based on the images and determine the appearance feature information of the objects to be grouped.
[0068] In step S203, the initial group photo image generation model can be trained by a diffusion model, a generative adversarial network, a variational autoencoder, etc. The trained initial group photo image generation model can generate an image based on a text description, that is, the initial group photo image generation model can output an image consistent with the text description information according to the input text description of the image to be generated. The initial group photo image generation model can be pre-stored in the cloud, and the electronic device calls the initial group photo image generation model through an application program or other means. If the computing performance of the electronic device is strong, the initial group photo image generation model can also be pre-stored in the electronic device.
[0069] Since the group photo image description information includes the objects to be grouped in each group of images, in order to ensure that the finally generated group photo image is consistent with the appearance characteristics of the objects to be grouped in each group of images, the initial group photo image generation model can be further trained using the appearance characteristic information of the objects to be grouped to generate a target group photo image generation model; alternatively, a new neural network model can be trained using the appearance characteristic information of the objects to be grouped in each group of images, and the neural network model trained based on the appearance characteristic information of the objects to be grouped is fused with the initial group photo image generation model to generate a target group photo image generation model.
[0070] In step S205, the group photo image description information is input into the target group photo image generation model, and the target group photo image generation model performs inference operations based on the group photo image description information to generate a group photo image adapted to the group photo image description information.
[0071] In some embodiments, the group photo image description information in step S204 includes scene information.
[0072] In order to make the content of the group photo image more rich, the group photo image description information may include scene information, that is, the background of the group photo image is defined, so that the user can create group photo images in more scenes and increase the authenticity of the group photo image. For example, the user inputs the group photo image description information in the form of text, and the specific content is: the objects to be grouped include three people A, B, and C. A, B, and C stand side by side at the entrance of the Temple of Heaven Park. Among them, B is in the middle, A stands on the left of B, and C stands on the right of B. Then the target group photo image generation model can output a group photo image with the background of the Temple of Heaven Park, the objects including three people A, B, and C, and A, B, and C standing side by side.
[0073] In the present disclosure, based on the extraction of the appearance characteristics of each group of images, the appearance characteristic information of the objects to be grouped in each group of images can be accurately obtained, so that the target group photo image generation model obtained based on the appearance characteristic information and the initial group photo image generation model can output a group photo image adapted to the group photo image description information. Based on the change of the group photo image description information by the user, the content of the group photo image can be controlled, and the visual effect of the group photo image is the same as that of the group photo image taken by the imaging device, which increases the convenience and authenticity of the group photo image generation.
[0074] According to an exemplary embodiment, as Figure 3 shown, the image generation method in this embodiment includes:
[0075] S301. Obtain multiple groups of images, each group of images includes at least one first sample image containing an object to be grouped, and the objects to be grouped included in the first sample images of different groups of images are different;
[0076] S302, fine-tuning the initial image generation model based on each group of images to obtain an appearance feature model of the object to be photographed contained in each group of images;
[0077] S303, fine-tuning the initial image generation model with the second training sample to obtain a group photo posture feature model;
[0078] S304, fusing the group photo posture feature model with the initial image generation model to obtain an initial group photo image generation model;
[0079] S305, fusing the appearance feature models of the objects to be photographed contained in each group of images with the initial group photo image generation model to obtain a target group photo image generation model;
[0080] S306, obtaining group photo image description information, where the group photo image description information includes objects to be photographed included in each group of images, postures of the objects to be photographed included in each group of images, and / or positional relationships between the objects to be photographed included in each group of images;
[0081] S307 : Generate a group photo image based on the group photo image description information and the target group photo image generation model.
[0082] Among them, the implementation methods of steps S301, S306, and S307 are the same as those of steps S201, S204, and S205 in the above embodiment, and are not repeated here.
[0083] In step S302, the initial image generation model is a diffusion model. Fine-tuning training based on the diffusion model is a large language model fine-tuning technology LoRA (Low-Rank Adaptation of Large Language Models). The weight parameters of the diffusion model are frozen, and then the trainable layer is injected into each attention module of the model for training. Since the gradient of the weight parameters of the diffusion model does not need to be recalculated, the amount of calculation required for training is reduced. Moreover, since the total amount of training parameters is small, the training of these parameters can be completed with very little data, achieving a good training effect. By fine-tuning the initial image generation model using a group of images corresponding to each object to be photographed, the appearance feature model of the object to be photographed contained in each group of images can be obtained, that is, one object to be photographed corresponds to one appearance feature model. When the user expects that there are other objects to be photographed in the group photo image, it is necessary to obtain a group of images of a new object to be photographed for fine-tuning training to obtain the appearance feature model corresponding to the new object to be photographed. For example, there are three objects to be photographed, A, B, and C, each of which has a set of images. The initial image generation model is fine-tuned using a set of images corresponding to the object A to be photographed, and the appearance feature model A is obtained; the initial image generation model is fine-tuned using a set of images corresponding to the object B to be photographed, and the appearance feature model B is obtained; the initial image generation model is fine-tuned using a set of images corresponding to the object C to be photographed, and the appearance feature model C is obtained. If the group photo images are expected to include the object D to be photographed, it is necessary to obtain another set of images of the object D to be photographed, and the diffusion model is fine-tuned using the set of images to obtain the appearance feature model D. In the process of fine-tuning the initial image generation model, the set of images corresponding to the object to be photographed can be directly input into the initial image generation model for fine-tuning to ensure the effect of fine-tuning; the set of images corresponding to the object to be photographed can also be input into the initial image generation model for fine-tuning in multiple times. If the fine-tuning of the initial image generation model is completed, the remaining images in the set of images corresponding to the object to be photographed are stopped from being input, so as to reduce the training time and improve the efficiency of fine-tuning.
[0084] In some embodiments, step S302 performs the following operations for each group of images:
[0085] Crop each first sample image according to the position of the object to be photographed contained therein to obtain a first sub-sample image;
[0086] Segmenting at least a portion of the first sub-sample images according to body parts of the subject to be photographed, to obtain a plurality of second sub-sample images;
[0087] The first sub-sample image and a plurality of second sub-sample images are used as first training samples, and the initial image generation model is fine-tuned to obtain an appearance feature model.
[0088] To improve the fine-tuning effect of the initial image generation model, it is necessary to process each first sample image corresponding to each object to be grouped in a photo, ensuring that the first sample image can clearly reflect the appearance features of the object to be grouped in a photo. When the first sample image contains the object to be grouped in a photo and scene content, if the object to be grouped in a photo only occupies a small position or an edge position in the first sample image, the first sample image is cropped according to the position of the object to be grouped in a photo to obtain a first sub-sample image, that is, to ensure that the object to be grouped in a photo occupies a larger area of the image and is not at the edge position of the image. For example, if the first sample image is a panoramic image including the object to be grouped in a photo and the object to be grouped in a photo only occupies one-twelfth of the area of the image, then it is cropped according to the position of the object to be grouped in a photo, so that the object to be grouped in a photo in the obtained first sub-sample image occupies one-half of the area of the image. It should be noted that the cropping is not to perform matte extraction on the object to be grouped in a photo in the first sample image, only retaining the object to be grouped in a photo and removing the background. If the size of the obtained first sub-sample image is too small, the first sub-sample image can be appropriately enlarged and processed for clarity.
[0089] Since the appearance feature information of the object to be grouped in a photo includes the facial features of the object to be grouped in a photo and the body features other than the face, at least some of the first sub-sample images can be segmented according to the body parts of the object to be grouped in a photo to obtain multiple second sub-sample images. For example, if the first sub-sample image is a full-body photo of the object to be grouped in a photo, the first sub-sample image can be segmented with the shoulders of the object to be grouped in a photo as the boundary to obtain two second sub-sample images, one of which is the facial image of the object to be grouped in a photo and the other is the body image of the object to be grouped in a photo.
[0090] Taking the first sub-sample image obtained through cropping and the second sub-sample image obtained through segmentation as the first training samples, the initial image generation model is fine-tuned and trained. Compared with directly using the first sample image for fine-tuning and training, the training effect of the appearance feature model can be effectively improved.
[0091] In step S303, the initial image generation model is fine-tuned and trained using the second training samples and the LoRA technology. The trained initial image generation model is the group photo pose feature model. The second training samples include multiple second sample images with annotations. The second sample images are group photo images, and the annotations are used to indicate the group photo objects in the second sample images, as well as the poses of each group photo object and / or the positional relationships between each group photo object. For example, the second training samples include ten second sample images with annotations. One of the second sample images is annotated as "In front of a Chinese-style building, A and B are sitting on a mahogany chair, and C is standing behind A and B", and another second sample image is annotated as "In a forest, A and C are running in the woods one after the other, and B is sitting on the grass".
[0092] In step S304, since the initial image generation model is a generalized image generation model and is not a generation model specifically used to generate group photos, it does not have the ability to accurately locate the characters in the text description. The group photo posture feature model that has been fine-tuned and trained can accurately locate the characters in the text description. Therefore, it is necessary to merge the group photo posture feature model with the initial image generation model to obtain the initial group photo image generation model, so that the initial group photo image generation model can understand the text description and generate an image that matches the characters and postures in the text description.
[0093] The fusion process mentioned above is a process of model stacking, which means that the initial image generation model is added to the corresponding training layer in the group photo posture feature model for locating the person based on the text description to form an initial group photo image generation model, wherein this training layer is in parallel with other training layers in the initial image generation model, that is, the parameters input into the initial group photo image generation model will be processed by this training layer and other training layers at the same time.
[0094] In one example, if Figure 5 As shown, a plurality of group photo images containing the objects to be photographed are accurately marked and used as second sample images to constitute second training samples. The second training samples are used to fine-tune the initial image generation model to obtain a group photo posture feature model, and the group photo posture feature model is fused with the initial image generation model to obtain an initial group photo generation model.
[0095] In step S305, the initial group photo image generation model is usually pre-stored in the cloud. The initial group photo image generation model is a general model, that is, the initial group photo image generation model can output a group photo object containing multiple people based on the text description. Since the initial group photo image generation model does not have the appearance feature information of the group photo object, the appearance feature model of the group photo object contained in each group of images can be fused with the initial group photo image generation model in the cloud to obtain the target group photo image generation model. Since the appearance feature models of the group photo objects are independent of each other, when the group photo image description information includes any group photo object, the corresponding appearance feature model in the target group photo image generation model can be independently awakened to avoid the appearance of portrait pollution problems such as the combination of the face of the group photo object A and the body of the group photo object B. In step S307, the target group photo image generation model can generate a group photo image that matches the group photo image description information based on the input group photo image description information.
[0096] In one example, if Figure 4As shown, three groups of images of objects A, B, and C to be in a group photo are obtained. The three groups of images are input into an image preprocessing model for preprocessing such as cropping and splitting. Each group of preprocessed images is used to fine-tune and train an initial image generation model to obtain appearance feature models A, B, and C. The appearance feature models A, B, and C are fused with the initial group photo generation model to obtain a target group photo generation model. In this way, by inputting group photo image description information into the target group photo generation model, a group photo image adapted to the group photo image description information can be output.
[0097] In the present disclosure, based on the fusion of a group photo pose feature model and an initial image generation model, and the fusion of the appearance feature models of the objects to be in a group photo included in each group of images with the initial group photo image generation model, a target group photo image generation model is obtained, which can effectively avoid the problem of portrait contamination in the generated group photo image, ensure the generation of a group photo image consistent with the specified group photo image description information, and the content of the group photo image is rich and the portraits are vivid.
[0098] An exemplary embodiment of the present disclosure provides an image generation device, as Figure 6 shown, a block diagram of an image generation device shown in the present disclosure.
[0099] The block diagram includes: an acquisition module 61 and an execution module 62. The acquisition module 61 is used to acquire multiple groups of images, each group of images includes at least one first sample image containing an object to be in a group photo, and the objects to be in a group photo included in the first sample images of different groups of images are different; the acquisition module 61 is further used to acquire group photo image description information, and the group photo image description information is used to describe the group photo generation effect of the objects to be in a group photo included in each group of images; the execution module 62 is used to generate a group photo image containing the objects to be in a group photo included in each group of images and matching the group photo image description information based on the multiple groups of images and the group photo image description information.
[0100] In an exemplary embodiment of the present disclosure, the execution module 62 is further used to: extract appearance feature information of the objects to be in a group photo included in each group of images respectively to obtain the appearance feature information of the objects to be in a group photo included in each group of images; obtain a target group photo image generation model based on the appearance feature information of the objects to be in a group photo included in each group of images and the initial group photo image generation model; generate a group photo image based on the group photo image description information and the target group photo image generation model.
[0101] In an exemplary embodiment of the present disclosure, the execution module 62 is further configured to: respectively fine-tune the initial image generation model based on each group of images to obtain an appearance feature model of the object to be grouped in each group of images, and use the appearance feature model of the object to be grouped in each group of images as the appearance feature information of the object to be grouped in each group of images; fuse the appearance feature models of the objects to be grouped in each group of images with the initial group photo image generation model to obtain a target group photo image generation model; and generate a group photo image based on the group photo image description information and the target group photo image generation model.
[0102] In an exemplary embodiment of the present disclosure, the execution module 62 is further configured to: crop each first sample image according to the position of the object to be grouped therein to obtain a first sub-sample image; segment at least some of the first sub-sample images according to the body parts of the object to be grouped to obtain a plurality of second sub-sample images; and use the first sub-sample images and the plurality of second sub-sample images as first training samples to fine-tune and train the initial image generation model to obtain an appearance feature model.
[0103] In an exemplary embodiment of the present disclosure, the group photo image description information includes the objects to be grouped in each group of images, and the poses of the objects to be grouped in each group of images and / or the positional relationships between the objects to be grouped in each group of images.
[0104] In an exemplary embodiment of the present disclosure, the execution module 62 is further configured to: fine-tune and train the initial image generation model with second training samples to obtain a group photo pose feature model, where the second training samples include a plurality of annotated second sample images, the second sample images are group photo images, and the annotations are used to indicate the group photo objects in the second sample images, and the poses of each group photo object and / or the positional relationships between each group photo object; and fuse the group photo pose feature model with the initial image generation model to obtain an initial group photo image generation model.
[0105] In an exemplary embodiment of the present disclosure, the group photo image description information further includes scene information.
[0106] Regarding the image generation device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0107] Figure 7 is a block diagram of an electronic device 700 shown according to an exemplary embodiment. For example, the electronic device 700 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0108] Refer to Figure 7, the electronic device 700 may include one or more of the following components: a processing component 702, a memory 704, a power component 706, a multimedia component 708, an audio component 710, an input / output (I / O) interface 712, a sensor component 714, and a communication component 716.
[0109] The processing component 702 generally controls the overall operation of the electronic device 700, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations. The processing component 702 may include one or more processors 720 to execute instructions to complete all or part of the steps of the above methods. In addition, the processing component 702 may include one or more modules to facilitate the interaction between the processing component 702 and other components. For example, the processing component 702 may include a multimedia module to facilitate the interaction between the multimedia component 708 and the processing component 702.
[0110] The memory 704 is configured to store various types of data to support the operation of the electronic device 700. Examples of such data include instructions for any application or method operating on the electronic device 700, contact data, phone book data, messages, pictures, videos, etc. The memory 704 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0111] The power component 706 provides power to the various components of the electronic device 700. The power component 706 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the electronic device 700.
[0112] The multimedia component 708 includes a screen that provides an output interface between the electronic device 700 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 708 includes a front camera and / or a rear camera. When the electronic device 700 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.
[0113] The audio component 710 is configured to output and / or input audio signals. For example, the audio component 710 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 700 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 704 or transmitted via the communication component 716. In some embodiments, the audio component 710 further includes a speaker for outputting audio signals.
[0114] The I / O interface 712 provides an interface between the processing component 702 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.
[0115] The sensor component 714 includes one or more sensors for providing status assessments of various aspects of the electronic device 700. For example, the sensor component 714 can detect the on / off state of the electronic device 700, the relative positioning of components, such as the display and the keypad of the electronic device 700. The sensor component 714 can also detect a change in the position of the electronic device 700 or a component of the electronic device 700, the presence or absence of user contact with the electronic device 700, the orientation or acceleration / deceleration of the electronic device 700, and the temperature change of the electronic device 700. The sensor component 714 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 714 can also include a light sensor, such as a CMOS or a CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 714 can further include an acceleration sensor, a gyro sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0116] The communication component 716 is configured to facilitate communication between the electronic device 700 and other devices in a wired or wireless manner. The electronic device 700 can access a communication standard-based wireless network, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 716 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 716 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0117] In an exemplary embodiment, the electronic device 700 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0118] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 704 including instructions, and the above instructions can be executed by the processor 720 of the electronic device 700 to complete the above image generation method. For example, the non-transitory computer-readable storage medium can be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0119] A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by the processor of the electronic device, enables the processing device of the electronic device to execute the image generation method provided by the exemplary embodiments of the present disclosure.
[0120] Those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present invention are pointed out by the following claims.
[0121] It should be understood that the present invention is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.
Claims
1. A method for generating an image, characterized in that, The method includes: Obtaining multiple groups of images, each group of images including at least one first sample image containing the object to be group-photographed, and the objects to be group-photographed included in the first sample images of different groups of images being different; Obtaining group-photo image description information, where the group-photo image description information is used to describe the group-photo generation effect of the objects to be group-photographed included in each group of images; Based on the multiple groups of images and the group-photo image description information, generating a group-photo image that includes the objects to be group-photographed included in each group of images and matches the group-photo image description information.
2. The method for generating an image according to claim 1, characterized in that, The generating, based on the multiple groups of images and the group-photo image description information, a group-photo image that includes the objects to be group-photographed included in each group of images and matches the group-photo image description information includes: Respectively performing appearance feature extraction on each group of images to obtain the appearance feature information of the objects to be group-photographed included in each group of images; Based on the appearance feature information of the objects to be group-photographed included in each group of images and an initial group-photo image generation model, obtaining a target group-photo image generation model; Based on the group-photo image description information and the target group-photo image generation model, generating the group-photo image.
3. The method for generating an image according to claim 2, characterized in that, The respectively performing appearance feature extraction on each group of images to obtain the appearance feature information of the objects to be group-photographed included in each group of images includes: Respectively fine-tuning the initial image generation model based on each group of images to obtain the appearance feature models of the objects to be group-photographed included in each group of images, and using the appearance feature models of the objects to be group-photographed included in each group of images as the appearance feature information of the objects to be group-photographed included in each group of images; The obtaining, based on the appearance feature information of the objects to be group-photographed included in each group of images and an initial group-photo image generation model, a target group-photo image generation model includes: Fusing the appearance feature models of the objects to be group-photographed included in each group of images with the initial group-photo image generation model to obtain the target group-photo image generation model; Based on the group-photo image description information and the target group-photo image generation model, generating the group-photo image.
4. The method for generating an image according to claim 3, characterized in that, The respectively fine-tuning the initial image generation model based on each group of images to obtain the appearance feature models of the objects to be group-photographed included in each group of images includes: Performing the following operations for each group of images: Cropping each of the first sample images according to the position of the object to be group-photographed included therein to obtain first sub-sample images; Dividing at least some of the first sub-sample images according to the body parts of the object to be group-photographed to obtain a plurality of second sub-sample images; Using the first sub-sample images and the plurality of second sub-sample images as first training samples to perform fine-tuning training on the initial image generation model to obtain the appearance feature model.
5. The method for generating an image according to claim 1, characterized in that, The group-photo image description information includes the objects to be group-photographed included in each group of images, as well as the poses of the objects to be group-photographed included in each group of images and / or the positional relationships between the objects to be group-photographed included in each group of images.
6. The method for generating an image according to claim 5, characterized in that, The initial group-photo image generation model is obtained by the following method: Fine-tune and train the initial image generation model with the second training sample to obtain a group photo pose feature model. The second training sample includes multiple second sample images with annotations. The second sample images are group photo images, and the annotations are used to indicate the group photo objects in the second sample images, as well as the poses of each group photo object and / or the positional relationships between each group photo object. Fuse the group photo pose feature model with the initial image generation model to obtain the initial group photo image generation model.
7. The method for generating an image according to any one of claims 1 to 6, characterized in that, The group photo image description information includes scene information.
8. An apparatus for generating an image, characterized in that, The device includes: An acquisition module, configured to acquire multiple groups of images, each group of images including at least one first sample image containing a to-be-group-photo object, and the to-be-group-photo objects included in the first sample images of different groups of images are different; The acquisition module is further configured to acquire group photo image description information, where the group photo image description information is used to describe the group photo generation effect of the to-be-group-photo objects included in each group of images; An execution module, configured to generate a group photo image that includes the to-be-group-photo objects included in each group of images and matches the group photo image description information based on the multiple groups of images and the group photo image description information.
9. An electronic device, characterized in that, The electronic device includes: A processor; A memory for storing executable instructions executable by the processor; Wherein, the processor is configured to execute the executable instructions in the memory to implement the image generation method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having executable instructions stored thereon, characterized in that, When the executable instructions are executed by the processor, the image generation method according to any one of claims 1 to 7 is implemented.