Group photo image generation method and device, electronic equipment and storage medium

By acquiring user images and determining group photo templates, and generating group photo images for multiple people, the problem of high site and time requirements for on-site shooting is solved, and high-quality group photo generation is achieved.

CN120219183APending Publication Date: 2025-06-27BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510286787.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, multiple people need to take photos on the spot in a group photo scene, resulting in high requirements for venue and time, and affected by the status of the person being photographed, it is easy to have defects in the group photo, affecting the shooting effect.

Method used

By acquiring user images of at least two invited users and determining a target group photo template based on the number of users and the group photo control information, a group photo image is generated based on the template and the user image, including the image background and facial features.

Benefits of technology

No on-site shooting is required, avoid facial expression defects of the group subject, reduce production costs, and improve the imaging quality of the group image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219183A_ABST
    Figure CN120219183A_ABST
Patent Text Reader

Abstract

Embodiments of the invention provide a group photo image generation method and apparatus, an electronic device and a storage medium. The method comprises the steps of obtaining user images corresponding to at least two invited users; according to the number of invited users and group photo control information, a target group photo template is determined, and the group photo control information represents a group photo posture and / or a group photo formation; a group photo image is generated based on the target group photo template and the user image, the group photo image comprises at least two group photo objects, the target group photo template is used for generating the image background of the group photo image, and the user image is at least used for generating the object face of the group photo object. A target group photo template is determined according to the number of invited users and group photo control information, and then an object body and an object face are constructed in combination with a user image and the target group photo template, so that a group photo image is generated. Multiple persons do not need to photograph on site, the facial expression flaws of the group photo object in the group photo image can be avoided, and the imaging quality of the group photo image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of artificial intelligence technology, and in particular, to a method, apparatus, electronic device, and storage medium for generating a group photo image. Background Art

[0002] Currently, for the scenario of group photos of multiple people, multiple people need to be photographed on site, which has relatively high requirements for the venue and time. At the same time, affected by the states of the people being photographed, the resulting group photos may have defects, affecting the shooting effect of the photos.

[0003] In the prior art, for the problems in the above-mentioned scenario of group photos of multiple people, usually only by repeatedly shooting multiple times can the shooting quality of the final group photo image be improved, resulting in high shooting costs, single scenarios, and unstable imaging quality. Summary of the Invention

[0004] Embodiments of the present disclosure provide a method, apparatus, electronic device, and storage medium for generating a group photo image to overcome the problems of high shooting costs, single scenarios, and unstable quality of group photo images.

[0005] In a first aspect, embodiments of the present disclosure provide a method for generating a group photo image, including:

[0006] Obtaining user images corresponding to at least two invited users; determining a target group photo template according to the number of the invited users and group photo control information, where the group photo control information represents a group photo pose and / or a group photo formation; generating a group photo image based on the target group photo template and the user images, where the group photo image includes at least two group photo objects, the target group photo template is used to generate the image background of the group photo image, and the user images are at least used to generate the object faces of the group photo objects.

[0007] In a second aspect, embodiments of the present disclosure provide a device for generating a group photo image, including:

[0008] A creation module, configured to obtain user images corresponding to at least two invited users;

[0009] A processing module, configured to determine a target group photo template according to the number of the invited users and group photo control information, where the group photo control information represents a group photo pose and / or a group photo formation;

[0010] A generation module, configured to generate a group photo image based on the target group photo template and the user images, where the group photo image includes at least two group photo objects, the target group photo template is used to generate the image background of the group photo image, and the user images are at least used to generate the object faces of the group photo objects.

[0011] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: a processor and a memory;

[0012] The memory stores computer-executable instructions;

[0013] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the group photo image generation method described in the first aspect and various possible designs of the first aspect above.

[0014] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored. When the processor executes the computer-executable instructions, the group photo image generation method described in the first aspect and various possible designs of the first aspect above is implemented.

[0015] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, and when the computer program is executed by a processor, the group photo image generation method described in the first aspect and various possible designs of the first aspect above is implemented.

[0016] The group photo image generation method, device, electronic device, and storage medium provided in this embodiment obtain user images corresponding to at least two invited users; determine a target group photo template according to the number of the invited users and group photo control information, where the group photo control information characterizes a group photo pose and / or a group photo formation; generate a group photo image based on the target group photo template and the user images, where the group photo image includes at least two group photo objects, the target group photo template is used to generate an image background of the group photo image, and the user images are at least used to generate object faces of the group photo objects. By determining a target group photo template according to the number of the invited users and the group photo control information, and then combining the user images and the target group photo template to construct an object body and an object face, a group photo image is generated. It is not necessary to take group photos on site for multiple people, and it can avoid facial expression defects of the group photo objects in the group photo image, reduce the production cost of the group photo image, and improve the imaging quality of the group photo image. Description of the Drawings

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 It is an application scenario diagram of the group photo image generation method provided by an embodiment of the present disclosure;

[0019] Figure 2 Flow chart of the group photo image generation method provided by the embodiments of the present disclosure Figure 1 ;

[0020] Figure 3 Schematic diagram of an interaction interface of a group photo group provided by the embodiments of the present disclosure;

[0021] Figure 4 For Figure 2 Flow chart of the specific implementation manner of step S103 in the illustrated embodiment;

[0022] Figure 5 For Figure 4 Flow chart of the specific implementation manner of step S1031 in the illustrated embodiment;

[0023] Figure 6 Schematic diagram of a process for generating a group photo image provided by the embodiments of the present disclosure;

[0024] Figure 7 Flow chart of the group photo image generation method provided by the embodiments of the present disclosure Figure 2 ;

[0025] Figure 8 For Figure 7 Flow chart of the specific implementation manner of step S203 in the illustrated embodiment;

[0026] Figure 9 Schematic diagram of a process for generating location information provided by the embodiments of the present disclosure;

[0027] Figure 10 For Figure 7 Flow chart of the specific implementation manner of step S204 in the illustrated embodiment;

[0028] Figure 11 Another schematic diagram of a process for generating a group photo image provided by the embodiments of the present disclosure;

[0029] Figure 12 Block diagram of the structure of the group photo image generation device provided by the embodiments of the present disclosure;

[0030] Figure 13 Schematic diagram of the structure of an electronic device provided by the embodiments of the present disclosure;

[0031] Figure 14 Schematic diagram of the hardware structure of the electronic device provided by the embodiments of the present disclosure. Detailed implementation manners

[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present disclosure with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.

[0033] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present disclosure are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.

[0034] The following explains the application scenarios of the embodiments of the present disclosure:

[0035] The group photo image generation method provided in the embodiments of the present disclosure can be applied to application programs (APPs, Applications) with image generation functions, such as video editing application programs, AI assistant application programs, etc. More specifically, it can be applied to application scenarios of generating group photo images based on Artificial Intelligence Generated Content (AIGC) technology. The execution subject of this embodiment can be a terminal device running the above-mentioned application program with image generation function, or a server deploying the server side corresponding to the above-mentioned application program, or other electronic devices with similar functions. Among them, when the execution subject is a terminal device, the terminal device executes the method provided in this embodiment by running the above-mentioned application program; when the execution subject is a server, the server side of the above-mentioned application program with image generation function can run partially or entirely on the server and execute the method provided in this embodiment on the server side, while the terminal device runs the client of the application program. Based on the server-client communication between the server and the terminal device, the terminal device can obtain the execution result of the method provided in this embodiment and display it as needed.

[0036] Among them, in some embodiments, the terminal device or the server can implement the group photo image generation method provided by the embodiments of the present disclosure by running various computer-executable instructions or computer programs. For example, the computer-executable instructions can be program-level commands, machine instructions, or software instructions. The computer program can be a native program or a software module in the operating system; it can be a local application, that is, a program that needs to be installed in the operating system to run, or it can be a small program embedded in any APP, that is, a program that runs based on the browser environment. In summary, the above computer-executable instructions can be instructions in any form, and the above computer programs can be application programs, modules, or plug-ins in any form, and the specific implementation form can be configured according to needs. Further, in the process of implementing the group photo image generation method provided by the embodiments of the present disclosure, the terminal device can execute the method by running the computer-executable instructions or computer programs set locally, or can execute the method by calling the computer-executable instructions or computer programs set in an external server. In some embodiments, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud storage, cloud communication, cloud databases, cloud computing, cloud functions, network services, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. Among them, the cloud service can be an interactive processing service for the terminal device to call.

[0037] Figure 1 A schematic diagram of an application scenario of the group photo image generation method provided by the embodiments of the present disclosure. Refer to Figure 1As shown in the figure, taking a terminal device as an example, a target application with an image generation function is running in the terminal device. The initiating user (User_1) triggers the group photo generation function of the target application, for example, clicks the "Take Group Photo" control as shown in the figure, creates a group photo group, and notifies the invited users (User_2, User_3, etc.) in the group photo group to upload corresponding user images to the server respectively, such as the images P1 and P2 shown in the figure. Additionally, optionally, the initiating user can also upload a user image in the group photo group. For example, after clicking the "Take a Photo" control as shown in the figure, takes and uploads the image P3 shown in the figure to the group photo group. After that, trigger the image generation function on the terminal device side, for example, click the "Generate Group Photo" control as shown in the figure, to generate a group photo image composed of different user images, thus achieving the effect of synthesizing a group photo image based on user images taken remotely. Then, according to the specific user needs, as shown in the figure, the synthesized image can be saved by triggering the "Save" control; or, by triggering the "Back" control, return to the previous page to make modifications again. Among them, the user images uploaded by the invited users (through the corresponding terminal devices) can be directly sent to the terminal device or sent to the server corresponding to the target application; the above process of generating the group photo image can be executed locally on the terminal device or executed on the server corresponding to the target application, and can be specifically configured according to needs.

[0038] In the prior art, for the scenario of group photos of multiple people, on-site shooting by multiple people is required, which has relatively high requirements for the venue and time. At the same time, affected by the states of the people being photographed, the resulting group photo may have defects, affecting the shooting effect of the photo. In some related technical solutions, unused user photos can be spliced into one picture by manual cropping and splicing to achieve the effect of a group photo image. However, the operation process of such solutions is complex, the interaction efficiency is low, and in the generated group photo image, the consistency between the group photo objects (such as people) and the image background is poor, affecting the imaging quality of the group photo image.

[0039] The embodiments of the present disclosure provide a method for generating a group photo image to solve the above problems.

[0040] Refer to Figure 2 , Figure 2 is a flowchart of the method for generating a group photo image provided by the embodiments of the present disclosure Figure 1 . The method of this embodiment can be applied to a terminal device or a server. Among them, the server can be used to deploy a function service implemented based on the method for generating a group photo image. The terminal device can access the above server and call the corresponding function service to implement the method for generating a group photo image provided by this embodiment. Exemplarily, the method for generating a group photo image provided by this embodiment includes:

[0041] Step S101: Obtain user images corresponding to at least two invited users.

[0042] Step S102: Determine a target group photo template according to the number of invited users and the group photo control information, where the group photo control information characterizes the group photo pose and / or the group photo formation.

[0043] Reference Figure 1 Referring to the schematic diagram of the application scenario shown, in this embodiment, the terminal device is used as the execution subject to introduce the provided group photo image generation method. Exemplarily, the terminal device provides a group photo generation function to the user by running a target application program. The user performs a trigger operation on the terminal device for the above-mentioned group photo generation function through the interaction interface of the target application program, so as to obtain user images corresponding to at least two invited users. In a possible implementation manner, after receiving a group creation instruction, the terminal device responds to it and creates a group photo group. The group photo group is a specific implementation manner of the user group, which is used to provide a user space and carry the invited users who join the group photo group. Specifically, the group photo group can be a chat group, and the users invited or joined in the chat group can at least upload user images in the group photo group. Optionally, based on the specific implementation manner of the group photo group, the invited users can also send text, voice, video and other contents in the group photo group. The group data corresponding to the group photo group can be stored on the server side of the target application program, that is, the terminal device creates a group photo group on the server side. In a possible implementation manner, after the group photo group is created on the server side, it includes at least two invited users. Then, the server side notifies the invited users to send user images to the server side by sending a notification message to the client corresponding to the invited users. After the invited users upload user images in the group photo group, the purpose of the terminal device to obtain user images corresponding to at least two invited users is realized; in another possible implementation manner, the terminal device obtains the user images corresponding to the invited users through other channels, such as the images shared by other invited users or the AI avatars of each invited user, etc. Specifically, the terminal device can obtain user images by accessing the server, receiving data transmission requests sent by other terminal devices, etc., which will not be elaborated here one by one.

[0044] Furthermore, after obtaining the user images, the terminal device determines a target group photo template according to the number of invited users and the group photo control information. In a possible implementation manner, the terminal device automatically determines the target group photo template according to the above information. Exemplarily, the specific implementation steps of step S102 include:

[0045] Step S1021: Obtain the description text and / or type identifier corresponding to the group photo control information;

[0046] Step S1022: Generate a template prompt based on the description text or type identifier and the number of invited users.

[0047] Step S1023: Generate a target group photo template based on the template prompt.

[0048] Exemplarily, the group photo control information can be a description text characterizing the group photo pose and / or formation, or a type identifier characterizing the group photo pose and / or formation, or a combination of both. By detecting the number of previously obtained user images and combining the above description text / type description, a template prompt is generated. Then, a pre-trained template generation large model is used to process the above template prompt to generate a target group photo template. Or, in another possible implementation, the above template prompt can be directly used as the target group photo template. In the subsequent process of generating group photo images, an image generation large model is used, and the template prompt is used as a control parameter for image generation, and combined with user images to generate group photo images that match the number of invited users and the group photo control information.

[0049] Furthermore, in another possible implementation, the target group photo template can be generated in combination with user operations. Exemplarily, before obtaining user images, it further includes: responding to a group creation instruction to create a group photo group, where the group photo group is used to carry at least two invited users; displaying user images uploaded by the invited users in the interaction interface of the group photo group; specifically, after creating the group photo group, the corresponding interaction interface of the group photo group is displayed, and then as the invited users send user images in the group photo group, on the terminal device side, in the interaction interface of the group photo group, the user images uploaded by each invited user are displayed. Specifically, for example, the user names of each invited user and the corresponding user images are displayed in the interaction interface. The user image can be a selfie of the invited user or an image of other content, such as a pet photo, etc. In the steps of this embodiment, the content of the user images uploaded by the invited users is not restricted. Figure 3 FIG. is a schematic diagram of an interaction interface of a group photo group provided by an embodiment of the present disclosure. The following will be combined with Figure 3 to introduce the above steps. As Figure 3 shown, in the interaction interface of the group photo group, different invited users upload corresponding user pictures at different times. For example, as shown in the figure, at time t1, the invited user User_1 uploads picture P1, at time t2, the invited user User_2 uploads picture P2, and at time t3, the invited user User_3 uploads picture P3. Through the group photo group, the collection of user images of different invited users is realized. Then, based on multiple user images in the group photo group, group photo images are generated, so as to achieve the purpose of generating multiple group photos based on images taken remotely.

[0050] After that, based on the above group photo group, the specific implementation of step S102 includes:

[0051] According to the number of user images displayed in the group photo group and the group photo control information, at least two alternative group photo templates are displayed. The alternative group photo templates include the image background of the group photo image and the body appearance of the object body. In response to a user operation, the target group photo template corresponding to the group photo group is determined from at least two alternative group photo templates. Specifically, first, the number of the above user images of the invited users in the group photo group is used as the first constraint condition. Among them, it matches the number of user images, that is, it matches the number of invited users (in the multi-person group photo scenario, one invited user can only upload one user image). Using the group photo control information as the second constraint condition, alternative group photo templates that meet the above first constraint condition and second constraint condition are selected from the group photo template library. After that, in response to a user instruction, the target group photo template is determined from the alternative group photo templates.

[0052] In a possible implementation manner, more specifically, before consuming the target group photo template, first, according to the number of user images uploaded by the previous invited users, a matching alternative group photo template is generated or obtained from the group photo template library, that is, the number of group photo objects in the alternative group photo template should be the same as the number of user images. After that, for the above alternative group photo template, according to the group photo poses and / or group photo formations of the person objects in the user images uploaded by the previous invited users, a matching group photo template is generated or obtained from the alternative group photo template as the screened alternative group photo template. The group photo poses and / or group photo formations of the group photo objects in the target group photo template (determined by the upload order of the user images) are the same as or similar to the group photo poses and / or group photo formations of the person objects in the user images. The above process of generating the target group photo template can be achieved by generating corresponding description features according to the number of user images and the group photo control information, and then inputting the description features into the template generation model to generate the corresponding target group photo template. The above process of selecting alternative group photo templates from the group photo template library can be realized by calculating the similarity between the group photo templates in the group photo template library and the corresponding number of user images and the postures of the person objects. Through the above steps in this embodiment, a target group photo template that matches the number of user images, the poses and formations of the person objects in the user images can be obtained, so that the group photo image generated based on the target group photo template and the user images is more realistic, reducing the sense of disconnection between the group photo objects and the image background, and thus improving the quality of the group photo image. Step S103: Generate a group photo image based on the target group photo template and the user images, where the group photo image includes at least two group photo objects. The target group photo template is used to generate the image background of the group photo image, and the user images are at least used to generate the object faces of the group photo objects.

[0053] Exemplarily, after collecting user images uploaded by multiple invited users through a group photo group, a target group photo template corresponding to the group photo group is obtained, and multiple user images in the group photo group are fused based on the target group photo template to generate a group photo image. Specifically, for the image background, in one possible implementation, the target group photo template contains relevant information describing the content of the image background. The content of the image background can be determined through the target group photo template, and then the image background is generated. Another possible image background can be fixed content determined based on user operations, such as a solid color background, that is, it has nothing to do with the target group photo template. In this case, the target group photo template is only used to generate the object bodies of the group photo objects. For the object bodies of the group photo objects, in one possible implementation, the target group photo template contains relevant information about the styles of the object bodies, specifically, for example, the clothing styles of the human objects. Multiple object bodies can be directly generated through the target group photo template. In another possible implementation, the object bodies are jointly determined by the target group photo template and the content of the user images. For example, the target group photo template determines the clothing color of the human object, and the user image determines the clothing shape of the human object, and so on. In still another possible implementation, the object bodies can also be independently determined by the user images. For example, according to the clothing of the human object in the user image, the object bodies are generated, and the object bodies are the result of style conversion of the clothing of the human object in the user image. That is, in this case, the target group photo target is only used to determine the image background.

[0054] In this embodiment, exemplarily, the target group photo template is used to generate the image background of the group photo image and the object bodies of the group photo objects, while the user images are at least used to generate the object faces of the group photo objects. The group photo image includes at least two group photo objects, that is, the group photo image is synthesized from at least two user images.

[0055] Further, in one possible implementation, the target group photo template is manually selected and determined by the user. Specifically, before obtaining the user images corresponding to at least two invited users, it further includes:

[0056] Further, in one possible implementation, the target group photo template includes a template prompt word, and the template prompt word is used to characterize the image content features of the group photo image. When the target group photo template is implemented in the form of a prompt word text, the alternative group photo templates are also correspondingly prompt word texts. After the user selects the target group photo template from multiple alternative group photo templates (that is, multiple prompt word texts), the target prompt word and the user images are directly input into an image generation model (that is, an AIGC model) to generate a group photo image with the image content features characterized by the template prompt word.

[0057] Further, in one possible implementation, as Figure 4As shown, the specific implementation of step S103 includes:

[0058] Step S1031: Based on the template prompt and the user image, generate a template base map, which includes an image background with image content features and an object body.

[0059] Step S1032: Extract facial features from the user image to obtain the corresponding facial image.

[0060] Step S1033: Blend the facial image into the object face corresponding to the object body in the template base map to generate a group photo image.

[0061] Exemplarily, the image generation model first generates a template base map according to the template prompt and the user image, based on the image content features described by the template prompt and the number of user images. The template base map includes an image background with image content features and an object body. Specifically, the content of the template prompt is, for example, "A group photo of 6 students in a classroom scene". Based on the semantics of the above model prompt, the template base map generated by the image generation model is an image of a "classroom scene", that is, the image background; in this image background, there are 6 human object figures in student costumes, that is, the object body. Here, the object body refers to the body part of the group photo object. Taking a human object as an example, the object body is the part of the human body excluding the head or face. After that, facial features are extracted from the user image to generate the corresponding facial image, and the facial image is blended into the object face corresponding to the object body in the template base map to generate a group photo image. The above process can be implemented based on the image generation ability of the image generation model, which is obtained during the model training process, and will not be elaborated here.

[0062] In the steps of this embodiment, a template base map with an image background and an object body having image content features is generated through the template prompt and the user image. Then, a facial image is generated in combination with the user image and blended into the object face corresponding to the object body in the template base map to generate a group photo image. This makes the target group photo template no longer a fixed style, but dynamically determines the image background, as well as the number and position of the group photo objects in combination with the situation of the user image, realizing the dynamic generation of the target group photo template, thereby avoiding the problems of limited number of people and position in the fixed-style group photo template and improving the content richness and imaging quality of the finally generated group photo image.

[0063] Further, in a possible implementation manner, as Figure 5 shown, the specific implementation of step S1031 includes:

[0064] Step S1031-1: Through the feature extraction model, obtain the image semantics corresponding to each user image, and the image semantics are at least used to characterize the clothing features of the human object in the user image.

[0065] Step S1031-2: Generate corresponding object bodies according to the image semantics of each user image.

[0066] Step S1031-3: Generate a template base map according to each object body and the template prompt words.

[0067] Exemplarily, the feature extraction model can be an independent model used in conjunction with the image generation model or a sub-model in the image generation model. In the step of generating the template base map, first, the image semantics corresponding to each user image are obtained through the feature extraction model. The image semantics are at least used to characterize the clothing features of the human object in the user image. Then, by utilizing the capabilities of the image generation model, corresponding object bodies are generated according to the image semantics of each user image, so that different user images have object bodies of different styles. Finally, a template base map is generated according to each object body and the template prompt words. The template base map consists of two parts: the image background and the object body. In this embodiment, during the process of generating the object body, the image semantics corresponding to the user image are combined to generate an object body that matches the clothing features of the human object in the user image. Then, in combination with the image background described by the template prompt words, the template base map is generated, thus making more full use of the effective information in the user image and making the group photo objects in the template base map and the finally generated group photo image more similar to the human objects in the user image, improving the authenticity of the group photo image.

[0068] Figure 6 FIG. is a schematic diagram of a process for generating a group photo image provided by an embodiment of the present disclosure. The following will be combined with Figure 6 , and a more detailed introduction to the above process will be given. As Figure 6 shown, after triggering the group photo image generation function, first, in the template page, a target group photo template is selected from multiple alternative group photo templates. The target group photo template includes template prompt words T1, T2, T3, etc. Among them, the content of the template prompt word T1 is, for example, [Group photo under the mountains]. After obtaining the user images P1, P2, and P3 through the group photo group, the template prompt word T1 and the user images P1, P2, and P3 are input into the image generation model. After the image generation model processes the above user images using the semantic feature extraction sub-model (not shown in the figure) included therein, in combination with the template prompt words, a template base map Q is generated. The template base map Q includes an image background and object bodies q1, q2, and q3 (shown as q1, q2, and q3 in the figure). Finally, facial images are generated using the user images P1, P2, and P3 and fused with the template base map Q to generate the final group photo image.

[0069] In this embodiment, user images corresponding to at least two invited users are obtained; a target group photo template is determined according to the number of invited users and the group photo control information, where the group photo control information characterizes the group photo pose and / or the group photo formation; a group photo image is generated based on the target group photo template and the user images, where the group photo image includes at least two group photo subjects, the target group photo template is used to generate the image background of the group photo image, and the user images are at least used to generate the faces of the group photo subjects. By determining the target group photo template according to the number of invited users and the group photo control information, and then combining the user images and the target group photo template to construct the body and face of the subject, a group photo image is generated. It is not necessary to take photos of multiple people on site, and it is possible to avoid facial expression defects of the group photo subjects in the group photo image, reduce the production cost of the group photo image, and improve the imaging quality of the group photo image.

[0070] Reference Figure 7 , Figure 7 is a schematic flowchart of a method for generating a group photo image provided by an embodiment of the present disclosure Figure 2 . This embodiment further refines step S103 on the basis of the embodiment shown in Figure 2 . The method for generating a group photo image includes:

[0071] Step S201: In response to a group creation instruction, a group photo group is created, and the user images uploaded by the invited users are displayed in the interaction interface of the group photo group. The group photo group is used to carry at least two invited users.

[0072] Step S202: Determine a target group photo template according to the number of invited users and the group photo control information.

[0073] Step S203: Configure position information for the user images based on the target group photo template. The position information is used to characterize the position of the group photo subject generated based on the user image in the group photo image.

[0074] Exemplarily, after creating a group photo group and obtaining the user images uploaded by the invited users, by responding to a configuration instruction, the position of the group photo subject generated based on the user image in the group photo image is determined, that is, the position information corresponding to the user image is determined. In a possible implementation manner, the position information is an image serial number corresponding to the user image, and a mapping is established between the image serial number and the body serial number of the object body in the target group photo template to determine the position of the group photo subject. Among them, the image serial number can be manually configured by the invited user or the initiating user, or can be determined based on the upload order of the user images, without limitation.

[0075] Furthermore, in addition to determining the position information of the user image based on the image serial number as described above, the position information of the user image can also be determined through the following operation interaction steps. In a possible implementation manner, as Figure 8As shown, the specific implementation of step S203 includes:

[0076] Step S2031: Display the template background image corresponding to the target group photo template, where the template background image includes the image background of the group photo image and at least two object bodies, and the at least two object bodies are located at preset positions in the image background, and the preset positions correspond to the target group photo template.

[0077] Step S2032: In response to a configuration instruction for the user image and the template background image, configure the mapping relationship between the user image and the object bodies in the template background image.

[0078] Step S2033: Generate the position information of the user image according to the mapping relationship between the user image and the object bodies.

[0079] Exemplarily, the target group photo template includes or corresponds to a template background image. The template background image is a kind of image data. The template background image includes the image background of the group photo image and at least two object bodies. The at least two object bodies are located at preset positions in the image background, and the preset positions correspond to the target group photo template, that is, in different target group photo templates, the positions of the group photo objects and the positional relationship between the group photo objects are different. After obtaining the template background image corresponding to the target group photo template, display the template background image, and by responding to the configuration instruction for the user image and the template background image, establish the mapping relationship between the user image and the object bodies in the template background image, so as to generate the position information of the user image. For example, pairwise click on the user image and the object body to establish the mapping relationship between the user image and the object body in the template background image, and then record the mapping relationship between the user image and the object body, so as to obtain the position information of each user image.

[0080] Figure 9 It is a schematic diagram of a process for generating position information provided by an embodiment of the present disclosure. The following will further introduce the above process in conjunction with Figure 9 as Figure 9As shown, first, within the template base map page, the template base map corresponding to the target group photo template and the user images uploaded in the current group photo group are displayed. Subsequently, the initiating user (i.e., the operating user of the terminal device) can click on the object bodies in the user images and the template base map to establish a mapping relationship between the two. For example, first click on user image P1 to make user image P1 in a highlighted state, then click on object body q2 to make object body q2 in a highlighted state, and then click the "Establish Mapping" control to establish the mapping relationship between user image P1 and object body q2. This mapping relationship is represented by a mapping identifier, more specifically, for example, as a mapping key-value pair or a mapping array [P1, q2], and this mapping relationship is the position information of user image P1. Repeat the above process to sequentially establish the mappings between image P2, image P3 and object bodies q1, q3 until all user images have established mapping relationships with object bodies. Based on the set of the above mapping relationships, the position information of each user image is generated.

[0081] In this embodiment, by loading the target group photo template for determining the positions of group photo objects, then displaying the template base map corresponding to the target group photo template, and establishing a mapping relationship with the user images, the rapid and efficient generation of the position information corresponding to the user images is realized. Subsequently, the object faces generated from the user images can be accurately configured to the corresponding positions, thereby generating a group photo image based on the group photo style provided by the target group photo template.

[0082] Further, optionally, before step S205, it further includes:

[0083] Step S204: Based on an image generation model, process the user image to generate a corrected user image, where the object face of the person object in the user image corresponds to a side shooting perspective, and the object face of the person object in the corrected user image corresponds to a front shooting perspective; or, the object face of the person object in the user image is partially blocked, and the object face of the person object in the corrected user image is unblocked.

[0084] Exemplarily, in the steps of this embodiment, after receiving the user image, the user image can be processed by an image generation model deployed locally or in the cloud on the terminal device to generate a corrected user image. Among them, the corrected user image is equivalent to the original user image and has better image quality. For example, noise reduction, increasing image sharpness, and super-resolution processing based on image generation technology are performed on the original user image.

[0085] In another possible implementation, the corrected user image is equivalent to the original user image, with a better viewing angle or visual effect. For example, in the original user image, the person object is not facing the camera directly, resulting in a side shooting angle for the person object, that is, only a part of the object's face is shown; or, a part of the object's face of the person object is blocked by (clothes, body, limbs), and only a part of the object's face is shown. In this case, the image generation ability of the image generation model is used to reconstruct the object's face of the person object to generate an object's face that can be displayed completely and frontally, so as to avoid facial defects in the subsequent process of using the object's face to generate a group photo object, and improve the image quality and authenticity of the group photo image.

[0086] Of course, it can be understood that the above step S204 is an optional step. In another possible implementation, step S204 may not be executed, and the subsequent steps may be directly performed using the user image, which will not be elaborated here.

[0087] Step S205: Generate a group photo image based on the target group photo template, the corrected user image, and the position information.

[0088] Exemplarily, after obtaining the target group photo template, the corrected user image, and the position information based on the above steps, the target group photo template is used to determine the image background of the group photo image and the object body of the group photo object, the corrected user image is used to determine the object face of the group photo object, and the position information is used to determine the position of the group photo object in the image background. Based on the above three groups of information, a complete description of the group photo image can be completed. Then, using a pre-trained image generation model and based on the AIGC technology for image generation, a group photo image that meets the above content description can be obtained.

[0089] Further, in a possible implementation, the position information includes a mapping identifier representing the mapping relationship between the user image and the object body, and the number of object bodies in the template base map is greater than the number of user images corresponding to the invited users. As Figure 10 shown, the specific implementation manner of step S204 includes:

[0090] Step S2041: Determine the target object body in the template base map according to the object body in the template base map and the mapping identifier corresponding to each user image. The target object body is the object body that has a mapping relationship with the user image.

[0091] Step S2042: Adjust the target positions of each target object body in the image background according to the number of target object bodies.

[0092] Step S2043: Based on the user image and the corresponding target object body, fuse them into the corresponding group photo object and display it at the target position corresponding to the target object body.

[0093] Exemplarily, when the number of object bodies in the template background image is greater than the number of user images corresponding to the invited users, only some of the object bodies are mapped. At this time, first, based on the object bodies in the template background image and the mapping identifiers corresponding to each user image, the effective object bodies that are mapped are determined, that is, the target object bodies. Subsequently, the group photo objects are constructed based on the target object bodies; while the object bodies that are not mapped by the user images (i.e., the position information does not indicate) cannot be matched with the user images and thus cannot generate group photo objects by fusing the user images, that is, invalid object bodies, and do not need to be displayed in the group photo image. After that, since the number of effective object bodies changes, the position arrangement of the group photo objects constructed based on the target object bodies may be unreasonable, such as uneven arrangement. Therefore, in this embodiment, after determining the target object bodies, according to the number of the target object bodies, the target positions of each target object body in the image background are adjusted. Then, the user images corresponding to each target object body are obtained, and object faces are generated based on the user images and then fused with the corresponding target object bodies and displayed at the target positions corresponding to the target object bodies, so that the position distribution of the group photo objects in the generated group photo image is more reasonable.

[0094] Figure 11 Another schematic diagram of the process of generating a group photo image provided by an embodiment of the present disclosure will be described below in conjunction with Figure 11 the above process for further introduction. As Figure 11 shown, the template background image includes four object bodies, namely object body q1 at addr_1, object body q2 at addr_2, object body q3 at addr_3, and object body q4 at addr_4 (shown as q1, q2, q3, and q4 in the figure). Then, based on the mapping information corresponding to user image P1, user image P2, and user image P3 (shown as P1, P2, and P3 in the figure) respectively, it is determined that object body q1, object body q2, and object body q4 are the target object bodies, that is, the object bodies that have a mapping relationship with the above user images. After that, based on the number of the above target object bodies, after reallocating and arranging object body q1, object body q2, and object body q4, the position of object body q1 is addr_5, the position of object body q2 is addr_6, and the position of object body q4 is addr_7. Then, based on the mapping relationship between user image P1, user image P2, and user image P3 and object body q1, object body q2, and object body q4, group photo object obj_1 is generated at addr_5, group photo object obj_2 is generated at addr_6, and group photo object obj_3 is generated at addr_7, thereby generating group photo objects.

[0095] In the steps of this embodiment, when the number of object bodies in the template base map is greater than the number of user images corresponding to the invited users, the target positions of each target object body in the image background are adjusted based on the number of target object bodies, so as to make the position distribution of each group photo object in the group photo image more reasonable and improve the quality of the group photo image.

[0096] Further, in another possible implementation manner, the position information includes position prompt words, and the position prompt words are used to describe the position arrangement characteristics between the group photo objects corresponding to each user image. The specific implementation manner of step S204 includes:

[0097] Step S204A: Process the target group photo template and the position prompt words through an image generation model, determine the target positions of each object body in the image background, and generate group photo objects at the corresponding target positions based on the user images and the corresponding object bodies to generate a group photo image.

[0098] Exemplarily, in another possible implementation manner, the position information is implemented through position prompt words, where the position prompt words are used to describe the position arrangement characteristics between the group photo objects corresponding to each user image. Exemplarily, the position prompt words can be input by the initiating user before or after creating the group photo group. Specifically, the content of the position prompt words is, for example, "stand in two rows one in front of the other" or "divide into multiple rows according to the specific number of people, with no more than 8 people in each row". It can be seen that the position prompt words are texts used to describe the position arrangement characteristics between the group photo objects corresponding to each user image and to constrain the position relationship between the group photo objects. The target group photo template is only used to determine the image background of the group photo image and no longer constrains the position relationship between the group photo objects. Then, the target group photo template and the position prompt words are processed through an image generation model. Based on the semantic understanding of the position prompt words, the image generation model determines the position relationship and arrangement method between the group photo objects, that is, determines the target positions of each object body in the image background, and fuses the user images and the corresponding object bodies to generate group photo objects and display them at the corresponding positions, and finally generates a group photo image.

[0099] In this embodiment, the position information of the user images is configured in the form of position prompt words, so that in the process of generating the group photo image, the positions and arrangement methods of the group photo objects are no longer restricted by the target group photo template, thereby greatly improving the diversity and flexibility of the generated group photo image.

[0100] In this embodiment, the implementation manners of steps S201 - S202 are the same as those of steps S101 - S102 in the embodiment Figure 2 shown in the present disclosure, and will not be elaborated here one by one.

[0101] Corresponding to the group photo image generation method in the above embodimentFigure 12 The structural block diagram of the group photo image generation device provided by the embodiments of the present disclosure. The methods introduced in the above embodiments can be executed by this group photo image generation device, which can be implemented in software and / or hardware, and can be integrated into an electronic device with certain data processing capabilities. Among them, the electronic device may include, but is not limited to, a mobile terminal with big data processing capabilities, and a fixed terminal with big data processing capabilities such as a desktop computer and a supercomputer.

[0102] For the sake of convenience of description, only parts related to the embodiments of the present disclosure are shown. Refer to Figure 12 , the group photo image generation device 3 includes:

[0103] A creation module 31, configured to obtain user images corresponding to at least two invited users;

[0104] A processing module 32, configured to determine a target group photo template according to the number of the invited users and the group photo control information, where the group photo control information characterizes a group photo pose and / or a group photo formation;

[0105] A generation module 33, configured to generate a group photo image based on the target group photo template and the user images, where the group photo image includes at least two group photo objects, the target group photo template is used to generate the image background of the group photo image, and the user images are at least used to generate the object faces of the group photo objects.

[0106] According to one or more embodiments of the present disclosure, the processing module 32 is specifically configured to: obtain a description text and / or a type identifier corresponding to the group photo control information; generate a template prompt word based on the description text or the type identifier and the number of the invited users; generate the target group photo template based on the template prompt word.

[0107] According to one or more embodiments of the present disclosure, after determining the target group photo template, the generation module 33 is further configured to: configure position information for the user images, where the position information is used to characterize the positions of the group photo objects generated based on the user images in the group photo image; when generating the group photo image based on the target group photo template and the user images, the generation module 33 is specifically configured to: generate the group photo image based on the target group photo template, the user images, and the position information.

[0108] According to one or more embodiments of the present disclosure, when generating the position information for the user image based on the target group photo template, the generating module 33 is specifically configured to: display the template background image corresponding to the target group photo template, where the template background image includes the image background of the group photo image and at least two object bodies, and the at least two object bodies are located at preset positions in the image background, and the preset positions correspond to the target group photo template; in response to the configuration instruction for the user image and the template background image, configure the mapping relationship between the user image and the object bodies in the template background image; and generate the position information of the user image according to the mapping relationship between the user image and the object bodies.

[0109] According to one or more embodiments of the present disclosure, the position information includes a mapping identifier characterizing the mapping relationship between the user image and the object body, and the number of object bodies in the template background image is greater than the number of user images corresponding to the invited users; when generating the group photo image based on the target group photo template, the user image, and the position information, the generating module 33 is specifically configured to: determine the target object bodies in the template background image according to the object bodies in the template background image and the mapping identifiers corresponding to each user image, where the target object bodies are the object bodies having a mapping relationship with the user image; adjust the target positions of the target object bodies in the image background according to the number of the target object bodies; and fuse the user image and the corresponding target object body into a corresponding group photo object and display it at the target position corresponding to the target object body.

[0110] According to one or more embodiments of the present disclosure, the position information includes position prompt words for describing the position arrangement characteristics between the group photo objects corresponding to each user image; when generating the group photo image based on the target group photo template, the user image, and the position information, the generating module 33 is specifically configured to: process the target group photo template and the position prompt words through an image generation model to determine the target positions of the object bodies in the image background, and generate group photo objects at the corresponding target positions based on the user image and the corresponding object bodies to generate the group photo image.

[0111] According to one or more embodiments of the present disclosure, before obtaining the user images corresponding to at least two invited users, the creating module 31 is further configured to:

[0112] In response to the group creation instruction, create a group photo group for hosting at least two invited users; display the user images uploaded by the invited users in the interaction interface of the group photo group; the processing module 32 is specifically configured to: display at least two alternative group photo templates according to the number of user images displayed in the group photo group and the group photo control information, where the alternative group photo templates include the image background of the group photo image and the body appearance of the object bodies; and determine the target group photo template corresponding to the group photo group in response to the user operation from the at least two alternative group photo templates.

[0113] According to one or more embodiments of the present disclosure, before generating a group photo image based on a target group photo template and a user image, the generating module 33 is further configured to: process the user image based on an image generation model to generate a corrected user image, where the object face of the person object in the user image corresponds to a side shooting angle, and the object face of the person object in the corrected user image corresponds to a front shooting angle; or, the object face of the person object in the user image is partially blocked, and the object face of the person object in the corrected user image is unblocked; when generating a group photo image based on the target group photo template and the user image, the generating module 33 is specifically configured to: generate a group photo image based on the target group photo template and the corrected user image.

[0114] According to one or more embodiments of the present disclosure, the target group photo template includes template prompt words, and the template prompt words are used to characterize the image content features of the group photo image; when generating a group photo image based on the target group photo template and the user image, the generating module 33 is specifically configured to: process the template prompt words and the user image through an image generation model to generate a group photo image.

[0115] According to one or more embodiments of the present disclosure, when the generating module 33 processes the template prompt words and the user image through an image generation model to generate a group photo image, the generating module 33 is specifically configured to: generate a template base map based on the template prompt words and the user image, where the template base map includes an image background and an object body with image content features; extract facial features from the user image to obtain a corresponding facial image; fuse the facial image into the object face corresponding to the object body in the template base map to generate a group photo image.

[0116] According to one or more embodiments of the present disclosure, when the generating module 33 generates a template base map based on the template prompt words and the user image, the generating module 33 is specifically configured to: obtain the image semantics corresponding to each user image through a feature extraction model, and the image semantics are at least used to characterize the clothing features of the person object in the user image; generate a corresponding object body according to the image semantics of each user image; generate a template base map according to each object body and the template prompt words.

[0117] Wherein, the creating module 31, the processing module 32, and the generating module 33 are connected in sequence. The group photo image generating device 3 provided in this embodiment can execute the technical solutions of the above method embodiments, and the implementation principles and technical effects are similar, which will not be elaborated here.

[0118] Figure 13 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. As Figure 13 shown, the electronic device 4 includes:

[0119] a processor 41 and a memory 42 communicatively connected to the processor 41;

[0120] The memory 42 stores computer-executable instructions;

[0121] The processor 41 executes the computer-executable instructions stored in the memory 42 to implement the group photo image generation method in the embodiments as Figures 2 - 11 shown in the embodiments.

[0122] Optionally, the processor 41 and the memory 42 are connected through a bus 43.

[0123] For relevant descriptions, reference may be made to Figures 2 - 11 the relevant descriptions and effects corresponding to the steps in the corresponding embodiments, and details are not elaborated here.

[0124] An embodiment of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, which are used to implement the group photo image generation method provided in any one of the corresponding embodiments of the present disclosure when executed by a processor. Figures 2 - 11 in the corresponding embodiments of the present disclosure.

[0125] An embodiment of the present disclosure provides a computer program product including a computer program, which implements the group photo image generation method provided in any one of the corresponding embodiments of the present disclosure when executed by a processor. Figures 2 - 11 in the corresponding embodiments of the present disclosure.

[0126] To implement the above embodiments, an embodiment of the present disclosure further provides an electronic device.

[0127] Referring to Figure 14 , which shows a schematic structural diagram of an electronic device 900 suitable for implementing the embodiments of the present disclosure. The electronic device 900 may be a terminal device or a server. Among them, the terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers, portable media players (PMPs), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs and desktop computers. Figure 14 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0128] As Figure 14As shown, the electronic device 900 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 901, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 are also stored. The processing device 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0129] Generally, the following devices may be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 908 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 may allow the electronic device 900 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 14 the electronic device 900 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0130] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 909, or installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above functions defined in the method of the embodiment of the present disclosure are executed.

[0131] It should be noted that the above-mentioned computer-readable medium in the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0132] The above-mentioned computer-readable medium may be included in the above-mentioned electronic device; or it may exist separately and not be assembled into the electronic device.

[0133] The above-mentioned computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to execute the method shown in the above embodiments.

[0134] Computer program code for performing the operations of this disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0135] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0136] The units or modules involved in the embodiments described in this disclosure may be implemented in software or in hardware. Wherein, the name of the unit or module does not constitute a limitation on the unit itself in some cases.

[0137] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.

[0138] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0139] In a first aspect, according to one or more embodiments of the present disclosure, there is provided a group photo image generation method, including:

[0140] Obtaining user images corresponding to at least two invited users; determining a target group photo template according to the number of the invited users and group photo control information, where the group photo control information characterizes a group photo pose and / or a group photo formation; generating a group photo image based on the target group photo template and the user images, where the group photo image includes at least two group photo objects, the target group photo template is used to generate an image background of the group photo image, and the user images are at least used to generate object faces of the group photo objects.

[0141] According to one or more embodiments of the present disclosure, the determining a target group photo template according to the number of the invited users and group photo control information includes: obtaining a description text and / or a type identifier corresponding to the group photo control information; generating a template prompt word based on the description text or the type identifier, and the number of the invited users; and generating the target group photo template based on the template prompt word.

[0142] According to one or more embodiments of the present disclosure, after determining the target group photo template, it further includes: configuring position information for the user images based on the target group photo template, where the position information is used to characterize positions of group photo objects generated based on the user images in the group photo image; and the generating a group photo image based on the target group photo template and the user images includes: generating a group photo image based on the target group photo template, the user images, and the position information.

[0143] According to one or more embodiments of the present disclosure, configuring the position information for the user image based on the target group photo template includes: displaying a template background image corresponding to the target group photo template, where the template background image includes an image background of the group photo image and at least two object bodies, and the at least two object bodies are located at preset positions in the image background, and the preset positions correspond to the target group photo template; in response to a configuration instruction for the user image and the template background image, configuring a mapping relationship between the user image and the object bodies in the template background image; and generating the position information of the user image according to the mapping relationship between the user image and the object bodies.

[0144] According to one or more embodiments of the present disclosure, the position information includes a mapping identifier characterizing the mapping relationship between the user image and the object body, and the number of object bodies in the template background image is greater than the number of user images corresponding to the invited users; generating a group photo image based on the target group photo template, the user image, and the position information includes: determining target object bodies in the template background image according to the object bodies in the template background image and the mapping identifiers corresponding to each user image, where the target object bodies are the object bodies having a mapping relationship with the user image; adjusting the target positions of each target object body in the image background according to the number of the target object bodies; and fusing the user image and the corresponding target object body into a corresponding group photo object and displaying it at the target position corresponding to the target object body.

[0145] According to one or more embodiments of the present disclosure, the position information includes position prompt words, and the position prompt words are used to describe the position arrangement characteristics between the group photo objects corresponding to each user image; generating a group photo image based on the target group photo template, the user image, and the position information includes: processing the target group photo template and the position prompt words through an image generation model to determine the target positions of each object body in the image background, and generating the group photo object at the corresponding target position based on the user image and the corresponding object body to generate the group photo image.

[0146] According to one or more embodiments of the present disclosure, before obtaining user images corresponding to at least two invited users, the method further includes: in response to a group creation instruction, creating a group photo group for hosting at least two invited users; displaying user images uploaded by the invited users within an interaction interface of the group photo group; determining a target group photo template according to the number of the invited users and group photo control information, including: displaying at least two alternative group photo templates according to the number of user images displayed in the group photo group and the group photo control information, where the alternative group photo templates include an image background of the group photo image and a body appearance of the object body; determining the target group photo template corresponding to the group photo group from the at least two alternative group photo templates in response to a user operation.

[0147] According to one or more embodiments of the present disclosure, before generating a group photo image based on the target group photo template and the user images, the method further includes: processing the user images based on an image generation model to generate corrected user images, where an object face of a person object in the user images corresponds to a side shooting perspective, and an object face of the person object in the corrected user images corresponds to a front shooting perspective; or, the object face of the person object in the user images is partially blocked, and the object face of the person object in the corrected user images is unblocked; generating the group photo image based on the target group photo template and the user images includes: generating the group photo image based on the target group photo template and the corrected user images.

[0148] According to one or more embodiments of the present disclosure, the target group photo template includes a template prompt word for characterizing image content features of the group photo image; generating the group photo image based on the target group photo template and the user images includes: processing the template prompt word and the user images through an image generation model to generate the group photo image.

[0149] According to one or more embodiments of the present disclosure, processing the template prompt word and the user images through an image generation model to generate the group photo image includes: generating a template base image based on the template prompt word and the user images, where the template base image includes an image background with the image content features and an object body; extracting facial features from the user images to obtain corresponding facial images; and fusing the facial images into object faces corresponding to the object body in the template base image to generate the group photo image.

[0150] According to one or more embodiments of the present disclosure, generating a template base map based on the template prompt and the user image includes: obtaining, through a feature extraction model, the image semantics corresponding to each of the user images, where the image semantics are at least used to characterize the clothing features of the human object in the user image; generating a corresponding object body according to the image semantics of each of the user images; and generating the template base map according to each of the object bodies and the template prompt.

[0151] In a second aspect, according to one or more embodiments of the present disclosure, there is provided a group photo image generation device, including:

[0152] A creation module, configured to obtain user images corresponding to at least two invited users;

[0153] A processing module, configured to determine a target group photo template according to the number of the invited users and the group photo control information, where the group photo control information characterizes a group photo pose and / or a group photo formation;

[0154] A generation module, configured to generate a group photo image based on the target group photo template and the user images, where the group photo image includes at least two group photo objects, the target group photo template is used to generate the image background of the group photo image, and the user images are at least used to generate the object faces of the group photo objects.

[0155] According to one or more embodiments of the present disclosure, the processing module is specifically configured to: obtain a description text and / or a type identifier corresponding to the group photo control information; generate a template prompt based on the description text or the type identifier and the number of the invited users; and generate the target group photo template based on the template prompt.

[0156] According to one or more embodiments of the present disclosure, after determining the target group photo template, the generation module is further configured to: configure position information for the user images based on the target group photo template, where the position information is used to characterize the position of the group photo object generated based on the user image in the group photo image; when the generation module generates a group photo image based on the target group photo template and the user images, it is specifically configured to: generate a group photo image based on the target group photo template, the user images, and the position information.

[0157] According to one or more embodiments of the present disclosure, when generating the position information for the user image based on the target group photo template, the generating module is specifically configured to: display the template background image corresponding to the target group photo template, where the template background image includes the image background of the group photo image and at least two object bodies, and the at least two object bodies are located at preset positions in the image background, and the preset positions correspond to the target group photo template; in response to the configuration instruction for the user image and the template background image, configure the mapping relationship between the user image and the object bodies in the template background image; and generate the position information of the user image according to the mapping relationship between the user image and the object bodies.

[0158] According to one or more embodiments of the present disclosure, the position information includes a mapping identifier representing the mapping relationship between the user image and the object body, and the number of object bodies in the template background image is greater than the number of user images corresponding to the invited users; when generating the group photo image based on the target group photo template, the user images, and the position information, the generating module is specifically configured to: determine the target object bodies in the template background image according to the object bodies in the template background image and the mapping identifiers corresponding to the respective user images, where the target object bodies are the object bodies having a mapping relationship with the user images; adjust the target positions of the respective target object bodies in the image background according to the number of the target object bodies; and fuse the user images and the corresponding target object bodies into corresponding group photo objects and display them at the target positions corresponding to the target object bodies.

[0159] According to one or more embodiments of the present disclosure, the position information includes a position prompt word, and the position prompt word is used to describe the position arrangement characteristics between the group photo objects corresponding to the respective user images; when generating the group photo image based on the target group photo template, the user images, and the position information, the generating module is specifically configured to: process the target group photo template and the position prompt word through an image generation model to determine the target positions of the respective object bodies in the image background, and generate group photo objects at the corresponding target positions based on the user images and the corresponding object bodies to generate the group photo image.

[0160] According to one or more embodiments of the present disclosure, before obtaining the user images corresponding to at least two invited users, the creating module is further configured to:

[0161] in response to the group creation instruction, create a group photo group for hosting at least two invited users; display the user images uploaded by the invited users in the interaction interface of the group photo group; the processing module is specifically configured to: display at least two alternative group photo templates according to the number of user images displayed in the group photo group and the group photo control information, where the alternative group photo templates include the image background of the group photo image and the body appearances of the object bodies; and determine the target group photo template corresponding to the group photo group in response to a user operation from the at least two alternative group photo templates.

[0162] According to one or more embodiments of the present disclosure, before generating a group photo image based on a target group photo template and a user image, the generating module is further configured to: process the user image based on an image generation model to generate a corrected user image, wherein the object face of the person object in the user image corresponds to a side shooting perspective, and the object face of the person object in the corrected user image corresponds to a front shooting perspective; or, the object face of the person object in the user image is partially blocked, and the object face of the person object in the corrected user image is unblocked; when generating a group photo image based on the target group photo template and the user image, the generating module is specifically configured to: generate a group photo image based on the target group photo template and the corrected user image.

[0163] According to one or more embodiments of the present disclosure, the target group photo template includes a template prompt, and the template prompt is used to characterize the image content features of the group photo image; when generating a group photo image based on the target group photo template and the user image, the generating module is specifically configured to: process the template prompt and the user image through an image generation model to generate a group photo image.

[0164] According to one or more embodiments of the present disclosure, when the generating module processes the template prompt and the user image through an image generation model to generate a group photo image, the generating module is specifically configured to: generate a template base map based on the template prompt and the user image, where the template base map includes an image background with image content features and an object body; extract facial features from the user image to obtain a corresponding facial image; and fuse the facial image into the object face corresponding to the object body in the template base map to generate a group photo image.

[0165] According to one or more embodiments of the present disclosure, when the generating module generates a template base map based on the template prompt and the user image, the generating module is specifically configured to: obtain the image semantics corresponding to each user image through a feature extraction model, where the image semantics is at least used to characterize the clothing features of the person object in the user image; generate a corresponding object body according to the image semantics of each user image; and generate a template base map according to each object body and the template prompt.

[0166] In a third aspect, according to one or more embodiments of the present disclosure, an electronic device is provided, including: at least one processor and a memory;

[0167] The memory stores computer-executable instructions;

[0168] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the group photo image generation method described in the first aspect above and various possible designs of the first aspect.

[0169] Fourthly, according to one or more embodiments of the present disclosure, there is provided a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the group photo image generation method as described in the first aspect above and various possible designs of the first aspect.

[0170] Fifthly, according to one or more embodiments of the present disclosure, there is provided a computer program product including a computer program, which, when executed by a processor, implements the group photo image generation method as described in the first aspect above and various possible designs of the first aspect.

[0171] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, a technical solution formed by mutually replacing the above features with technical features having similar functions (but not limited to) disclosed in the present disclosure.

[0172] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although a number of specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0173] Although the subject matter has been described in language specific to structural features and / or methodological act logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A method for generating a group photo image, characterized in that: include: Obtain user images corresponding to at least two invited users; determining a target group photo template according to the number of invited users and group photo control information, wherein the group photo control information represents a group photo posture and / or a group photo formation; A group photo image is generated based on the target group photo template and the user image, wherein the group photo image includes at least two group photo objects, the target group photo template is used to generate an image background of the group photo image, and the user image is used to generate at least object faces of the group photo objects.

2. The method according to claim 1, characterized in that The determining a target group photo template according to the number of invited users and the group photo control information includes: Obtaining a description text and / or a type identifier corresponding to the group photo control information; Generate a template prompt word based on the description text or type identifier and the number of invited users; The target group photo template is generated based on the template prompt word.

3. The method according to claim 1, characterized in that After determining the target photo template, the method further includes: configuring position information for the user image based on the target group photo template, where the position information is used to represent a position of a group photo object generated based on the user image in the group photo image; The generating a group photo image based on the target group photo template and the user image includes: A group photo image is generated based on the target group photo template, the user image and the location information.

4. The method according to claim 3, characterized in that The configuring the location information of the user image based on the target group photo template includes: Displaying a template base image corresponding to the target group photo template, wherein the template base image includes an image background of the group photo image and at least two object bodies, and the at least two object bodies are located at preset positions in the image background, and the preset positions correspond to the target group photo template; In response to a configuration instruction for the user image and the template base map, configuring a mapping relationship between the user image and an object body in the template base map; The position information of the user image is generated according to the mapping relationship between the user image and the object body.

5. The method according to claim 4, characterized in that The location information includes a mapping identifier representing a mapping relationship between a user image and an object body, and the number of object bodies in the template base map is greater than the number of user images corresponding to the invited users; The generating a group photo image based on the target group photo template, the user image and the location information includes: Determine a target object body in the template base map according to the object body in the template base map and the mapping identifiers corresponding to the user images, wherein the target object body is an object body that has a mapping relationship with the user image; According to the number of the target object bodies, adjusting the target position of each target object body in the image background; Based on the user image and the corresponding target object body, they are fused into a corresponding photo object and displayed at a target position corresponding to the target object body.

6. The method according to claim 3, characterized in that The position information includes a position prompt word, and the position prompt word is used to describe the position arrangement characteristics between the group photo objects corresponding to each of the user images; The generating a group photo image based on the target group photo template, the user image and the location information includes: The target group photo template and the position prompt word are processed by an image generation model to determine the target position of each object body in the image background, and the group photo object is generated at the corresponding target position based on the user image and the corresponding object body to generate the group photo image.

7. The method according to claim 1, characterized in that Before acquiring user images corresponding to at least two invited users, the method further includes: In response to the group creation instruction, create a group for taking a photo, where the group for taking a photo is used to carry at least two invited users; Displaying the user image uploaded by the invited user in the interactive interface of the photo group; The determining a target group photo template according to the number of invited users and the group photo control information includes: displaying at least two candidate group photo templates according to the number of user images displayed in the group photo group and the group photo control information, wherein the candidate group photo templates include the image background of the group photo image and the body appearance of the subject's body; In response to a user operation, a target group photo template corresponding to the group photo group is determined from the at least two candidate group photo templates.

8. The method according to claim 1, characterized in that Before generating the group photo image based on the target group photo template and the user image, the method further includes: Based on the image generation model, the user image is processed to generate a corrected user image, wherein the object face of the person object in the user image corresponds to a side shooting angle, and the object face of the person object in the corrected user image corresponds to a front shooting angle; or the object face of the person object in the user image is partially blocked, and the object face of the person object in the corrected user image is not blocked; The generating a group photo image based on the target group photo template and the user image includes: A group photo image is generated based on the target group photo template and the corrected user image.

9. The method according to claim 1, characterized in that: The target group photo template includes template prompt words, and the template prompt words are used to characterize image content features of the group photo image; The generating a group photo image based on the target group photo template and the user image includes: The template prompt word and the user image are processed by an image generation model to generate the group photo image.

10. The method according to claim 9, characterized in that The step of processing the template prompt word and the user image through an image generation model to generate the group photo image includes: Based on the template prompt word and the user image, a template base image is generated, wherein the template base image includes an image background having the image content features and a body of the object; Extracting facial features from the user image to obtain a corresponding facial image; The facial image is fused to the object face corresponding to the object body in the template base image to generate the group photo image.

11. The method according to claim 10, characterized in that The step of generating a template base map based on the template prompt word and the user image includes: Acquire image semantics corresponding to each of the user images through a feature extraction model, wherein the image semantics is at least used to characterize clothing features of a person object in the user image; Generating a corresponding object body according to the image semantics of each of the user images; The template base map is generated according to the body of each object and the template prompt words.

12. A group photo image generating device, characterized in that: include: Creating a module for obtaining user images corresponding to at least two invited users; a processing module, configured to determine a target group photo template according to the number of invited users and group photo control information, wherein the group photo control information represents a group photo posture and / or a group photo formation; A generating module is used to generate a group photo image based on the target group photo template and the user image, wherein the group photo image includes at least two group photo objects, the target group photo template is used to generate an image background of the group photo image, and the user image is used to generate at least an object face of the group photo object.

13. An electronic device, characterized in that: include: Processor and memory; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method for generating a group photo image according to any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions. When the processor executes the computer-executable instructions, the method for generating a group photo image according to any one of claims 1 to 11 is implemented.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for generating a group photo image according to any one of claims 1 to 11 is implemented.