A group photo generation method

By adjusting the portrait size based on the user's height information and using portrait segmentation and key point detection models, the problem of poor group photo results when users are at different distances from the camera has been solved, resulting in more realistic and aesthetically pleasing group photos.

CN115375593BActive Publication Date: 2026-07-28HISENSE GRP HLDG CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HISENSE GRP HLDG CO LTD
Filing Date
2021-05-20
Publication Date
2026-07-28

AI Technical Summary

Technical Problem

When generating virtual group photos, the distance between users and the camera varies, affecting the realism and aesthetics of the photos.

Method used

By adjusting the size ratio of the portrait based on the user's height information, and using a pre-trained portrait segmentation and key point detection model, the baseline portrait key points are determined and adjusted to ensure the correct display position of the portrait in the group photo.

Benefits of technology

It improves the realism and aesthetics of group photos, ensuring that the proportions and positions of people in the photos match the user's height, thus enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115375593B_ABST
    Figure CN115375593B_ABST
Patent Text Reader

Abstract

The application discloses a group photo generation method, a display device, an apparatus, an electronic device and a medium, which are used to ensure the authenticity of the group photo effect and improve user experience. According to the received group photo request instruction, the application can determine target portraits of target users in the video frames to be generated, which meet the set requirements, and can determine the height information corresponding to each target user according to the correspondence between the saved user and height information. According to the height information corresponding to each target user, the height ratio between each target user is determined. The size of the target portrait of each target user can be adjusted according to the height ratio between each target user. Since the size ratio of each target portrait after size adjustment is the same as the height ratio between each user, the authenticity of the generated group photo effect can be ensured when the group photo is generated according to the target portrait after size adjustment, and the user experience can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of group photo generation technology, and in particular to a group photo generation method, display device, apparatus, electronic device and medium. Background Technology

[0002] Virtual group photos allow users to connect with friends, family, colleagues, and clients in different locations, creating group photos that include multiple users and bringing convenience and fun to their work and lives. These virtual group photos are primarily achieved using image segmentation and image compositing techniques. For example, for each user currently in a video call, the video frames containing that user's video stream are first segmented to obtain their portrait (an image containing only the user without a background is called a portrait). Then, an image compositing algorithm is used to combine the portraits of multiple users onto a preset background image, generating a new image—a group photo containing the portraits of the users in the video call.

[0003] However, if each user is at a different distance from the camera, directly compositing the portraits obtained from each video frame into the group photo will usually affect the realism of the group photo. Summary of the Invention

[0004] This application provides a method, display device, apparatus, electronic device, and medium for generating group photos, in order to ensure the authenticity of the group photo effect and improve the user experience.

[0005] Firstly, this application provides a method for generating group photos, the method comprising:

[0006] Based on the received group photo request instruction, determine the target human image of the target user that meets the set requirements contained in the video frame of the group photo to be generated;

[0007] Based on the saved correspondence between user and height information, determine the height information corresponding to the target user, and based on the height information, determine the height ratio among the target users;

[0008] Based on the height ratio, the size of the target portrait is adjusted, and a group photo is generated based on the adjusted target portrait.

[0009] Secondly, this application also provides a method for generating group photos, the method comprising:

[0010] Based on the received group photo request instruction, the video frame to be generated is input into the pre-trained portrait segmentation model to obtain the portrait of the target user contained in the video frame.

[0011] The video frame is input into a pre-trained portrait key point detection model to determine the portrait key points contained in the portrait of the video frame;

[0012] Determine the key points of the baseline portrait, adjust the portrait according to the key points of the baseline portrait, and use the adjusted portrait as the target portrait that meets the set requirements; wherein the key points of the portrait contained in the adjusted portrait are the baseline portrait key points;

[0013] Generate a group photo based on the target portrait.

[0014] Thirdly, this application also provides a method for generating group photos, the method comprising:

[0015] Based on the received group photo request instruction, determine the target human image of the target user that meets the set requirements contained in the video frame of the group photo to be generated;

[0016] Based on the vertical coordinate of the display position of the target user in the group photo as set for the target user, determine the vertical coordinate of the display position of the target user's portrait in the group photo.

[0017] A group photo is generated based on the displayed location.

[0018] Fourthly, this application provides a display device, the display device comprising:

[0019] A display screen for displaying group photos;

[0020] Controller, the controller being used to perform:

[0021] Based on the height ratio between the target users to be photographed, the size of the target portraits of the target users in the video frame is adjusted, and the group photo is generated based on the adjusted target portraits.

[0022] Fifthly, this application also provides a display device, the display device comprising:

[0023] A display screen for displaying group photos;

[0024] Controller, the controller being used to perform:

[0025] Determine the reference portrait key points, and adjust the portrait of the target user contained in the video frame of the group photo to be generated based on the reference portrait key points; use the adjusted portrait as the target portrait that meets the set requirements; wherein the portrait key points contained in the adjusted portrait are the reference portrait key points;

[0026] Generate a group photo based on the target portrait.

[0027] Sixthly, this application also provides a display device, the display device comprising:

[0028] A display screen for displaying group photos;

[0029] Controller, the controller being used to perform:

[0030] Based on the vertical coordinate of the display position of the target user in the group photo set for the target user to be generated, determine the vertical coordinate of the display position of the target image of the target user in the group photo in the video frame.

[0031] A group photo is generated based on the displayed location.

[0032] Seventhly, this application provides a group photo generation apparatus, the apparatus comprising:

[0033] The first determining module is used to determine the target image of the target user that meets the set requirements in the video frame of the group photo to be generated, based on the received group photo request instruction;

[0034] The second determining module is used to determine the height information corresponding to the target user based on the stored correspondence between user and height information, and to determine the height ratio between the target users based on the height information.

[0035] The first generation module is used to adjust the size of the target portrait according to the height ratio, and generate a group photo based on the target portrait after size adjustment.

[0036] Eighthly, this application also provides a group photo generation apparatus, the apparatus comprising:

[0037] The first input module is used to input the video frame of the group photo to be generated into the pre-trained portrait segmentation model according to the received group photo request instruction, so as to obtain the portrait of the target user contained in the video frame.

[0038] The second input module is used to input the video frame into a pre-trained portrait key point detection model to determine the portrait key points contained in the portrait of the video frame.

[0039] An adjustment module is used to determine the key points of a reference portrait, adjust the portrait according to the key points of the reference portrait, and use the adjusted portrait as the target portrait that meets the set requirements; wherein the key points of the portrait contained in the adjusted portrait are the key points of the reference portrait.

[0040] The second generation module is used to generate a group photo based on the target portrait.

[0041] Ninthly, this application also provides a group photo generation apparatus, the apparatus comprising:

[0042] The third determining module is used to determine the target image of the target user that meets the set requirements in the video frame of the group photo to be generated, based on the received group photo request instruction;

[0043] The fourth determining module is used to determine the vertical coordinate of the display position of the target image of the target user in the group photo based on the vertical coordinate of the display position of the target user in the group photo set for the target user;

[0044] The third generation module is used to generate a group photo based on the display position.

[0045] In a tenth aspect, this application provides an electronic device comprising at least a processor and a memory, wherein the processor is configured to execute a computer program stored in the memory to implement the steps of any of the above-described group photo generation methods.

[0046] In another aspect, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of any of the above-described group photo generation methods.

[0047] This application can determine the target human images of the target users that meet the set requirements within the video frame containing the group photo to be generated, based on the received group photo request instruction. It can also determine the height information corresponding to each target user based on the saved correspondence between user and height information, and determine the height ratio between each target user based on their height information. Furthermore, it can adjust the size of each target user's image based on the height ratio between target users, and then generate a group photo based on each adjusted target human image. Since the size ratio of each adjusted target human image in this application is the same as the height ratio between each user, the generation of group photos based on the adjusted target human images can guarantee the realism of the generated group photo effect and improve the user experience. Attached Figure Description

[0048] To more clearly illustrate the implementation methods in the embodiments of this application or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings.

[0049] Figure 1 A schematic diagram of a first group photo generation process provided by some embodiments is shown;

[0050] Figure 2 A schematic diagram of a second group photo generation process provided by some embodiments is shown;

[0051] Figure 3 A schematic diagram of a third group photo generation process provided in some embodiments is shown;

[0052] Figure 4 A schematic diagram of a fourth group photo generation process provided in some embodiments is shown;

[0053] Figure 5 A schematic diagram of an electronic device provided in some embodiments is shown;

[0054] Figure 6 A schematic diagram of a fifth group photo generation process provided in some embodiments is shown;

[0055] Figure 7 A schematic diagram of a first display device provided in some embodiments is shown;

[0056] Figure 8 A schematic diagram of a first group photo generation apparatus provided in some embodiments is shown;

[0057] Figure 9 A schematic diagram of a sixth group photo generation process provided in some embodiments is shown;

[0058] Figure 10 A schematic diagram of a second display device provided in some embodiments is shown;

[0059] Figure 11 A schematic diagram of a second group photo generation apparatus provided in some embodiments is shown;

[0060] Figure 12 A schematic diagram of the seventh group photo generation process provided in some embodiments is shown;

[0061] Figure 13 A schematic diagram of a third display device provided in some embodiments is shown;

[0062] Figure 14 A schematic diagram of a third group photo generation apparatus provided in some embodiments is shown;

[0063] Figure 15 A schematic diagram of an electronic device structure provided by some embodiments is shown. Detailed Implementation

[0064] To ensure the authenticity of group photos and improve user experience, this application provides a group photo generation method, display device, apparatus, electronic device, and medium.

[0065] To make the objectives and implementation methods of this application clearer, the exemplary implementation methods of this application will be clearly and completely described below with reference to the accompanying drawings of the exemplary embodiments of this application. Obviously, the exemplary embodiments described are only some embodiments of this application, and not all embodiments.

[0066] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.

[0067] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.

[0068] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.

[0069] The term "module" refers to any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.

[0070] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0071] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the described embodiments and various different variations of embodiments suitable for specific use considerations.

[0072] In practical use, the electronic device can determine the target user's image that meets the set requirements within the video frame containing the group photo to be generated, based on the received group photo request instruction. It can also determine the height information corresponding to each target user based on the saved correspondence between user and height information, and determine the height ratio between each target user. Furthermore, it can adjust the size of each target user's image based on the height ratio between target users, and then generate a group photo based on each adjusted target image. Since the size ratio of each adjusted target image in this application is the same as the height ratio between each user, generating a group photo based on the adjusted target images can ensure the authenticity of the generated group photo and improve the user experience.

[0073] Figure 1 The diagram illustrates a group photo generation process according to some embodiments, which includes the following steps:

[0074] S101: Based on the received group photo request instruction, determine the target image of the target user that meets the set requirements contained in the video frame of the group photo to be generated.

[0075] The group photo generation method provided in this application is applied to electronic devices, such as PCs, mobile terminals, servers, home automation systems, social TVs, etc.

[0076] In one possible implementation, the video frame to be used to generate the group photo can be a video frame from the video stream of the user currently making a video call, or any video frame provided by the user. This application does not impose specific limitations on this, and the choice can be made flexibly according to needs. Regardless of the form of the video frame, a group photo can be generated based on the group photo generation method provided in the embodiments of this application. For ease of understanding, the following example illustrates the group photo generation process provided in the embodiments of this application, using a video frame from each video stream of each user currently making a video call as the example.

[0077] In one possible implementation, the electronic device may include an image acquisition module such as a camera, allowing users to conduct video calls. During a video call, any user can trigger a group photo request by clicking a "Invite Group Photo" button, and other users can accept the invitation by clicking an "OK" button. This application does not specifically limit the number of users participating in the video call; for example, there can be one or more other users, which can be flexibly set according to needs. For ease of description, the user who clicks the "Invite Group Photo" button (the initiator) is referred to as the first user, and the user who clicks the "OK" button to accept the invitation (the invitee) is referred to as the second user. The electronic device used by the first user is referred to as the first electronic device, and the electronic device used by the second user is referred to as the second electronic device. In one possible implementation, after receiving the group photo request instruction, the electronic device (such as the first electronic device) can obtain the video stream of each user participating in the video call from a video cloud server. For example, the first electronic device can obtain the video streams corresponding to the first and second users from the video cloud server based on the user identification information of the first user who clicked the "Invite Group Photo" button and the user identification information of the second user who clicked the "OK" button. Each user corresponds to a video stream; for example, the first user corresponds to the first video stream, and the second user corresponds to the second video stream.

[0078] After acquiring the video stream of each user participating in a video call, the electronic device can determine the target portrait of the user (target user) that meets the set requirements within the video frames of each video stream. In one possible implementation, when determining the target portrait of the target user that meets the set requirements, the video frames of each video stream of the video call to be generated can be input into a pre-trained portrait segmentation model, and each portrait obtained can be identified as the target portrait that meets the set requirements.

[0079] S102: Based on the saved correspondence between user and height information, determine the height information corresponding to the target user, and based on the height information, determine the height ratio between the target users conducting the video call.

[0080] In one possible implementation, to obtain a user's height information, the electronic device can pre-store the correspondence between users and their height information. For example, when a user first uses an application on the electronic device that can generate group photos, the device can use facial recognition to identify the user (user identification information) and display a prompt to input height information. The user can then input their height, and the electronic device can store the correspondence between the user (user identification information) and their height. In another possible implementation, the electronic device (such as a second electronic device) can upload its stored correspondence between users and their height to a video cloud server. The electronic device (such as a first electronic device) can obtain and store the correspondence between users and their height from other electronic devices (such as a second electronic device) through the video cloud server. When a group photo request instruction is received, the electronic device (such as the first electronic device) can use image recognition to determine which user (target user) is included in the video frame where the group photo is to be generated (user identification information). Based on the stored correspondence between users and their height, the height information corresponding to each target user can be determined. After determining the height information of each target user, the electronic device (such as the first electronic device) can determine the height ratio between each target user in the video call. For example, if we denote the height of the first user as A and the height of the second user as B, then the height ratio between the first user and the second user is A / B.

[0081] S103: Adjust the size of the target portrait according to the height ratio, and generate a group photo based on the adjusted target portrait.

[0082] To ensure the authenticity of the group photo, an electronic device (such as a first electronic device) can adjust the size of the target portraits according to the height ratio between the target users. The adjusted size of each target portrait matches the height ratio between the target users. For example, see [link to relevant documentation]. Figure 2 , Figure 2 The diagram illustrates a second group photo generation process provided by some embodiments, such as... Figure 2As shown, taking the above embodiment as an example, if the height of the first user A is represented by A, and the height of the second user B is represented by B, for example, the height ratio A / B of the first user and the second user is 2:1, assuming that the distances of the first user and the second user from the camera are different, the size ratio of the first user's portrait (target portrait) A' to the second user's portrait (target portrait) B' is not 2:1, that is, there is a distortion in the portrait size ratio and the user height ratio. If this portrait is directly composited into a group photo, it will usually affect the realism of the group photo effect. In order to improve the realism of the group photo effect, this application can adjust the size of each target portrait according to the height ratio between the target users in the video call. For example, the ratio of the target portrait A' to B' after the size adjustment is the same as the height ratio A / B, which is also 2:1. Since the size ratio of each target portrait after the size adjustment is the same as the height ratio between each user, when generating a virtual group photo based on the target portrait after the size adjustment, the realism of the generated group photo effect can be guaranteed, and the user experience can be improved.

[0083] Specifically, the electronic device may also include a display terminal such as a monitor, and the generated group photo can be displayed on the display terminal, allowing users to view the generated group photo. In one possible implementation, the electronic device (such as the first electronic device) can also send the generated group photo to the terminal (second electronic device) corresponding to other target users (second users) among the target users. For example, the first electronic device can send the generated group photo to a video cloud server, and the second electronic device can obtain the group photo from the video cloud server, allowing the second user to view the group photo through the second electronic device. It is understood that the other target users (second users) are users who are having a video call with the currently logged-in target user (first user).

[0084] This application can determine the target human images of the target users that meet the set requirements within the video frame containing the group photo to be generated, based on the received group photo request instruction. It can also determine the height information corresponding to each target user based on the saved correspondence between user and height information, and determine the height ratio between each target user based on their height information. Furthermore, it can adjust the size of each target user's image based on the height ratio between target users, and then generate a group photo based on each adjusted target human image. Since the size ratio of each adjusted target human image in this application is the same as the height ratio between each user, the generation of group photos based on the adjusted target human images can guarantee the realism of the generated group photo effect and improve the user experience.

[0085] In practical use, taking a portrait containing five key facial features (head, chest, abdomen, knees, and feet) as an example, if each user in a video call is at a different distance from the camera, the number of key facial features in each user's portrait may differ. For example, a user closer to the camera may have fewer key facial features, perhaps only the head, chest, and abdomen (upper body), while a user farther from the camera may have more, perhaps four (head, chest, abdomen, and knees). Generating a group photo based on portraits with different key facial features may affect the aesthetics of the photo. To improve the aesthetics of the group photo, based on the above embodiments, in this embodiment, determining the target portrait of the target user in the video frame containing the group photo to meet the set requirements includes:

[0086] The video frame to be generated is input into a pre-trained portrait segmentation model to obtain the portrait of the target user contained in the video frame.

[0087] The video frame is input into a pre-trained portrait key point detection model to determine the portrait key points contained in the portrait of the video frame;

[0088] Determine the key points of the baseline portrait, adjust the portrait according to the key points of the baseline portrait, and use the adjusted portrait as the target portrait that meets the set requirements; wherein the key points of the portrait contained in the adjusted portrait are the key points of the baseline portrait.

[0089] In one possible implementation, when an electronic device (such as a first electronic device) determines a target human image that meets the set requirements, for each video stream in a video call, it can first input the video frames of that video stream into a pre-trained human image segmentation model to obtain the human image of the target user contained in the video frames. Alternatively, the video frames can also be input into a pre-trained human keypoint detection model to determine the human keypoints contained in the human image of that video frame.

[0090] In one possible implementation, the process of training the human image segmentation model includes:

[0091] Obtain any first sample image containing a human portrait from the sample set. Each pixel of the first sample image corresponds to a sample category label indicating whether the pixel is a human portrait.

[0092] The original human portrait segmentation model is used to determine the recognition category label of each pixel in the first sample image;

[0093] The original portrait segmentation model is trained based on the sample category label and the recognition category label to obtain the trained portrait segmentation model.

[0094] In this embodiment of the application, the sample set includes multiple first sample images. For each pixel in each first sample image, the pixel has a corresponding sample category label, which is used to identify whether the pixel is a human portrait pixel.

[0095] When training the original portrait segmentation model, a first sample image containing a human figure can be obtained from the sample set. Each pixel of this first sample image corresponds to a sample category label. This obtained first sample image is then input into the original portrait segmentation model, which uses the model to obtain the recognition category label for each pixel of the first sample image.

[0096] In practice, after determining the recognition category label for each pixel of the first input sample image, since the sample category labels for each pixel of the first sample image are pre-saved, the accuracy of the portrait segmentation model's recognition result can be determined by whether the sample category labels match the recognition category labels. If they do not match, it indicates that the portrait segmentation model's recognition result is inaccurate, and the parameters of the portrait segmentation model need to be adjusted to train the model.

[0097] In practice, when adjusting the parameters in the portrait segmentation model, the gradient descent algorithm can be used to backpropagate the gradient of the parameters of the portrait segmentation model, thereby training the portrait segmentation model.

[0098] In one possible implementation, the above operation can be performed on each first sample image in the sample set, and when the preset convergence condition is met, the training of the portrait segmentation model is determined to be complete.

[0099] The preset convergence condition can be that the number of first sample images correctly identified by the original portrait segmentation model in the sample set is greater than a set number, or the number of iterations for training the portrait segmentation model reaches the set maximum number of iterations. These conditions can be flexibly set in practice and are not specifically limited here.

[0100] In one possible implementation, when training the original portrait segmentation model, the first sample image in the sample set can be divided into a training first sample image and a test first sample image. The original portrait segmentation model is first trained based on the training first sample image, and then the reliability of the trained portrait segmentation model is verified based on the test first sample image.

[0101] In one possible implementation, the process of training the portrait keypoint detection model includes:

[0102] Obtain any second sample image in the sample set that contains human key points. The second sample image corresponds to a human key point sample category label and the sample position information of the human key points in the second sample image corresponding to the human key point sample category label. The human key point sample category label is used to identify the category of the human key points contained in the second sample image.

[0103] Using the original portrait key point detection model, the portrait key point recognition category label and corresponding recognition location information of the portrait key points contained in the second sample image are determined;

[0104] The original portrait key point detection model is trained based on the portrait key point sample category label, the portrait key point recognition category label, the sample location information, and the recognition location information to obtain the trained portrait key point detection model.

[0105] In order to accurately determine the key points of a person's image, in this application, the sample set contains multiple second sample images. Each second sample image containing a key point of a person's image corresponds to a key point of a person's image sample category label. The key point of a person's image sample category label is used to identify the category of the key point of the person's image contained in the second sample image. The category of the key point of a person's image can be head, chest, abdomen, knee, foot, etc.

[0106] To obtain the location information of facial key points, the second sample image also includes the sample location information of each facial key point corresponding to its sample category label within the second sample image. This sample location information may include the coordinates of the pixels in the second sample image of the lower left, lower right, upper left, upper right, or center point of the target bounding box of the facial key point.

[0107] When training the original portrait keypoint detection model, a second sample image containing any portrait keypoint can be acquired from the sample set. This second sample image corresponds to a portrait keypoint sample category label and the sample location information of the portrait keypoint corresponding to the portrait keypoint sample category label within the second sample image. This acquired second sample image is then input into the original portrait keypoint detection model. Through the original portrait keypoint detection model, the portrait keypoint recognition category label and corresponding recognition location information of the portrait keypoint corresponding to the second sample image are obtained.

[0108] In practice, after determining the facial keypoint recognition category label and corresponding recognition location information of the input second sample image, since the facial keypoint sample category label and the sample location information of the corresponding facial keypoint in the second sample image are pre-saved, the accuracy of the facial keypoint detection model's recognition result can be determined by whether the facial keypoint sample category label and the facial keypoint recognition category label are consistent, and whether the sample location information and the recognition location information are consistent. In practice, if they are inconsistent, it indicates that the recognition result of the facial keypoint detection model is inaccurate, and the parameters of the facial keypoint detection model need to be adjusted to train the model.

[0109] In practice, when adjusting the parameters in the portrait key point detection model, the gradient descent algorithm can be used to backpropagate the gradient of the parameters of the portrait key point detection model, thereby training the portrait key point detection model.

[0110] In one possible implementation, the above operation can be performed on each second sample image in the sample set, and when the preset convergence condition is met, it is determined that the training of the portrait key point detection model is complete.

[0111] The preset convergence condition can be that the number of second sample images correctly identified by the original portrait keypoint detection model in the sample set is greater than a set number, or the number of iterations for training the portrait keypoint detection model reaches the set maximum number of iterations, etc. These conditions can be flexibly set in practice and are not specifically limited here.

[0112] In one possible implementation, when training the original portrait key point detection model, the second sample images in the sample set can be divided into training second sample images and test second sample images. The original portrait key point detection model is first trained based on the training second sample images, and then the reliability of the trained portrait key point detection model is verified based on the test second sample images.

[0113] To improve the aesthetics of group photos, an electronic device (such as a first electronic device) can determine a reference portrait key point. Based on this reference portrait key point, the portrait obtained from the portrait segmentation model is adjusted so that the portrait key points in each adjusted portrait are the reference portrait key points. Then, the adjusted portrait is used as the target portrait that meets the set requirements. Based on the adjusted portrait (the target portrait that meets the set requirements), the subsequent steps of adjusting the size of the target portrait according to the height ratio between the target users are performed.

[0114] To accurately determine the key points of the reference portrait, based on the above embodiments, in this application, determining the key points of the reference portrait includes:

[0115] Based on the established portrait key points, determine the reference portrait key points; or,

[0116] The minimum number of facial key points in the image of the target user making the video call is determined as the baseline facial key point.

[0117] In one possible implementation, to quickly determine reference facial key points, one or more facial key points can be pre-set and designated as reference facial key points. For example, the head, chest, and other parts of the upper body can be pre-set as reference facial key points. Based on these set reference facial key points, the facial image of each user (target user) participating in a video call can be quickly adjusted.

[0118] In one possible implementation, to improve the flexibility of determining the baseline facial key points, the baseline facial key points suitable for each user's current facial image can be flexibly determined based on the facial key points contained in the facial image of each user currently participating in the video call. Specifically, the minimum number of facial key points contained in the facial image of each target user participating in the video call can be determined as the baseline facial key points. Based on these baseline facial key points, the facial images of each user (target user) participating in the video call can be flexibly adjusted.

[0119] To facilitate understanding, the process of determining the benchmark portrait key points provided in this application will be described below through a specific embodiment. Taking a pre-set portrait key point system including five key points such as head, chest, abdomen, knees, and feet as an example, if the portraits of the first user and the second user contain the same five key points, for example, both of them are the same five key points, then the benchmark portrait key points are the five key points of the head, chest, abdomen, knees, and feet. It can be considered that the portraits of the first user and the second user are both portraits containing complete portrait key points, and thus the portrait can be directly used as the target portrait that meets the set requirements.

[0120] If the portraits of the first user and the second user contain different key points, refer to... Figure 3 , Figure 3 The diagram illustrates a third group photo generation process provided by some embodiments, such as... Figure 3As shown, for example, the portrait of the first user (User A) contains three key points: head, chest, and abdomen, while the portrait of the second user (User B) contains four key points: head, chest, abdomen, and knees. The head, chest, and abdomen can be used as the baseline portrait key points. Portrait key point matching is then performed on the second user's portrait. Specifically, based on these baseline key points, the portion of the second user's portrait below the abdomen key point and above the knee key point is discarded. Adjustments are then made to the second user's portrait, resulting in the target portrait that meets the set requirements. Both the adjusted second user's portrait and the first user's portrait contain the three baseline portrait key points: head, chest, and abdomen. Taking a 1:1 height ratio between the first and second users as an example, the size ratio of the target portrait (meeting the requirements) can be adjusted to 1:1 based on this ratio. When generating a group photo using this adjusted target portrait, not only is the realism of the group photo guaranteed, but its aesthetics are also improved.

[0121] To ensure the realism and aesthetics of the group photo, based on the above embodiments, in this application, the step of generating a group photo based on the target portrait after size adjustment includes:

[0122] Determine the display position of the resized target portrait in the group photo;

[0123] A group photo is generated based on the displayed location.

[0124] Specifically, when generating a group photo, the electronic device can first determine the display position of each target image in the group photo after resizing, and then place each target image in its corresponding display position to generate the group photo. In one possible implementation, determining the display position of a target image in the group photo can be based on the horizontal and vertical coordinates of the target image within the video frame to which it belongs. For example, the horizontal and vertical coordinates of the target image can be determined based on the horizontal and vertical coordinates of the pixels at the bottom left, bottom right, top left, top right, and center point of the target frame within the target image's frame within the video frame. It is understood that when determining the horizontal and vertical coordinates of a target image within the group photo based on its horizontal and vertical coordinates within the video frame, the user can control the vertical coordinate of the target image by moving it forward and backward, and simultaneously control its horizontal coordinate by moving it left and right.

[0125] In one possible implementation, when determining the x and y coordinates of the target person in the group photo based on the x and y coordinates of the target person in the video frame to which the target person belongs, if the positions of each user's target person in the video frame are cluttered, the display position of each user's target person in the group photo may also be cluttered, thus affecting the aesthetics and realism of the group photo. To improve the aesthetics and realism of the group photo, based on the above embodiments, in this application, determining the display position of the resized target person in the group photo includes:

[0126] Based on the vertical coordinate of the target user in the group photo as set for the target user, determine the vertical coordinate of the target user's sized portrait in the group photo.

[0127] In one possible implementation, the vertical coordinates of the pixels representing the bottom left, bottom right, top left, top right, and center points of the target frame in the group photo can be pre-set for each target user. For each target user's portrait, the vertical coordinates of that portrait in the group photo can be determined based on the pre-set vertical coordinates. The vertical coordinates set for each target user can be the same or different, and can be flexibly set according to needs. For example, the same vertical coordinates can be set for each target user, so that the portraits of each target user are arranged side-by-side in the group photo. See also... Figure 4 , Figure 4 The diagram illustrates a fourth group photo generation process provided in some embodiments, such as... Figure 4 As shown in the diagram, the left side (left and right in the diagram) illustrates the group photo effect, where the horizontal and vertical coordinates (display positions) of each target user in the group photo are determined based on their horizontal and vertical coordinates within the video frame to which they belong. It can be seen that the vertical coordinates (display positions) of each target user are rather cluttered, resulting in a less realistic and aesthetically pleasing group photo. The right side (left and right in the diagram) illustrates the group photo effect, where each target user has the same vertical coordinate, causing their images to be arranged side-by-side in the group photo. It can be seen that the display positions of each target user are more aligned, resulting in a more realistic and aesthetically pleasing group photo.

[0128] The vertical coordinate of each target user in the group photo can be flexibly set according to needs, and this application does not impose specific limitations on it. When setting the vertical coordinate of each target user in the group photo, one target user can be set first, such as setting the vertical coordinate of the first user in the group photo first, and then setting the vertical coordinate of the second user in the group photo to be the same as the vertical coordinate of the first user, or to be at a set distance from the vertical coordinate of the first user, thereby completing the setting of the vertical coordinate of the second user in the group photo.

[0129] In one possible implementation, the horizontal coordinate of the target person in the group photo can be determined based on the horizontal coordinate of the target person's image within the video frame to which the target person's image belongs; this will not be elaborated further here. To ensure the aesthetics and realism of the group photo, based on the above embodiments, in this embodiment, determining the display position of the resized target person's image in the group photo includes:

[0130] Determine whether there is overlap between pixels of target portraits with the same vertical coordinate. If so, adjust the horizontal coordinate of the target portrait with the same vertical coordinate according to the set horizontal spacing. Based on the adjusted horizontal coordinate, determine the display position of the target portrait in the group photo.

[0131] In one possible implementation, when determining the display position of the resized target portrait in the group photo, it can be determined whether there is overlap between pixels of target portraits with the same vertical coordinate. If there is overlap, it can be assumed that there may be occlusion between target portraits of target users located in the same row (with the same vertical coordinate) in the group photo. To ensure the aesthetics and realism of the group photo, the horizontal coordinates of the target portraits with the same vertical coordinate can be adjusted according to a set horizontal spacing, so that the horizontal spacing between the adjusted target portraits is the set horizontal spacing, thereby preventing target portraits with the same vertical coordinate from occluding each other. For example, see [reference needed]. Figure 4 In the right image, the person on the left is referred to as the first person, and the person on the right is referred to as the second person. The horizontal coordinates of the bottom-right corner pixel of the target frame containing the first person and the bottom-left corner pixel of the target frame containing the second person can be determined first. The distance between the adjusted horizontal coordinates of the bottom-right corner pixel of the target frame containing the first person and the bottom-left corner pixel of the target frame containing the second person is the set horizontal spacing. The horizontal spacing can be flexibly set according to requirements; for example, it can be any value not less than 0. This application does not impose specific limitations on this.

[0132] To facilitate understanding, the group photo generation process provided in this application will be described below through a specific embodiment. (See reference...) Figure 5 , Figure 5 The diagram illustrates a schematic representation of an electronic device structure provided in some embodiments, such as... Figure 5As shown, the electronic device can be a social TV. The image acquisition module, such as a camera, within the electronic device can capture video streams from users making video calls. After receiving a group photo request instruction, the home host in the electronic device can retrieve each video stream from the video cloud server and determine the target user's image that meets the set requirements within the video frame containing the group photo. The home host can determine the height information of each target user in the video call based on the stored correspondence between user and height information. Based on this height information, it can determine the height ratio between each target user in the video call and adjust the size of each target image according to this height ratio. Then, it generates a group photo based on the adjusted target image and controls the display (display terminal) in the electronic device to display (play) the generated group photo. In one possible implementation, the home host in the electronic device can send the generated group photo to the terminals corresponding to other target users (users currently logged into the electronic device who are making video calls with the target user) via the video cloud server.

[0133] To facilitate understanding, the group photo generation process provided in this application will be described below through a specific embodiment. (See reference...) Figure 6 , Figure 6 The diagram illustrates a fifth group photo generation process provided in some embodiments, such as... Figure 6 As shown, the process includes the following steps:

[0134] S601: When a user uses the photo-generating application on an electronic device for the first time, the device can identify the user's identity information through facial recognition and display a prompt to input height information. The user can then input their height, and the electronic device can save the mapping between the user (identifier information) and their height. Simultaneously, the electronic device can upload this saved mapping to a video cloud server. Furthermore, the electronic device can retrieve and save the mappings between users and their heights from other electronic devices via the video cloud server.

[0135] S602: When a group photo request instruction is received, the electronic device can determine the user identification information of each user (target user) in the video call, and thus determine the height information of each target user according to the correspondence between the saved user and height information, and determine the height ratio between each target user in the video call according to the height information of each target user.

[0136] S603: The electronic device inputs the video frames of each video stream of the video call to be generated into a pre-trained portrait segmentation model, which is obtained from the video cloud server, to obtain the portrait of the target user contained in the video frames of each video stream; and inputs the video frames of each video stream of the video call to be generated into a pre-trained portrait key point detection model to determine the portrait key points contained in the portrait of the target user contained in the video frames of each video stream; determines the reference portrait key points, and adjusts each portrait according to the reference portrait key points, and uses the adjusted portrait as the target portrait that meets the set requirements, wherein the portrait key points contained in the adjusted portrait are the reference portrait key points.

[0137] S604: The electronic device adjusts the size of each target image that meets the set requirements according to the height ratio between the target users making the video call.

[0138] S605: The electronic device determines the display position of each target portrait after size adjustment in the group photo; based on the display position, it generates a group photo containing each target portrait after size adjustment.

[0139] S606: The electronic device sends the generated group photo to the terminals of other target users among the target users via a video cloud server. These other target users are users who are having a video call with the currently logged-in target user.

[0140] Based on the same technical concept, this application also provides a display device. Figure 7 A schematic diagram of a first type of display device provided in some embodiments is shown, such as Figure 7 As shown, the display device includes:

[0141] Display 71, the display 71 being used to display the group photo;

[0142] Controller 72, the controller 72 being configured to perform:

[0143] Based on the height ratio between the target users to be photographed, the size of the target portraits of the target users in the video frame is adjusted, and the group photo is generated based on the adjusted target portraits.

[0144] In one possible implementation, the display device can perform the steps of the electronic device performing the corresponding function in the above method, which will not be described in detail here.

[0145] Based on the same technical concept, this application also provides a group photo generation device. Figure 8 Schematic diagrams of a first type of group photo generation apparatus provided in some embodiments are shown, such as Figure 8 As shown, the device includes:

[0146] The first determining module 81 is used to determine the target image of the target user that meets the set requirements in the video frame of the group photo to be generated, based on the received group photo request instruction.

[0147] The second determining module 82 is used to determine the height information corresponding to the target user based on the stored correspondence between user and height information, and to determine the height ratio between the target users based on the height information.

[0148] The first generation module 83 is used to adjust the size of the target portrait according to the height ratio, and generate a group photo based on the target portrait after size adjustment.

[0149] In one possible implementation, the first determining module 81 is specifically used to input the video frame of the group photo to be generated into the pre-trained portrait segmentation model, and determine the obtained portrait as the target portrait that meets the set requirements.

[0150] In one possible implementation, the first determining module 81 is specifically used to input the video frame of the group photo to be generated into a pre-trained portrait segmentation model to obtain the portrait of the target user contained in the video frame.

[0151] The video frame is input into a pre-trained portrait key point detection model to determine the portrait key points contained in the portrait of the video frame;

[0152] Determine the key points of the baseline portrait, adjust the portrait according to the key points of the baseline portrait, and use the adjusted portrait as the target portrait that meets the set requirements; wherein the key points of the portrait contained in the adjusted portrait are the key points of the baseline portrait.

[0153] In one possible implementation, the first determining module 81 is specifically configured to determine the reference portrait key points based on the set portrait key points; or,

[0154] The minimum number of facial key points in the image of the target user making the video call is determined as the baseline facial key point.

[0155] In one possible implementation, the first generation module 83 is specifically used to determine the display position of the resized target portrait in the group photo;

[0156] A group photo is generated based on the displayed location.

[0157] In one possible implementation, the first generation module 83 is specifically used to determine the vertical coordinate of the target user's sized target image in the group photo based on the vertical coordinate set for the target user in the group photo.

[0158] In one possible implementation, the first generation module 83 is specifically used to determine whether there is overlap between the pixels of the target portrait with the same vertical coordinate. If so, the horizontal coordinate of the target portrait with the same vertical coordinate is adjusted according to the set horizontal spacing, and the display position of the target portrait in the group photo is determined based on the adjusted horizontal coordinate.

[0159] In one possible implementation, the device further includes:

[0160] The first sending module is used to send the generated group photo to the terminals corresponding to other target users among the target users, wherein the other target users are users who are having a video call with the currently logged-in target user.

[0161] In one possible implementation, based on the above embodiments, this application also provides a method for generating group photos, see below. Figure 9 , Figure 9 The diagram illustrates a sixth group photo generation process provided in some embodiments, such as... Figure 9 As shown, the process includes the following steps:

[0162] S901: Based on the received group photo request instruction, the video frame of the group photo to be generated is input into the pre-trained portrait segmentation model to obtain the portrait of the target user contained in the video frame.

[0163] S902: Input the video frame into a pre-trained portrait key point detection model to determine the portrait key points contained in the portrait of the video frame;

[0164] S903: Determine the reference portrait key points, adjust the portrait according to the reference portrait key points, and use the adjusted portrait as the target portrait that meets the set requirements; wherein the portrait key points contained in the adjusted portrait are the reference portrait key points;

[0165] S904: Generate a group photo based on the target portrait.

[0166] The group photo generation method provided in this application is applied to electronic devices, such as PCs, mobile terminals, servers, home automation systems, social TVs, etc.

[0167] Since this application can adjust the portrait of the target user based on the determined reference portrait key points, and the portrait key points contained in the adjusted portrait are all reference portrait key points, when generating a group photo based on the adjusted portrait (target portrait), the aesthetics and realism of the group photo effect can be improved compared with generating a group photo based on the unadjusted portrait.

[0168] In one possible implementation, determining the key points of the reference portrait includes:

[0169] Based on the established portrait key points, determine the reference portrait key points; or,

[0170] The minimum number of facial key points in the image of the target user making the video call is determined as the baseline facial key point.

[0171] In one possible implementation, generating a group photo based on the target portrait includes:

[0172] Determine the display position of the target portrait in the group photo;

[0173] A group photo is generated based on the displayed location.

[0174] In one possible implementation, determining the display position of the target portrait in the group photo includes:

[0175] Based on the vertical coordinate of the target user in the group photo, determine the vertical coordinate of the target user's portrait in the group photo.

[0176] In one possible implementation, determining the display position of the target portrait in the group photo includes:

[0177] Determine whether there is overlap between pixels of target portraits with the same vertical coordinate. If so, adjust the horizontal coordinate of the target portrait with the same vertical coordinate according to the set horizontal spacing. Based on the adjusted horizontal coordinate, determine the display position of the target portrait in the group photo.

[0178] In one possible implementation, the method further includes:

[0179] The generated group photo is sent to the terminals of other target users among the target users, wherein the other target users are users who are having a video call with the currently logged-in target user.

[0180] Based on the same technical concept, this application also provides a display device. Figure 10 The diagram illustrates a second display device provided in some embodiments, such as... Figure 10 As shown, the display device includes:

[0181] Display 101, the display 101 being used to display the group photo;

[0182] Controller 102, the controller 102 being configured to perform:

[0183] Determine the reference portrait key points, and adjust the portrait of the target user contained in the video frame of the group photo to be generated based on the reference portrait key points; use the adjusted portrait as the target portrait that meets the set requirements; wherein the portrait key points contained in the adjusted portrait are the reference portrait key points;

[0184] Generate a group photo based on the target portrait.

[0185] In one possible implementation, the display device can perform the steps of the electronic device performing the corresponding function in the above method, which will not be described in detail here.

[0186] Based on the same technical concept, this application also provides a group photo generation device. Figure 11 Schematic diagrams of a second group photo generation apparatus provided in some embodiments are shown, such as Figure 11 As shown, the device includes:

[0187] The first input module 111 is used to input the video frame of the group photo to be generated into the pre-trained portrait segmentation model according to the received group photo request instruction, so as to obtain the portrait of the target user contained in the video frame.

[0188] The second input module 112 is used to input the video frame into a pre-trained portrait key point detection model to determine the portrait key points contained in the portrait of the video frame.

[0189] The adjustment module 113 is used to determine the key points of the reference portrait, adjust the portrait according to the key points of the reference portrait, and use the adjusted portrait as the target portrait that meets the set requirements; wherein the key points of the portrait contained in the adjusted portrait are the key points of the reference portrait.

[0190] The second generation module 114 is used to generate a group photo based on the target portrait.

[0191] In one possible implementation, the adjustment module 113 is specifically used to determine the reference portrait key points based on the set portrait key points; or,

[0192] The minimum number of facial key points in the image of the target user making the video call is determined as the baseline facial key point.

[0193] In one possible implementation, the second generation module 114 is specifically used to determine the display position of the target portrait in the group photo; and generate the group photo based on the display position.

[0194] In one possible implementation, the second generation module 114 is specifically used to determine the vertical coordinate of the target user's portrait in the group photo based on the vertical coordinate set for the target user in the group photo.

[0195] In one possible implementation, the second generation module 114 is specifically used to determine whether there is overlap between the pixels of the target portrait with the same vertical coordinate. If so, the horizontal coordinate of the target portrait with the same vertical coordinate is adjusted according to the set horizontal spacing, and the display position of the target portrait in the group photo is determined based on the adjusted horizontal coordinate.

[0196] In one possible implementation, the device further includes:

[0197] The second sending module is used to send the generated group photo to the terminals corresponding to other target users among the target users, wherein the other target users are users who are having a video call with the currently logged-in target user.

[0198] In one possible implementation, based on the above embodiments, this application also provides a method for generating group photos, see below. Figure 12 , Figure 12 The diagram illustrates a seventh group photo generation process provided in some embodiments, such as... Figure 12 As shown, the process includes the following steps:

[0199] S1201: Based on the received group photo request instruction, determine the target image of the target user that meets the set requirements contained in the video frame of the group photo to be generated;

[0200] S1202: Determine the vertical coordinate of the display position of the target image of the target user in the group photo based on the vertical coordinate of the display position of the target user in the group photo as set for the target user;

[0201] S1203: Generate a group photo based on the displayed position.

[0202] The group photo generation method provided in this application is applied to electronic devices, such as PCs, mobile terminals, servers, home automation systems, social TVs, etc.

[0203] Since this application can determine the vertical coordinate of the target user's image in the group photo based on the vertical coordinate of the display position set for the target user in the group photo, when generating the group photo based on the display position, it can improve the aesthetics and realism of the group photo effect compared to related technologies that do not set the vertical coordinate of the display position of the target user in the group photo.

[0204] In one possible implementation, the method further includes:

[0205] Determine whether there is overlap between pixels of target portraits with the same vertical coordinate. If so, adjust the horizontal coordinate of the target portrait with the same vertical coordinate according to the set horizontal spacing. Based on the adjusted horizontal coordinate, determine the display position of the target portrait in the group photo.

[0206] In one possible implementation, the method further includes:

[0207] The generated group photo is sent to the terminals of other target users among the target users, wherein the other target users are users who are having a video call with the currently logged-in target user.

[0208] Based on the same technical concept, this application also provides a display device. Figure 13 The diagram illustrates a third display device provided in some embodiments, such as... Figure 13 As shown, the display device includes:

[0209] Display 131, the display 131 being used to display the group photo;

[0210] Controller 132, the controller 132 being configured to perform:

[0211] Based on the vertical coordinate of the display position of the target user in the group photo set for the target user to be generated, determine the vertical coordinate of the display position of the target image of the target user in the group photo in the video frame.

[0212] A group photo is generated based on the displayed location.

[0213] In one possible implementation, the display device can perform the steps of the electronic device performing the corresponding function in the above method, which will not be described in detail here.

[0214] Based on the same technical concept, this application also provides a group photo generation device. Figure 14 Schematic diagrams of a third group photo generation apparatus provided in some embodiments are shown, such as Figure 14 As shown, the device includes:

[0215] The third determining module 141 is used to determine the target image of the target user that meets the set requirements in the video frame of the group photo to be generated, based on the received group photo request instruction.

[0216] The fourth determining module 142 is used to determine the vertical coordinate of the display position of the target image of the target user in the group photo based on the vertical coordinate of the display position of the target user in the group photo set for the target user;

[0217] The third generation module 143 is used to generate a group photo based on the display position.

[0218] In one possible implementation, the device further includes:

[0219] The judgment module is used to determine whether there is overlap between the pixels of the target portrait with the same vertical coordinate. If so, the horizontal coordinate of the target portrait with the same vertical coordinate is adjusted according to the set horizontal spacing. Based on the adjusted horizontal coordinate, the display position of the target portrait in the group photo is determined.

[0220] In one possible implementation, the device further includes:

[0221] The third sending module is used to send the generated group photo to the terminals corresponding to other target users among the target users, wherein the other target users are users who are having a video call with the currently logged-in target user.

[0222] Based on the same technical concept, this application also provides an electronic device. Figure 15 The diagram illustrates a schematic representation of an electronic device structure provided in some embodiments, such as... Figure 15 As shown, it includes: processor 151, communication interface 152, memory 153 and communication bus 154, wherein processor 151, communication interface 152 and memory 153 communicate with each other through communication bus 154.

[0223] The memory 153 stores a computer program, which, when executed by the processor 151, causes the processor 151 to perform the steps of the electronic device performing the corresponding function in the above method.

[0224] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0225] Communication interface 152 is used for communication between the above-mentioned electronic device and other devices.

[0226] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0227] The processors mentioned above can be general-purpose processors, including central processing units, network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits, field-programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0228] Based on the same technical concept, this application also provides a computer-readable storage medium storing a computer program executable by an electronic device, wherein computer-executable instructions are used to cause a computer to perform the process executed in the aforementioned method section.

[0229] The aforementioned computer-readable storage medium can be any available medium or data storage device that can be accessed by the processor in an electronic device, including but not limited to magnetic storage such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), optical storage such as CDs, DVDs, BDs, HVDs, etc., and semiconductor storage such as ROMs, EPROMs, EEPROMs, non-volatile memory (NAND flash), solid-state drives (SSDs), etc.

[0230] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0231] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0232] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0233] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0234] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for generating group photos, characterized in that, The method includes: Based on the received group photo request instruction, determine the target human image of the target user that meets the set requirements contained in the video frame of the group photo to be generated; Based on the saved correspondence between user and height information, the height information corresponding to the target user is determined, and based on the height information, the height ratio between the target users is determined; wherein, when a user uses the service for the first time, the user is identified by facial recognition, and a prompt message for inputting height information is displayed, the input height information is received, and the correspondence between the identified user and the input height information is saved; Based on the height ratio, the size of the target portrait is adjusted, and a group photo is generated based on the adjusted target portrait. The target portraits of the target user that meet the set requirements are included in the video frame containing the group photo to be generated. The video frame to be generated is input into a pre-trained portrait segmentation model to obtain the portrait of the target user contained in the video frame. The video frame is input into a pre-trained portrait key point detection model to determine the portrait key points contained in the portrait of the video frame; Determine the key points of the baseline portrait, adjust the portrait according to the key points of the baseline portrait, and use the adjusted portrait as the target portrait that meets the set requirements; wherein the key points of the portrait contained in the adjusted portrait are the key points of the baseline portrait.

2. The method according to claim 1, characterized in that, The target portraits of the target user that meet the set requirements are included in the video frame containing the group photo to be generated. The video frames of the group photo to be generated are input into the pre-trained portrait segmentation model, and the resulting portraits are identified as target portraits that meet the set requirements.

3. The method according to claim 1, characterized in that, The key points for determining the baseline portrait include: Based on the established portrait key points, determine the reference portrait key points; or, The minimum number of facial key points in the image of the target user making the video call is determined as the baseline facial key point.

4. The method according to any one of claims 1-3, characterized in that, The process of generating a group photo based on the target portrait after size adjustment includes: Determine the display position of the resized target portrait in the group photo; A group photo is generated based on the displayed location.

5. The method according to claim 4, characterized in that, The determination of the display position of the target portrait after size adjustment in the group photo includes: Based on the vertical coordinate of the target user in the group photo as set for the target user, determine the vertical coordinate of the target user's sized portrait in the group photo.

6. The method according to claim 4, characterized in that, The determination of the display position of the target portrait after size adjustment in the group photo includes: Determine whether there is overlap between pixels of target portraits with the same vertical coordinate. If so, adjust the horizontal coordinate of the target portrait with the same vertical coordinate according to the set horizontal spacing. Based on the adjusted horizontal coordinate, determine the display position of the target portrait in the group photo.

7. The method according to claim 1, characterized in that, The method further includes: The generated group photo is sent to the terminals of other target users among the target users, wherein the other target users are users who are having a video call with the currently logged-in target user.