Virtual group photo generation method and device

By using a diffusion model in virtual group photo technology to integrate characters and backgrounds naturally, the limitations of virtual group photo unnatural and user manual location selection in the existing technology are solved, and a more natural and convenient virtual group photo generation is achieved.

CN119963403APending Publication Date: 2025-05-09HISENSE GRP HLDG CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311476812.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-08
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

When the existing virtual photo technology deals with inconsistent colors of portraits and backgrounds, the generated virtual photo has a sense of abruptness and a floating portrait, and requires the user to manually select the character position, which limits the limitations of the application.

Method used

By extracting the area where the character is located in the image containing the character, input the sub-image and background image of the region into the pre-trained diffusion model, and obtaining a virtual group photo of the characters output from the diffusion model fused in the background image.

Benefits of technology

It realizes the natural integration of characters and background, eliminates the need for users to manually select locations, and improves the naturalness of virtual photos and the convenience of application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963403A_ABST
    Figure CN119963403A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a virtual group photo generation method and equipment, which are used for solving the problems that a person and a background are not natural enough and a user needs to manually select the position of the person when a virtual group photo is generated in the prior art. According to the embodiment of the invention, the electronic equipment extracts the area where the figure is located in the image containing the figure; and inputting the sub-image of the area and the background image into a pre-trained diffusion model, and obtaining a virtual group photo after the character output by the diffusion model is fused in the background image. Due to the fact that fusion of the character and the background image can be well achieved through the diffusion model, the character and the background in the generated virtual group photo are natural, and a user does not need to manually select the position where the character is located.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image generation or image synthesis, and in particular to a method and device for generating a virtual group photo. Background Art

[0002] Virtual group photo is a technology that virtually combines the characters in one image with the background in another image. Virtual group photo can be used in entertainment works such as movies, TV series, games, etc. to make virtual characters more realistic and enhance the audience experience. At the same time, it also provides creators with more freedom and creative possibilities to create unprecedented character images. In addition, virtual group photo can also be applied to other scenarios.

[0003] The synthesis process commonly used in virtual group photos in related technologies can be expressed by the formula: I = α*F + (1-α) * B, where F is the pixel value of the pixel point of the portrait, α is the corresponding weight, and B is the pixel value of the pixel point of the background image. The current virtual group photo technology has the following problems: First, when the color of the portrait and the background image are inconsistent, there will be an obvious sense of abruptness and floating of the portrait after synthesis, that is, the person and the background in the virtual group photo are abrupt and unnatural; second, it can be seen from the above formula that the size of the portrait and the background image must be consistent. Although the image size can be ignored by using layer overlay, this also requires the user to manually select the position of the person, which also increases the limitations of the application. Summary of the invention

[0004] The embodiments of the present application provide a method and device for generating a virtual group photo, which are used to solve the problem in the prior art that when generating a virtual group photo, the characters and the background are not natural enough and the user needs to manually select the position of the characters.

[0005] In a first aspect, an embodiment of the present application provides a method for generating a virtual group photo, the method comprising:

[0006] Extracting the area where the person is located in the image containing the person;

[0007] The sub-image of the region and the background image are input into a pre-trained diffusion model, and a virtual photo of the person output by the diffusion model after being fused with the background image is obtained.

[0008] In a second aspect, an embodiment of the present application further provides an electronic device, which includes at least a processor and a memory, and the processor is used to implement the steps of the virtual group photo generation method as described in any one of the above items when executing a computer program stored in the memory.

[0009] In the embodiment of the present application, the electronic device extracts the area where the person is located in the image containing the person; the sub-image of the area and the background image are input into the pre-trained diffusion model to obtain a virtual group photo of the person fused with the background image output by the diffusion model. Since the diffusion model can well realize the fusion of the person and the background image, the person and the background in the generated virtual group photo are more natural, and there is no need for the user to manually select the position of the person. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0011] Figure 1 A schematic diagram of a process of generating a virtual group photo provided in an embodiment of the present application;

[0012] Figure 2 A schematic diagram of a process of combining a portrait area and a background image provided in an embodiment of the present application;

[0013] Figure 3 A schematic diagram of a process of combining a sub-image of a regular area with a background image provided by an embodiment of the present application;

[0014] Figure 4 A schematic diagram of a process of generating a virtual group photo provided in an embodiment of the present application;

[0015] Figure 5 A schematic diagram of a process for determining a portrait area provided in an embodiment of the present application;

[0016] Figure 6 A schematic diagram of a pixel shift process provided in an embodiment of the present application;

[0017] Figure 7 A schematic diagram of a process for determining a region where a person is located based on a portrait region provided in an embodiment of the present application;

[0018] Figure 8a A schematic diagram of a virtual group photo generated using relevant technologies provided in an embodiment of the present application;

[0019] Figure 8b A schematic diagram of a virtual group photo provided in an embodiment of the present application;

[0020] Fig. 9 A schematic diagram of a process for generating a virtual group photo provided in an embodiment of the present application;

[0021] Fig.10 A schematic diagram of the structure of a virtual group photo generating device provided in an embodiment of the present application;

[0022] Fig.11 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0023] The present application will be further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present application.

[0024] In order to generate a virtual photo with a natural background and a person, the embodiment of the present application provides a method and device for generating a virtual photo.

[0025] The virtual group photo generation method comprises: extracting the area where the person is located in the image containing the person; inputting the sub-image of the area and the background image into a pre-trained diffusion model, and obtaining a virtual group photo after the person is fused into the background image output by the diffusion model.

[0026] Figure 1 A schematic diagram of a process of generating a virtual group photo provided in an embodiment of the present application, the process comprising the following steps:

[0027] S101: extracting the area where the person is located in the image containing the person.

[0028] The virtual group photo generation method provided in the embodiment of the present application is applied to an electronic device, which may be a smart device such as a PC or a server.

[0029] In order to achieve the fusion of the person and the background, the electronic device may first extract the area where the person is located in the image of the task, wherein the area may be a rectangular area.

[0030] Specifically, in order to extract the area where the person is located, the electronic device may locally store a pre-trained person extraction model, and the electronic device may input an image containing the person into the person extraction model and obtain the output of the person extraction model, and the output of the person extraction model is an image with the area where the person is located marked. That is, it is equivalent to extracting the area where the person is located in the image containing the person.

[0031] S102: Inputting the sub-image of the region and the background image into a pre-trained diffusion model, and obtaining a virtual photo output by the diffusion model in which the person is fused with the background image.

[0032] In order to accurately generate a virtual group photo, the electronic device locally stores a pre-trained diffusion model. After obtaining the area where the person is located, the electronic device can input the sub-image of the area where the person is located and the background image into the pre-trained diffusion model and obtain the output of the diffusion model. The output of the diffusion model is the virtual group photo after the person and the background image are fused.

[0033] Among them, the diffusion model can be understood as an algorithm that can generate images. In addition, the electronic device can input the vector corresponding to the background image together with the sub-image of the region into the diffusion model. The vector corresponding to the background image can be a vector representing the description information of the background image obtained by processing the background image using the Contrastive Language-Image Pretraining (CLIP) algorithm.

[0034] In one possible implementation, the electronic device may replace areas of the acquired image except for areas containing people with preset pixel values, for example, with 0, and then input the image together with the background image into a diffusion model to regenerate the background area and obtain a virtual photo.

[0035] The embodiment of the present application is equivalent to regenerating the background area of ​​the image based on the diffusion model to achieve the synthesis of the portrait and the new background image. The diffusion model is a generative algorithm. The electronic device can also input the image containing the person into the diffusion model. The embodiment of the present application is equivalent to providing a virtual photo method based on a generative algorithm, which is a method for synthesizing the foreground portrait with the new background image.

[0036] In the embodiment of the present application, the electronic device extracts the area where the person is located in the image containing the person; wherein the area is an irregular area; the sub-image of the area and the background image are input into the pre-trained diffusion model, and a virtual group photo of the person fused with the background image output by the diffusion model is obtained. Since the area where the person is located is an irregular area, and the fusion of the person and the background image can be well achieved through the diffusion model, the person and the background in the generated virtual group photo are more natural, and there is no need for the user to manually select the position of the person.

[0037] In order to accurately obtain the area where the person is located, based on the above embodiment, in the embodiment of the present application, extracting the area where the person is located in the image containing the person includes:

[0038] Using a human body detection algorithm, extracting a portrait area in an image containing a person; wherein the portrait area is a rectangular area;

[0039] A preset number of pixel points are obtained at each edge of the portrait area; each obtained pixel point is offset; and each offset pixel point is connected to obtain the area where the person is located.

[0040] In order to accurately obtain the area where the person is located, the electronic device can use a human body detection algorithm to extract a portrait area in an image containing the person; wherein the portrait area is a rectangular area, and it should be noted that the portrait area completely contains the person and part of the background.

[0041] After acquiring the portrait area, the electronic device may acquire a preset number of pixel points at each edge of the portrait area. In a possible implementation, the number of points acquired by the electronic device at the left and right boundaries of the portrait area (the left and right described here are the left and right in the image) is appropriately greater than that at the upper and lower boundaries (the upper and lower described here are the upper and lower in the image). Three pixel points may be acquired at the upper and lower boundaries (the upper and lower described here are the upper and lower in the image), and four points may be acquired at the left and right boundaries (the left and right described here are the left and right in the image).

[0042] After obtaining each pixel point, each pixel point can be offset. Specifically, any direction can be randomly selected and offset in that direction. After the offset, each offset pixel point is connected according to the original connection mode of each pixel point to obtain the area where the person is located. For example, if pixel point 1 is originally connected to pixel point 2, then pixel point 1 after offset is still connected to pixel point 2 after offset.

[0043] Figure 2 A schematic diagram of a process of separating a portrait area and a background image provided in an embodiment of the present application.

[0044] Depend on Figure 2 It can be seen that if Figure 2 The background image and the image containing the person shown in are used to generate a virtual group photo, and the portrait area in the image containing the person is extracted, and the virtual group photo is generated based on the portrait area and the background image. When the virtual group photo is generated by the diffusion model, Figure 2 The area around the character's arms shown cannot be generated.

[0045] It should be noted that the current portrait segmentation is responsible for separating the portrait in the image from the original background, and image synthesis is responsible for synthesizing the segmented portrait with the new background. With the popularity of AI painting technology in recent years, background generation or replacement of images based on diffusion models has also become a new method for realizing virtual group photos. Virtual group photos based on diffusion models are a method of modifying or editing part of an image, because the portrait in the original image cannot be changed, and only the background area can be regenerated. At present, there is a problem with this method. When there is a changeable background area in the portrait area, the pixels in the background area are often difficult to generate. For example, the area between the legs of the portrait, the area between the arms and the body, the pixels generated in these areas are difficult to associate with the content generated outside the portrait area, such as Figure 2 shown. Figure 2 The pixels around the area where the character's arm is located are all unchangeable portrait areas. The generation algorithm generates the target background while the foreground portrait remains unchanged. Therefore, for the algorithm, the pixels around this area are "unreferenceable". It is difficult to cross the domain portrait area and connect it with the content generated in the external background area, resulting in this part of the pixels being monochrome pixels in most cases or unable to generate new content.

[0046] Figure 3 A schematic diagram of a process of combining a sub-image of a regular area with a background image provided in an embodiment of the present application.

[0047] Depend on Figure 3 It can be seen that if Figure 3 The background image and the image containing the person shown in the figure are used to generate a virtual group photo, and the area where the person is located in the image containing the person is extracted. If the area is a regular area, that is, a standard rectangular area, when the virtual group photo is generated by the diffusion model, it can solve the problem Figure 2 The problem of the elbow not being able to be generated in the middle, but when the gap between the background and the foreground is too large, the boundary problem may appear. Figure 3 Middle right side (left and right described here are Figure 3 The reason is that the regular rectangular area is the same as the image, and the diffusion model can easily learn a simple implementation method, that is, directly embedding the area into another image to generate a group photo. It only needs to ensure that the embedded image is similar in style to the target background image. Therefore, if the area is not changed to a standard rectangular area, the composite photo may have the phenomenon that the foreground portrait area is superimposed on the new background image.

[0048] In order to solve the defect of the above-mentioned virtual group photo method based on the diffusion model that when there is an independent background area inside the portrait area, the content generated by the independent background area cannot be linked to the content generated by the outer background area, the original image containing the person in this application is first subjected to portrait segmentation or human body detection to obtain a standard rectangular area, and then the irregular area is obtained based on the rectangular area. Finally, the virtual group photo obtained is the unchanged area, that is, the content of the area where the person is located remains unchanged, and the changed area, that is, the background area of ​​the original image is regenerated according to the style of the target background through the diffusion model.

[0049] Figure 4 A schematic diagram of a process for generating a virtual group photo provided in an embodiment of the present application.

[0050] Depend on Figure 4 It can be seen that the electronic device can first obtain the portrait area in the image containing the person, and construct an irregular area based on the portrait area, and input the sub-image of the irregular area and the background image into the diffusion model to obtain a virtual photo. Figure 2 and Figure 3 By comparison, it can be seen that the virtual group photo generated by using the sub-image of the irregular area is more natural.

[0051] In addition, in the present application, the changed area is designed to be an irregular area, and the interior of the area is connected, and there is no changed area, that is, the two areas do not contain each other. In order to obtain this irregular area, the electronic device can also use a portrait segmentation algorithm, obtain an accurate portrait mask through portrait segmentation, and then obtain the coordinates of the outermost boundary pixels of the four sides of the area through "connected domain" calculation or "contour" calculation, and further obtain a portrait area of ​​a standard rectangle circumscribed to the portrait mask area, and finally, based on the boundary pixel points of the upper, lower, left and right sides of the portrait area, connect the boundary pixel points in a certain order to construct an irregular area.

[0052] Figure 5 A schematic diagram of a process for determining a portrait area provided in an embodiment of the present application.

[0053] Depend on Figure 5 It can be seen that after obtaining the portrait mask, the electronic device can calculate the coordinates of the boundary pixels through the connected domain or the contour, and calculate the circumscribed standard rectangle of the mask area according to the coordinates of the boundary pixels to obtain the portrait area.

[0054] In order to accurately obtain the area where the person is located, based on the above embodiments, in the embodiment of the present application, the offset of each pixel point obtained includes:

[0055] For each pixel point obtained, a target offset number is randomly obtained from each offset number of pre-saved pixel points, and the pixel points are offset in any direction by the target offset number.

[0056] In order to accurately obtain the area where the person is located, the electronic device can randomly obtain an offset number from each offset number of the pre-saved pixel point for each pixel point obtained, and the randomly obtained offset number is the target offset number. For example, each offset number can be 5-10. After obtaining the target offset number, the target offset number of pixels can be offset in any direction. That is, the offset of the pixel point is completed.

[0057] Specifically, taking the human body detection algorithm to obtain the portrait area as an example, the human body detection algorithm can obtain the left vertex coordinates x, y and the width and height (w, h) of the portrait area. Then, considering that the human body is a rectangle with a height greater than the width when standing, the coordinates of each point are randomly offset by 5-10 pixels, and the offset direction is offset in the four directions of up, down, left and right. Finally, these points can be connected in sequence to construct an irregular area. Among them, points can be selected from the four boundaries of the obtained rectangular area, up, down, left and right (the up, down, left and right described here are the up, down, left and right in the image), and a point is selected every 1 / 4 of the distance on the upper and lower boundaries, and three points are selected on each side; a point is selected every 1 / 5 of the distance on the left and right sides, and four points are selected on each side. All points are randomly offset by 5-10 pixels up, down, left and right. The points after the offset are connected in sequence to successfully construct an irregular area, which is called the unchanged area here.

[0058] Figure 6 A schematic diagram of a pixel shift process provided in an embodiment of the present application.

[0059] Figure 6 The straight line in is the boundary of the portrait area. Figure 6 The original pixel point in the image is shifted in any of the four directions: up, down, left, and right. The shifted pixel point is located at Figure 6 Middle right side (left and right described here are Figure 6 The positions indicated in the figure are as follows.

[0060] In an embodiment of the present application, the electronic device obtains the portrait area of ​​the foreground through a portrait segmentation or human body detection algorithm, then selects a certain number of points at the boundary of the portrait area, and randomly shifts the positions of these points, and then connects the shifted points to make the portrait area an irregularly shaped area, the content of the area is not changed, and the outside of the area is regenerated according to the style of the background image through a diffusion model. The method provided in an embodiment of the present application avoids the problem that the diffusion model cannot generate new content for the background inside the portrait area when there is a background in the portrait area. The method can obtain the portrait area based on a human body detection algorithm or a portrait segmentation algorithm, and has low requirements on the algorithm accuracy. The portrait area and the background area do not contain each other. When the diffusion model regenerates the background area, it can effectively connect the contents of the two areas, so that the generated image is more natural. Whether it is a portrait segmentation algorithm or a human body detection algorithm, for this application, it is to obtain a portrait area, that is, a standard rectangular area.

[0061] Figure 7 A schematic diagram of a process for determining an area where a person is located based on a portrait area provided in an embodiment of the present application.

[0062] Figure 7 Left side of center (left and right described here are Figure 7 The left and right parts shown are the portrait area. Figure 7 Each point in the middle is the position of each pixel obtained after random offset. Figure 7 Middle right side (left and right described here are Figure 7 The left and right shown in the figure are the areas containing people obtained after sequentially connecting each offset pixel point. Figure 7 It can be seen that the area containing the person is an irregular area.

[0063] In order to accurately generate a virtual group photo, based on the above embodiments, in an embodiment of the present application, obtaining a virtual group photo output by the diffusion model in which the characters are fused with the background image includes:

[0064] The network layer of the diffusion model adds noise to the sub-image to generate a noise image; obtains the remaining number of repetitions saved in advance; and loops through the following steps until there is no remaining number of repetitions, and outputs the noise image as a virtual group photo:

[0065] Inputting the noise image, the background image, the sub-image and the remaining number of repetitions into a diffusion network Unet pre-trained in the diffusion model, and the Unet outputs the predicted noise;

[0066] The network layer of the diffusion model filters the predicted noise in the noise image to obtain a target image, uses the target image to replace the noise image, and updates the remaining number of repetitions.

[0067] In order to accurately generate a virtual group photo, in an embodiment of the present application, noise can be first added to the sub-image of the area where the characters are located through the network layer of the diffusion model to generate a noise image, wherein the added noise is usually Gaussian noise. Specifically, the network layer of the diffusion model can randomly generate a Gaussian noise image with the same size as the sub-image, and the Gaussian noise image and the sub-image are combined into a noise image. And obtain the pre-saved remaining number of repetitions, wherein the larger the remaining number of repetitions, the more iterations are required. Generally speaking, the quality of the generated virtual group photo is also higher, but it also makes the virtual group photo generation more time-consuming. Under the premise of comprehensively considering the effect and time, the value range of the remaining number of repetitions is set to 50-100.

[0068] The network layer of the diffusion model loops through the following steps until there are no more repetitions left:

[0069] The network layer of the diffusion model inputs the noise image, the background image, the sub-image of the area where the character is located, and the remaining number of repetitions into the Unet pre-trained in the diffusion model, wherein the specific input into the Unet may be a vector representing the remaining number of repetitions and an embedding vector representing the characteristics of the background image, and obtains the noise predicted by the Unet output; the network layer of the diffusion model filters the predicted noise in the noise image, obtains the target image, replaces the noise image with the target image, and updates the remaining number of repetitions. The remaining number of repetitions can also be referred to as a moment. Specifically, the electronic device uses the CLIP algorithm to process the background image and generate an embedding vector. The embedding algorithm is used to obtain the vector corresponding to the remaining number of repetitions. For example, when the remaining number of repetitions is 14, the corresponding vector is [0.2, 0.4, ..., 0.02], and the length of the vector is generally 1280.

[0070] When there are no remaining repetitions, the diffusion model outputs the acquired target image as a virtual group photo. The final result is generated as a composite virtual group photo of the same size as the original image containing the person after image decoding.

[0071] It should be noted that Unet is the core network of the diffusion model, which is used to predict the noise in the image.

[0072] In order to accurately obtain the virtual group photo, based on the above embodiments, in the embodiment of the present application, the Unet is trained in the following manner:

[0073] Obtain a pre-saved character image of any character, and add any preset noise to the character image; input the character image after adding noise, any number of times and a background image into Unet, and obtain the predicted noise output by the Unet;

[0074] The Unet is trained according to the predicted noise and the preset noise.

[0075] In order to realize the training of Unet, the electronic device can first obtain a pre-saved character image of any character, add any preset noise to the character image, and input the character image after adding noise, any number and a background image into Unet, wherein the input into Unet can be a vector corresponding to the number, and the background image input into Unet can also be an embedded vector of the background image, and obtain the predicted noise output by the Unet, and train the Unet according to the predicted noise and the preset noise. Specifically, the electronic device can calculate the loss according to the cross entropy of the predicted noise distribution and the actual added noise distribution, and train the Unet according to the loss. Specifically, the electronic device can continuously change the character image, iterate the training, and minimize the loss.

[0076] The bottom layer of the diffusion model is a Unet, whose main task is to predict the noise in the image. In the actual scenario, during the training phase, the diffusion model adds noise to the image x0 at time t0 to obtain x1, and lets Unet predict the noise added in the process of obtaining x1 from x0, and then continuously transforms the sample x0-x t , until the loss between the predicted output noise and the actual added noise is minimized, and the network converges. In the inference phase, given the noise sample x at time t, t , predict the Gaussian noise added at time t through the Unet network, and then let x t Subtract the predicted noise to get x t-1 Repeat the above process to continue x t-1 Get x t-2 , and so on, we finally get the image x0.

[0077] In order to obtain a trained diffusion model, based on the above embodiments, in the embodiment of the present application, the diffusion model is trained in the following manner:

[0078] Obtain any sample group photo in the sample set, and a sample portrait and a sample background image saved for the sample group photo; wherein the sample portrait is an irregular image;

[0079] Inputting the sample portrait and the sample background image into the original diffusion model to obtain a predicted group photo output by the original diffusion model;

[0080] The original diffusion model is trained according to the predicted group photo and the sample group photo.

[0081] In order to train the diffusion model, a sample set for training is stored in the embodiment of the present application, and the sample set includes a sample group photo, a sample portrait containing a person, and a sample background image of the background stored for the sample group photo, wherein the sample portrait is an irregular image. In order to train the diffusion model, the electronic device can obtain any sample group photo in the sample set, and obtain a sample background image containing a sample portrait containing a person and the background stored for the sample group photo.

[0082] In an embodiment of the present application, after obtaining any sample group photo in the sample set and the sample portrait and sample background image corresponding to the sample group photo, the electronic device can input the sample portrait set and the sample background image into the original diffusion model, and the original diffusion model outputs the predicted group photo.

[0083] After the original diffusion model outputs the predicted group photo, the original diffusion model is trained according to the sample group photo and the predicted group photo to obtain a trained diffusion model.

[0084] In order to accurately generate a virtual group photo, based on the above embodiments, in the embodiment of the present application, a sample portrait and a sample background image corresponding to a sample group photo are obtained in the following manner:

[0085] Extracting the area where the person in the sample group photo is located, and segmenting the area in the sample group photo;

[0086] The image inpainting algorithm is used to process the segmented sample group photo to generate a sample background image.

[0087] The electronic device may extract the area where the person is located in the sample group photo, segment the area in the sample group photo, and use an image repair algorithm to process the sample group photo after segmenting the area to generate a sample background image.

[0088] The image inpainting algorithm can repair the missing content (or holes) in the image. The image inpainting algorithm in the embodiment of the present application can be an open source image inpainting (Lama) algorithm, the input of which is a sample group photo and a sample portrait obtained by human body detection, and the output is an image in which the portrait in the sample group photo is removed and the removed area is filled, and the image is a pure background image without a portrait.

[0089] In order to accurately generate a virtual group photo, based on the above embodiments, in the embodiment of the present application, the method further includes:

[0090] Performing generalization processing on the sample portrait, and replacing the sample portrait with the generalized image; and / or,

[0091] The sample background image is generalized and the processed image is used to replace the sample background image.

[0092] Since in actual scenes, the portrait and background images input to the diffusion model are usually images collected from different scenes, different lighting, etc., and the sample images are processed by the method provided in the embodiment of the present application to obtain the sample portrait and the sample background image as images collected from the same scene, etc. If the diffusion model is trained only by such sample images and sample background images, it is impossible to effectively obtain a diffusion model that can generate a more natural virtual photo. Therefore, in the embodiment of the present application, the electronic device can generalize the sample portrait after obtaining the sample portrait, and replace the sample portrait with the generalized image. Thereby, the sample portrait and the sample background image are not images collected from the same scene, and the training of the diffusion model is further realized. The electronic device can also generalize the sample background image after obtaining the sample background image, and replace the sample background image with the generalized image. Thereby, the sample portrait and the sample background image are not images collected from the same scene, and the training of the diffusion model is further realized, and the generalization ability of the model is improved.

[0093] In order to accurately implement the training of the diffusion model, on the basis of the above embodiments, in the embodiment of the present application, the generalization processing includes at least one of the following processing: color transformation processing, gamma transformation processing, image enhancement processing, blurring processing, image flipping processing and image cropping processing.

[0094] The generalized processing described in the embodiments of the present application includes at least one of the following processing: color transformation processing, gamma transformation processing, image enhancement processing, blurring processing, image flipping processing and image capture processing.

[0095] The goal of a virtual group photo is to generate a photo of a person in a new background that is realistic and harmonious. To this end, this application uses sample group photos of different tasks in different backgrounds taken in real life as labels for algorithm training, and then uses methods such as human body detection and image inpainting to create sample people and sample background images required for diffusion model training. Then, methods such as image capture and data enhancement are used to simulate a variety of real-life scenarios in virtual group photos, thereby enhancing the generalization ability of the model.

[0096] By enhancing the data of the foreground portrait, the diffusion model has the ability to harmonize the image. In the generated image, the foreground portrait and the background will be color adjusted, making the new image more natural and realistic.

[0097] Usually, in the newly generated virtual group photo, the background and the background image have the same style, but the content is not completely the same, and the size of the image containing the person and the background image uploaded by the user are also different. Therefore, during the training process, the electronic device obtains multiple background images by randomly cutting out areas of different sizes on the sample background image, and also collects some background images with similar styles. Further data augmentation is performed by color transformation, blurring, and left-right flipping, so that the new background image is consistent with the background in the image containing the person in style and inconsistent with the background in the image containing the person in content, so that enough samples for training can be constructed. For the foreground portrait, considering that in actual applications, the image containing the person uploaded by the user and the background image have inconsistent colors and lighting, the foreground portrait is subjected to color transformation and gamma transformation to simulate the situation of the diversity of user portrait lighting. By performing data augmentation on the foreground portrait and background, the diffusion model can adapt to the scene with inconsistent colors between the foreground portrait and the background and diverse background images, so that the generated group photo is real and natural.

[0098] As mentioned above, for the diffusion model, the new sample background images obtained through data augmentation are similar in style to the original sample background images, but these newly generated sample background images and the original sample background images are used as labels in the training process. After the sample portraits are processed by color adjustment and gamma transformation, their training label data does not change, and the labels of the original sample portraits can also be used as labels in the training process. Therefore, the labels of the multiple input data obtained through data augmentation have not changed, and they are still the original real-life images, and do not need to be generated separately, which also simplifies the complexity of training data annotation during network training.

[0099] The electronic device inputs the sample background image and the sample portrait into the original diffusion model to obtain the predicted group photo output by the diffusion model. The original diffusion model is trained based on the predicted group photo and the sample group photo.

[0100] When there is an independent background area inside the accurate portrait mask, it is difficult for existing methods to generate effective content for the independent area inside. To this end, in this solution, we make the changed area and the unchanged area fully connected, that is, there is no other area inside a certain area. At the same time, we construct an irregular unchanged area. By using the irregular shape, the algorithm can avoid the misunderstanding of copying multiple learned images together when generating content. Irregular shapes can also be generated based on the boundary for reference. The generated content has a natural transition with the boundary of the portrait area and is similar to the target background style. We no longer use the precise portrait area boundary as the unchanged area. This solution constructs an irregular area based on the results of human body detection or portrait segmentation. It can solve the problem of poor edge effect of the final virtual group photo portrait when there are mis-segmentation and wrong segmentation of the boundary segmented by the portrait segmentation algorithm, making the virtual group photo more natural and realistic.

[0101] Figure 8a A schematic diagram of a virtual group photo generated using relevant technologies provided in an embodiment of the present application.

[0102] Depend on Figure 8a It can be seen that there may be some areas (i.e. Figure 8a The circled area in the figure is not completely integrated with the background image.

[0103] Figure 8b A schematic diagram of a virtual group photo provided in an embodiment of the present application.

[0104] in, Figure 8b A virtual group photo is generated by the method provided in the embodiment of the present application. Figure 8a and Figure 8b It can be seen that the virtual group photo generated by the method provided in the embodiment of the present application can solve the problem of poor edge effects of portraits, and is more natural and realistic.

[0105] Fig. 9 A schematic diagram of a process for generating a virtual group photo provided in an embodiment of the present application.

[0106] Depend on Fig. 9 It can be seen that the electronic device can first obtain the original image containing the person, and obtain the portrait area in the image, take points at the boundary of the portrait area, offset the taken points, sequentially connect the offset points, generate an irregular area, and obtain a sub-image corresponding to the irregular area. The embedding vector obtained by processing the sub-image and the background image with the CLIP algorithm is input into the diffusion model to obtain a picture of the person in the new background output by the diffusion model, that is, a virtual photo.

[0107] Fig.10 A schematic diagram of a virtual group photo generation device provided in an embodiment of the present application, the device comprising:

[0108] An extraction module 1001 is used to extract the area where the person is located in the image containing the person;

[0109] The processing module 1002 is used to input the sub-image of the area and the background image into a pre-trained diffusion model, and obtain a virtual photo of the person fused with the background image output by the diffusion model.

[0110] Furthermore, the processing module 1002 is specifically used to use a human body detection algorithm to extract a portrait area in an image containing a person; wherein the portrait area is a rectangular area; a preset number of pixels are obtained at each edge of the portrait area; each obtained pixel is offset; and each pixel after offset is connected to obtain the area where the person is located.

[0111] Furthermore, the processing module 1002 is specifically configured to randomly obtain a target offset number from each offset number of pre-saved pixel points for each acquired pixel point, and offset the target offset number of pixels in any direction.

[0112] Furthermore, the processing module 1002 is specifically used for the network layer of the diffusion model to add noise to the sub-image to generate a noise image; obtain the remaining number of repetitions saved in advance; loop the following steps until there are no remaining repetitions, and output the noise image as a virtual photo: input the noise image, the background image, the sub-image and the remaining number of repetitions into the diffusion network Unet pre-trained in the diffusion model, and the Unet outputs the predicted noise; the network layer of the diffusion model filters the predicted noise in the noise image to obtain a target image, replaces the noise image with the target image, and updates the remaining number of repetitions.

[0113] Furthermore, the processing module 1002 is also used to train the Unet in the following manner: obtain a pre-saved character image of any character and add any preset noise to the character image; input the character image after adding noise, any number of times and a background image into the Unet to obtain the predicted noise output by the Unet; and train the Unet according to the predicted noise and the preset noise.

[0114] Furthermore, the processing module 1002 is also used to train the diffusion model in the following manner: obtain any sample group photo in the sample set, and a sample portrait and a sample background image saved for the sample group photo; wherein the sample portrait is an irregular image; input the sample portrait and the sample background image into the original diffusion model to obtain a predicted group photo output by the original diffusion model; and train the original diffusion model based on the predicted group photo and the sample group photo.

[0115] Furthermore, the processing module 1002 is further used to extract the area where the characters in the sample group photo are located, and segment the area in the sample group photo; and use an image repair algorithm to process the segmented sample group photo to generate a sample background image.

[0116] Furthermore, the processing module 1002 is also used to perform generalization processing on the sample portrait and replace the sample portrait with the generalized image; and / or perform generalization processing on the sample background image and replace the sample background image with the processed image.

[0117] Fig.11 The present invention provides a schematic diagram of an electronic device structure according to an embodiment of the present invention. Based on the above embodiments, the present invention further provides an electronic device, such as Fig.11 As shown, it includes: a processor 1101, a communication interface 1102, a memory 1103 and a communication bus 1104, wherein the processor 1101, the communication interface 1102, and the memory 1103 communicate with each other through the communication bus 1104;

[0118] The memory 1103 stores a computer program. When the program is executed by the processor 1101, the processor 1101 performs the following steps:

[0119] Extracting the area where the person is located in the image containing the person;

[0120] The sub-image of the region and the background image are input into a pre-trained diffusion model, and a virtual photo of the person fused with the background image output by the diffusion model is obtained.

[0121] Further, the processor 1101 is specifically configured to extract a portrait region from an image containing a person using a human body detection algorithm; wherein the portrait region is a rectangular region;

[0122] A preset number of pixel points are obtained at each edge of the portrait area; each obtained pixel point is offset; and each offset pixel point is connected to obtain the area where the person is located.

[0123] Furthermore, the processor 1101 is specifically configured to randomly obtain a target offset number from each offset number of pre-saved pixel points for each acquired pixel point, and offset the target offset number of pixels in any direction.

[0124] Further, the processor 1101 is specifically configured to add noise to the sub-image by the network layer of the diffusion model to generate a noise image; obtain the remaining number of repetitions saved in advance; and loop through the following steps until there is no remaining number of repetitions, and output the noise image as a virtual group photo:

[0125] Inputting the noise image, the background image, the sub-image and the remaining number of repetitions into a diffusion network Unet pre-trained in the diffusion model, and the Unet outputs the predicted noise;

[0126] The network layer of the diffusion model filters the predicted noise in the noise image to obtain a target image, uses the target image to replace the noise image, and updates the remaining number of repetitions.

[0127] Furthermore, the processor 1101 is specifically configured to train the Unet in the following manner:

[0128] Obtain a pre-saved character image of any character, and add any preset noise to the character image; input the character image after adding noise, any number of times and a background image into Unet, and obtain the predicted noise output by the Unet;

[0129] The Unet is trained according to the predicted noise and the preset noise.

[0130] Further, the processor 1101 is specifically configured to train the diffusion model in the following manner:

[0131] Obtain any sample group photo in the sample set, and a sample portrait and a sample background image saved for the sample group photo; wherein the sample portrait is an irregular image;

[0132] Inputting the sample portrait and the sample background image into the original diffusion model to obtain a predicted group photo output by the original diffusion model;

[0133] The original diffusion model is trained according to the predicted group photo and the sample group photo.

[0134] Furthermore, the processor 1101 is specifically configured to extract the area where the person in the sample group photo is located, and segment the area in the sample group photo;

[0135] The image inpainting algorithm is used to process the segmented sample group photo to generate a sample background image.

[0136] Further, the processor 1101 is specifically configured to perform generalization processing on the sample portrait, and replace the sample portrait with the generalized image; and / or,

[0137] The sample background image is generalized and the processed image is used to replace the sample background image.

[0138] The communication bus mentioned in the above server can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0139] The communication interface is used for communication between the above electronic device and other devices.

[0140] The memory may include a random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0141] The above-mentioned processor can be a general-purpose processor, including a central processing unit, a network processor (Network Processor, NP), etc.; it can also be a digital signal processing processor (Digital Signal Processing, DSP), an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, etc.

[0142] On the basis of the above embodiments, an embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program executable by an electronic device, and when the program is run on the electronic device, the electronic device implements the following steps when executing:

[0143] The memory stores a computer program, and when the program is executed by the processor, the processor performs the following steps:

[0144] Extracting the area where the person is located in the image containing the person;

[0145] The sub-image of the region and the background image are input into a pre-trained diffusion model, and a virtual photo of the person fused with the background image output by the diffusion model is obtained.

[0146] In a possible implementation, extracting the area where the person is located in the image containing the person includes:

[0147] Using a human body detection algorithm, extracting a portrait area in an image containing a person; wherein the portrait area is a rectangular area;

[0148] A preset number of pixel points are obtained at each edge of the portrait area; each obtained pixel point is offset; and each offset pixel point is connected to obtain the area where the person is located.

[0149] In a possible implementation manner, offsetting each acquired pixel point includes:

[0150] For each pixel point obtained, a target offset number is randomly obtained from each offset number of pre-saved pixel points, and the pixel points are offset in any direction by the target offset number.

[0151] In a possible implementation manner, obtaining a virtual photo of the person output by the diffusion model and fused with the background image includes:

[0152] The network layer of the diffusion model adds noise to the sub-image to generate a noise image; obtains the remaining number of repetitions saved in advance; and loops through the following steps until there is no remaining number of repetitions, and outputs the noise image as a virtual group photo:

[0153] Inputting the noise image, the background image, the sub-image and the remaining number of repetitions into a diffusion network Unet pre-trained in the diffusion model, and the Unet outputs the predicted noise;

[0154] The network layer of the diffusion model filters the predicted noise in the noise image to obtain a target image, uses the target image to replace the noise image, and updates the remaining number of repetitions.

[0155] In one possible implementation, the Unet is trained in the following manner:

[0156] Obtain a pre-saved character image of any character, and add any preset noise to the character image; input the character image after adding noise, any number of times and a background image into Unet, and obtain the predicted noise output by the Unet;

[0157] The Unet is trained according to the predicted noise and the preset noise.

[0158] In a possible implementation, the diffusion model is trained in the following manner:

[0159] Obtain any sample group photo in the sample set, and a sample portrait and a sample background image saved for the sample group photo; wherein the sample portrait is an irregular image;

[0160] Inputting the sample portrait and the sample background image into the original diffusion model to obtain a predicted group photo output by the original diffusion model;

[0161] The original diffusion model is trained according to the predicted group photo and the sample group photo.

[0162] In a possible implementation, a sample portrait and a sample background image corresponding to a sample group photo are obtained in the following manner:

[0163] Extracting the area where the person in the sample group photo is located, and segmenting the area in the sample group photo;

[0164] The image inpainting algorithm is used to process the segmented sample group photo to generate a sample background image.

[0165] In a possible implementation, the method further includes:

[0166] Performing generalization processing on the sample portrait, and replacing the sample portrait with the generalized image; and / or,

[0167] The sample background image is generalized and the processed image is used to replace the sample background image.

[0168] In a possible implementation, the generalization processing includes at least one of the following processing: color transformation processing, gamma transformation processing, image enhancement processing, blurring processing, image flipping processing, and image capture processing.

[0169] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0170] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0171] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0172] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0173] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A method for generating a virtual group photo, characterized in that: The method comprises: Extracting the area where the person is located in the image containing the person; The sub-image of the region and the background image are input into a pre-trained diffusion model, and a virtual photo of the person fused with the background image output by the diffusion model is obtained.

2. The method according to claim 1, characterized in that The extracting of the area where the person is located in the image containing the person comprises: Using a human body detection algorithm, extracting a portrait area in an image containing a person; wherein the portrait area is a rectangular area; A preset number of pixel points are obtained at each edge of the portrait area; each obtained pixel point is offset; and each offset pixel point is connected to obtain the area where the person is located.

3. The method according to claim 2, characterized in that The offset of each pixel point obtained includes: For each pixel point obtained, a target offset number is randomly obtained from each offset number of pre-saved pixel points, and the pixel points are offset in any direction by the target offset number.

4. The method according to claim 1, characterized in that: The step of obtaining a virtual group photo of the person output by the diffusion model and fused with the background image comprises: The network layer of the diffusion model adds noise to the sub-image to generate a noise image; obtains the remaining number of repetitions saved in advance; and loops through the following steps until there is no remaining number of repetitions, and outputs the noise image as a virtual group photo: Inputting the noise image, the background image, the sub-image and the remaining number of repetitions into a diffusion network Unet pre-trained in the diffusion model, and the Unet outputs the predicted noise; The network layer of the diffusion model filters the predicted noise in the noise image to obtain a target image, uses the target image to replace the noise image, and updates the remaining number of repetitions.

5. The method according to claim 4, characterized in that The Unet is trained in the following way: Obtain a pre-saved character image of any character, and add any preset noise to the character image; input the character image after adding noise, any number of times and a background image into Unet, and obtain the predicted noise output by the Unet; The Unet is trained according to the predicted noise and the preset noise.

6. The method according to claim 1, characterized in that The diffusion model is trained in the following way: Obtain any sample group photo in the sample set, and a sample portrait and a sample background image saved for the sample group photo; wherein the sample portrait is an irregular image; Inputting the sample portrait and the sample background image into the original diffusion model to obtain a predicted group photo output by the original diffusion model; The original diffusion model is trained according to the predicted group photo and the sample group photo.

7. The method according to claim 6, characterized in that The sample portrait and sample background image corresponding to a sample group photo are obtained in the following way: Extracting the area where the person in the sample group photo is located, and segmenting the area in the sample group photo; The image inpainting algorithm is used to process the segmented sample group photo to generate a sample background image.

8. The method according to claim 7, characterized in that The method further comprises: Performing generalization processing on the sample portrait, and replacing the sample portrait with the generalized image; and / or, The sample background image is generalized and the processed image is used to replace the sample background image.

9. The method according to claim 8, characterized in that The generalization processing includes at least one of the following processing: color transformation processing, gamma transformation processing, image enhancement processing, blurring processing, image flipping processing and image interception processing.

10. An electronic device, characterized in that: The electronic device comprises at least a processor and a memory, and the processor is used to implement the steps of the virtual group photo generation method as described in any one of claims 1 to 9 when executing the computer program stored in the memory.