Intelligent generation method for stylized unified fusion character group photo

Through the adversarial generation network of integrated position-color prediction model and content-style perception model, the problem of difficulty in dealing with character portraits and background synthesis in the existing technology is solved, and the unified style, harmonious and natural image generation is achieved, which is suitable for virtual social scenes.

CN119919516APending Publication Date: 2025-05-02ZHEJIANG UNIVERSITY OF MEDIA AND COMMUNICATIONS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411937212.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

Existing image synthesis technologies are difficult to solve the synthesis of character portraits and backgrounds at the same time, especially in terms of appearance and geometric inconsistency. Most technologies only focus on single or a few sub-problems, and fail to fully meet the needs of virtual social scenarios.

Method used

A stylized unified and fusion intelligent generation method for character group photos is proposed, integrating portrait and background fusion, portrait position prediction in the background, and style unified functions of portrait and background. Through the adversarial generation network, combining the position-color prediction model and the content-style perception model, the image synthesis of character images and backgrounds under various conditions is realized, so that the generated images can use the style of the background uniformly, and the picture is harmonious, real and natural.

Benefits of technology

It realizes high-quality integration of character portraits and backgrounds, unify styles, enhances the authenticity and nature of images, and is suitable for image generation needs in virtual social scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919516A_ABST
    Figure CN119919516A_ABST
Patent Text Reader

Abstract

The invention discloses a stylized unified fusion figure group photo intelligent generation method, which comprises the following steps: through a defined adversarial generative network, predicting the position and color of a portrait in a background through a position color prediction model and a content style perception model to obtain a fusion image, finely adjusting the content and style of the fused image through a content style perception model to obtain a generated image; and inputting a triple to the defined generative adversarial network to alternately train the generative model and the discrimination model to obtain a final generative model, and inputting a background image and a portrait image to the generative model to generate a fused generated image. The method integrates the functions of portrait and background fusion, portrait position prediction in the background and portrait and background style unification. The synthesized picture uniformly uses the style of the background picture, and the picture is harmonious, real and natural.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to an image fusion intelligent generation method, in particular to a stylized unified fusion group photo intelligent generation method. Background Art

[0002] In recent years, with the rise of digital media, new visual effects have been created by integrating character images with different backgrounds, which are widely used in many fields such as character background replacement, virtual design, artistic creation, and automatic generation of advertising pictures. Therefore, the technical demand for intelligent generation of character image fusion background is getting higher and higher, and it can show great application potential in many fields.

[0003] In order to obtain an ideal image, image synthesis and image generation are generally used together. Image generation is responsible for creating something from nothing, while image synthesis is responsible for improving something from something to an excellent one. Image generation has limited controllability. Even if a large amount of conditional information is provided, it may not be possible to generate a picture that fully meets expectations. Image synthesis is better at fine control, splicing completely matching visual elements to obtain a real and reasonable image that better meets people's expectations, making image synthesis have great application value.

[0004] Image synthesis refers to cutting the foreground of a picture and pasting it onto another background picture to obtain a composite image. Because the foreground and background in the composite image are originally real, the inconsistency between the foreground and background after forming a whole makes the image unrealistic. This problem can be further divided into appearance inconsistency and geometric inconsistency.

[0005] The appearance inconsistency is mainly manifested in unnatural boundaries between foreground and background, lighting mismatch between foreground and background, missing / unreasonable shadows or reflections in the foreground, and mismatched textures between foreground and background.

[0006] Geometric inconsistency mainly manifests itself in the foreground being too large / too small, the foreground not being supported (floating in the air), the unreasonable occlusion relationship between foreground and background objects, and the inconsistent perspective of the foreground and background;

[0007] Appearance inconsistency and geometric inconsistency can be further divided into many sub-problems, each of which is very challenging and worth studying. According to the tasks of solving different sub-problems, they are divided into image fusion, image harmonization, image placement and spatial deformation, shadow generation, etc.

[0008] Existing image synthesis technologies are mainly divided into traditional synthesis methods and deep learning synthesis methods.

[0009] Traditional methods for composite image synthesis are mostly rule-based methods, which are based on the manual features that need to be matched and extracted, such as Poisson image fusion method and multi-scale image harmonization. Their limitations lie in the limited nature of manual features, most of which focus on color gradients and color statistics.

[0010] With the advancement of computer deep learning research, various deep learning models have been introduced into image fusion. The features of fusion based on deep learning are richer than those of traditional methods and more realistic in visual perception.

[0011] Although some current technologies have achieved some results, they still have some defects, such as:

[0012] 1. Only focus on solving general domains, without paying too much attention to portrait and background synthesis, which is the most common editing task. As virtual social gatherings and meetings become more popular in our lives, the demand for more realistic images using portraits and backgrounds is increasing.

[0013] 2. In the field of image fusion, most technologies only focus on one or two sub-problems, and few technologies focus on multiple sub-problems at the same time. Summary of the invention

[0014] In order to solve the above problems, the present invention provides a method for intelligently generating group photos of people with unified style fusion. It integrates the functions of portrait and background fusion, portrait position prediction in the background, and style unification of portrait and background; from the perspective of character image synthesis, the method combines the tasks related to portrait image position prediction, character image color prediction, and content style perception model optimization of image content and style, and can synthesize images with people as foreground and various backgrounds taken under various conditions, so that the synthesized images uniformly use the style of the background image, and the pictures are harmonious, real and natural.

[0015] The present invention comprises the following steps:

[0016] A method for intelligently generating a stylized unified and integrated group photo of people comprises the following steps:

[0017] Step 1: Collect character-background image materials, perform preprocessing, and obtain preprocessed character-background images;

[0018] Step 2: The preprocessed person-background image is subjected to a cutout model to obtain a person image without background, and the person image without background is subjected to value processing to obtain a person image mask; the preprocessed person-background image, the person image without background, and the person image mask are combined into a triplet;

[0019] Step 3: define a generative adversarial network, including a generative model and a discriminative model, wherein the generative model includes a position-color prediction model, a content-style perception model, and a generative loss module, and the discriminative model includes a series-connected discriminator and a discriminative loss module;

[0020] Input the triplet obtained in step 2 into the generative model to obtain a generated image, and the generated image is passed through the generative loss module to obtain a generated loss value;

[0021] The generated image is input into the discriminant model to obtain the probability value of the false value. The preprocessed person-background image obtained in step 1 is input into the discriminant model to obtain the probability value of the true value. The discriminant loss module obtains the discriminant loss value by comparing the difference between the probability value of the false value and the value 0 and the difference between the probability value of the true value and the value 1. The generated loss value and the discriminant loss value constitute the loss value of the adversarial generative network.

[0022] First, fix the generative model, maximize the loss value of the adversarial generative network, and train the discriminative model; then fix the discriminative model, minimize the loss value of the adversarial generative network, and train the generative model. By alternately training the generative model and the discriminative model, an optimized generative model is obtained.

[0023] Step 4: Input the background image and portrait image to be fused into the optimized generation model obtained in step 3 to obtain a generated image of the fusion of the person and the background.

[0024] In step 1, character-background image materials are collected and preprocessed to obtain preprocessed character-background images, which specifically include:

[0025] 1.1) Collect pictures with people and backgrounds, and select pictures with obvious people and obvious backgrounds;

[0026] 1.2) Scale the images selected in step 1.1) to a uniform size, denoted as I, where the i-th person-background image is denoted as I i .

[0027] In step 2, the preprocessed person-background image is subjected to a cutout model to obtain a person image without background, and the person image without background is subjected to value processing to obtain a person image mask; the preprocessed person-background image, the person image without background, and the person image mask are combined into a triplet, specifically including:

[0028] 2.1) Use MODNet to cut out I and get a character image without background, denoted as J, where the jth character image is denoted as J j ;

[0029] 2.2) Use PIL to process the value of J to generate a character image mask with black transparent part and white color part, denoted as K. The jth character image mask is denoted as K j .

[0030] In step 3, the position-color prediction model, content-style perception model and generation model loss module are connected in series.

[0031] In step 3, the position-color prediction model includes a person image position prediction model and a person image color prediction model.

[0032] In step 3, the content-style perception model includes a perception model and a content-style loss module, the perception model includes an encoder and a decoder, and the content-style loss module adopts a picture classification model VGG-16 encoder.

[0033] In step 3, the generation loss module uses the mean square error loss function, the L1 norm function, the Euclidean distance calculated by the 7th layer of the classification model VGG-16 encoder, and the square of the Frobenius norm between the Gram matrices calculated by the 2nd, 4th, 7th, and 10th layers of the classification model VGG-16 encoder.

[0034] In step 3, the discriminant loss module is calculated using a binary cross loss function.

[0035] Step 4: Input the background image and the portrait image to be fused into the optimized generation model obtained in step 3 to obtain a generated image of the fusion of the person and the background, which specifically includes:

[0036] 4.1) Use MODNet to cut out the portrait image to be fused to obtain an image without the background;

[0037] 4.2) The image of the person without background and the background image to be fused are simultaneously input into the generative model in step 3, and a generated image of the fusion of the person and the background is obtained.

[0038] Compared with the prior art, the present invention has the following advantages:

[0039] This proposal integrates the functions of portrait and background fusion, portrait position prediction in the background, and portrait and background style unification. It can synthesize images with people as the foreground and various backgrounds taken under various conditions, so that the generated synthetic images use the style of the background in a unified manner, with a harmonious picture, a real, natural, and beautiful effect.

[0040] 1. The proposed method for intelligently generating stylized unified and fused group photos of people integrates the functions of portrait and background fusion, portrait position prediction in the background, and style unification of portrait and background, through a comprehensive position color prediction model and content style perception model; the position color prediction model realizes the prediction of the position of the character image and the color of the character image to obtain the first-stage fused image; the content style perception model realizes the fine-tuning of the content and style of the fused image, and achieves the effect of unified style of the character image and background, harmonious picture, real and natural, and beautiful appearance.

[0041] 2. The content-style perception model in the proposed method for intelligently generating stylized unified and integrated group photos of people can more meticulously unify the content of the portrait and the original image, and integrate the image with consistent style and background.

[0042] 3. The loss network of the content-style perception model in the proposed method for intelligently generating stylized unified and fused group photos of people uses a pre-trained VGG-6 model, in which the style loss is composed of its 2nd, 4th, 7th, and 10th layers. Since it uses multi-layer features to calculate the style loss, the model ensures the style unity of the fused image from more dimensions; the content loss is composed of its 7th layer; the relatively high-dimensional features are used to ensure the unity of the portrait in the fused image at a higher dimension. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 The figure is a schematic diagram of the overall process of the method for intelligently generating a group photo of people with stylized unified integration according to the present invention. DETAILED DESCRIPTION

[0044] The overall method for intelligently generating a unified and integrated group photo of people is further explained below with reference to the accompanying drawings.

[0045] like Figure 1 As shown, the unified and integrated group photo intelligent generation method proposed by the present invention includes the following steps:

[0046] Step 1: Collect character-background image materials, perform preprocessing, and obtain preprocessed character-background images;

[0047] Step 2: The preprocessed person-background picture is subjected to a cutout model to obtain a person image without background, and the person image without background is subjected to value processing to obtain a person image mask; a person-background picture, a person image, and a person image mask are formed into a triplet;

[0048] Step 3: define a generative adversarial network, including a generative model and a discriminative model, wherein the generative model includes a position-color prediction model, a content-style perception model, and a generative loss module; the position-color prediction model and the content-style perception model are connected in series to the generative model loss module; the triplet in step 2 is input into the generative model to obtain a generated image, and the generated image is passed through the generative model loss module to obtain a generated loss value; the discriminative model includes a discriminator and a discriminative loss module; the generated image and the character background image are input into the discriminative model to obtain probability values ​​respectively, and the discriminative loss module obtains the discriminative loss value by comparing the difference between the probability value and the true value, and the generated loss value and the discriminative loss value constitute the loss value of the generative adversarial network; first fix the generative model, maximize the loss value, and train the discriminative model; then fix the discriminative model, minimize the loss value, and train the generative model; and obtain an optimized generative model by alternately training the generative model and the discriminative model;

[0049] Step 4: For the model generated in step 4, input the background image and the portrait image to obtain the generated image that is a fusion of the person and the background;.

[0050] In step 1, character-background image materials are collected and preprocessed to obtain preprocessed character-background images; specifically, the following steps are included:

[0051] 1.1) Collect pictures with people and backgrounds, and manually select pictures with obvious people and obvious backgrounds;

[0052] 1.2) Scale the selected images to 512*512, denoted as I, where the i-th person-background image is denoted as I i ;

[0053] In step 2, the preprocessed person-background picture is subjected to a cutout model to obtain a person image without background, and the person image without background is processed to obtain a person image mask; a person-background picture, a person image, and a person image mask are formed into a triplet, specifically including:

[0054] 2.1) Use the deep image segmentation model MODNet to cut out I and obtain an image without background characters, denoted as J, where the jth character image is denoted as J j ;

[0055] 2.2) Use Python image library PIL to process the value of J and generate a mask of the person image with the transparent part as black (0) and the color part as white (255), denoted as K. The jth person image mask is denoted as K j ;

[0056] In step 3, a generative adversarial network is defined, which includes a generative model and a discriminative model. The generative model includes a position-color prediction model, a content-style perception model, and a generative model loss module; the position-color prediction model and the content-style perception model are connected in series to the generative model loss module; the triplet in step 2 is input into the generative model to obtain a generated image, and the generated image is passed through the generative model loss module to obtain a loss value. The discriminative model includes a discriminator and a discriminative loss module, which specifically includes:

[0057] 3.1) The generative model of the adversarial generative network defined, the generative model is denoted as G, which includes a position-color prediction model, a content-style perception model and a generative model loss module;

[0058] 3.2) The position-color prediction model includes a person image position prediction model and a person image color prediction model, and the person-background image and the person image are input to predict the position and color of the person image, specifically:

[0059] 3.2.1) Person image position prediction model: Use a multi-layer convolutional layer as the model, the model is denoted as A, and the model input is a person-background picture I with i≠j i 、Character ImageJ j Get the prediction of the portrait position (i.e. generate the character image mask, denoted as K j ′), K j ′ and the real person image mask K i For comparison, the mean square error loss function is used to calculate the loss, denoted as L g , K j The generation formula is as follows:

[0060] K j ′=A(I i J j ) (1)

[0061] This formula represents the relationship between the character image position prediction model A and the input and output; i A picture representing a person and a background, J j represents a person image, and i≠j;

[0062] L g The generation formula is as follows;

[0063]

[0064] This formula represents the generation of the character image mask K j ′ and the real person image mask K i The mean square error loss function is: where N represents the number of training data;

[0065] 3.2.2) Character image color prediction model: Use a multi-layer convolutional layer as the model, the model is denoted as C, and the model input is a character-background picture I with i≠j i 、Character ImageJ j Get person image J j in I i Color prediction in (i.e. generating the color of the character image, denoted by J j ′), J j ' and real portrait J i For comparison, the loss is calculated using the L1 normal function, denoted as L c , J j ′, L c The formula is as follows:

[0066] J j ′=C(I i J j ) (3)

[0067] This formula represents the relationship between the character image color prediction model C and the input and output; i A picture representing a person and a background, J j represents a person image, and i≠j;

[0068] The Lc generation formula is as follows;

[0069]

[0070] This formula represents the generated character image color J j ' and real portrait J i The loss function is calculated by the L1 norm function; where N represents the number of training data;

[0071] 3.2.3) Mask the generated character image K j ′ and character image color J j ′ to merge and retain I i The background image of J j Fusion to I i The fused image generated in is denoted as I′ j_comp ; Formula table

[0072] The following are shown:

[0073]

[0074] This formula represents the generated character image mask K j ′ and character image color J j 'To fuse;

[0075] Among them I i A picture representing a person and background, Kj ′ represents the generated character image mask, J j ′ represents the color of the generated character image;

[0076] in represents the Hadamard product, which is the element-by-element product of matrices of the same shape.

[0077] 3.3) The content-style perception model includes a perception model and a content-style loss module, for the fused image I′ generated in 3.2.3) j_comp The input content-style perception model is fine-tuned to fine-tune the boundary problem and generate a more refined fused image; denoted as I j_comp ;Specifically:

[0078] 3.3.1) Content-style loss module, denoted as φ, uses the pre-trained image classification model VGG-16 encoder as a fixed loss network;

[0079] VGG-16 contains 13 convolutional layers and 3 fully connected layers. 13 of the convolutional layers are used as the loss network, where the style loss is composed of the 2nd, 4th, 7th, and 10th layers, denoted as φ style ; The content loss consists of the 7th layer, denoted as φ feature ;

[0080] 3.3.2) Perception model, denoted as Net; includes an encoder and a decoder. The encoder uses a downsampled residual network and implementation to encode the image to obtain image coding features, and the decoder uses an upsampled convolutional network to decode the image coding features to obtain the image. Specifically:

[0081] 3.3.2.1) Input I′ to the perception model j_comp , get the image O j_comp , O j_comp By φ feature Get content feature y j_comp , J j By φ feature Get content feature y j , calculate y by Euclidean distance and normalization j_comp and the target value y j The difference between them is taken as loss, recorded as loss feature ;

[0082] The formula is as follows:

[0083]

[0084] This formula represents the generated image O of the computational perception model. j_comp With real person imagesJ jThe loss function of the content feature difference between them;

[0085] where φ j j=7 represents the 7th layer of VGG-16;

[0086] Among them C j H j W j Represents the three-dimensional array corresponding to the 7th layer of VGG-16, C j is the number of channels of the feature map (Channel), H j is the height of the feature map, W j is the width of the feature map;

[0087] where y' i That is, the input value O j_comp ,y i That is, the input value J j ;

[0088] Where N represents the number of training data;

[0089] 3.3.2.2)O j_comp By φ style Get style feature z j_comp , I i By φ style Get style feature z j , by calculating z j_comp and z j The difference in the square of the Frobenius Norm between the Gram Matrices of the images is taken as the loss, denoted as loss style ;

[0090]

[0091] This formula represents the generated image O of the computational perception model. j_comp With real people - background image I i The loss function of the difference between the segmentation features;

[0092] Where j∈{2,4,7,10} represents the 2nd, 4th, 7th, and 10th layers of VGG-16;

[0093] Among them G φ j Represents the Gram matrix of the j-th layer features of VGG-16;

[0094] Where F represents the Frobenius norm;

[0095] where y' i That is, the input value Oj_comp ,y i That is, the input value I i ;

[0096] Where N represents the number of training data;

[0097] 3.3.2.3) The total loss of the perceptual model is denoted as L f ; The formula is as follows:

[0098] L f =loss feature +loss style (8)

[0099] This formula indicates that the perceptual model loss is composed of content feature loss loss feature And style feature loss loss style composition

[0100] 3.3.3) The perception model generates the fused image I′ in 3.2.3) j_comp Fine-tune the content and style, expressed as the following formula:

[0101] I comp =Net(I′ comp ) (9)

[0102] This formula represents the fused image I′ j_comp The fused image I′ is j_comp Fine-tune the content and style to obtain the generated image I comp ;

[0103] 3.4) The discriminant model consists of a discriminant network and a loss module; specifically:

[0104] 3.4.1) The discriminant network consists of three convolutional layers, a spectral normalization layer (to limit the range of variation of the weight matrix), and a Sigmoid function, denoted as D;

[0105] 3.4.2) Input the generated image I obtained in step 2 into the discriminant network comp and the real person-background image, and get the probability value of the image being false / true, denoted as p i , the difference between the probability value and the true label is calculated by the binary cross loss function to obtain the discriminant loss value, denoted as L a ;

[0106]

[0107] This formula represents the loss value calculated by the binary cross loss function;

[0108] y iis the true label of the i-th sample, which takes the value of 0 or 1. comp When the input is a real person-background image, it is 0;

[0109] p i is the predicted probability of the i-th sample, that is, the probability value output by the model, between 0 and 1;

[0110] Where N represents the number of training data;

[0111] 3.5) The final loss function is denoted as L, and the formula is as follows:

[0112]

[0113] This formula represents the composition of the loss function of the adversarial generation network;

[0114] Where L a Represents the discriminant loss value;

[0115] L c Indicates the color loss value corresponding to the character image color prediction model;

[0116] L g Represents the position loss value corresponding to the character image position prediction model;

[0117] L f Represents the sum of the content value and style loss value corresponding to the content-style perception model;

[0118] in Indicates optimizing the discriminant network by maximizing the loss value;

[0119] Represents minimizing the loss value to optimize the generated network;

[0120] 3.6) By alternately training the generator and the discriminator, the final optimized generation model is obtained.

[0121] Step 4: Input the background image and the portrait image into the model generated in step 3 to obtain a generated image that is a fusion of the person and the background, which specifically includes:

[0122] 4.1) Use the deep image segmentation model MODNet to cut out the portrait image to obtain an image without the background;

[0123] 4.2) Input the image of the person without background and the background image into the generative model in step 3 at the same time, and obtain the generated image of the fusion of the person and the background.

[0124] The above is only an illustration of the present invention, and is not intended to limit the present invention. A person skilled in the art should recognize that any changes or modifications made to the present invention will fall within the protection scope of the present invention.

Claims

1. A method for intelligently generating a stylized, unified and integrated group photo of people, characterized in that: The following steps are involved: Step 1: Collect character-background image materials, perform preprocessing, and obtain preprocessed character-background images; Step 2: The preprocessed person-background image is subjected to a cutout model to obtain a person image without background, and the person image without background is subjected to value processing to obtain a person image mask; the preprocessed person-background image, the person image without background, and the person image mask are combined into a triplet; Step 3: define a generative adversarial network, including a generative model and a discriminative model, wherein the generative model includes a position-color prediction model, a content-style perception model, and a generative loss module, and the discriminative model includes a series-connected discriminator and a discriminative loss module; Input the triplet obtained in step 2 into the generative model to obtain a generated image, and the generated image is passed through the generative loss module to obtain a generated loss value; The generated image is input into the discriminant model to obtain the probability value of the false value. The preprocessed person-background image obtained in step 1 is input into the discriminant model to obtain the probability value of the true value. The discriminant loss module obtains the discriminant loss value by comparing the difference between the probability value of the false value and the value 0 and the difference between the probability value of the true value and the value 1. The generated loss value and the discriminant loss value constitute the loss value of the adversarial generative network. First, fix the generative model, maximize the loss value of the adversarial generative network, and train the discriminative model; then fix the discriminative model, minimize the loss value of the adversarial generative network, and train the generative model. By alternately training the generative model and the discriminative model, an optimized generative model is obtained. Step 4: Input the background image and portrait image to be fused into the optimized generation model obtained in step 3 to obtain a generated image of the fusion of the person and the background.

2. The method for intelligently generating a stylized, unified and integrated group photo of people according to claim 1, characterized in that: In step 1, character-background image materials are collected and preprocessed to obtain preprocessed character-background images, specifically including: 1.1) Collect pictures with people and backgrounds, and select pictures with obvious people and obvious backgrounds; 1.2) Scale the images selected in step 1.1) to a uniform size, denoted as I, where the i-th person-background image is denoted as I i .

3. The method for intelligently generating a stylized, unified and integrated group photo of people according to claim 1, characterized in that: In step 2, the preprocessed person-background image is subjected to a cutout model to obtain a person image without background, and the person image without background is processed to obtain a person image mask; The preprocessed person-background image, the person image without background, and the person image mask are combined into a triplet, specifically including: 2.1) Use MODNet to cut out I and get a character image without background, denoted as J, where the jth character image is denoted as J j ; 2.2) Use PIL to process the value of J to generate a character image mask with black transparent part and white color part, denoted as K. The jth character image mask is denoted as K j .

4. The method for intelligently generating a stylized, unified and integrated group photo of people according to claim 3, characterized in that: In step 3, the position-color prediction model, content-style perception model and generation model loss module are connected in series.

5. The method for intelligently generating a stylized, unified and integrated group photo of people according to claim 4, characterized in that: In step 3, the position-color prediction model includes a person image position prediction model and a person image color prediction model.

6. The method for intelligently generating a stylized, unified and integrated group photo of people according to claim 5, characterized in that: In step 3, the content-style perception model includes a perception model and a content-style loss module, the perception model includes an encoder and a decoder, and the content-style loss module adopts a picture classification model VGG-16 encoder.

7. The method for intelligently generating a stylized, unified and integrated group photo of people according to claim 1, characterized in that: In step 3, the generation loss module uses the mean square error loss function, the L1 norm function, the Euclidean distance calculated by the 7th layer of the classification model VGG-16 encoder, and the square of the Frobenius norm between the Gram matrices calculated by the 2nd, 4th, 7th, and 10th layers of the classification model VGG-16 encoder.

8. The method for intelligently generating a stylized, unified and integrated group photo of people according to claim 1, characterized in that: In step 3, the discriminative loss module is calculated using a binary cross loss function.

9. The method for intelligently generating a stylized, unified and integrated group photo of people according to claim 1, characterized in that: Step 4: Input the background image and the portrait image to be fused into the optimized generation model obtained in step 3 to obtain a generated image of the fusion of the person and the background, which specifically includes: 4.1) Use MODNet to cut out the portrait image to be fused to obtain an image without the background; 4.2) The image of the person without background and the background image to be fused are simultaneously input into the generative model in step 3, and a generated image of the fusion of the person and the background is obtained.