Method and device for manufacturing model by combining diffusion network and adversarial network

By combining diffusion networks and adversarial networks, high-resolution real model images are generated using adversarial neural networks and face recognition models, the limitations and authorization problems of generated images in the prior art are solved, and high-quality and low-cost real model generation is achieved.

CN120070625APending Publication Date: 2025-05-30ZIXUN TECHNOLOGY (FUJIAN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510108257.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has limitations when generating real model images, making it difficult to generate high-resolution images, and requires art retouching, which has authorization problems, resulting in increased costs.

Method used

Combining the diffusion network and the adversarial network, the initial image is generated through the adversarial neural network, and high-dimensional feature vectors are extracted using the face recognition model, and the diffusion model is input for final generation, realizing multi-view batch generation of non-infringing real models.

Benefits of technology

It effectively improves the quality and authenticity of real model production, realizes multi-perspective batch generation of non-infringement real models, and reduces the cost of merchants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070625A_ABST
    Figure CN120070625A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for manufacturing a model by combining a diffusion network and an adversarial network, and the method comprises the steps: screening out a needed picture from a face image library, and taking the needed picture as training image data; processing the training image data to obtain an image with a set resolution as training image processing data; the adversarial neural network performs main framework construction based on a StyleGAN2 architecture, and the training image processing data is input into the adversarial neural network for training to obtain a generative network model; the method comprises the following steps: generating a reference image through a generative network model, then performing feature extraction on the reference image by using a face recognition model FaceNet, and extracting a high-dimensional feature vector of 16 * 768 dimensions after switching of a full connection layer; inputting the high-dimensional feature vector into a diffusion model to generate a final real model portrait image; and the generated model can be directly used for commercial use, so that the enterprise cost is greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image generation, and particularly to a method and device for creating models by combining a diffusion network and a generative adversarial network. Background Art

[0002] In the field of creating real models, traditional technologies have limitations in generating model images, and can only generate low-resolution facial images; in addition, if the images are directly obtained by photographing a person, they need to be retouched by an artist before being used in the diffusion model. Without retouching, due to the flaws in the directly photographed images, when used in the diffusion model, the flaws will be magnified and cannot be used in practice; and if using the face of this person, authorization from the person himself / herself is also required, resulting in a further increase in the cost of the enterprise. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a method and device for creating models by combining a diffusion network and a generative adversarial network. Utilizing the high-performance generation ability of the diffusion network and combining the randomness of human faces of the generative adversarial network, through the synergistic effect of the two, the quality and realism of creating real models are effectively improved, and multi-view batch generation of non-infringing real models is achieved.

[0004] In a first aspect, the present invention provides a method for creating models by combining a diffusion network and a generative adversarial network, including the following steps:

[0005] Step 1: Screen out the required pictures from the face image library as training image data;

[0006] Step 2: Process the training image data to obtain an image with a set resolution as the processed training image data;

[0007] Step 3: Build the main framework of the adversarial neural network based on the StyleGAN2 architecture, input the processed training image data into the adversarial neural network for training to obtain a generation network model;

[0008] Step 4: Generate a reference image through the generation network model, then use the face recognition model FaceNet to extract features from the reference image, and after passing through a fully connected layer for transfer, extract a high-dimensional feature vector of 16 * 768 dimensions;

[0009] Step 5: Input the high-dimensional feature vector into the diffusion model to generate the final real model portrait.

[0010] In a second aspect, the present invention provides a device for creating models by combining a diffusion network and a generative adversarial network, including:

[0011] A data screening module that screens out the required pictures from the face image library as training image data;

[0012] A data processing module processes the training image data to obtain an image with a set resolution as the training image processing data.

[0013] A model generation training module builds a main framework for the adversarial neural network based on the StyleGAN2 architecture, inputs the training image processing data into the adversarial neural network for training, and obtains a generation network model.

[0014] A feature vector acquisition module generates a reference image through the generation network model, then uses the face recognition model FaceNet to extract features from the reference image, and after passing through a fully connected layer, extracts a high-dimensional feature vector with a dimension of 16 * 768.

[0015] A portrait generation module inputs the high-dimensional feature vector into a diffusion model to generate a final real model portrait.

[0016] One or more technical solutions provided by the present invention have at least the following technical effects or advantages:

[0017] The present invention utilizes the high-performance generation ability of the diffusion network, combines the randomness of the face in the generative adversarial network, and through the synergistic effect of the two, effectively improves the quality and realism of the production of real models, and realizes the multi-view batch generation of non-infringing real models; the models generated in this way do not have authorization problems and are guaranteed to be different from all existing faces; after the real model pictures are generated, they can be used as display pictures for merchants in the later stage, greatly reducing the costs of merchants.

[0018] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically gives the specific embodiments of the present invention. Brief Description of the Drawings

[0019] The present invention will be further described below with reference to the accompanying drawings in conjunction with embodiments.

[0020] Figure 1 It is a flowchart of the method in Embodiment 1 of the present invention;

[0021] Figure 2 It is a structural schematic diagram of the device in Embodiment 2 of the present invention. Detailed Embodiments

[0022] The overall idea of the technical solution in the embodiments of the present application is as follows:

[0023] First, an adversarial neural network needs to be trained for the generation of real model faces, and then the face is input into a diffusion model for the full-body generation of the model.

[0024] The adversarial neural network prepares the training dataset with the following steps: First, it executes the detection function using the yolov5face algorithm to identify and filter out the parts of frontal face images. For the eligible face images, it executes the calculate_face_area() function to ensure that the face area occupies 60%-80% of the central area of the image. It uses the calculate_rotation() function to check whether the side rotation angle is within ±5°. Then, it performs color histogram analysis through the analyze_histogram() function to eliminate abnormal exposure samples, ensuring that the standard deviation of the RGB three channels of the image is less than 45 and the brightness is concentrated in the range of 0.4-0.8.

[0025] After preparing the head portraits, it uses the select_high_quality_images() function to perform the second screening on the images. This function first uses cv2.GaussianBlur to perform Gaussian blur denoising on the images (kernel_size = 3) to avoid the influence of excessive noise on the final image sharpness score; then it uses cv2.cvtColor to convert the images to grayscale, and then uses cv2.Laplacian(gray, cv2.CV_64F) to calculate the Laplacian operator, takes the absolute value of the result and converts it to the uint8 type, and finally calculates the mean value of the result to obtain the sharpness score. The threshold is set to 6. When the sharpness score is greater than 6, it is marked as a high-quality sample, and other samples are not marked. In this way, 10,000 eligible high-quality training images are screened out from the original dataset.

[0026] After obtaining the screened training images, preprocessing preparation before training is required, and only after processing can they be fed into the neural network for training. Before training, the following methods are used for preprocessing operations. All images are processed to a resolution of 512×512 using the cv2.resize() function and the INTER_AREA interpolation method. First, the images are scaled proportionally so that the shorter side reaches 1024 pixels. Then, the crop_image_center() function is used to crop the images with the face center point as the center. The pixel values are normalized from [0, 255] to [-1, 1] through the normalize_pixels() function, and random horizontal flipping and color jittering are performed using the random_flip() and random_color_jitter() functions to enhance data diversity. Finally, the data required before neural network training is obtained.

[0027] The adversarial neural network is built based on the StyleGAN2 architecture. The mapping network of the generator uses 8 fully connected layers (each layer with 512 dimensions) and the LReLU activation function (α = 0.2). The discriminator uses the ResNet architecture, including 6 residual blocks, and each residual block uses a 3×3 convolutional kernel and the LeakyReLU activation function (α = 0.2) as the building blocks.

[0028] Based on the data prepared above, training the adversarial neural network can obtain the model of the StyleGAN2 generation network. Using this network model, arbitrary model avatars of 512*512 can be finally generated, which are not infringing and can be used as the input images before the next step of using the diffusion network.

[0029] To further enhance the facial expression details of the models in the generated images, we use the face recognition model FaceNet to extract features from each training image. After being transferred through the output fully connected layer, a high-dimensional feature vector of 16*768 dimensions is extracted, which accurately encodes key information such as the geometric structure of the facial contour and facial features, skin texture and micro-expression details, and identity recognition features. This high-dimensional feature vector, as the basic feature representation for enhancing expression details, will be fed into the diffusion model for inference in the next step.

[0030] After obtaining the high-dimensional feature vector, it is input into the diffusion model for the final generation of models. During the denoising process, the DDIM (Denoising Diffusion Implicit Model) sampler is used, with the number of sampling steps set to 50 steps and the time step size for each step set to 0.02. The Euler ancestral sampling method is used during the sampling process, and the CFG scale is set to 7 to balance the generation quality and diversity. At the same time, attention guidance is introduced during the sampling, set to 1.5, to enhance the attention to the key facial regions.

[0031] Through the iterative denoising process, the model will continuously generate models that meet the corresponding features based on the high-dimensional feature vector. Finally, after the generation process of the diffusion model, a high-definition image with a resolution of 1024×1024 is output, and then it is enlarged to a resolution of 2048×2048 through a 2x upsampling network, realizing the overall generation of high-definition non-infringing model images.

[0032] Finally, the generated model image is cropped to a size with an aspect ratio of 2:3 to obtain the final real model portrait.

[0033] Embodiment 1

[0034] As Figure 1 shown, this embodiment provides a method for making models by combining a diffusion network and an adversarial network, including the following steps:

[0035] Step 1: Screen out the required pictures from the face image library as training image data;

[0036] Step 2: Process the training image data to obtain an image with a set resolution as the training image processing data;

[0037] Step 3: Build the main framework of the adversarial neural network based on the StyleGAN2 architecture, input the training image processing data into the adversarial neural network for training to obtain a generation network model;

[0038] Step 4: Generate a reference image through the generation network model, then use the face recognition model FaceNet to extract features from the reference image. After passing through the fully connected layer, extract a high-dimensional feature vector of 16 * 768 dimensions;

[0039] Step 5: Input the high-dimensional feature vector into the diffusion model to generate the final real model portrait.

[0040] In this embodiment, preferably, the specific content of Step 1 is as follows: First, use the yolov5face algorithm to execute the detection function to identify and screen out the frontal face part; for the screened face images, use the calculate_face_area function to ensure that the face area occupies 60%-80% of the central area of the image; use the calculate_rotation function to check whether the side rotation angle is within ±5°; then, perform color histogram analysis through the analyze_histogram function to eliminate abnormal exposure samples, ensure that the standard deviation of the RGB three channels of the image is less than 45, and the brightness is concentrated in the range of 0.4-0.8 to obtain the preprocessed image data; use the select_high_quality_images function to screen the preprocessed data. The select_high_quality_images function first uses cv2.GaussianBlur to perform Gaussian blur denoising on the image; then uses cv2.cvtColor to convert the image to a grayscale image, then uses cv2.Laplacian(gray, cv2.CV_64F) to calculate the Laplacian operator, take the absolute value of the result, and convert it to the uint8 type. Finally, calculate the average value of all results to obtain the sharpness score. Set the threshold to 6, and screen out 10,000 images with a sharpness score greater than 6 from the preprocessed image data to obtain the training image data.

[0041] In this embodiment, preferably, step 2 is specifically as follows: All images in the training image data are processed using the cv2.resize function and the INTER_AREA interpolation method, and each image is processed to a resolution of 512×512; that is, the image is scaled proportionally so that the shorter side reaches 1024 pixels, and then the crop_image_center function is used to crop with the center point of the face as the center to obtain an image with a resolution of 512×512.

[0042] The images with a resolution of 512×512 are randomly horizontally flipped and color jittered using the random_flip function and the random_color_jitter function, and then the pixel values are normalized from [0, 255] to [-1, 1] through the normalize_pixels function, and finally the training image processing data is obtained.

[0043] In this embodiment, preferably, step 3 is specifically as follows: The adversarial neural network is built with the StyleGAN2 architecture as the main framework. Among them, the mapping network of the generator adopts 8 fully connected layers, each fully connected layer is 512-dimensional, and the LReLU activation function is used; the discriminator adopts the ResNet architecture, and the ResNet architecture includes 6 residual blocks, and each residual block includes a 3×3 convolution kernel and a LeakyReLU activation function; the training image processing data is input into the adversarial neural network for training to obtain the StyleGAN2 generation network model.

[0044] In this embodiment, preferably, step 5 is specifically as follows: The high-dimensional feature vector is input into the diffusion model.

[0045] In the diffusion model, the denoising process uses the DDIM sampler, and the number of sampling steps is set to 50 steps; Euler ancestral sampling is used, the CFG scale is set to 7, and attention guidance is introduced, set to 1.5.

[0046] An image with a resolution of 1024×1024 is generated and output through the diffusion model, the image is enlarged to a resolution of 2048×2048 through a 2x upsampling network, and then cropped to an aspect ratio of 2:3 to obtain the final real model portrait.

[0047] Based on the same inventive concept, the present application also provides an apparatus corresponding to the method in Embodiment 1, as detailed in Embodiment 2.

[0048] Embodiment 2

[0049] As Figure 2 shown, in this embodiment, an apparatus for making a model by combining a diffusion network and an adversarial network is provided, including:

[0050] A data screening module that screens out the required pictures from the face image library as training image data;

[0051] A data processing module that processes the training image data to obtain an image with a set resolution as processed training image data;

[0052] A model generation and training module that builds the main framework of the adversarial neural network based on the StyleGAN2 architecture, inputs the processed training image data into the adversarial neural network for training, and obtains a generation network model;

[0053] A feature vector acquisition module that generates a reference image through the generation network model, then uses the face recognition model FaceNet to extract features from the reference image, and after passing through a fully connected layer transfer, extracts a high-dimensional feature vector of 16 * 768 dimensions;

[0054] A portrait generation module that inputs the high-dimensional feature vector into a diffusion model to generate a final real model portrait.

[0055] In this embodiment, preferably, the data screening module is specifically as follows: First, use the yolov5face algorithm to execute the detection function to identify and screen out the frontal view face part; for the screened face images, use the calculate_face_area function to ensure that the face area occupies 60%-80% of the central area of the image; use the calculate_rotation function to check whether the side rotation angle is within ±5°; then, perform color histogram analysis through the analyze_histogram function to eliminate abnormal exposure samples, ensure that the standard deviation of the RGB three channels of the image is less than 45, and the brightness is concentrated in the range of 0.4-0.8 to obtain preprocessed image data; use the select_high_quality_images function to screen the preprocessed data. The select_high_quality_images function first uses cv2.GaussianBlur to perform Gaussian blur denoising on the image; then uses cv2.cvtColor to convert the image to a grayscale image, then uses cv2.Laplacian(gray, cv2.CV_64F) to calculate the Laplacian operator, takes the absolute value of the result, and converts it to the uint8 type. Finally, calculate the average value of all results to obtain the sharpness score. Set the threshold to 6, and screen out 10,000 images with a sharpness score greater than 6 from the preprocessed image data to obtain the training image data.

[0056] In this embodiment, preferably, the data processing module is specifically configured to: process all the images in the training image data using the cv2.resize function and the INTER_AREA interpolation method, and process each image to a resolution of 512×512; that is, scale the image proportionally so that the shorter side reaches 1024 pixels, and then use the crop_image_center function to crop with the face center point as the center to obtain an image with a resolution of 512×512.

[0057] Randomly horizontally flip and perform color jitter on the 512×512 resolution images using the random_flip function and the random_color_jitter function, and then normalize the pixel values from [0, 255] to [-1, 1] through the normalize_pixels function to finally obtain the training image processing data.

[0058] In this embodiment, preferably, the model generation training module is specifically configured to: build the main framework of the adversarial neural network based on the StyleGAN2 architecture, where the mapping network of the generator uses 8 fully connected layers, each fully connected layer is 512-dimensional, and the LReLU activation function is used; the discriminator uses the ResNet architecture, and the ResNet architecture includes 6 residual blocks, each residual block includes a 3×3 convolutional kernel and the LeakyReLU activation function; input the training image processing data into the adversarial neural network for training to obtain the StyleGAN2 generation network model.

[0059] In this embodiment, preferably, the generated portrait module is specifically configured to: input the high-dimensional feature vector into the diffusion model;

[0060] The denoising process in the diffusion model uses the DDIM sampler, sets the number of sampling steps to 50 steps; uses Euler ancestral sampling, sets the CFG scale to 7, and introduces attention guidance, set to 1.5;

[0061] Generate and output an image with a resolution of 1024×1024 through the diffusion model, magnify the image to a resolution of 2048×2048 through a 2x upsampling network, and then crop it to an aspect ratio of 2:3 to obtain the final real model portrait.

[0062] Since the device introduced in the second embodiment of the present invention is the device used to implement the method of the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and variations of the device, so it will not be elaborated here. Any device used in the method of the first embodiment of the present invention belongs to the scope of protection of the present invention.

[0063] The technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0064] Through the method of this embodiment, a real face portrait can be randomly generated. This real face portrait can be directly used by users, who can use it in a diffusion model to generate the commercial exhibition pictures they need for display in combination with products.

[0065] Although the specific embodiments of the present invention have been described above, those skilled in the art should understand that the specific embodiments we described are illustrative rather than limiting the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should be covered by the scope of the claims of the present invention.

Claims

1. A method for making a model by combining a diffusion network and an adversarial network, characterized in that: The steps include: Step 1: Filter the required pictures from the face library as training image data; Step 2: Process the training image data to obtain an image with a set resolution as training image processing data; Step 3: The main framework of the adversarial neural network is built based on the StyleGAN2 architecture, and the training image processing data is input into the adversarial neural network for training to obtain a generated network model; Step 4: Generate a reference image by generating a network model, and then use the face recognition model FaceNet to extract features from the reference image. After being transferred through the fully connected layer, a high-dimensional feature vector of 16*768 dimensions is extracted; Step 5: Input the high-dimensional feature vector into the diffusion model to generate the final real model portrait.

2. The method for making a model by combining a diffusion network and an adversarial network according to claim 1, characterized in that: The step 1 is specifically as follows: first, the detection function is executed using the yolov5face algorithm to identify and screen out the face portion of the front view; the screened face image is used to use the calculate_face_area function to ensure that the face area occupies 60%-80% of the center area of ​​the image; the calculate_rotation function is used to check whether the side rotation angle is within ±5°; then, the color histogram analysis is performed through the analyze_histogram function to eliminate abnormal exposure samples, ensure that the standard deviation of the three RGB channels of the image is less than 45, and the brightness is concentrated in the range of 0.4-0.8, and obtain the preprocessed image data; The preprocessed image data is screened using the select_high_quality_images function, which first uses cv2.GaussianBlur to perform Gaussian blur denoising on the image; then uses cv2.cvtColor to convert the image into a grayscale image, and then uses cv2.Laplacian(gray,cv2.CV_64F) to calculate the Laplacian operator, takes the absolute value of the result, and converts it to uint8 type, and finally calculates the average of all results to obtain a clarity score, sets the threshold to 6, and screens out 10,000 images with a clarity score greater than 6 from the preprocessed image data to obtain training image data.

3. The method for making a model by combining a diffusion network and an adversarial network according to claim 1, characterized in that: The step 2 is specifically as follows: all images in the training image data are processed using the cv2.resize function and the INTER_AREA interpolation method, and each image is processed to a resolution of 512×512; that is, the image is proportionally scaled so that the short side reaches 1024 pixels, and then the crop_image_center function is used to crop the image with the center point of the face as the center to obtain an image with a resolution of 512×512; The 512×512 resolution image is randomly flipped horizontally and color jittered using the random_flip function and random_color_jitter function, and then the pixel value is normalized from [0,255] to [-1,1] using the normalize_pixels function to obtain the training image processing data.

4. The method for making a model by combining a diffusion network and an adversarial network according to claim 1, characterized in that: The step 3 is specifically as follows: the adversarial neural network is built based on the StyleGAN2 architecture for the main framework, wherein the mapping network of the generator adopts 8 fully connected layers, each fully connected layer is 512-dimensional, and uses the LReLU activation function; the discriminator adopts the ResNet architecture, and the ResNet architecture includes 6 residual blocks, each residual block includes a 3×3 convolution kernel and a LeakyReLU activation function; the training image processing data is input into the adversarial neural network for training to obtain the StyleGAN2 generation network model.

5. The method for making a model by combining a diffusion network and an adversarial network according to claim 1, characterized in that: The step 5 specifically includes: inputting the high-dimensional feature vector into the diffusion model, In the diffusion model, the denoising process uses the DDIM sampler and sets the sampling step to 50 steps. Eulerancestral sampling is used, the CFG scale is set to 7, and attention guidance is introduced and set to 1.

5. The diffusion model generates and outputs an image with a resolution of 1024×1024, which is then enlarged to a resolution of 2048×2048 through a 2x upsampling network and then cropped to an aspect ratio of 2:3 to obtain the final real model portrait.

6. A device for making a model by combining a diffusion network and an adversarial network, characterized in that: include: The data screening module selects the required images from the face library as training image data; A data processing module processes the training image data to obtain an image of a set resolution as training image processing data; Model generation training module, the adversarial neural network is built based on the StyleGAN2 architecture as the main framework, the training image processing data is input into the adversarial neural network for training, and the generated network model is obtained; The feature vector acquisition module generates a reference image by generating a network model, and then uses the face recognition model FaceNet to extract features from the reference image. After being transferred through the fully connected layer, a high-dimensional feature vector of 16*768 dimensions is extracted; Generate portrait module, input high-dimensional feature vector into diffusion model to generate the final real model portrait.

7. The device for making a model by combining a diffusion network and an adversarial network according to claim 6, characterized in that: The data screening module is specifically as follows: first, the detection function is executed using the yolov5face algorithm to identify and screen out the face part of the front view; the screened face image is used to use the calculate_face_area function to ensure that the face area occupies 60%-80% of the center area of ​​the image; the calculate_rotation function is used to check whether the side rotation angle is within ±5°; then, the color histogram analysis is performed through the analyze_histogram function to eliminate abnormal exposure samples, ensure that the standard deviation of the three RGB channels of the image is less than 45, and the brightness is concentrated in the range of 0.4-0.8, and obtain the pre-processed image data; The preprocessed image data is screened using the select_high_quality_images function, which first uses cv2.GaussianBlur to perform Gaussian blur denoising on the image; then uses cv2.cvtColor to convert the image into a grayscale image, and then uses cv2.Laplacian(gray,cv2.CV_64F) to calculate the Laplacian operator, takes the absolute value of the result, and converts it to uint8 type, and finally calculates the average of all results to obtain a clarity score, sets the threshold to 6, and screens out 10,000 images with a clarity score greater than 6 from the preprocessed image data to obtain training image data.

8. The device for making a model by combining a diffusion network and an adversarial network according to claim 6, characterized in that: The data processing module specifically includes: processing all images in the training image data using the cv2.resize function and the INTER_AREA interpolation method, processing each image to a resolution of 512×512; that is, scaling the image in proportion so that the short side reaches 1024 pixels, and then using the crop_image_center function to crop the image with the center point of the face as the center to obtain an image with a resolution of 512×512; The 512×512 resolution image is randomly flipped horizontally and color jittered using the random_flip function and random_color_jitter function, and then the pixel value is normalized from [0,255] to [-1,1] using the normalize_pixels function to obtain the training image processing data.

9. The device for making a model by combining a diffusion network and an adversarial network according to claim 6, characterized in that: The model generation training module is specifically as follows: the adversarial neural network is built based on the StyleGAN2 architecture as the main framework, wherein the mapping network of the generator adopts 8 fully connected layers, each fully connected layer is 512 dimensions, and the LReLU activation function is used; the discriminator adopts the ResNet architecture, and the ResNet architecture includes 6 residual blocks, each residual block includes a 3×3 convolution kernel and a LeakyReLU activation function; the training image processing data is input into the adversarial neural network for training to obtain the StyleGAN2 generation network model.

10. The device for making a model by combining a diffusion network and an adversarial network according to claim 6, characterized in that: The portrait generation module specifically includes: inputting the high-dimensional feature vector into the diffusion model; In the diffusion model, the denoising process uses the DDIM sampler and sets the sampling step to 50 steps. Eulerancestral sampling is used, the CFG scale is set to 7, and attention guidance is introduced and set to 1.

5. The diffusion model generates and outputs an image with a resolution of 1024×1024, which is then enlarged to a resolution of 2048×2048 through a 2x upsampling network and then cropped to an aspect ratio of 2:3 to obtain the final real model portrait.