AIGC-based photographed image generation method and system, and photographing device

By extracting user pose skeletons and facial feature vectors, and performing foreground subject generation based on lighting priors and background fusion guided by depth, the problem of foreground and background matching in AIGC technology is solved, generating photographic images with consistent lighting and spatial representation and accurate identity.

CN120897043AActive Publication Date: 2025-11-04浙江天怀数智科技有限公司

Patent Information

Application Number
CN202511440000.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-10
Publication Date
2025-11-04
Estimated Expiration
2045-10-10

AI Technical Summary

Technical Problem

Existing AIGC technology struggles to achieve a natural match between foreground figures and background lighting environment and spatial layout in the generation of photographic images, leading to problems such as conflicting lighting directions, missing ambient light reflections, and reduced fidelity of user identity features.

Method used

By acquiring user images and style template images, extracting user pose skeletons and facial feature vectors, performing foreground subject generation based on lighting priors, and combining depth-guided background fusion and high-fidelity detail restoration, an image with consistent lighting and spatial representation and preserved user identity features is generated.

Benefits of technology

It achieves image generation with a high degree of consistency between foreground and background in terms of lighting and space, improving the image fusion quality and visual realism, while preserving the user's identity characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120897043A_ABST
    Figure CN120897043A_ABST
Patent Text Reader

Abstract

The invention discloses an AIGC-based photographed image generation method and system and photographing equipment, and relates to the field of photographed image generation, and the method comprises the steps: carrying out the parallel deep analysis of a user image and a style template image, synchronously extracting the high-fidelity constraint conditions, such as the identity and posture of a user, and obtaining a high-fidelity image; and multi-dimensional stylized guidance information including clothing, background and illumination. And the illumination prior information is used as a preposition constraint, an illumination environment of a target scene is injected in a foreground main body generation stage, and space depth information of a foreground figure is synchronously constructed. And finally, providing accurate spatial constraint and guidance for a background fusion process by using the depth information, and performing detail enhancement in combination with a high-fidelity repair network, thereby generating a vivid image in which a foreground and a background are highly consistent in illumination and space and user identity features are reserved. In this way, self-adaptive regulation and control of the illumination and space relation in the generation process can be achieved, and the fusion quality and the visual reality sense of the final image are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of photograph image generation, and more specifically, to an AIGC-based photograph image generation method and system, and a photographing device. BACKGROUND

[0002] With the rapid development of digital media and social networks, users' demand for personalized and high-quality visual content is growing. Although traditional photography modes can produce high-quality images, the process often involves high equipment costs, professional personnel configuration, specific site selection, and time-consuming post-production, which greatly limits the freedom and efficiency of ordinary users in creative image creation. In this context, the breakthrough of generative artificial intelligence (AIGC) technology, especially diffusion models in the field of image generation, provides a new technical path for virtual photography and personalized image creation. By using AIGC technology, users only need to provide simple personal photos and style references, and it is expected to generate artistic works comparable to professional photography, greatly reducing the threshold of high-quality image creation, and showing broad application prospects and commercial value.

[0003] However, in the practice of applying AIGC technology to high-fidelity photograph image generation, existing technologies still face many challenges. Specifically, some mainstream solutions tend to adopt a phased generation strategy. First, in an independent phase, the core foreground figure is generated based on the user-provided pose information and identity features in a relatively neutral or simplified background. Subsequently, in the second phase, the generated figure image is used as the input condition of the image-to-image conversion model, and the background and atmosphere elements in the target style template are combined for secondary rendering and fusion. Although it simplifies the control difficulty to some extent, its inherent decoupling characteristics also bring insurmountable defects. Since the generation process of the foreground figure lacks prior knowledge of the final scene lighting environment and spatial layout, the lighting effect of the figure and the subsequent background integration are difficult to match naturally, often causing problems such as contradictory lighting direction and missing environmental light reflection, making the final image appear obvious "paste feeling". In addition, the stylization process in the second phase may cause uncontrollable disturbance to the precisely generated facial features and pose details of the figure, thereby weakening the fidelity of the user's identity features, making it difficult to balance identity consistency and environmental integration authenticity.

[0004] Therefore, there is an urgent need for an optimized AIGC-based photograph image generation method, system, and photographing device. SUMMARY

[0005] To solve the above technical problems, the present application is proposed.

[0006] According to an aspect of the present application, a photographing image generation method based on AIGC is provided, which comprises: acquiring a user image input by a user and a style template image selected by the user; extracting a user posture skeleton and a face feature vector from the user image; performing parameter analysis on the style template image to obtain a clothing description prompt word, a background description prompt word and a lighting description prompt word; performing foreground subject generation with lighting prior on the user posture skeleton and the face feature vector based on the clothing description prompt word and the lighting description prompt word to obtain a foreground figure image, a foreground figure mask and a foreground figure depth map; performing background fusion based on depth guidance on the foreground figure image, the foreground figure mask and the foreground figure depth map based on the background description prompt word to obtain a fused image; and performing high-fidelity detail repair on the fused image to obtain an enhanced image.

[0007] According to another aspect of the present application, a photographing image generation system based on AIGC is provided, which comprises: a user input acquisition module for acquiring a user image input by a user and a style template image selected by the user; a user feature extraction module for extracting a user posture skeleton and a face feature vector from the user image; a parameter analysis module for performing parameter analysis on the style template image to obtain a clothing description prompt word, a background description prompt word and a lighting description prompt word; a foreground subject generation module for performing foreground subject generation with lighting prior on the user posture skeleton and the face feature vector based on the clothing description prompt word and the lighting description prompt word to obtain a foreground figure image, a foreground figure mask and a foreground figure depth map; a background fusion module for performing background fusion based on depth guidance on the foreground figure image, the foreground figure mask and the foreground figure depth map based on the background description prompt word to obtain a fused image; and a high-fidelity detail repair module for performing high-fidelity detail repair on the fused image to obtain an enhanced image.

[0008] According to still another aspect of the present application, a photographing device is provided, which can execute the photographing image generation method based on AIGC as described above.

[0009] Compared with the prior art, the AIGC-based photographed image generation method, system and photographed device provided by the application can realize adaptive regulation of the light and space relationship in the generation process, effectively improve the fusion quality and visual reality of the final image. BRIEF DESCRIPTION OF DRAWINGS

[0010] The above and other objects, features and advantages of the present application will become more apparent from the following detailed description of embodiments of the present application, when taken in conjunction with the accompanying drawings. The drawings provided in the present application are used to provide further understanding of the embodiments of the present application, and constitute a part of the specification, and are used to explain the present application together with the embodiments of the present application, and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0011] Figure 1 The flowchart of the AIGC-based photographed image generation method according to the embodiments of the present application is shown.

[0012] Figure 2 The data flowchart of the AIGC-based photographed image generation method according to the embodiments of the present application is shown.

[0013] Figure 3 The flowchart of the sub-step S2 of the AIGC-based photographed image generation method according to the embodiments of the present application is shown.

[0014] Figure 4 The flowchart of the sub-step S3 of the AIGC-based photographed image generation method according to the embodiments of the present application is shown.

[0015] Figure 5 The flowchart of the sub-step S33 of the AIGC-based photographed image generation method according to the embodiments of the present application is shown.

[0016] Figure 6 The flowchart of the sub-step S4 of the AIGC-based photographed image generation method according to the embodiments of the present application is shown.

[0017] Figure 7 The flowchart of the sub-step S42 of the AIGC-based photographed image generation method according to the embodiments of the present application is shown.

[0018] Figure 8 Flowchart of sub-step S5 of the AIGC-based photographed image generation method according to the embodiment of the present application.

[0019] Figure 9 Flowchart of sub-step S6 of the AIGC-based photographed image generation method according to the embodiment of the present application.

[0020] Figure 10 Block diagram of the AIGC-based photographed image generation system according to the embodiment of the present application. DETAILED DESCRIPTION

[0021] Embodiments of the present disclosure will be described in more detail with reference to the drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood.

[0022] To solve the problems in the background art, the present application provides an AIGC-based photographed image generation method. Figure 1 Flowchart of the AIGC-based photographed image generation method according to the embodiment of the present application. Figure 2 Data flowchart of the AIGC-based photographed image generation method according to the embodiment of the present application. As shown in Figure 1 and Figure 2 The AIGC-based photographed image generation method includes the following steps: S1, obtaining a user image input by a user and a style template image selected by the user; S2, extracting a user posture skeleton and a face feature vector from the user image; S3, performing parameter analysis on the style template image to obtain a clothing description prompt word, a background description prompt word, and a lighting description prompt word; S4, based on the clothing description prompt word and the lighting description prompt word, performing foreground subject generation with lighting prior on the user posture skeleton and the face feature vector to obtain a foreground figure image, a foreground figure mask, and a foreground figure depth map; S5, based on the background description prompt word, performing background fusion based on depth guidance on the foreground figure image, the foreground figure mask, and the foreground figure depth map to obtain a fused image; and S6, performing high-fidelity detail repair on the fused image to obtain an enhanced image.

[0023] In the above-mentioned AIGC-based photographed image generation method, in step S1, a user image input by a user and a style template image selected by the user are obtained. It should be understood that, in order to make the AIGC-driven photographed image generation process have a clear content subject and style guide, and avoid the generated result deviating from the user's expectation due to single image input, the present application obtains the image input by the user carrying core content and the template image selected by the user defining artistic style, to provide the subsequent model with clear content feature source and style feature reference. In this way, it can be ensured that the generated image not only retains the key content structure in the user image, but also accurately matches the style quality expected by the user, providing accurate data support for subsequent model processing, and avoiding the generated result from having content deviation or style inconsistency. Specifically, the user can take a photo or select a user image from local storage, and the system performs resolution correction, denoising and other preprocessing on the user image. At the same time, the user can select a style template image from the style library provided by the system (such as classified by realism, animation, etc.), and the system reads and caches the template image data.

[0024] In the above-mentioned AIGC-based photographed image generation method, in step S2, a user posture skeleton and a face feature vector are extracted from the user image. It should be understood that, since the AIGC-based photographed image generation needs to ensure the restoration degree of the user's posture and the uniqueness of the identity, if the posture and face features are not accurately extracted, the subsequent generation process is prone to problems such as posture deviation from the original action and low recognition of the face from the user. Therefore, the present application further extracts a posture skeleton representing the human action structure and a face feature vector uniquely identifying the user's identity from the user image, to provide a core constraint basis for the subsequent AIGC model generation process. In this way, it can be ensured that the subsequent foreground subject generation strictly follows the user's original posture, avoids limb misplacement, accurately retains the user's unique facial features, and prevents face distortion, laying a foundation for generating an image with accurate posture and identity fidelity.

[0025] In particular, in one specific embodiment, Figure 3 A flowchart of the sub-step S2 of the AIGC-based photographed image generation method according to the embodiment of the present application. As shown in Figure 3 S2, the step S2 includes: S21, inputting the user image into a posture extraction module to obtain a user posture skeleton, the posture extraction module being an Open Pose model; S22, inputting the user image into a face feature extraction module to obtain a face feature vector, the face feature extraction module being an Insight Face model.

[0026] Specifically, the step S21 inputs the user image into a pose extraction module to obtain a user pose skeleton, and the pose extraction module is an Open Pose model. Specifically, the user image is further input into the Open Pose model, which can efficiently locate key joints such as eyes, nose, shoulders and elbows of a human body through a multi-stage convolutional neural network, calculate two-dimensional coordinates and confidence thereof, and finally output structured data containing joint coordinates and skeleton connection relationships, thereby providing accurate pose constraints for subsequent generation.

[0027] Specifically, the step S22 inputs the user image into a face feature extraction module to obtain a face feature vector, and the face feature extraction module is an Insight Face model. It should be understood that, in order to avoid identity confusion caused by insufficient feature dimension or noise interference, the user image is further input into the Insight Face model, which can effectively capture subtle unique features of a face through deep metric learning. Specifically, a face image is first detected and cropped through an algorithm such as MTCNN, and then input into the Insight Face model after alignment and normalization, so as to be converted into a high-dimensional (such as 512-dimensional) face feature vector that can uniquely represent the identity of the user, thereby ensuring the identity consistency of the generated face.

[0028] In the above-mentioned AIGC-based photograph image generation method, the step S3 analyzes the style template image to obtain a clothing description prompt word, a background description prompt word and a light description prompt word. It should be understood that, since the AIGC-based photograph image generation needs to obtain multi-dimensional style information from the style template, directly using the template as a whole may cause mutual interference of clothing, background and light features, so that the image style generated subsequently is inconsistent. Therefore, the style template image is disassembled and analyzed in the present application, and clothing, background and light features are extracted and converted into corresponding clothing description prompt words, background description prompt words and light description prompt words, so as to provide multi-dimensional style guidance for the fusion of the foreground and the background.

[0029] In particular, in one specific embodiment, Figure 4 The flowchart of the sub-step S3 of the AIGC-based photograph image generation method according to the embodiment of the present application is shown in FIG. 3. As shown in FIG. 3, Figure 4 The step S3 includes: S31 decouples the foreground and the background of the style template image to obtain a person mask, a segmented person image and a segmented background image; S32 inputs the segmented person image and the segmented background image into an image-text conversion model respectively to obtain a clothing description prompt word and a background description prompt word; and S33 analyzes and quantizes the global light environment of the style template image to obtain a light description prompt word.

[0030] Specifically, the step S31, the style template image is foreground-background decoupled to obtain a person mask, a segmented person image and a segmented background image. It should be understood that in order to avoid subsequent extraction of garment features being disturbed by the background, extraction of background features containing person information, and reduction of feature accuracy, the present application further decouples the foreground and background of the style template image. For example, the style template image can be input into a general segmentation model (such as Segment Anything Model, SAM), which can automatically identify and accurately outline the main person contour in the image, and output a binary person mask. Based on this mask, the original template image can be separated into a segmented person image containing only the main body of the person and a segmented background image with the background removed (or filled with a neutral color) through pixel-level operation, realizing independent processing of the person and the background. In this way, it can be ensured that there is no background interference when extracting garment features subsequently, and there is no person residue when extracting background features, and at the same time, the person mask can be used for main body protection when the background is fused, improving the processing accuracy of each link and laying a foundation for subsequent feature extraction.

[0031] Specifically, the step S32, the segmented person image and the segmented background image are input into an image-text conversion model to obtain garment description prompt words and background description prompt words. Specifically, to realize accurate textual description of image content, the present application further inputs the segmented image into an image-text conversion model based on an encoder-decoder architecture. The model first extracts the visual content of the input image into a series of high-level feature embeddings through its visual encoder (such as Vision Transformer, ViT), and then its language decoder (such as a structure based on GPT) decodes these visual features as context conditions in a self-recursive manner, thereby generating descriptive natural language text. For the segmented person image, the model is guided to focus on the details of the material, cutting, color and decoration of the garment, and generate garment description prompt words. And for the segmented background image, the model generates background description prompt words describing the scene type, core elements, environmental atmosphere and other information.

[0032] Specifically, the step S33, the global light environment analysis and quantification are performed on the style template image to obtain the light description prompt word. It can be understood that, since light is a core factor affecting the visual authenticity of an image, it is easy to have ambiguous description of the light of the style template by visual observation only, resulting in a large deviation between the generated light effect and the template subsequently. Therefore, the global light environment analysis and quantification are further performed on the style template image, and the light parameters are converted into accurate text prompt words, so as to provide a clear basis for the light simulation in the subsequent foreground generation. In this way, it can be ensured that the light (such as the highlight position, the shadow length, and the color tendency) of the generated foreground figure is completely matched with the global light of the template, avoiding the contradiction of the light direction and the color temperature, and the quantitative analysis can improve the accuracy of the light description, so that the AIGC model accurately reproduces the light atmosphere of the template, and enhances the visual authenticity of the generated image.

[0033] In particular, in one specific embodiment, Figure 5 A flowchart of the sub-step S33 of the AIGC-based photograph image generation method according to the embodiment of the present application is shown. As shown in the figure, Figure 5 The step S33 includes: S331, inputting the style template image into a pre-trained light estimation neural network to obtain a light feature vector; and S332, inputting the light feature vector into a preset mapping function to obtain the light description prompt word.

[0034] More specifically, the step S331, the style template image is input into a pre-trained light estimation neural network to obtain a light feature vector. Specifically, the style template image is further input into a pre-trained light estimation model constructed based on a deep convolutional neural network. The model learns to regress the light information of a three-dimensional scene from a single two-dimensional image by training on a dataset containing a large number of images and their corresponding real-world light parameters. The output light feature vector can be a set of spherical harmonic function coefficients in one possible embodiment. This set of coefficients can compactly and accurately encode the ambient light distribution from various directions in the scene, including the main light source direction, intensity, and complex light effects such as ambient diffuse reflection. In this way, it can be ensured that the extracted light features cover the subtle differences of the light source, avoid the subjective bias of manual analysis, and the high-dimensional feature vector can accurately represent the global distribution of the light, laying a data foundation for the subsequent mapping into text prompt words.

[0035] More specifically, in step S332, the illumination feature vector is input into a preset mapping function to obtain the illumination description prompts. It should be understood that since the illumination feature vector is high-dimensional numerical data, the AIGC model cannot directly recognize it and use it for style guidance; it needs to be converted into structured text prompts to play a guiding role. Therefore, this application further inputs the illumination feature vector into a preset, rule-based mapping function. This mapping function contains a set of quantization-to-text conversion rules to establish a precise correspondence between numerical parameters and text descriptions. Specifically, the mapping function first parses the input illumination feature vector into multiple numerical components, including light source direction, light source intensity, and light source color temperature. Then, a lookup table and matching are performed for each component: the light source direction component is looked up in a preset angle-direction dictionary; the light source intensity component is mapped to the description "bright" or "intense" based on its numerical range (e.g., 0.8-1.0); and the light source color temperature component is mapped to cool tones, warm tones, or specific color words, such as "sunset orange light," based on its numerical value. Finally, the text fragments obtained from the conversion of each component are combined to form a coherent and structured lighting description prompt, thereby realizing the conversion of lighting quantization data into AIGC-recognizable text.

[0036] In the aforementioned AIGC-based image generation method, step S4 involves generating a foreground subject with lighting priors based on clothing description prompts and lighting description prompts, using the user's pose skeleton and facial feature vectors to obtain a foreground subject image, a foreground subject mask, and a foreground subject depth map. It should be understood that in traditional AIGC image generation, the foreground subject's clothing and lighting are often disconnected from the subsequent background, easily leading to pose deviations and low facial recognition. Furthermore, background fusion requires subject boundary and spatial depth information; lacking these can result in fusion distortion. Therefore, this application further uses clothing prompts to define the style and lighting prompts to annotate the lighting priors, combining pose skeletons and facial feature vectors to generate a foreground image and its accompanying mask and depth map, thereby achieving uniformity in the foreground subject's style, pose, identity, and lighting. This ensures that clothing fits the template, lighting matches the target scene, the mask protects the subject, and the depth map aids spatial adaptation, enhancing the realism of the final image.

[0037] In particular, in one specific embodiment, Figure 6 This is a flowchart of sub-step S4 of the AIGC-based image generation method according to an embodiment of this application. Figure 6 As shown, step S4 includes: S41, constructing a first-stage aggregated prompt based on clothing description prompts and lighting description prompts; S42, performing multi-condition guided foreground image synthesis on user pose skeleton and face feature vector based on the first-stage aggregated prompts to obtain a foreground person image; S43, inputting the foreground person image into a person segmentation model and a monocular depth estimation model to obtain a foreground person mask and a foreground person depth map.

[0038] Specifically, the step S41, based on the clothing description prompt word and the light description prompt word, constructs a first-stage aggregated prompt word. It can be understood that since the clothing and light prompt words are independently input into the AIGC model, it is easy to distract attention, resulting in missing clothing details or mismatching light and material, and the first stage needs to avoid complex background interference in foreground generation and waste of model resources. Therefore, the application further integrates the clothing and light prompt words, adds a neutral background constraint to construct a first-stage aggregated prompt word, thereby guiding the model to focus on the generation of the character, clothing, and light in cooperation, and excluding background interference. In this way, the model can take into account the clothing details and light effects, so that the light naturally acts on the clothing, the neutral background avoids resource waste, and the foreground generation efficiency and quality are improved.

[0039] Specifically, the step S42, based on the first-stage aggregated prompt word, performs multi-condition guided foreground image synthesis on the user pose skeleton and the face feature vector to obtain a foreground character image. It can be understood that since the foreground image is generated only by relying on the aggregated prompt word, it is easy to have problems such as pose deviation from the original action or large difference between the face and the user, and the pose skeleton and the face feature vector are core constraints for accurate action and identity preservation, and lack of which will make the generated result deviate from the expectation. Therefore, the application further adds strong constraints of the pose skeleton and the face feature vector to the multi-condition foreground synthesis based on the aggregated prompt word as the style guide, so as to balance the stylization and the preservation of real features. In this way, the foreground character can be ensured to follow the style requirements and restore the unique features of the user's pose and face, avoiding the situation that the style is in place but the identity and pose are distorted.

[0040] In particular, in one specific embodiment, Figure 7 The flowchart of the sub-step S42 of the AIGC-based photographed image generation method according to the embodiment of the application is shown in FIG. 4B. Figure 7 As shown in FIG. 4B, the step S42 includes: S421, initializing a latent diffusion model, the latent diffusion model including a ControlNet module and an IP-Adapter module; S422, inputting the user pose skeleton into the ControlNet module to obtain a pose condition; S423, inputting the face feature vector into the IP-Adapter module to obtain an identity condition; S424, inputting the first-stage aggregated prompt word into a text encoder to obtain a text condition; S425, based on the pose condition, the identity condition, and the text condition, performing diffusion denoising on an initial noise tensor to obtain a pure latent representation of the foreground character; and S426, inputting the pure latent representation of the foreground character into a decoder of a variational autoencoder to obtain the foreground character image.

[0041] More specifically, the step S421 initializes the latent diffusion model, which includes a ControlNet module and an IP-Adapter module. It should be understood that, since the general latent diffusion model can only be guided to generate by a text prompt word, it lacks precise control of user posture and exclusive constraint of face identity, and cannot meet the core demand of preserving posture and identity in AIGC photo generation, which is easy to cause posture misplacement or face celebrity. Therefore, the present application further initializes the latent diffusion model integrated with the ControlNet module and the IP-Adapter module, so as to endow the model with the ability to process multi-dimensional constraint conditions, and take into account the advantages of stylized generation and real feature preservation. In this way, it can be ensured that the model controls the limb structure through ControlNet, locks the identity through IP-Adapter, while maintaining the advantage of stylization, and provides support for multi-condition guided foreground synthesis.

[0042] More specifically, the steps S422, S423 and S424, the user posture skeleton is input into the ControlNet module to obtain the posture condition; the face feature vector is input into the IP-Adapter module to obtain the identity condition; and the first stage aggregated prompt word is input into the text encoder to obtain the text condition. It should be understood that, since the user posture skeleton (joint node data), the face feature vector (high-dimensional numerical value), and the first stage aggregated prompt word (text) are all non-model native recognizable formats, if directly input, the model cannot be parsed, and effective constraint cannot be formed. Therefore, the present application further converts the three types of original data into model processable feature conditions through the corresponding modules, i.e. ControlNet converts posture, IP-Adapter converts identity, and text encoder converts text condition, so as to build a multi-dimensional generation constraint system. In this way, it can be ensured that the posture condition accurately reflects the joint correlation (such as shoulder-elbow-wrist connection), the identity condition uniquely matches the user's facial features (such as eye distance, nose type), and the text condition completely transmits the style information (such as clothing material, light direction), and the three types of conditions work together to provide clear guidance for subsequent diffusion denoising.

[0043] In particular, in one possible embodiment, the steps S422, S423 and S424 are implemented as follows: first, the user pose skeleton is input into the OpenPose branch of the ControlNet, and the module extracts joint spatial correlation features through a convolutional network, outputting a pose condition feature map matching the intermediate feature dimension of the UNet. To ensure the strong constraint of the pose, its control weight can be set to a high value, for example, between 0.8 and 1.0. Then the face feature vector is input into the IP-Adapter module, which converts it into an identity condition vector of the same dimension as the text feature through a feature mapping network, and completes the preliminary fusion with the text feature. To ensure the identity fidelity, the control weight of the IP-Adapter can also be set in a high range, such as 0.7 to 0.9. Finally, the first-stage aggregated prompt word is input into the CLIP ViT-L / 14 text encoder, which extracts text semantic features through an attention mechanism, outputs standardized text condition features, and finally stores three types of conditions for subsequent calling. By configuring the weights of different conditions, the balance between the pose, identity and stylization of the generated results can be flexibly adjusted.

[0044] More specifically, the step S425 diffuses and denoises the initial noise tensor based on the pose condition, the identity condition and the text condition to obtain a pure latent representation of the foreground character. It should be understood that since the initial noise tensor is random Gaussian distribution data and does not contain any foreground character related features, if only a single condition is used for denoising, it is easy to generate a latent representation deviating from the user's demand (such as incorrect pose, inconsistent style). Therefore, the application further uses the three conditions of pose, identity and text to guide the diffusion and denoising process of the initial noise tensor, so as to gradually shape the foreground character features that meet all constraints in the latent space. In this way, it can ensure that the limb deviation is corrected, the facial features are anchored, and the clothing lighting is optimized during denoising, avoiding generation deviation, and finally obtaining a pure latent representation containing complete user features and target style.

[0045] More specifically, the step S426 inputs the foreground figure pure latent representation into the decoder of the variational autoencoder to obtain the foreground figure image. It should be understood that the foreground figure pure latent representation is a feature tensor in a low-dimensional latent space, rather than a pixel image that can be directly visualized. Therefore, the present application utilizes the decoder part of an independently pre-trained variational autoencoder to complete the mapping conversion from the latent space to the pixel space. The VAE itself is independently trained on a large image dataset, and the goal is to learn an efficient image compression and reconstruction function; the encoder compresses the high-resolution image into a compact latent representation, and the decoder can recover the original image from the representation with high quality. In the framework of the present application, the entire diffusion denoising process is carried out in the latent space defined by the VAE, so as to greatly reduce the computational complexity. Therefore, the decoder receives the final pure latent representation generated by the diffusion model as input, performs upsampling through a series of transpose convolutions, and combines residual connections to finely reconstruct image details, and finally outputs a foreground figure image with the same resolution as the original image, clear content, and compliance with all previous constraints.

[0046] Specifically, the step S43 inputs the foreground figure image into the figure segmentation model and the monocular depth estimation model to obtain the foreground figure mask and the foreground figure depth map. It should be understood that since the subsequent background fusion stage needs to accurately distinguish the foreground figure and the background region to avoid the background generation covering the figure subject, and needs to adjust the background light and shadow according to the spatial depth information of the foreground figure (for example, the background elements behind the figure need to conform to the perspective relationship), the foreground figure image itself does not contain these structured information. Therefore, the present application further inputs the foreground figure image into the figure segmentation model and the monocular depth estimation model respectively, extracts the foreground figure mask (defines the boundary) and the foreground figure depth map (defines the space), so as to provide accurate spatial constraints and subject protection basis for the background fusion. In this way, it can be ensured that during the subsequent background fusion, the model only generates content in the background region outside the mask, avoiding the figure subject being tampered. At the same time, the depth map can guide the background light and shadow to be naturally projected to the figure surface (for example, the shadow formed by the background light source on the back of the figure conforms to the depth relationship), solving the discomfort of the figure floating on the background, and improving the scene spatial consistency.

[0047] In particular, in one possible embodiment, the step S43 is implemented as follows: first, the foreground figure image is preprocessed, the resolution is adjusted to 512x512 pixels, and the brightness deviation is corrected to eliminate the interference of light on model recognition. The preprocessed image is input into the pre-trained U-2-Net figure segmentation model, the model locates the figure region through multi-scale feature extraction, and outputs a binary foreground figure mask (the pixel value of the figure region is 255, and the background is 0). At the same time, the same preprocessed image is input into the MiDaSv3 monocular depth estimation model, the model analyzes the three-dimensional structure of the figure in the image (such as the close-up hand and the distant shoulder), and outputs a grayscale foreground figure depth map (the higher the pixel brightness, the closer the distance). Finally, the mask edge is subjected to Gaussian blur processing to eliminate the hard edge, and the depth map is subjected to normalization processing to adapt to the input range of the subsequent fusion model.

[0048] In the above-mentioned AIGC-based photograph image generation method, the step S5 is to perform background fusion based on depth guidance on the foreground figure image, the foreground figure mask, and the foreground figure depth map based on the background description prompt word to obtain a fused image. It can be understood that traditional background fusion is mostly simple superposition or figure-to-figure, which is easy to cause the foreground to be "floating" on the background, resulting in light and shadow conflicts, spatial confusion, and even damage to the posture and identity features of the foreground. Therefore, the scene style is further determined based on the background description prompt word in this application, the foreground image main body, mask boundary protection, and depth map space guidance are combined, and fusion is completed through depth guidance, taking into account the stylization and spatial consistency. In this way, it can be ensured that the background matches the prompt word elements, the correct level is formed according to the depth map, the foreground is protected from being covered by the mask, and finally a fused image with unified light and shadow and space is obtained.

[0049] In particular, in one specific embodiment, Figure 8 The flowchart of the sub-step S5 of the AIGC-based photograph image generation method according to the embodiment of the application is shown in FIG. 5. Figure 8 As shown in FIG. 5, the step S5 includes: S51, inputting the foreground figure image into the encoder of the variational autoencoder to obtain a pure latent representation of the foreground; S52, adding Gaussian noise corresponding to a denoising time step to the pure latent representation of the foreground to obtain a noisy initial latent tensor; S53, inputting the foreground figure mask and the foreground figure depth map into the ControlNet module to obtain a mask control condition and a depth control condition; S54, inputting the background description prompt word into the text encoder to obtain a text semantic condition; S55, based on the text semantic condition, the mask control condition, and the depth control condition, performing multi-control network guided conditional denoising generation on the noisy initial latent tensor to obtain a fused pure latent tensor; and S56, inputting the fused pure latent tensor into the decoder of the variational autoencoder to obtain the fused image.

[0050] Specifically, the step S51 inputs the foreground figure image into the encoder of the variational autoencoder to obtain a pure latent representation of the foreground. It should be understood that the foreground figure image is pixel-level data, which is difficult to directly use for fusion and is incompatible with the latent diffusion model process, and cannot efficiently inject background features. Therefore, the present application further inputs it into the encoder of the variational autoencoder, and obtains a pure latent representation through low-dimensional compression, balances feature preservation and efficiency, and adapts to the diffusion model logic. In this way, the subsequent calculation amount can be greatly reduced, the key features of the foreground are completely preserved, the loss of details is avoided, and the diffusion model is adapted, which lays a foundation for noise injection and background generation.

[0051] Specifically, the step S52 adds Gaussian noise corresponding to a denoising time step to the pure latent representation of the foreground to obtain a noisy initial latent tensor. It should be understood that the present application provides a starting point for the subsequent diffusion denoising process, which contains both the high-level features of the foreground figure and sufficient randomness for the model to inject new background information. Specifically, the initial denoising time step t needs to be determined, which is directly related to a user-controllable denoising strength parameter (the value range is usually 0 to 1). A higher denoising strength (such as 0.75) corresponds to a larger initial time step t, such as t can be set to 750 in a total of 1000 time steps, which means that more Gaussian noise is added to the pure latent representation, giving the model more space to generate a background that is fused with the foreground. Conversely, a lower strength corresponds to a smaller time step and less noise, and more of the original foreground image structure is preserved. According to the selected time step t, a corresponding proportion of Gaussian noise is added to the pure latent representation of the foreground using a noise scheduler, thereby generating a noisy initial latent tensor, ensuring that the core features of the foreground are preserved during denoising, and gradually injecting background information to achieve a smooth fusion of the background generated based on the foreground.

[0052] Specifically, the step S53 inputs the foreground figure mask and the foreground figure depth map into the ControlNet module to obtain a mask control condition and a depth control condition. It should be understood that the foreground mask and the depth map are pixel-level data, which cannot be directly analyzed by the diffusion model, making it difficult for the model to distinguish between foreground and background and to understand spatial structure. Therefore, the present application further inputs the two into the corresponding branches of ControlNet to convert them into control conditions that can be processed by the model, and constructs a spatial constraint and subject protection mechanism. In this way, the mask condition can mark the figure region, and the model can only generate a background outside; the depth condition can pass on the spatial hierarchy to prevent the background from penetrating the figure and protect the spatial logic.

[0053] Specifically, the step S54, the background description prompt word is input into the text encoder to obtain the text semantic condition. It should be understood that the background prompt word is a natural language, and the scene information contained therein is difficult for the diffusion model to directly utilize, and is easy to deviate from the expected background due to ambiguous expression, such as generating an autumn day for “autumn day”. Therefore, the application further inputs the prompt word into the pre-trained text encoder to convert it into a high-dimensional semantic condition, accurately conveying scene details and style. In this way, the core information of the prompt word can be completely preserved, language ambiguity can be avoided to cause deviation, and clear guidance is provided for the model to generate a background that meets the expectations.

[0054] Specifically, the step S55, based on the text semantic condition, the mask control condition and the depth control condition, the noise-bearing initial latent tensor is guided by the multi-control network to generate a conditionally denoised to obtain a fused pure latent tensor. It should be understood that single-condition denoising is easy to deviate, such as covering the foreground only by text, and generating a spatially disordered background only by mask, which is difficult to meet multiple requirements. Therefore, the application further cooperatively guides the denoising by text style, mask foreground preservation, and depth level, and constructs a fused feature in the latent space. In this way, the background can be generated according to the prompt word during denoising, the foreground can be preserved, the correct spatial relationship can be formed according to the depth, and finally a pure latent representation of the fused foreground and background is obtained.

[0055] Specifically, the step S56, the fused pure latent tensor is input into the decoder of the variational autoencoder to obtain the fused image. It should be understood that the fused pure latent tensor is a low-dimensional feature and cannot be directly visualized, and the details of the foreground and background contained therein need to be converted at the pixel level to be perceived. Therefore, the application further inputs it into the VAE decoder to obtain the fused image by feature restoration, realizing the conversion of latent features to visualized images. In this way, the decoder can completely restore the details to avoid blurring and missing, and finally output an image with natural fusion of foreground and background and no synthetic traces.

[0056] In particular, in one possible embodiment, the implementation process of the step S56 is as follows: first, load the pre-trained VAE decoder matched with the encoder, initialize and switch to inference mode; input the fused pure latent tensor; the decoder enlarges the dimension through multiple rounds of upsampling, combines residual convolution to refine details, and performs batch normalization after each round of upsampling to stabilize the features; the activation function maps the pixel value to the visualizable range, converts to the RGB space, and outputs the fused image; after verifying the details and fusion degree, it is confirmed that there are no blur traces, and it is stored in the result database.

[0057] In the aforementioned AIGC-based image generation method, step S6 involves high-fidelity detail restoration of the fused image to obtain an enhanced image. It should be understood that because the background fusion stage model focuses on the spatial and lighting coordination between the foreground and background, it often prioritizes resource allocation to background generation, resulting in blurring and artifacts in foreground facial details (such as eye wrinkles and lip lines) and clothing textures (such as fabric wrinkles), reducing image realism. Therefore, this application further performs high-fidelity detail restoration on the fused image, precisely optimizing flawed areas to restore the core details of the person and the texture of the clothing. This ensures that the enhanced image retains the consistency of foreground and background fusion while possessing clear facial features and clothing textures, avoiding "overall harmony but rough details," significantly improving image visual appeal and user acceptance.

[0058] In particular, in one specific embodiment, Figure 9 This is a flowchart of sub-step S6 of the AIGC-based image generation method according to an embodiment of this application. Figure 9 As shown, step S6 includes: S61, performing facial region detection and extraction on the fused image to obtain facial slices to be repaired and facial bounding boxes; S62, inputting the facial slices to be repaired into a facial repair model based on a generative adversarial network to obtain a repaired facial image; S63, seamlessly fusing the repaired facial image with the fused image based on the facial bounding boxes to obtain the enhanced image.

[0059] Specifically, in step S61, facial region detection and extraction are performed on the fused image to obtain the facial slice to be repaired and the facial bounding box. It should be understood that since the face is the core area for user identification, background fusion easily leads to distortion of facial details, such as blurred eyes, abnormal eyebrow shape, and uneven skin tone. Furthermore, the repair needs to be precisely applied to the face to avoid affecting the background and clothing areas. Therefore, this application further performs facial region detection and extraction on the fused image to obtain the facial slice to be repaired and its corresponding bounding box coordinates, thereby providing accurate region positioning for subsequent targeted repair. This ensures that subsequent repair focuses only on facial blemishes without interfering with the optimized background and clothing areas. Simultaneously, the bounding box coordinates provide a positional basis for the re-attaching of the face after repair, avoiding problems such as facial displacement and proportional imbalance, and ensuring the accuracy of the repair and the overall image harmony.

[0060] In particular, in one possible embodiment, the implementation process of step S61 is as follows: first, pre-process the fused image to adjust the brightness and contrast to improve the distinction between the face and the background; load the MTCNN face detection model to locate the face through a three-layer cascaded network and output the boundary box coordinates containing the key points; crop the face region according to the coordinates and adjust the resolution to adapt to the repair model to obtain the to-be-repaired slice. Record the pixel precision information of the boundary box to ensure that the face can be accurately pasted in the subsequent process, and complete the output and storage of the slice and the boundary box.

[0061] Specifically, in step S62, the to-be-repaired face slice is input into a face repair model based on a generative adversarial network to obtain a repaired face image. It should be understood that, since traditional face repair methods (such as interpolation filling) are prone to cause the repaired region to be blurred and lack real texture, and cannot restore the unique features of the user's face, such as eyebrow arch height and eye shape, the repair model based on the generative adversarial network can generate high-fidelity details through the adversarial training of the generator and the discriminator. Therefore, the application further inputs the to-be-repaired face slice into such a model to restore the fine features and natural texture of the face. In this way, it can ensure that the repaired image eliminates blurring, artifacts and other defects, accurately retains the unique features of the user's face, avoids "looking like someone else" or "becoming a net celebrity" after repair, and guarantees the fidelity of the identity and the sense of reality of the details.

[0062] In particular, in one possible embodiment, the implementation process of step S62 is as follows: first, pre-process the to-be-repaired slice to align the face based on the line connecting the two eyes to eliminate the tilt, and normalize the pixel value to the range suitable for the model; load the GFPGAN repair model, whose U-Net module first removes blurring artifacts, and then injects face features into the StyleGAN2 generator through the channel separation spatial feature transformation layer; the generator outputs the repair result combined with the user's face features, inversely normalizes the slice to obtain the repaired face image, and checks the consistency with the original face features for backup.

[0063] Specifically, in step S63, based on the face boundary box, the repaired face image is seamlessly fused with the fused image to obtain the enhanced image. It should be understood that, since directly replacing the repaired face is prone to cause synthetic marks such as hard transition of the face and hair due to differences in brightness and color, which destroys the overall coordination of the image and affects the visual reality. Therefore, the application further determines the pasting position according to the face boundary box, and uses image fusion technology to achieve seamless connection between the repaired face and the original image, so as to eliminate the synthetic marks and guarantee the overall visual consistency. In this way, it can ensure that the repaired face is completely matched with the background, hair and neck in brightness and color temperature, and the edge transition is natural, without the uncomfortable feeling of "being pasted on", and finally outputs an enhanced image that is overall coordinated and has realistic details.

[0064] In particular, in one possible embodiment, the step S63 is implemented as follows: first, the repaired face is adjusted to the size and angle consistent with the original face according to the boundary box coordinates; a Poisson blending algorithm is used to define a blending area with the original face edge as a reference, and gradient smoothing transition is achieved by solving the Poisson equation; real-time verification of brightness and color temperature matching degree is performed, and the repaired face tone is fine-tuned at the difference; after fusion, the overall image is slightly sharpened, and the enhanced image without synthetic traces is output and stored in the final result library.

[0065] In particular, in another possible preferred embodiment, the step S63 includes: generating a semantic importance map of the repaired face image, the semantic importance map being used to represent the importance degree of different regions of the face to identity fidelity; based on the semantic importance map, a weighted Poisson equation is solved for the gradient field of the repaired face image in the face boundary box region of the fused image, so as to make the non-key region smoothly transition with the fused image while retaining the key features of the identity, thereby obtaining the enhanced image.

[0066] When the repaired face image and the fused image are Poisson blended, the low-frequency information such as illumination and color gradient can be well matched, but the semantic information is not concerned, for example, the Poisson blending is the same as the smooth transition of the cheek and the sharp transition of the eye edge. Here, the AI-generated stylized output image, i.e., the fused image and the high-fidelity image of the user, i.e., the repaired face image, can express the semantic importance of different regions of the face, so a semantic-guided Poisson blending method can be used, that is, by introducing a semantic importance map, the gradient field of the source face image is uniformly forced to be applied to the target region. Here, the semantic importance map assigns higher weights to those pixels that are crucial to maintaining the user's identity (e.g., eyes, corners of the mouth, unique wrinkles), and lower weights to those regions that are smooth and less critical to identity features (e.g., cheeks, forehead). In this way, the fusion process is constructed as a weighted optimization problem, in which the influence of the source image gradient will be proportional to its semantic importance.

[0067] Specifically, the AIGC output image has the desired target style and illumination consistency, but its face structure is the AI's stylized interpretation of the user's face, and although the IP-Adapter module can roughly retain the identity, those fine, high-frequency details that define the individual's uniqueness may be softened, stylized, or slightly changed, that is, its semantic distribution is coherent in style, but may lack in identity fidelity. While the repaired face image has high identity fidelity, it contains the accurate and true geometric structure of the user's unique facial features, but its local texture and micro-illumination may not perfectly match the stylized synthetic environment in the fused image.

[0068] Thus, when all gradients from the restored facial image are imposed onto the fused image, conflicts arise. In smooth areas like the cheeks, the subtle textures of the original photo may clash with the painting / rendering style of the AI ​​image, even if the fusion itself is seamless. Furthermore, to minimize overall energy, the solver may unconsciously average out very sharp, identity-defining gradients, resulting in a loss of sharpness in the final image. Introducing a semantic graph guides the solver to be strict in critical areas and lenient in non-critical areas. For example, it aims to actively preserve gradients that define the structure of the eyes, lips, and nose, while allowing gradients in skin areas to be more relaxed to better adapt to the stylized textures of the surrounding fused image.

[0069] For standard Poisson fusion, let This refers to the target region in the target image, i.e., the fused image. Inside, we need to solve for the unknown pixel value function. Let... The source image represents the reconstructed facial image. represent For the target image outside the region, i.e., the fused image, the goal of standard Poisson fusion is to find the function... , so that its gradient With a guiding vector field (i.e., source image) gradient field Minimize the difference: ;in, For the minimization operator of the fusion function, For gradient operators, For the target integration area, For the boundary of the target area, Let be the function for fusing pixel values ​​to be determined. A function for known pixel values ​​outside the boundary. This is a function for the pixel values ​​of the source image.

[0070] This minimization problem is equivalent to finding the solution domain under given boundary conditions. Poisson's equation on: ;in, For divergence operators, For the Laplace operator.

[0071] Specifically, a semantic importance map of the repaired facial image is generated, which characterizes the importance of different facial regions to identity fidelity. The semantic importance map is generated in the process of generating the semantic importance map. hour, (If a stronger retention effect is needed, it can be greater than 1). The larger the value, the higher the semantic importance of that point. This map can be constructed by combining multiple pieces of information related to facial identity. First, facial landmark analysis is performed, and a heatmap is created using a facial landmark detector (such as dlib's 68-point model) on a reference facial slice or a reconstructed facial image. Its intensity is highest at key points (eyes, eyebrows, bridge of the nose, nostrils, and lip contours), and diffuses outward in a Gaussian decay manner. Then, based on identity information, it is also conveyed through high-frequency details such as wrinkles, pores, and hair, and high-frequency detail analysis is performed. For example, this information is captured by calculating the amplitude of the Laplacian operator response of the source face image. ;in, Source image Second-order partial derivatives in the direction, Source image Second-order partial derivatives in the direction, The function is the pixel value of the source image. This is a high-frequency detail map of the face.

[0072] Then, Normalized to the [0,1] interval, and combined with the graph, the final semantic importance graph is obtained by weighted combination of the keypoint graph and the high-frequency detail graph: ;in, A heatmap of key facial features. This is a hyperparameter (e.g., set to 0.7) used to balance the importance between structural keypoints and fine textures. This is a semantic importance map. This ensures that during subsequent fusion, the repaired features of core areas (such as the eyes and lip line) are prioritized for preservation, while non-critical areas (such as the cheeks) can flexibly adapt to the lighting and skin tone of the fused image. This avoids identity distortion and allows for adjustments to ensure a smooth transition, improving the overall naturalness of the fusion. For example, when a user selects an astronaut-style template, it ensures that the repaired eye shape, brow bone, and other core features are highly consistent with the user, avoiding the distortion of a "generic astronaut face" and meeting the core requirement of AIGC's photography equipment to generate personalized style photos.

[0073] Then, based on the semantic importance map, the gradient field of the restored facial image is weighted to solve a weighted Poisson equation within the facial bounding box region of the fused image. This preserves key identity features while ensuring a smooth transition between non-key regions and the fused image, resulting in the enhanced image. Therefore, the semantic importance map is incorporated into the objective function of the Poisson fusion. Inserted as a weighting factor into the integral term, the new objective function becomes: Thus, in Output gradient where the value is very high (e.g., at the corner of the eye). With source gradient Any deviation between them will be severely penalized, forcing the solver to almost exactly replicate the source gradient; while Where the value is very low (e.g., on a smooth cheek), the penalty is small. The solver has greater freedom to find a value that might deviate from the expected value. gradient This allows for better matching of boundary conditions and a smoother transition with the texture of the surrounding blended image. Its gradient field can flexibly adapt to the lighting and skin tone of the blended image, ultimately achieving enhancement without distortion of key features or artifacts in non-key areas, avoiding both a synthetic look and identity distortion. For example, after blending against a cyberpunk-style background, key features such as the user's lips and eye wrinkles in a heart-shaped gesture are fully preserved, while the cheek area naturally adapts to the neon cool-toned lighting of the background, with no obvious stitching marks, meeting the practical requirements of AIGC's camera-generated high-quality artistic photos.

[0074] As a result, key facial features from the user's original photo can be transferred with higher fidelity, effectively preventing the "AI face" phenomenon that loses subtle uniqueness. Meanwhile, in secondary areas, the fusion process is more like a gentle diffusion, allowing AI-generated styles and textures to permeate the face, making the final result look like a holistic rendered image rather than a cut-and-paste product. Furthermore, by reducing the weight of gradient transfer in smooth areas, the risk of introducing unnatural-looking textures that clash with the target style is reduced, thus minimizing artifacts.

[0075] In summary, the AIGC-based image generation method based on the embodiments of this application is explained. It performs parallel deep analysis on user images and style template images, simultaneously extracting high-fidelity constraints such as user identity and pose, as well as multi-dimensional stylized guidance information including clothing, background, and lighting. Furthermore, prior lighting information is used as a pre-constraint, injecting the lighting environment of the target scene during the foreground subject generation stage, and simultaneously constructing the spatial depth information of the foreground figure. Finally, this depth information is used to provide precise spatial constraints and guidance for the background fusion process, and combined with a high-fidelity inpainting network for detail enhancement, thereby generating a realistic image where the foreground and background are highly consistent in lighting and space, and user identity features are preserved. This enables adaptive control of lighting and spatial relationships during the generation process, effectively improving the fusion quality and visual realism of the final image.

[0076] Figure 10 This is a block diagram of an AIGC-based image generation system according to an embodiment of this application. Figure 10As shown, the AIGC-based photographed image generation system 100 according to the embodiment of the present application comprises: a user input collection module 110, configured to acquire a user image input by a user and a style template image selected by the user; a user feature extraction module 120, configured to extract a user posture skeleton and a face feature vector from the user image; a parameter analysis module 130, configured to perform parameter analysis on the style template image to obtain a clothing description prompt word, a background description prompt word and a lighting description prompt word; a foreground subject generation module 140, configured to perform foreground subject generation with lighting priori based on the clothing description prompt word and the lighting description prompt word on the user posture skeleton and the face feature vector to obtain a foreground figure image, a foreground figure mask and a foreground figure depth map; a background fusion module 150, configured to perform depth-guided background fusion based on the background description prompt word on the foreground figure image, the foreground figure mask and the foreground figure depth map to obtain a fused image; and a high-fidelity detail repair module 160, configured to perform high-fidelity detail repair on the fused image to obtain an enhanced image.

[0077] Here, those skilled in the art can understand that the specific operations of each step in the above AIGC-based photographed image generation system have been described in detail above with reference to the description of the AIGC-based photographed image generation method of Figures 1 to 9 Therefore, the repeated description thereof will be omitted.

[0078] As described above, the AIGC-based photographed image generation system 100 according to the embodiment of the present application can be implemented in various wireless terminals, such as a server with a gas composition adaptive thermal fluid generation control algorithm, etc. In one possible implementation, the AIGC-based photographed image generation system 100 according to the embodiment of the present application can be integrated into a wireless terminal as a software module and / or a hardware module. For example, the AIGC-based photographed image generation system 100 can be a software module in the operating system of the wireless terminal, or can be an application program developed for the wireless terminal; of course, the AIGC-based photographed image generation system 100 can also be one of the many hardware modules of the wireless terminal.

[0079] The embodiment of the present application also provides a photographed device, which can perform the AIGC-based photographed image generation method as described above.

Claims

1. A method for generating photographic images based on AIGC, characterized in that, include: Retrieve the user image input by the user and the style template image selected by the user; Extract user pose skeleton and facial feature vector from user images; The style template image is parsed to obtain clothing description prompts, background description prompts, and lighting description prompts; based on the clothing description prompts and lighting description prompts, a foreground subject is generated with lighting prior on the user's pose skeleton and facial feature vector to obtain a foreground person image, a foreground person mask, and a foreground person depth map; Based on background description prompts, depth-guided background fusion is performed on the foreground person image, foreground person mask, and foreground person depth map to obtain the fused image. High-fidelity detail restoration is performed on the fused image to obtain the enhanced image.

2. The method for generating images based on AIGC according to claim 1, characterized in that, Extracting user pose skeleton and facial feature vector from user image includes: inputting user image into pose extraction module to obtain user pose skeleton, wherein pose extraction module is Open Pose model; and inputting user image into facial feature extraction module to obtain facial feature vector, wherein facial feature extraction module is Insight Face model.

3. The method for generating images based on AIGC according to claim 1, characterized in that, The style template image is subjected to parameter parsing to obtain clothing description prompts, background description prompts, and lighting description prompts. This includes: decoupling the foreground and background of the style template image to obtain a character mask, a segmented character image, and a segmented background image; inputting the segmented character image and the segmented background image into an image-to-text conversion model to obtain clothing description prompts and background description prompts; and performing global lighting environment analysis and quantization on the style template image to obtain lighting description prompts.

4. The AIGC-based image generation method according to claim 3, characterized in that, The process of performing global illumination environment analysis and quantification on a style template image to obtain illumination description prompts includes: inputting the style template image into a pre-trained illumination estimation neural network to obtain an illumination feature vector; and inputting the illumination feature vector into a preset mapping function to obtain the illumination description prompts.

5. The method for generating images based on AIGC according to claim 1, characterized in that, Based on clothing description prompts and lighting description prompts, a foreground subject generation method with lighting prior is performed on the user's pose skeleton and facial feature vector to obtain a foreground person image, a foreground person mask, and a foreground person depth map. This includes: constructing a first-stage aggregated prompt based on clothing description prompts and lighting description prompts; performing multi-condition guided foreground image synthesis on the user's pose skeleton and facial feature vector based on the first-stage aggregated prompts to obtain a foreground person image; and inputting the foreground person image into a person segmentation model and a monocular depth estimation model to obtain a foreground person mask and a foreground person depth map.

6. The AIGC-based image generation method according to claim 5, characterized in that, Based on the first-stage aggregated prompts, a foreground image is synthesized using multi-condition guidance based on the user pose skeleton and facial feature vector to obtain a foreground person image. This includes: initializing a latent diffusion model, which includes a ControlNet module and an IP-Adapter module; inputting the user pose skeleton into the ControlNet module to obtain pose conditions; inputting the facial feature vector into the IP-Adapter module to obtain identity conditions; inputting the first-stage aggregated prompts into a text encoder to obtain text conditions; performing diffusion denoising on the initial noise tensor based on the pose conditions, identity conditions, and text conditions to obtain a clean latent representation of the foreground person; and inputting the clean latent representation of the foreground person into the decoder of a variational autoencoder to obtain the foreground person image.

7. The method for generating images based on AIGC according to claim 1, characterized in that, Based on background description cues, a depth-guided background fusion is performed on the foreground person image, foreground person mask, and foreground person depth map to obtain a fused image. This includes: inputting the foreground person image into the encoder of a variational autoencoder to obtain a clean latent representation of the foreground; adding Gaussian noise corresponding to the denoising time step to the clean latent representation of the foreground to obtain a noisy initial latent tensor; inputting the foreground person mask and foreground person depth map into a ControlNet module to obtain mask control conditions and depth control conditions; inputting the background description cues into a text encoder to obtain text semantic conditions; based on the text semantic conditions, mask control conditions, and depth control conditions, performing conditional denoising on the noisy initial latent tensor guided by a multi-control network to obtain the fused clean latent tensor; and inputting the fused clean latent tensor into the decoder of the variational autoencoder to obtain the fused image.

8. The method for generating images based on AIGC according to claim 1, characterized in that, The process of performing high-fidelity detail restoration on the fused image to obtain an enhanced image includes: performing facial region detection and extraction on the fused image to obtain facial slices and facial bounding boxes to be restored; inputting the facial slices to be restored into a facial restoration model based on a generative adversarial network to obtain a restored facial image; and seamlessly fusing the restored facial image with the fused image based on the facial bounding boxes to obtain the enhanced image.

9. An AIGC-based image generation system, characterized in that, include: The user input acquisition module is used to acquire user images input by the user and style template images selected by the user. The user feature extraction module is used to extract user pose skeletons and facial feature vectors from user images; The parameter parsing module is used to parse the style template image to obtain clothing description prompts, background description prompts, and lighting description prompts. The foreground subject generation module is used to generate a foreground subject with lighting prior based on clothing description prompts and lighting description prompts, using the user's pose skeleton and facial feature vector to obtain a foreground person image, a foreground person mask, and a foreground person depth map. The background fusion module is used to perform depth-guided background fusion of the foreground person image, foreground person mask, and foreground person depth map based on background description prompts to obtain the fused image. The high-fidelity detail restoration module is used to perform high-fidelity detail restoration on the fused image to obtain an enhanced image.

10. A photographic device, characterized in that, include: The photographing device is capable of executing the AIGC-based image generation method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Automatic face image illumination editing method under complex background

    CN104463181A

  • Figure image reloading method and device, storage medium and computer equipment

    CN117670656A

  • Image condition redrawing method with illumination perception and illumination reality sense

    CN118429530A

  • Lora model training-based commodity graph-to-background method, apparatus and device, and medium

    CN120070663A

  • Methods and Systems for Automatically Generating Backdrop Imagery for a Graphical User Interface

    US20230064723A1

Cited By

  • Character and background fusion method and system based on adaptive skin color protection

    CN121213428A

  • On-site personalized invitation letter generation method and system based on artificial intelligence

    CN121725099A