Scene image generation method and device, electronic equipment and readable storage medium

Through the text encoder and image feature extraction module combined with the character face feature fusion module and diffusion model, the problem of difficulty in accurately shaping the personalized characteristics of the characters in the existing technology is solved, and scene images with consistent style, coordinated background, and clear character identities are generated.

CN120339762APending Publication Date: 2025-07-18SHENZHEN XUMI YUNTU SPACE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510182672.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing scene image generation methods are difficult to accurately shape the personalized characteristics of the characters. Especially in complex scenes where multiple images coexist, it is difficult to accurately distinguish the human details, clothing and background information of each character, resulting in the generated image background inconsistent and the character movements are incoherent.

Method used

By obtaining text description information and reference images, the text encoder and image feature extraction module extract features, combined with the character face feature fusion module and diffusion model, the target scene character images are generated, so as to achieve coordinated and coherent characters and backgrounds, and accurately distinguish between character identities and clothing.

Benefits of technology

The generated target scene character images maintain the consistency of the picture style, the background information is coordinated and coherent with the character's movements, accurately distinguish the character's identity and clothing, improve the accuracy of the character's facial expressions, and generate higher quality personalized scene images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339762A_ABST
    Figure CN120339762A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image generation, and provides a scene image generation method and device, electronic equipment and a readable storage medium. The method comprises the following steps: processing text description information through a text encoder to obtain a first text feature; processing the reference image through an image feature extraction module to obtain a global image feature, an image background feature, a skeleton image, a figure semantic image fusion feature and a face semantic image fusion feature; performing cross fusion on the first text feature, the character semantic image fusion feature and the face semantic image fusion feature through a character face feature fusion module to obtain a character face fusion feature; and through a diffusion model, processing the features after the preset noise graph and the skeleton graph are fused, the global image features, the image background features, the first text features and the human face fusion features, and generating a target scene human image of the object to be generated. The technical problem that it is difficult to precisely shape the personalized features of the figure in the related technology is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image generation technology, and in particular, to a method, apparatus, electronic device, and readable storage medium for generating a scene image. Background Art

[0002] With the development of artificial intelligence technology, users are increasingly inclined to use application programs to generate personalized images. For example, when a user inputs a reference picture and a text instruction, the application program will generate a dynamic image according to the text instruction and the reference picture.

[0003] However, in complex scenarios involving the coexistence of multiple people's images, such as scenarios with occlusion, where there is overlap between the body parts of the people in the picture, the related technology at this time cannot accurately distinguish information such as the body details, clothing, and background of each person, and it is difficult to accurately extract the personalized features of the whole body of the people. Therefore, problems such as uncoordinated backgrounds and discontinuous actions of the people in the generated images will occur.

[0004] Therefore, the existing methods for generating scene images have the technical problem of being difficult to accurately shape the personalized features of the whole body of the people. Summary of the Invention

[0005] In view of this, the embodiments of the present application provide a method, apparatus, electronic device, and readable storage medium for generating a scene image to solve the problem that the existing methods for generating scene images are difficult to accurately shape the personalized features of the whole body of the people.

[0006] In the first aspect of the embodiments of the present application, a method for generating a scene image is provided, including:

[0007] Obtain the text description information of the object to be generated and the reference image of the object to be generated;

[0008] Process the text description information through a text encoder to obtain the first text feature of the text description information;

[0009] Process the reference image through an image feature extraction module to obtain the target image feature of the reference image, where the target image feature includes a global image feature, an image background feature, a skeleton graph, a person semantic image fusion feature, and a face semantic image fusion feature; where the person semantic image fusion feature is the feature after the fusion of the person identity feature and the person semantic feature in the reference image, and the face semantic image fusion feature is the feature after the fusion of the face identity feature and the face semantic feature in the reference image;

[0010] Cross-fuse the first text feature, the person semantic image fusion feature, and the face semantic image fusion feature through a person face feature fusion module to obtain a person face fusion feature;

[0011] Through a diffusion model, the features after fusing a preset noise map and a skeleton map, global image features, image background features, first text features, and human face fusion features are processed to generate a target scene human image of the object to be generated.

[0012] In a second aspect of the embodiments of the present application, there is provided a device for generating a scene image, including:

[0013] An acquisition module, configured to acquire text description information of the object to be generated and a reference image of the object to be generated;

[0014] A first processing module, configured to process the text description information through a text encoder to obtain first text features of the text description information;

[0015] A second processing module, configured to process the reference image through an image feature extraction module to obtain target image features of the reference image, where the target image features include global image features, image background features, a skeleton map, human semantic image fusion features, and face semantic image fusion features; where the human semantic image fusion features are features after fusing human identity features and human semantic features in the reference image, and the face semantic image fusion features are features after fusing face identity features and face semantic features in the reference image;

[0016] A cross-fusion module, configured to perform cross-fusion on the first text features, human semantic image fusion features, and face semantic image fusion features through a human face feature fusion module to obtain human face fusion features;

[0017] A generation module, configured to process the features after fusing a preset noise map and a skeleton map, global image features, image background features, first text features, and human face fusion features through a diffusion model to generate a target scene human image of the object to be generated.

[0018] In a third aspect of the embodiments of the present application, there is provided an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, where when the processor executes the computer program, the steps of the above method are implemented.

[0019] In a fourth aspect of the embodiments of the present application, there is provided a readable storage medium, where the readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0020] The beneficial effects of the embodiments of the present application compared with the prior art are:

[0021] By obtaining the text description information of the object to be generated and the reference image of the object to be generated, the scene materials, character materials and user requirements of the object to be generated are obtained; by processing the text description information through a text encoder, the first text feature of the text description information is obtained, providing data support for subsequent feature fusion, thereby improving the accuracy of the fused feature and generating a target scene character image that meets the user's expectations; by processing the reference image through an image feature extraction module, the target image feature of the reference image is obtained, providing a data basis for the target scene character image, enabling the target scene character image to maintain the consistency of the picture style and the coordination and coherence between the background information and the character actions, avoiding the problem that the overall style of the images before and after the character actions in the target scene character image varies too much, and being able to accurately distinguish the identities, clothing of the characters in a multi-person scene and generate a more accurate facial expression of the characters; by cross-fusing the first text feature, the character semantic image fusion feature and the face semantic image fusion feature through a character face feature fusion module, the character face fusion feature is obtained, which can learn the correlation relationship between the first text feature, the character semantic image fusion feature and the face semantic image fusion feature, thereby specifically focusing on the user's requirements for the characters or faces in the target scene character image, effectively fusing the text feature with the character feature and the face feature, and increasing the accuracy of the character face fusion feature; by a diffusion model, processing the features after fusing the preset noise map and the skeleton map, the global image feature, the image background feature, the first text feature and the character face fusion feature to generate the target scene character image of the object to be generated, which can perform multiple cross-fusions on the global image feature, the image background feature, the first text feature and the character face fusion feature, learn the complex correlations and dependencies between the features, and thus obtain a more accurate fused feature. It can solve the technical problem that the existing methods for generating scene images are difficult to accurately shape the personalized features of the whole body of the characters. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0023] Figure 1 is a schematic flowchart of a method for generating a scene image provided by an embodiment of the present application;

[0024] Figure 2 is a schematic block diagram of a module structure of a method for generating a scene image provided by an embodiment of the present application;

[0025] Figure 3It is a schematic diagram of the module structure of another method for generating a scene image provided by an embodiment of the present application;

[0026] Figure 4 It is a schematic diagram of the module structure of yet another method for generating a scene image provided by an embodiment of the present application;

[0027] Figure 5 It is a schematic diagram of the structure of a device for generating a scene image provided by an embodiment of the present application;

[0028] Figure 6 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0029] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0030] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same type and do not limit the number of objects. For example, the first object can be one or multiple.

[0031] In addition, it should be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusively, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, the elements defined by the statement "including..." do not exclude the existence of additional identical elements in the process, method, article or device including the elements.

[0032] Next, a method and a device for generating a scene image according to an embodiment of the present application will be described in detail with reference to the accompanying drawings.

[0033] Figure 1 It is a schematic flowchart of a method for generating a scene image provided by an embodiment of the present application. Figure 1 The method for generating a scene image can be executed by a terminal device. As Figure 1As shown, the method for generating the scene image includes:

[0034] S101, obtaining the text description information of the object to be generated and the reference image of the object to be generated.

[0035] Specifically, the text description information is used to express the user's needs, so that the actions of the characters in the generated target scene character image can meet the user's expectations. For example: "A family is making dumplings together"; the reference image can be a static picture, such as a photo or drawing of "making dumplings with the family", providing scene materials and character materials for the object to be generated.

[0036] S102, processing the text description information through a text encoder to obtain the first text feature of the text description information.

[0037] Specifically, the text encoder can select a Contrastive Language-Image Pre-training (CLIP) text encoder. The text encoder converts the input text description information into a semantic vector, that is, the first text feature corresponding to the text description information, so as to facilitate subsequent fusion with the image feature, thereby improving the accuracy of the fusion feature and generating a target scene character image that meets the user's expectations.

[0038] S103, processing the reference image through an image feature extraction module to obtain the target image feature of the reference image.

[0039] The target image feature includes a global image feature, an image background feature, a skeleton diagram, a person semantic image fusion feature, and a face semantic image fusion feature; among them, the person semantic image fusion feature is the feature after the fusion of the person identity feature and the person semantic feature in the reference image, and the face semantic image fusion feature is the feature after the fusion of the face identity feature and the face semantic feature in the reference image.

[0040] Specifically, the global image feature can represent the overall style of the reference image, providing a data basis for the generation of the target scene character image, so as to maintain the consistency of the picture style and avoid the problem that the overall style of the images before and after the actions of the characters in the target scene character image varies too much.

[0041] Specifically, the image background feature can represent the background style of the reference image, providing a data basis for the target scene character image, ensuring the coordination, consistency, accuracy, and coherence between the background information and the character actions in the target scene character image.

[0042] Specifically, the skeleton diagram includes the skeleton information of each person in the reference image, which can provide the pose and structure information of the person for the generation of the target scene character image, so as to maintain the consistency of the appearance of the characters in the generated image.

[0043] Specifically, the person identity feature and the person semantic feature are obtained by extracting the features of the person image in the reference image, and the person image is obtained by segmenting the reference image. After the person semantic image fusion feature fuses the person identity feature and the person semantic feature, it can more accurately describe the details of the body, clothing, etc. of the person in the reference image, provide data support for distinguishing the person identity and clothing in a multi-person scenario, and reduce the interference of factors such as physical contact between people or the occlusion of the person's body by an occluder on the person image in the target scenario.

[0044] Specifically, the face identity feature and the face semantic feature are obtained by extracting the features of the face image in the reference image, and the face image is obtained by segmenting the person image. After the face semantic image fusion feature fuses the face identity feature and the face semantic feature, it can more accurately describe the details of the facial features, expressions, accessories, etc. of the face in the reference image, and provide data support for generating the facial expressions of the person in a multi-person scenario.

[0045] The reference image is processed by the image feature extraction module to obtain the target image feature of the reference image, which provides a data basis for the person image in the target scenario, enables the person image in the target scenario to maintain the consistency of the picture style, the coordination and coherence between the background information and the person's actions, avoids the problem of too large a difference in the overall style of the images before and after the person's actions in the person image in the target scenario, accurately distinguishes the person identity and clothing in a multi-person scenario, and generates a more accurate facial expression of the person.

[0046] S104. The person face feature fusion module cross-fuses the first text feature, the person semantic image fusion feature, and the face semantic image fusion feature to obtain the person face fusion feature.

[0047] Specifically, the person face feature fusion module can perform the cross-fusion of the first text feature, the person semantic image fusion feature, and the face semantic image fusion feature twice, and the specific fusion order is not specifically limited. For example, the first text feature and the person semantic image fusion feature can be first cross-fused, and then the fused feature and the face semantic image fusion feature can be second cross-fused; it can also be that the first text feature and the face semantic image fusion feature are first cross-fused, and then the fused feature and the person semantic image fusion feature are second cross-fused.

[0048] Through the method of cross-fusion, the person face feature fusion module can learn the correlation relationship between the first text feature, the person semantic image fusion feature, and the face semantic image fusion feature, so as to specifically focus on the user's needs for the person or face in the person image in the target scenario, effectively fuse the text feature with the person feature and the face feature, and increase the accuracy of the person face fusion feature.

[0049] S105. Process the features after fusing the preset noise map and the skeleton map, the global image features, the image background features, the first text features, and the fused features of the human face through a diffusion model to generate the target scene character image of the object to be generated.

[0050] Specifically, the diffusion model adopts the Stable Diffusion (SD) model, and its structure adopts the U-shaped Network (UNet) structure. The UNet structure is composed of multiple identical sub-modules. In addition, the diffusion model has also been specifically improved on the basis of the SD model by adding a cross-attention module in the sub-module to facilitate multiple cross-fusions of the global image features, the image background features, the first text features, and the fused features of the human face, learn the complex associations and dependencies between the features, so as to obtain more accurate fused features and generate the target scene character image.

[0051] Specifically, fuse the skeleton map and the preset noise map to obtain a feature map, that is, the features after fusing the preset noise map and the skeleton map, and input this feature map into the diffusion model as the overall input of the diffusion model.

[0052] In addition, to optimize the generated scene character image, a combined loss function is also used during the generation process to ensure that the target scene character image is consistent with the reference image and the text description information. The combined loss function can be a combination of two or more loss functions. For example, a combination of a reconstruction loss function and a text consistency loss function is adopted. Among them, the reconstruction loss function can use the mean-square error (MSE) to calculate the difference between the generated scene character image and the reference image, so as to improve the appearance similarity between the generated scene character image and the reference image. The text consistency loss function can use the Contrastive Loss (CL) to calculate the consistency between the generated scene character image and the text description information, so as to improve the semantic similarity between the generated scene character image and the text description information.

[0053] According to the technical solution provided by the embodiments of the present application, by obtaining the text description information of the object to be generated and the reference image of the object to be generated, the scene materials, character materials and user requirements of the object to be generated are obtained; by processing the text description information through a text encoder, the first text feature of the text description information is obtained, providing data support for subsequent feature fusion, thereby improving the accuracy of the fusion feature and generating a target scene character image that meets user expectations; by processing the reference image through an image feature extraction module, the target image feature of the reference image is obtained, providing a data basis for the target scene character image, enabling the target scene character image to maintain the consistency of the picture style, as well as the coordination and coherence between the background information and the character actions, avoiding the problem that the overall style of the images before and after the character actions in the target scene character image varies too much, and being able to accurately distinguish the identities, clothes of the characters in a multi-person scene and generate a more accurate character facial expression; by cross-fusing the first text feature, the character semantic image fusion feature, and the face semantic image fusion feature through a character face feature fusion module, the character face fusion feature is obtained, which can learn the correlation relationship between the first text feature, the character semantic image fusion feature, and the face semantic image fusion feature, thereby specifically focusing on the user's requirements for the characters or faces in the target scene character image, effectively fusing the text feature with the character feature and the face feature, and increasing the accuracy of the character face fusion feature; by a diffusion model, processing the features after fusing the preset noise map and the skeleton map, the global image feature, the image background feature, the first text feature, and the character face fusion feature to generate the target scene character image of the object to be generated, which can perform multiple cross-fusions on the global image feature, the image background feature, the first text feature, and the character face fusion feature, learn the complex correlations and dependencies between the features, and thus obtain a more accurate fusion feature. It can solve the technical problem that the existing methods for generating scene images are difficult to accurately shape the personalized features of the whole body of the character.

[0054] In some embodiments, the image feature extraction module includes a first image encoder, a second image encoder, a skeleton extraction module, a segmentation module, a face semantic image fusion module, and a character semantic image fusion module;

[0055] Processing the reference image through the image feature extraction module to obtain the target image feature of the reference image includes:

[0056] Processing the reference image through the first image encoder to obtain an initial feature, and performing a dimensionality transformation on the initial feature through a linear layer to obtain the global image feature of the reference image;

[0057] Processing the reference image through the skeleton extraction module to obtain the skeleton map of the reference image;

[0058] The reference image is processed by a segmentation module to obtain the background image, the person image, and the face image of the reference image;

[0059] The background image is processed by a second image encoder to obtain the image background features of the background image;

[0060] The face identity features and face semantic features of the face image are fused by a face semantic image fusion module to obtain the face semantic image fusion features of the face image;

[0061] The person identity features and person semantic features of the person image are fused by a person semantic image fusion module to obtain the person semantic image fusion features of the person image.

[0062] Specifically, as Figure 2 shown, the reference image is respectively input into a skeleton extraction module, a first image encoder, and a first segmentation module for processing.

[0063] Specifically, the first image encoder can be a Contrastive Language-Image Pre-training (CLIP) image encoder. The first image encoder is used to capture the global visual features of the reference image, output the initial features of the reference image, and provide a data basis for subsequent feature fusion processing. The initial features are input into a linear layer module, and the linear layer performs operations such as dimension transformation on the initial features and outputs the global image features of the reference image, so as to facilitate the fusion of the global image features with the subsequent text image features. The global image features can reflect the overall information of the reference image, including color, texture, and shape, etc., so as to maintain the consistency of the overall content and structure between the generated image and the reference image.

[0064] Specifically, the skeleton extraction module can use an OpenPose skeleton extraction model. The skeleton extraction module is used to extract the human skeleton information of the person in the reference picture, provide accurate data support for the pose structure of the person in the generated image, and facilitate maintaining the coordination and consistency of the appearance of the person in the target scene person image.

[0065] Specifically, the segmentation module includes a first segmentation module and a second segmentation module, and both the first segmentation module and the second segmentation module can use the Segment Anything Model (SAM). The first segmentation module performs segmentation based on the semantic information of the reference image, outputs a background image and a person image. By segmenting the background and the person, the background and the person can be processed separately, and the image background features and person features can be extracted more accurately, thereby controlling the details of the generated image. The second segmentation module performs segmentation based on the semantic information of the person image and outputs a face image. Separately processing the face and the person can not only reduce unnecessary computational burdens, make more efficient use of computing resources, and speed up the generation speed, but also be able to extract the facial contour features and the person's body clothing features more accurately.

[0066] Specifically, the second image encoder can use the Vision Transformer (VIT). The VIT image encoder can capture the global dependencies of the background image, thereby accurately extracting the image background features of the background image and providing data support for generating a coherent and consistent scene person image.

[0067] Specifically, as Figure 3 shown, the face semantic image fusion module includes a face feature extraction module, a third image encoder, a first resampling module, a second resampling module, a face feature fusion module, a first self-attention module, and a first multi-layer perceptron module. Among them, the face feature extraction module is used to extract face identity features, and the face identity features can be the facial features of the face for identity recognition. The third image encoder is used to extract face semantic features, and the face semantic features can be the information conveyed by facial expressions for the recognition and understanding of each semantic region in the face image, such as eyes, mouth, hair, etc. Fusing the face identity features and the face semantic features can obtain more comprehensive face information and improve the accuracy of the face semantic image fusion features.

[0068] Specifically, as Figure 4 shown, the person semantic image fusion module includes a person feature extraction module, a fourth image encoder, a third resampling module, a fourth resampling module, a person feature fusion module, a second self-attention module, and a second multi-layer perceptron module. Among them, the person feature extraction module is used to extract person identity features for identity recognition. The fourth image encoder is used to extract person semantic features for the recognition and understanding of each semantic region in the person image, such as limbs, torso, clothing, etc. Fusing the person identity features and the person semantic features can obtain more comprehensive person information and improve the accuracy of the person semantic image fusion features.

[0069] According to the technical solution provided by the embodiments of the present application, the reference image is processed by the first image encoder to obtain initial features, and the initial features are dimensionally transformed by a linear layer to obtain the global image features of the reference image, so as to maintain the consistency of the overall content and structure between the generated image and the reference image; the reference image is processed by the skeleton extraction module to obtain the skeleton map of the reference image, so as to maintain the coordination and consistency of the appearance of the characters in the target scene character image; the reference image is processed by the segmentation module to obtain the background image, character image and face image of the reference image, which can more efficiently utilize computing resources, accelerate the generation speed, and accurately extract features; the background image is processed by the second image encoder to obtain the image background features of the background image, providing data support for generating a coherent and consistent scene character image; the face semantic image fusion module fuses the face identity features and face semantic features of the face image to obtain the face semantic image fusion features of the face image, which can obtain more comprehensive face information and improve the accuracy of the face semantic image fusion features; the character semantic image fusion module fuses the character identity features and character semantic features of the character image to obtain the character semantic image fusion features of the character image, which can obtain more comprehensive character information and improve the accuracy of the character semantic image fusion features.

[0070] In some embodiments, fusing the face identity features and face semantic features of the face image by the face semantic image fusion module to obtain the face semantic image fusion features of the face image includes:

[0071] Processing the face image by the face feature extraction module to obtain the first face identity features of the face image;

[0072] Processing the first face identity features by the first resampling module to obtain the second face identity features of the character image;

[0073] Processing the face image by the third image encoder to obtain the first face semantic features of the face image;

[0074] Processing the first face semantic features by the second resampling module to obtain the second face semantic features of the face image;

[0075] Fusing and processing the second face identity features and the second face semantic features by the face feature fusion module to obtain face fusion features;

[0076] Processing the second face semantic features by the first self-attention module to obtain the third face semantic features;

[0077] Processing the face fusion features and the third face semantic features by the first multi-layer perceptron module to obtain the face semantic image fusion features of the face image.

[0078] Specifically, the face feature extraction module can use the Arcface face feature extraction model to extract the first face identity features, which may include appearance features, physiological features, expression features, etc. The second face identity features correspond to the first face identity features. The first resampling module uses dynamic resampling to unify the spatial resolution of the first face identity features, thereby ensuring data consistency and accelerating the generation speed.

[0079] Specifically, the third image encoder can use the Bootstrapping Language-Image Pretraining 2 (BLIP2) encoder to extract the first face semantic features. The second face semantic features correspond to the first face semantic features. The second resampling module is used to unify the spatial resolution of the first face semantic features and improve data consistency.

[0080] Specifically, the face feature fusion module can use the Querying Transformer (Q-former) technology. The face feature fusion module is used to fuse the second face identity features and the second face semantic features. The Q-former technology can effectively fuse different types of features, improve the expressiveness of the fused features, thereby enhancing the accuracy and generalization ability of face recognition, and making the target scene person image more vivid.

[0081] Specifically, the first self-attention module highlights the important parts in the second face semantic features by learning the internal dependencies in the second face semantic features, thereby obtaining a third face semantic feature with higher quality and expressiveness.

[0082] Specifically, the first Multilayer Perceptron (MLP) module is used to perform non-linear transformation and feature enhancement on the face fusion features and the third face semantic features, thereby improving the expressiveness and accuracy of the face semantic image fusion features.

[0083] According to the technical solution provided by the embodiments of the present application, the face feature extraction module processes the face image to obtain the first face identity feature of the face image; the first resampling module processes the first face identity feature to obtain the second face identity feature of the person image; the third image encoder processes the face image to obtain the first face semantic feature of the face image; the second resampling module processes the first face semantic feature to obtain the second face semantic feature of the face image; accurate and consistent face identity features and face semantic features are obtained, and the image generation speed is accelerated; the face feature fusion module fuses and processes the second face identity feature and the second face semantic feature to obtain the face fusion feature, improving the expressiveness of the fusion feature; the first self-attention module processes the second face semantic feature to obtain the third face semantic feature, improving the quality and expressiveness of the face semantic feature; the first multi-layer perceptron module processes the face fusion feature and the third face semantic feature to obtain the face semantic image fusion feature of the face image, improving the expressiveness and accuracy of the face semantic image fusion feature, so that the face depicted in the target scene person image is more vivid and accurate, and has a higher consistency with the face material of the reference image.

[0084] In some embodiments, the person semantic image fusion module fuses the person appearance feature and the person semantic feature of the person image to obtain the person semantic image fusion feature of the person image, including:

[0085] The person feature extraction module processes the person image to obtain the first person identity feature of the person image;

[0086] The third resampling module processes the first person identity feature to obtain the second person identity feature of the person image;

[0087] The fourth image encoder processes the person image to obtain the first person semantic feature of the person image;

[0088] The fourth resampling module processes the first person semantic feature to obtain the second person semantic feature of the person image;

[0089] The person feature fusion module fuses and processes the second person identity feature and the second person semantic feature to obtain the person fusion feature;

[0090] The second self-attention module processes the second person semantic feature to obtain the third person semantic feature;

[0091] The second multi-layer perceptron module processes the person fusion feature and the third person semantic feature to obtain the person semantic image fusion feature of the person image.

[0092] Specifically, the person feature extraction module can use the VIT image feature extraction model to extract the first person identity feature. The second person identity feature corresponds to the first person identity feature, and the third resampling module is used to unify the spatial resolution of the first person identity feature.

[0093] Specifically, the fourth image encoder can use the BLIP2 encoder to extract the first person semantic feature. The second person semantic feature corresponds to the first person semantic feature, and the fourth resampling module is used to unify the spatial resolution of the first person semantic feature.

[0094] Specifically, the person feature fusion module can use the Q-former technology to improve the expressiveness and accuracy of the person fusion feature and ensure the consistency of the person images in the target scene.

[0095] Specifically, the second self-attention module learns the internal dependencies in the second person semantic feature, highlights the important parts in the second person semantic feature, and obtains a third person semantic feature with higher quality and expressiveness.

[0096] Specifically, the second multi-layer perceptron module is used to perform non-linear transformation and feature enhancement on the person fusion feature and the third person semantic feature, so as to improve the expressiveness and accuracy of the person semantic image fusion feature, and facilitate the refinement of the person image in the target scene person image.

[0097] According to the technical solution provided by the embodiments of the present application, the person feature extraction module processes the person image to obtain the first person identity feature of the person image; the third resampling module processes the first person identity feature to obtain the second person identity feature of the person image; the fourth image encoder processes the person image to obtain the first person semantic feature of the person image; accurate and consistent face identity features and face semantic features are obtained, providing data support for face feature fusion; the fourth resampling module processes the first person semantic feature to obtain the second person semantic feature of the person image; the person feature fusion module fuses the second person identity feature and the second person semantic feature to obtain a person fusion feature; the expressiveness of the person fusion feature is improved, enabling the person fusion feature to convey person features more efficiently; the second self-attention module processes the second person semantic feature to obtain a third person semantic feature; the second multi-layer perceptron module processes the person fusion feature and the third person semantic feature to obtain the person semantic image fusion feature of the person image, which can improve the expressiveness and accuracy of the person semantic image fusion feature, facilitate the refinement of the person image in the target scene person image, accurately distinguish information such as person identity and clothing, and meet the personalized needs of users.

[0098] In some embodiments, the character face feature fusion module cross - fuses the first text feature, the character semantic image fusion feature, and the face semantic image fusion feature to obtain the character face fusion feature, including:

[0099] The first text feature is processed by an average pooling layer to obtain a second text feature of the text description information;

[0100] The second text feature and the face semantic image fusion feature are fused by a first cross - attention module to obtain a text - face fusion feature;

[0101] The text - face fusion feature and the character semantic image fusion feature are fused by a second cross - attention module to obtain the character face fusion feature.

[0102] Specifically, the character face fusion module includes an average pooling layer (average pooling, Avgpool), a first cross - attention module, and a second cross - attention module. Among them, Avgpool is used to aggregate and reduce the dimension of the first text feature, reduce the spatial size of the first text feature, thereby reducing the amount of calculation and the number of parameters, while maintaining important feature information, and outputting the second text feature.

[0103] Specifically, the first cross - attention module is used to learn the correlation between the second text feature and the face semantic image fusion feature, complement and enhance the matching degree of the image and text features, and can more accurately capture and reproduce the detailed texture of the face; the second cross - attention module is used to learn the correlation between the text - face fusion feature and the character semantic image fusion feature, control the overall style of the character face in the generated image through the text feature, increase the controllability of the generated character, and improve the overall quality of the generated image.

[0104] According to the technical solution provided by the embodiments of the present application, the first text feature is processed by an average pooling layer to obtain a second text feature of the text description information; the aggregation and dimension reduction operations of the first text feature are realized, the amount of calculation and the number of parameters are reduced, and the proportion of important feature information is increased; the second text feature and the face semantic image fusion feature are fused by a first cross - attention module to obtain a text - face fusion feature, so as to more accurately capture and reproduce the detailed texture of the face; the text - face fusion feature and the character semantic image fusion feature are fused by a second cross - attention module to obtain the character face fusion feature, realizing the effect of controlling the overall style of the character face in the generated image through the text feature, increasing the controllability of the generated character, being able to improve the overall quality of the generated image, and making the generated target scene character image more realistic and natural.

[0105] In some embodiments, the diffusion model includes a U-shaped network module; the U-shaped network module includes N sub-modules with the same structure, where N is greater than or equal to 2;

[0106] Through the diffusion model, the features after fusing the preset noise map and the skeleton map, the global image features, the image background features, the first text features, and the human face fusion features are processed to generate the target scene human image of the object to be generated, including:

[0107] The global image features, the image background features, the first text features, and the human face fusion features are fused through N identical sub-modules to obtain the final fusion features;

[0108] According to the features after fusing the preset noise map and the skeleton map and the final fusion features, the target scene human image of the object to be generated is generated.

[0109] Specifically, the sub-modules in the UNet structure are connected in a U-shaped structure, and the output of the previous sub-module is the input of the next sub-module. Each sub-module performs preprocessing, feature fusion, etc. on the input features, understands and learns the correlation between multi-modal features, adjusts the image features, improves the detail accuracy in the generated image, and finally obtains the final fusion features, generating a target scene human image that is consistent with the text description and has natural coherence in the background and characters.

[0110] Specifically, the preset noise map is a pre-selected image with noise. The diffusion model restores the image details by locally adding and removing noise to generate high-quality images. The skeleton map provides data such as human proportions and dynamics for the generation of the target scene human image, ensuring that the human image in the generated image is more vivid and accurate.

[0111] According to the technical solution provided by the embodiments of the present application, the global image features, the image background features, the first text features, and the human face fusion features are fused through N identical sub-modules to obtain the final fusion features, achieving the effect of continuously adjusting and fusing multi-modal features during the generation process, depicting image details, and improving the accuracy of the fusion features; according to the features after fusing the preset noise map and the skeleton map and the final fusion features, the target scene human image of the object to be generated is generated, achieving the generation of a high-quality personalized scene human image that meets the requirements, with the personalized features of the whole body of the character accurately shaped, the appearance of the character consistent with the background information, and the overall coherence.

[0112] In some embodiments, the sub-module includes a multi-head self-attention module, a text cross-attention module, a background cross-attention module, a human face cross-attention module, and a feed-forward network module;

[0113] Fusing the global image feature, the image background feature, the first text feature, and the human face fusion feature through N identical sub-modules to obtain the final fusion feature, including:

[0114] Processing the input feature through a multi-head self-attention module, where the input feature is the feature obtained by the previous sub-module through fusing the global image feature, the image background feature, the first text feature, and the human face fusion feature;

[0115] Fusing the feature obtained from the previous step and the first text feature through a text cross-attention module;

[0116] Fusing the feature obtained from the previous step and the image background feature through a background cross-attention module;

[0117] Fusing the feature obtained from the previous step and the human face fusion feature through a human face cross-attention module;

[0118] Fusing the feature obtained from the previous step through a feed-forward network module to obtain the output feature;

[0119] Outputting the output feature so that the subsequent sub-module can fuse the output feature to obtain the final fusion feature.

[0120] Specifically, the sub-module also includes multiple linear layer modules. Before processing the input feature through the multi-head self-attention module, the global image feature, the image background feature, the first text feature, and the human face fusion feature are respectively subjected to dimension adjustment and preprocessing through the linear layer module, so as to dynamically adjust the transformation parameters according to the feature distribution and importance degree, and output features with unified dimensions for subsequent fusion operations.

[0121] Specifically, the multi-head self-attention module receives the feature output by the linear layer for processing. The multi-head self-attention module can simultaneously focus on information in different positions and different feature dimensions, capture the diversity of features from different angles, thereby improving the expression ability and adaptability of the fusion feature.

[0122] Specifically, the text cross-attention module is used to adjust the details in the image generation process according to the semantic information of the text description. For example, the text description information is "a little girl in a red dress is chasing butterflies", and the semantic information such as color and style in it can help adjust the image feature and improve the accuracy of the generated image.

[0123] Specifically, the background cross-attention module is used to learn the relationship between the image background features and other features, so as to make the background and the characters in the generated image more harmonious and natural. The cross-attention module for human faces is used to improve the consistency of the generated image in terms of human facial features and overall appearance with the reference image and the text description information, so as to ensure that the human face and overall appearance in the generated image are accurate and natural.

[0124] Specifically, the feed-forward network module further fuses the features obtained from the previous step through non-linear transformation and feature enhancement operations, increases the flexibility of the fused features, removes noise and interference information at the same time, improves the visual effect of the image, and enhances the clarity of the image, so as to obtain a clearer and more realistic image.

[0125] According to the technical solution provided by the embodiment of the present application, by processing the input features through the multi-head self-attention module, the diversity of features can be captured from different perspectives, so as to improve the expression ability and adaptability of the fused features; by fusing the features obtained from the previous step and the first text features through the text cross-attention module, the details in the image generation process can be adjusted according to the semantic information of the text description, and the accuracy of the generated image can be improved; by fusing the features obtained from the previous step and the image background features through the background cross-attention module, the relationship between the image background features and other features can be learned, so as to make the background and the characters in the generated image more harmonious and natural; by fusing the features obtained from the previous step and the human face fused features through the cross-attention module for human faces, the consistency of the generated image in terms of human facial features and overall appearance with the reference image and the text description information can be improved, so as to ensure that the human face and overall appearance in the generated image are accurate and natural; by fusing the features obtained from the previous step through the feed-forward network module to obtain the output features, the flexibility of the fused features can be increased, thereby improving the visual effect of the image and enhancing the clarity of the image; output the output features, so that the subsequent sub-modules can fuse the output features to obtain the final fused features, thereby generating a target scene character image with high accuracy and clear and realistic.

[0126] Any combination of the above all optional technical solutions can form an optional embodiment of the present application, which will not be elaborated here one by one.

[0127] Figure 2 It is a schematic diagram of the module structure of a method for generating a scene image provided by an embodiment of the present application.

[0128] As Figure 2 shown, the method for generating a scene image includes:

[0129] Input the reference image into the skeleton extraction module, extract the skeleton map, and fuse the skeleton map with the preset noise map;

[0130] Input the reference image into the first image encoder, and input the features output by the first image encoder into the linear layer to obtain global image features;

[0131] Input the reference image into the first segmentation module to obtain the background image and the person image, and input the background image into the second image encoder to obtain image background features;

[0132] Input the person image into the person semantic image fusion extraction module to obtain person semantic image fusion features;

[0133] Input the person image into the second segmentation module to obtain the face image, and input the face image into the face semantic image fusion extraction module to obtain face semantic image fusion features;

[0134] Input the text description information into the text encoder to obtain the first text feature, and input the first text feature into the average pooling layer to obtain the second text feature;

[0135] Input the second text feature and the face semantic image fusion features into the first cross-attention module to obtain text-face fusion features;

[0136] Input the text-face fusion features and the person semantic image fusion features into the second cross-attention module to obtain person-face fusion features;

[0137] Input the features after fusing the skeleton map and the preset noise map, the global image features, the image background features, the person-face fusion features, and the first text feature into the diffusion model to obtain the target scene person image.

[0138] The following is an apparatus embodiment of the present application, which can be used to execute the method embodiment of the present application. For details not disclosed in the apparatus embodiment of the present application, please refer to the method embodiment of the present application.

[0139] Figure 5 It is a schematic diagram of a device for generating a scene image provided by an embodiment of the present application. As Figure 5 shown, the device for generating the scene image includes:

[0140] An acquisition module 501, configured to acquire the text description information of the object to be generated and the reference image of the object to be generated;

[0141] A first processing module 502, configured to process the text description information through a text encoder to obtain the first text feature of the text description information;

[0142] The second processing module 503 is used to process the reference image through the image feature extraction module to obtain the target image features of the reference image. The target image features include global image features, image background features, skeleton maps, person-semantic image fusion features, and face-semantic image fusion features. The person-semantic image fusion feature is the feature after fusing the person identity feature and the person semantic feature in the reference image, and the face-semantic image fusion feature is the feature after fusing the face identity feature and the face semantic feature in the reference image.

[0143] The cross-fusion module 504 is used to cross-fuse the first text feature, the person-semantic image fusion feature, and the face-semantic image fusion feature through the person-face feature fusion module to obtain the person-face fusion feature.

[0144] The generation module 505 is used to process the fused features of the preset noise map and the skeleton map, the global image features, the image background features, the first text feature, and the person-face fusion feature through the diffusion model to generate the target scene person image of the object to be generated.

[0145] According to the technical solution provided by the embodiments of the present application, the acquisition module 501 acquires the text description information of the object to be generated and the reference image of the object to be generated, obtaining the scene materials, character materials and user requirements of the object to be generated; the first processing module 502 processes the text description information through a text encoder to obtain the first text feature of the text description information, providing data support for subsequent feature fusion, thereby improving the accuracy of the fusion feature and generating a target scene character image that meets the user's expectations; the second processing module 503 processes the reference image through an image feature extraction module to obtain the target image feature of the reference image, providing a data basis for the target scene character image, enabling the target scene character image to maintain the consistency of the picture style and the coordination and coherence between the background information and the character actions, avoiding the problem of excessive overall style differences in the images before and after the character actions in the target scene character image, and being able to accurately distinguish the identities, clothes of the characters in a multi-person scene and generate more accurate character facial expressions; the cross-fusion module 504 cross-fuses the first text feature, the character semantic image fusion feature, and the face semantic image fusion feature through the character face feature fusion module to obtain the character face fusion feature, which can learn the correlation relationship between the first text feature, the character semantic image fusion feature, and the face semantic image fusion feature, thereby specifically focusing on the user's requirements for the characters or faces in the target scene character image, effectively fusing the text feature with the character feature and the face feature, and increasing the accuracy of the character face fusion feature; the generation module 505 processes the features after the fusion of the preset noise map and the skeleton map, the global image feature, the image background feature, the first text feature, and the character face fusion feature through a diffusion model to generate the target scene character image of the object to be generated, which can perform multiple cross-fusions on the global image feature, the image background feature, the first text feature, and the character face fusion feature, learning the complex associations and dependencies between the features, thereby obtaining more accurate fusion features. It can solve the technical problem that the existing methods for generating scene images are difficult to accurately shape the personalized features of the whole body of the character.

[0146] In some embodiments, the image feature extraction module includes a first image encoder, a second image encoder, a skeleton extraction module, a segmentation module, a face semantic image fusion module, and a person semantic image fusion module; the second processing module 503 is specifically configured to: process the reference image through the first image encoder to obtain initial features, and perform dimensional transformation on the initial features through a linear layer to obtain the global image features of the reference image; process the reference image through the skeleton extraction module to obtain the skeleton map of the reference image; process the reference image through the segmentation module to obtain the background image, the person image, and the face image of the reference image; process the background image through the second image encoder to obtain the image background features of the background image; fuse the face identity features and the face semantic features of the face image through the face semantic image fusion module to obtain the face semantic image fusion features of the face image; fuse the person identity features and the person semantic features of the person image through the person semantic image fusion module to obtain the person semantic image fusion features of the person image.

[0147] In some embodiments, the second processing module 503 is further configured to: process the face image through the face feature extraction module to obtain the first face identity features of the face image; process the first face identity features through the first resampling module to obtain the second face identity features of the person image; process the face image through the third image encoder to obtain the first face semantic features of the face image; process the first face semantic features through the second resampling module to obtain the second face semantic features of the face image; perform a fusion process on the second face identity features and the second face semantic features through the face feature fusion module to obtain face fusion features; process the second face semantic features through the first self-attention module to obtain the third face semantic features; process the face fusion features and the third face semantic features through the first multi-layer perceptron module to obtain the face semantic image fusion features of the face image.

[0148] In some embodiments, the second processing module 503 is further configured to: process the person image through the person feature extraction module to obtain the first person identity features of the person image; process the first person identity features through the third resampling module to obtain the second person identity features of the person image; process the person image through the fourth image encoder to obtain the first person semantic features of the person image; process the first person semantic features through the fourth resampling module to obtain the second person semantic features of the person image; perform a fusion process on the second person identity features and the second person semantic features through the person feature fusion module to obtain person fusion features; process the second person semantic features through the second self-attention module to obtain the third person semantic features; process the person fusion features and the third person semantic features through the second multi-layer perceptron module to obtain the person semantic image fusion features of the person image.

[0149] In some embodiments, the cross-fusion module 504 is specifically configured to: process the first text feature through an average pooling layer to obtain a second text feature of the text description information; fuse and process the second text feature and the face semantic image fusion feature through a first cross-attention module to obtain a text-face fusion feature; fuse and process the text-face fusion feature and the person semantic image fusion feature through a second cross-attention module to obtain a person-face fusion feature.

[0150] In some embodiments, the generation module 505 is specifically configured to: the diffusion model includes a U-shaped network module; the U-shaped network module includes N sub-modules with the same structure, and N is greater than or equal to 2; fuse and process the global image feature, the image background feature, the first text feature, and the person-face fusion feature through N identical sub-modules to obtain a final fusion feature; generate a target scene person image of the object to be generated according to the preset noise map and the fused feature of the skeleton map and the final fusion feature.

[0151] In some embodiments, the sub-module includes a multi-head self-attention module, a text cross-attention module, a background cross-attention module, a person-face cross-attention module, and a feed-forward network module; the generation module 505 is further configured to: process the input feature through the multi-head self-attention module, and the input feature is the feature obtained by the previous sub-module fusing and processing the global image feature, the image background feature, the first text feature, and the person-face fusion feature; fuse and process the feature obtained in the previous step and the first text feature through the text cross-attention module; fuse and process the feature obtained in the previous step and the image background feature through the background cross-attention module; fuse and process the feature obtained in the previous step and the person-face fusion feature through the person-face cross-attention module; fuse and process the feature obtained in the previous step through the feed-forward network module to obtain an output feature; output the output feature so that the subsequent sub-module fuses and processes the output feature to obtain a final fusion feature.

[0152] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0153] Figure 6 is a schematic diagram of the electronic device 6 provided by the embodiment of the present application. As Figure 6As shown, the electronic device 6 of this embodiment includes: a processor 601, a memory 602, and a computer program 603 stored in the memory 602 and executable on the processor 601. When the processor 601 executes the computer program 603, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor 601 executes the computer program 603, the functions of the various modules / units in the above-mentioned device embodiments are implemented.

[0154] The electronic device 6 may be a desktop computer, a notebook, a palm computer, a cloud server, or other electronic devices. The electronic device 6 may include, but is not limited to, the processor 601 and the memory 602. Those skilled in the art can understand that Figure 6 merely examples of the electronic device 6, which do not constitute a limitation on the electronic device 6, may include more or fewer components than shown in the figure, or different components.

[0155] The processor 601 may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0156] The memory 602 may be an internal storage unit of the electronic device 6, for example, the hard disk or memory of the electronic device 6. The memory 602 may also be an external storage device of the electronic device 6, for example, a plug-in hard disk equipped on the electronic device 6, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. The memory 602 may also include both an internal storage unit and an external storage device of the electronic device 6. The memory 602 is used to store computer programs and other programs and data required by the electronic device.

[0157] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0158] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium (such as a computer-readable storage medium). Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of this application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. The computer program can include computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable storage medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0159] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of this application, and should all be included in the protection scope of this application.

Claims

1. A method for generating a scene image, characterized in that, Including: Obtain the text description information of the object to be generated and the reference image of the object to be generated; Process the text description information through a text encoder to obtain the first text feature of the text description information; Process the reference image through an image feature extraction module to obtain the target image feature of the reference image, where the target image feature includes a global image feature, an image background feature, a skeleton graph, a person semantic image fusion feature, and a face semantic image fusion feature; Wherein the person semantic image fusion feature is the feature after fusing the person identity feature and the person semantic feature in the reference image, and the face semantic image fusion feature is the feature after fusing the face identity feature and the face semantic feature in the reference image; Cross-fuse the first text feature, the person semantic image fusion feature, and the face semantic image fusion feature through a person face feature fusion module to obtain a person face fusion feature; Process the fused feature of the preset noise map and the skeleton graph, the global image feature, the image background feature, the first text feature, and the person face fusion feature through a diffusion model to generate the target scene person image of the object to be generated.

2. The method according to claim 1, wherein The image feature extraction module includes a first image encoder, a second image encoder, a skeleton extraction module, a segmentation module, a face semantic image fusion module, and a person semantic image fusion module; The process of obtaining the target image feature of the reference image through the image feature extraction module includes: Process the reference image through the first image encoder to obtain an initial feature, and perform dimensional transformation on the initial feature through a linear layer to obtain the global image feature of the reference image; Process the reference image through the skeleton extraction module to obtain the skeleton graph of the reference image; Process the reference image through the segmentation module to obtain the background image, the person image, and the face image of the reference image; Process the background image through the second image encoder to obtain the image background feature of the background image; Fuse the face identity feature and the face semantic feature of the face image through the face semantic image fusion module to obtain the face semantic image fusion feature of the face image; Fuse the person identity feature and the person semantic feature of the person image through the person semantic image fusion module to obtain the person semantic image fusion feature of the person image.

3. The method according to claim 2, wherein The process of fusing the face identity feature and the face semantic feature of the face image through the face semantic image fusion module to obtain the face semantic image fusion feature of the face image includes: Process the face image through a face feature extraction module to obtain the first face identity feature of the face image; Process the first face identity feature through a first resampling module to obtain the second face identity feature of the person image; Process the face image through a third image encoder to obtain the first face semantic feature of the face image; The first face semantic feature is processed by a second resampling module to obtain a second face semantic feature of the face image; The second face identity feature and the second face semantic feature are fused by a face feature fusion module to obtain a face fusion feature; The second face semantic feature is processed by a first self-attention module to obtain a third face semantic feature; The face fusion feature and the third face semantic feature are processed by a first multi-layer perceptron module to obtain a face semantic image fusion feature of the face image.

4. The method according to claim 2, wherein The person appearance feature and the person semantic feature of the person image are fused by the person semantic image fusion module to obtain the person semantic image fusion feature of the person image, including: The person image is processed by a person feature extraction module to obtain a first person identity feature of the person image; The first person identity feature is processed by a third resampling module to obtain a second person identity feature of the person image; The person image is processed by a fourth image encoder to obtain a first person semantic feature of the person image; The first person semantic feature is processed by a fourth resampling module to obtain a second person semantic feature of the person image; The second person identity feature and the second person semantic feature are fused by a person feature fusion module to obtain a person fusion feature; The second person semantic feature is processed by a second self-attention module to obtain a third person semantic feature; The person fusion feature and the third person semantic feature are processed by a second multi-layer perceptron module to obtain the person semantic image fusion feature of the person image.

5. The method according to claim 1, characterized in that, The first text feature, the person semantic image fusion feature, and the face semantic image fusion feature are cross-fused by a person-face feature fusion module to obtain a person-face fusion feature, including: The first text feature is processed by an average pooling layer to obtain a second text feature of the text description information; The second text feature and the face semantic image fusion feature are fused by a first cross-attention module to obtain a text-face fusion feature; The text-face fusion feature and the person semantic image fusion feature are fused by a second cross-attention module to obtain a person-face fusion feature.

6. The method according to claim 1, wherein The diffusion model includes a U-shaped network module; the U-shaped network module includes N sub-modules with the same structure, where N is greater than or equal to 2; The preset noise map and the features after fusing the skeleton map, the global image feature, the image background feature, the first text feature, and the person-face fusion feature are processed by the diffusion model to generate a target scene person image of the object to be generated, including: The global image feature, the image background feature, the first text feature, and the person-face fusion feature are fused by the N identical sub-modules to obtain a final fusion feature; Generate the target scene character image of the object to be generated according to the features after fusing the preset noise map and the skeleton map and the final fused features.

7. The method according to claim 6, wherein The sub-module includes a multi-head self-attention module, a text cross-attention module, a background cross-attention module, a character face cross-attention module, and a feed-forward network module; The fusing the global image feature, the image background feature, the first text feature, and the character face fusion feature through the N identical sub-modules to obtain a final fused feature includes: Process the input feature through the multi-head self-attention module, where the input feature is the feature obtained by the previous sub-module fusing the global image feature, the image background feature, the first text feature, and the character face fusion feature; Fuse the feature obtained from the previous step and the first text feature through the text cross-attention module; Fuse the feature obtained from the previous step and the image background feature through the background cross-attention module; Fuse the feature obtained from the previous step and the character face fusion feature through the character face cross-attention module; Fuse the feature obtained from the previous step through the feed-forward network module to obtain an output feature; Output the output feature so that the subsequent sub-module fuses the output feature to obtain the final fused feature.

8. An apparatus for generating a scene image, characterized in that, Includes: An acquisition module for acquiring the text description information of the object to be generated and the reference image of the object to be generated; A first processing module for processing the text description information through a text encoder to obtain the first text feature of the text description information; A second processing module for processing the reference image through an image feature extraction module to obtain the target image feature of the reference image, where the target image feature includes a global image feature, an image background feature, a skeleton map, a character semantic image fusion feature, and a face semantic image fusion feature; Where the character semantic image fusion feature is the feature after fusing the character identity feature and the character semantic feature in the reference image, and the face semantic image fusion feature is the feature after fusing the face identity feature and the face semantic feature in the reference image; A cross-fusion module for cross-fusing the first text feature, the character semantic image fusion feature, and the face semantic image fusion feature through a character face feature fusion module to obtain a character face fusion feature; A generation module for processing the features after fusing the preset noise map and the skeleton map, the global image feature, the image background feature, the first text feature, and the character face fusion feature through a diffusion model to generate the target scene character image of the object to be generated.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.