Image generation method and device, electronic equipment and storage medium
By fusing features from original object images, style clothing images, and pose images, and combining them with an image generation model, the problem of insufficient understanding of image generation intent in existing technologies is solved, thereby improving the personalized image generation effect.
Patent Information
- Application Number
- CN202410610769.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-16
- Publication Date
- 2025-11-18
AI Technical Summary
Existing text-to-image solutions fail to accurately understand the user's image generation intent, resulting in generated images that do not match the user's needs and cannot maintain the personalized characteristics of the person, leading to poor image generation results.
By acquiring features from the original object image, the target style clothing image, and the target pose image, and combining them with the target image description text, the feature fusion process is performed and then input into a preset image generation model to generate the target object image, ensuring the accurate expression of the person's pose and clothing style.
It improves the accuracy of representing the main object during image generation, enhances personalized expression capabilities, ensures the matching of generated images with user intent, and improves the fine-grained feature learning performance of image generation.
Smart Images

Figure CN120976328A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to an image generation method and device, electronic equipment and a storage medium. BACKGROUND
[0002] Photography is an art and also a technology; after taking photos, many photography enthusiasts will also perform post-processing on the photos to enhance the beauty and expressiveness of the pictures, so that the pictures can better meet the creativity and style expression ability of the photographer, and can also make up for the shortcomings or errors during shooting.
[0003] With the development of artificial intelligence (AI) technology, text-to-image in AI technology is widely used in image secondary processing; in the existing text-to-image scheme, the original reference image of a person and an instruction text representing the image generation intention are input into a diffusion model to generate images of various styles and types desired. However, the existing text-to-image scheme only combines a short descriptive text to describe the image requirements such as style, and the large model cannot well understand the image generation intention of the user, and the personalized image of the person cannot be well maintained, resulting in that the finally generated image does not match the user's requirements, and the personalized features of the main subject cannot be maintained, and the image generation effect is poor. SUMMARY
[0004] The present application provides an image generation method, device, equipment, storage medium and computer program product, which can improve the understanding ability of the model to the image generation requirements on the basis of improving the representation accuracy of the generated image to the main subject, guarantee the matching between the generated image and the user's requirements, and further greatly improve the personalized image generation effect.
[0005] In one aspect, the present application provides an image generation method, which comprises:
[0006] obtaining an original object image, a target style clothing image, a target pose image and a target image description text, the target image description text being used to indicate a target object corresponding to the original object image, in a target pose corresponding to the target pose image, wearing a target style clothing corresponding to the target style clothing image;
[0007] determining a first object feature corresponding to the original object image, a first style feature corresponding to the target style clothing image and a first description feature corresponding to the target image description text;
[0008] fusing the first object feature and the first style feature to obtain a second style feature;
[0009] The first object feature and the first description feature are fused to obtain a second description feature;
[0010] The second style feature, the second description feature, and the target posture image are input into a preset image generation model for image generation processing to obtain a target object image corresponding to the target object.
[0011] Another aspect provides an image generation device, the device comprising:
[0012] An information acquisition module configured to perform acquisition of an original object image, a target style clothing image, a target posture image, and a target image description text, the target image description text being used to indicate generation of a target object corresponding to the original object image, in a target posture corresponding to the target posture image, wearing a target style clothing corresponding to the target style clothing image, an image;
[0013] A first feature determination module configured to perform determination of a first object feature corresponding to the original object image, a first style feature corresponding to the target style clothing image, and a first description feature corresponding to the target image description text;
[0014] A first fusion processing module configured to perform fusion processing of the first object feature and the first style feature to obtain a second style feature;
[0015] A second fusion processing module configured to perform fusion processing of the first object feature and the first description feature to obtain a second description feature;
[0016] A first image generation processing module configured to perform image generation processing of the second style feature, the second description feature, and the target posture image input into a preset image generation model to obtain a target object image corresponding to the target object.
[0017] Another aspect provides an electronic device comprising a processor;
[0018] A memory for storing processor-executable instructions;
[0019] The processor is configured to execute the instructions to implement any of the above image generation methods.
[0020] Another aspect provides a computer-readable storage medium, when instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform any of the above image generation methods.
[0021] The other aspect provides a computer program product or computer program, which comprises computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions, so that the computer device executes the image generation method provided in various optional implementations.
[0022] The image generation method, device, equipment, storage medium and computer program product provided by the application have the following technical effects:
[0023] In the image generation process, the original object image, the target style clothing image, the target pose image, and the target image description text for indicating an image of a target object corresponding to the original object image, wearing a target style clothing corresponding to the target style clothing image, in a target pose corresponding to the target pose image are obtained. Then, a first object feature corresponding to the original object image, a first style feature corresponding to the target style clothing image, and a first description feature corresponding to the target image description text are determined. The first object feature and the first style feature are fused to obtain a second style feature, and the first object feature and the first description feature are fused to obtain a second description feature. Then, the second style feature fused with the object feature, the second description feature fused with the object feature, and the target pose image are input into a preset image generation model for image generation processing to obtain a target object image corresponding to the target object. In this way, the representation accuracy of the model for the subject object in the image generation process can be improved by combining the second style feature fused with the object feature and the second description feature fused with the object feature. In addition, by combining the image pose on the basis of the style feature and the description feature of the image, the dressing and pose of the character can be accurately controlled, the personalized expression ability can be enhanced, the learning performance of the model for the pose, the subject object, and the style and other fine-grained features can be improved, the understanding ability of the model for the image generation requirement and the user intention can be improved, the matching between the generated image and the user intention can be ensured, and thus the personalized image generation effect can be greatly improved. BRIEF DESCRIPTION OF DRAWINGS
[0024] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the drawings needed in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0025] Figure 1 is a schematic diagram of an application environment of an image generation method provided by an embodiment of the present application;
[0026] Figure 2 is a flowchart of a method for generating an image according to an embodiment of the present application;
[0027] Figure 3 is a flowchart of a method for generating an image according to an embodiment of the present application;
[0028] Figure 4 is a flowchart of a method for generating an image according to an embodiment of the present application;
[0029] Figure 5 is a flowchart of a method for generating an image according to an embodiment of the present application;
[0030] Figure 6 is a flowchart of a method for generating an image according to an embodiment of the present application;
[0031] Figure 7 is a flowchart of a method for generating an image according to an embodiment of the present application;
[0032] Figure 8 is a flowchart of a method for generating an image according to an embodiment of the present application; DETAILED DESCRIPTION
[0033] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, any other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily mean a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0035] In the embodiments of this application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as processing circuitry or memory) or a combination thereof. Similarly, one processor (or multiple processors or memory) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an integral module or unit that includes the functions of the module or unit.
[0036] With the research and progress of artificial intelligence technology, artificial intelligence technology has been researched and applied in many fields, such as common smart home, smart wearable device, virtual assistant, smart speaker, smart marketing, unmanned driving, autonomous driving, unmanned aerial vehicle, digital twin, virtual human, robot, artificial intelligence generated content (AIGC, Artificial Intelligence Generated Content), conversational interaction, smart medical treatment, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0037] The scheme provided by the embodiments of the present application mainly relates to artificial intelligence generated content and other technologies. Artificial intelligence generated content is a new way of content creation, which uses artificial intelligence technology to assist or replace manual content generation, and generates content based on user input keywords or requirements. Specifically, the following embodiments are used for illustration:
[0038] Please refer to Figure 1 , Figure 1 is a schematic diagram of an application environment of an image generation method provided by the embodiments of the present application. The application environment can at least include a server 100 and a terminal 200.
[0039] In an optional embodiment, the server 100 can be used to pre-train a preset image generation model required by the image generation process and other models. The server 100 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0040] In an optional embodiment, the terminal 200 can be used to provide image generation services for users, for example, on some social media, an AI-based image generation function portal can be provided; accordingly, users can process newly uploaded images or historical images on social media based on the image generation function portal. Specifically, the terminal 200 can include, but is not limited to, a smart phone, a desktop computer, a tablet computer, a notebook computer, a smart speaker, a digital assistant, an augmented reality (AR) / virtual reality (VR) device, a smart wearable device, a vehicle-mounted terminal, a smart television, and the like. It can also be a software running on the above-mentioned electronic devices, such as an application program, a mini program, and the like. The operating system running on the electronic device in the embodiment of the present application can include, but is not limited to, an Android system, an IOS system, Linux, Windows, and the like.
[0041] In addition, it should be noted that, Figure 1 The above is only an application environment of the image generation method, and the embodiments of the present application are not limited to the above.
[0042] In the embodiments of the present application, the server 100 and the terminal 200 can be connected directly or indirectly through wired or wireless communication, which is not limited in the present application.
[0043] The following describes an image generation method of the present application, Figure 2 is a flowchart of an image generation method provided by an embodiment of the present application. The present application provides method operation steps as in the embodiments or flowcharts, but more or fewer operation steps can be included based on conventional or non-creative labor. The order of steps listed in the embodiments is only one of the many execution orders, and does not represent the only execution order. In actual terminal or server product execution, the method order shown in the embodiments or the drawings can be executed in sequence or in parallel (for example, in a parallel processor or multi-threaded processing environment). Specifically, as shown in Figure 2 The above method can include:
[0044] S201: Obtain an original object image, a target style clothing image, a target pose image, and a target image description text.
[0045] In a specific embodiment, the original object image can be an image containing a target object; specifically, the target object can be a person; specifically, the image containing the target object can be an image containing at least a person's face. Optionally, in addition to containing a person's face, the original object image can also include other parts of the person, etc.
[0046] In an optional embodiment, the above obtaining an original object image can include:
[0047] obtaining a preset object image containing the target object;
[0048] taking the preset object image as the original object image.
[0049] In a specific embodiment, the preset object image can be an initial image containing the target object provided by the user. Alternatively, the preset object image can be a newly uploaded object image of the user, or a historical object image on the social media. Specifically, the historical object image can be an object image uploaded or published to the social media within a historical time period. The historical time period can be set in combination with actual application, for example, a time period before the current login of the social media account.
[0050] In an optional embodiment, the above obtaining the original object image can include:
[0051] obtaining a preset object image containing the target object;
[0052] cropping an object region image from the preset object image;
[0053] performing enhancement processing on the object region image to obtain the original object image.
[0054] In a specific embodiment, the object region image can be an image of a region where the face of the target object is located in the preset object image. Specifically, the above enhancement processing on the object region image can include whitening, increasing image clarity, etc., which can improve image quality but do not affect image processing reflecting the features of the subject (object subject) such as facial features.
[0055] In the above embodiment, the preset object image is first cropped to obtain the object region image, and then the object region image is enhanced and processed to obtain the original object image, which is used to learn the features of the object subject in the subsequent image generation process. This can better extract the features of the object subject.
[0056] In a specific embodiment, the target style costume image can be an image containing a target style costume. The target style costume can be a costume with a target style. Specifically, the style of the costume (costume style) can be set in combination with actual application requirements, for example, Qin and Han Dynasties, Wei and Jin Dynasties, Sui and Tang Dynasties, Song Dynasty, Ming Dynasty, etc., and can also include modern styles, etc. Correspondingly, a plurality of preset style costume images corresponding to different costume styles can be pre-configured in combination with actual application requirements, and the plurality of preset style costume images include the target style costume image.
[0057] In a specific embodiment, the target posture image can be a human body skeleton figure image corresponding to the target posture (human body skeleton figure image in the target posture). In an optional embodiment, the target posture image is obtained in the following manner:
[0058] obtaining a target sample object image;
[0059] inputting the target sample object image into a preset posture generation model for posture image generation processing to obtain the target posture image.
[0060] In a specific embodiment, the target sample object image contains an object with a target posture. The preset posture generation model can be obtained by training a first preset model based on a plurality of sample object images and a labeled posture image corresponding to each sample object image. Specifically, the plurality of sample object images correspond to a plurality of postures, and the labeled posture image corresponding to each sample object image can be a human body skeleton figure image in the corresponding posture. Accordingly, the trained preset posture generation model can generate a posture image (human body skeleton figure image in the corresponding posture) corresponding to an object image based on the object image with a certain posture. Optionally, the model structure of the preset posture generation model can be set according to actual application requirements, for example, it can be a preprocessor of the openpose series.
[0061] In the above embodiment, the preset posture generation model is used to generate a posture image from an image containing an object with a target posture, which can effectively ensure the accuracy of the generated posture image in presenting the posture.
[0062] In a specific embodiment, the target image description text is used to indicate an image of a target object corresponding to the original object image, wearing a target style of clothing corresponding to the target style of clothing image, in a target posture corresponding to the target posture image; for example, an image of a character with the same face as the input character image (original object image), wearing the same style of clothing as the input style of clothing image, in the same posture as the input posture image. Optionally, the target image description text is also used to indicate background information of the finally generated image. Optionally, if the target image description text does not contain indication information of the background information of the finally generated image, the background information of the target style of clothing image can be used as the background information of the finally generated image. Specifically, the target image description text can be set by a user according to actual needs.
[0063] S203: determining a first object feature corresponding to the original object image, a first style feature corresponding to the target style of clothing image, and a first description feature corresponding to the target image description text.
[0064] In a specific embodiment, the first object feature is the representation information of the original object image, i.e., a feature vector; specifically, the first object feature can represent the object feature of the target object. Optionally, the original object image can be input into a pre-trained object feature extraction model for object feature extraction processing to obtain the first object feature. Specifically, the model structure of the object feature extraction model can be set in combination with actual application, for example, an ArcFace model (a kind of face recognition model, which is constructed based on a deep learning convolutional neural network).
[0065] In a specific embodiment, the first style feature can be the representation information of the target style clothing image, i.e., a feature vector; specifically, the first style feature can represent the clothing style feature of the target style clothing corresponding to the target style clothing image; optionally, the target style clothing image can be input into a pre-trained image feature extraction model for image feature extraction processing to obtain the first style feature. Specifically, the model structure of the image feature extraction model can be set in combination with actual application, for example, a vision transformer (a visual encoder).
[0066] In a specific embodiment, the first description feature can be the representation information of the target image description text, i.e., a feature vector; specifically, the first description feature can represent the image generation requirement feature described by the target image description text; optionally, the target image description text can be input into a pre-trained text feature extraction model for text feature extraction processing to obtain the first description feature; specifically, the model structure of the text feature extraction model can be set in combination with actual application; optionally, the text feature extraction model can include a text feature extraction network, and accordingly, the target image description text can be input into the text feature extraction network for text feature extraction processing, and the description feature output by the text feature extraction network can be taken as the first description feature. Optionally, the text feature extraction network can be set in combination with actual application, for example, Flan-UL2 (a large language model framework, i.e., a framework of Unifying Language Learning Paradigms).
[0067] In an optional embodiment, the text feature extraction model can include a text feature extraction network and a prior feature conversion network; accordingly, the first description feature is determined in the following manner:
[0068] The target image description text is input into the text feature extraction network for text feature extraction processing to obtain an initial description feature;
[0069] The initial description feature is input into the prior feature conversion network for processing to obtain the first description feature;
[0070] In a specific embodiment, the network structure of the prior feature conversion network can be set in combination with actual application; and the first description feature can be a prior feature in the image generation process of the preset image generation model for generating an image.
[0071] In the above embodiment, the text feature (initial description feature) corresponding to the target image description text is converted into the prior feature of the preset image generation model, which can facilitate subsequent further fusion with other features and optimize the image generation effect of the preset image generation model.
[0072] S205: The first object feature and the first style feature are fused to obtain a second style feature;
[0073] In a specific embodiment, the fusion of the first object feature and the first style feature can be splicing processing to obtain the second style feature. Specifically, the second style feature can be a style feature fused with the object feature.
[0074] S207: The first object feature and the first description feature are fused to obtain a second description feature;
[0075] In a specific embodiment, the fusion of the first object feature and the first description feature can be splicing processing to obtain the second description feature. Specifically, the second description feature can be a description feature fused with the object feature.
[0076] S209: The second style feature, the second description feature, and the target pose image are input into the preset image generation model for image generation processing to obtain a target object image corresponding to the target object.
[0077] In a specific embodiment, the target object image corresponding to the target object can be an image of the target object wearing target style clothes in a target pose.
[0078] In a specific embodiment, as shown in Figure 3 The preset image generation model can include a pose feature extraction network, a first feature fusion network, a linear projection network, and an image generation network; and correspondingly, the input of the second style feature, the second description feature, and the target pose image into the preset image generation model for image generation processing to obtain the target object image corresponding to the target object can include:
[0079] S2091: The target pose image is input into the pose feature extraction network for pose feature extraction processing to obtain a first pose feature;
[0080] S2093: input the first pose feature, the second style feature, and the second description feature into the first feature fusion network for fusion processing to obtain a first image fusion feature;
[0081] S2095: input the first image fusion feature into a linear projection network for linear projection processing to obtain a second image fusion feature;
[0082] S2097: input the second image fusion feature into an image generation network for image generation processing to obtain a target object image.
[0083] In one specific embodiment, the network structure of the pose feature extraction network can be set in combination with actual application. Optionally, the pose feature extraction network can be a text2Image Adaperts (text-to-image feature conversion adapter). Specifically, the pose feature extraction network can include, in sequence, a pixel-level back-migration network, one convolutional layer, two residual blocks, a down-sampling processing layer, one convolutional layer, and two residual blocks, thereby ensuring better extraction of pose features. The first pose feature can be a pose feature extracted by the pose feature extraction network from a target pose image.
[0084] In one specific embodiment, the above fusion processing of the first pose feature, the second style feature, and the second description feature into the first feature fusion network can include splicing processing of the first pose feature, the second style feature, and the second description feature to obtain the first image fusion feature. Specifically, the first image fusion feature can be a feature after splicing of the first pose feature, the second style feature, and the second description feature.
[0085] In one specific embodiment, the network structure of the linear projection network can be set in combination with actual application. In the linear projection network, the first image fusion feature can be linearly converted to obtain the second image fusion feature, thereby better learning the required image feature. Specifically, the second image fusion feature is a feature after linear conversion of the first image fusion feature.
[0086] In one specific embodiment, the network structure of the image generation network can be set in combination with actual application. Optionally, the image generation network can be a UNet network (denoising network) in a diffusion model.
[0087] In the above embodiments, before the fusion of multiple features, a pose feature extraction network is separately added for pose feature extraction processing, and the fused features are directly unified into the image generation network using a linear projection layer, which can increase the fine control of the pose feature and greatly improve the image generation performance of the model.
[0088] In an optional embodiment, the preset image generation model can further include a cross-attention network, and correspondingly, the method can further include:
[0089] The second image fusion feature is input into the cross-attention network for cross-attention learning to obtain an attention learning feature;
[0090] Correspondingly, the image generation processing of the second image fusion feature into the image generation network to obtain the target object image includes:
[0091] The attention learning feature is input into the image generation network for image generation processing to obtain the target object image.
[0092] In a specific embodiment, the image generation network can include a multi-layer network connected in sequence, and on the basis of taking the output of the previous layer network of the current layer network as the input of the current layer network, the second image fusion feature can also be taken as the input of the current layer network. Optionally, in the process of taking the second image fusion feature as the input of the current layer network, the cross-attention network can be combined to perform cross-attention learning on the second image fusion feature, and then the attention learning feature is taken as the input of the current layer network, so as to better improve the representation accuracy of the learned feature to the image generation requirement. Specifically, the attention learning feature can be the feature obtained by performing cross-attention learning on the second image fusion feature.
[0093] In a specific embodiment, the preset image generation model can be obtained by performing image generation training on a to-be-trained model based on a plurality of original sample object images and training data corresponding to each original sample object image. Optionally, as shown in Figure 4 The preset image generation model is trained in the following manner:
[0094] S401: Obtain a plurality of original sample object images and training data corresponding to each original sample object image;
[0095] S403: Determine a current sample object image from the plurality of original sample object images;
[0096] S405: Determine a first current object feature corresponding to the current sample object image, a first current style feature corresponding to the current sample object image, and a first current description feature corresponding to the current sample object image;
[0097] S407: Fuse the first current object feature and the first current style feature to obtain a second current style feature;
[0098] S409: Fuse the first current object feature and the first current description feature to obtain a second current description feature;
[0099] S411: input the second current style feature, the second current description feature, and the preset pose image corresponding to the current sample object image into the to-be-trained model for image generation processing to obtain a predicted object image corresponding to the current sample object image;
[0100] S413: train the to-be-trained model based on the predicted object image corresponding to the current sample object image and the preset object image corresponding to the current sample object image to obtain a preset image generation model.
[0101] In one specific embodiment, the training data corresponding to each original sample object image includes: a preset style clothing image corresponding to each original sample object image, a preset pose image corresponding to each original sample object image, a preset image description text corresponding to each original sample object image, and a preset object image corresponding to each original sample object image; a plurality of preset style clothing images corresponding to a plurality of original sample object images, a plurality of preset pose images, and a plurality of sample image description texts. Specifically, the plurality of preset style clothing images can include a target style clothing image; the plurality of preset pose images can include a target pose image.
[0102] In one specific embodiment, the preset image description text corresponding to each original sample object image is used to indicate an image of generating an object corresponding to the original sample object image, wearing style clothing (style clothing corresponding to the corresponding preset style clothing image) in a pose (pose corresponding to the corresponding preset pose image) corresponding to the original sample object image.
[0103] In one specific embodiment, the current sample object image can be a sample object image corresponding to a current training round; specifically, the current sample object image can include a plurality of sample object images. Specifically, the first current object feature corresponding to the current sample object image can be the representation information (object feature) of the preset object image corresponding to the current sample object image; the first current style feature corresponding to the current sample object image can be the representation information (style feature) of the preset style clothing image corresponding to the current sample object image; and the first current description feature corresponding to the current sample object image can be the representation information (description feature) of the sample image description text corresponding to the current sample object image.
[0104] In one specific embodiment, the specific refinement of determining the first current object feature corresponding to the current sample object image, the first current style feature corresponding to the current sample object image, and the first current description feature corresponding to the current sample object image can refer to the specific refinement of determining the first object feature corresponding to the original object image, the first style feature corresponding to the target style clothing image, and the first description feature corresponding to the target image description text, which will not be described here.
[0105] In a specific embodiment, the fusion processing of the first current object feature and the first current style feature can include splicing processing of the first current object feature and the first current style feature to obtain a second current style feature.
[0106] In a specific embodiment, the fusion processing of the first current object feature and the first current description feature can include splicing processing of the first current object feature and the first current description feature to obtain a second current description feature.
[0107] In a specific embodiment, the specific refinement of the inputting of the second current style feature, the second current description feature and the preset pose image corresponding to the current sample object image into the to-be-trained model for image generation processing to obtain the predicted object image corresponding to the current sample object image can refer to the specific refinement of the inputting of the second style feature, the second description feature and the target pose image into the preset image generation model for image generation processing to obtain the target object corresponding to the target object image, which will not be described here.
[0108] In a specific embodiment, the training of the to-be-trained model based on the predicted object image corresponding to the current sample object image and the preset object image corresponding to the current sample object image to obtain the preset image generation model can include: determining an image generation loss based on the predicted object image corresponding to the current sample object image and the preset object image corresponding to the current sample object image; updating the model parameters of the to-be-trained model based on the image generation loss; performing the next round of iterative training (i.e., repeating the steps of determining the current sample object image from the plurality of original sample object images to updating the model parameters of the to-be-trained model based on the image generation loss) based on the updated to-be-trained model until a preset convergence condition is met, and taking the to-be-trained model corresponding to the time when the preset convergence condition is met as the preset image generation model.
[0109] In a specific embodiment, the image generation loss can reflect the image generation performance of the preset image generation model. Specifically, the smaller the difference between the predicted object image corresponding to the current sample object image and the preset object image corresponding to the current sample object image, the smaller the image generation loss. Conversely, the larger the difference between the predicted object image corresponding to the current sample object image and the preset object image corresponding to the current sample object image, the larger the image generation loss. Specifically, the image generation loss can be determined in combination with a loss function that can calculate the difference between the predicted object image corresponding to the current sample object image and the preset object image corresponding to the current sample object image. Specifically, the loss function can be set in combination with actual application.
[0110] In a specific embodiment, the gradient descent method can be combined in the process of updating the model parameters of the to-be-trained model based on the image generation loss.
[0111] In a specific embodiment, the preset convergence condition can be set in combination with actual application, for example, the number of iterations of the execution of the training reaches a preset number, the image generation loss is less than a specified threshold, etc., which can be set in combination with the training speed and the model accuracy requirement.
[0112] In the above embodiment, in the preset image generation model training process, the second current style feature fused with the object feature, the second current description feature fused with the object feature, and the preset pose image corresponding to the current sample object image are input into the to-be-trained model for image generation processing, which can make the style feature and the description feature of the image incorporate the features of the object subject, and in the preset image generation model training process, the extraction of the style feature and the description feature is not required, and the model parameters in the extraction model of the style feature and the description feature do not need to be updated, which can effectively reduce the calculation amount, on the basis of effectively preserving the generation capability of the existing model, in combination with the image generation model training, on the basis of improving the representation accuracy of the generated image on the subject object, the understanding capability of the model on the image generation requirement is improved, the matching between the generated image and the user requirement is ensured, and thus the personalized image generation effect can also be greatly improved.
[0113] In a specific embodiment, as shown in Figure 5 , Figure 5 is a schematic diagram of image generation processing based on a preset image generation model provided by an embodiment of the present application; specifically, in combination with Figure 5 It can be seen that the target style clothing image can be input into the image feature extraction model for image feature extraction processing to obtain the first style feature, and the original object image can be input into the object feature extraction model for object feature extraction processing to obtain the first object feature; and the target image description text can be input into the text feature extraction model for text feature extraction processing to obtain the first description feature; then, the first object feature and the first style feature are input into the second feature fusion network for fusion processing to obtain the second style feature; and the first object feature and the first description feature are input into the third feature fusion network for fusion processing to obtain the second description feature; in addition, the target pose image is input into the pose feature extraction network in the preset image generation model for pose feature extraction processing to obtain the first pose feature; then, the first pose feature, the second style feature fused with the object feature, and the second description feature fused with the object feature are input into the first feature fusion network in the preset image generation model for fusion processing to obtain the first fusion feature; and the image generation processing is performed through the linear projection network in the preset image generation model and the image generation network in the preset image generation model to obtain the target object image personalized for the target object.
[0114] It can be seen from the technical solutions provided by the embodiments of the present specification that, in the image generation process, the present specification acquires an original object image, a target style clothing image, a target pose image, and a target image description text indicating an image of a target object corresponding to the original object image, wearing a target style clothing corresponding to the target style clothing image, in a target pose corresponding to the target pose image; and first object features corresponding to the original object image, first style features corresponding to the target style clothing image, and first description features corresponding to the target image description text are determined first; the first object features and the first style features are fused to obtain second style features, and the first object features and the first description features are fused to obtain second description features; then, the second style features fused with the object features, the second description features fused with the object features, and the target pose image are input into a preset image generation model for image generation processing to obtain a target object image corresponding to the target object, so that the second style features fused with the object features and the second description features fused with the object features can be combined to improve the accuracy of the model in representing the subject object in the image generation process, and the image pose can be combined on the basis of the style features and the description features of the image to accurately control the clothing and the pose of the person, enhance the expression ability of individualization, improve the learning performance of the model on the pose, the subject object, and the style, and further improve the understanding ability of the model on the image generation demand and the user intention, so as to ensure the matching between the generated image and the user intention, and further greatly improve the effect of individualized image generation.
[0115] In addition, the image generation method provided by the embodiments of the present application can well create personalized clothing pictures of the user's own personalized clothing and pose, the specific personalized pose and clothing pictures can enhance the user's ability to express individual emotions, meet the user's demand for individualization, better create the user's social network image, make it easier to find friends, expand social relationships, effectively increase the user stickiness of the social network and expand the influence range, and use the end-to-end image generation model to make the association between the user's demand description and the final result more direct, so that the image can be generated again, the image creation threshold and cost are reduced, the original object image can be presented in a new way, the personalized object image can be generated without special photo shooting and PS image processing, the interestingness of social entertainment is improved, the user stickiness to the social network is increased, and the diversity of image content distribution in the social network is increased.
[0116] The embodiments of the present application also provide an image generation device, as shown in Figure 6 The device includes:
[0117] The information acquisition module 610 is configured to acquire an original object image, a target style clothing image, a target pose image, and a target image description text, the target image description text being used to indicate a target object corresponding to the original object image, a target pose corresponding to the target pose image, and a target style clothing corresponding to the target style clothing image.
[0118] The first feature determination module 620 is configured to determine a first object feature corresponding to the original object image, a first style feature corresponding to the target style clothing image, and a first description feature corresponding to the target image description text.
[0119] The first fusion processing module 630 is configured to perform fusion processing on the first object feature and the first style feature to obtain a second style feature.
[0120] The second fusion processing module 640 is configured to perform fusion processing on the first object feature and the first description feature to obtain a second description feature.
[0121] The first image generation processing module 650 is configured to input the second style feature, the second description feature, and the target pose image into a preset image generation model to perform image generation processing to obtain a target object image corresponding to the target object.
[0122] In an optional embodiment, the preset image generation model includes a pose feature extraction network, a first feature fusion network, a linear projection network, and an image generation network.
[0123] The first image generation processing module 650 includes:
[0124] The pose feature extraction unit is configured to input the target pose image into the pose feature extraction network to perform pose feature extraction processing to obtain a first pose feature.
[0125] The fusion processing unit is configured to input the first pose feature, the second style feature, and the second description feature into the first feature fusion network to perform fusion processing to obtain a first image fusion feature.
[0126] The linear projection processing unit is configured to input the first image fusion feature into the linear projection network to perform linear projection processing to obtain a second image fusion feature.
[0127] The image generation processing unit is configured to input the second image fusion feature into the image generation network to perform image generation processing to obtain the target object image.
[0128] In an optional embodiment, the preset image generation model further comprises a cross-attention network; and the device further comprises:
[0129] a cross-attention learning module configured to perform cross-attention learning on the second image fusion feature input into the cross-attention network, to obtain an attention learning feature;
[0130] the image generation processing unit is further configured to perform image generation processing on the attention learning feature input into the image generation network, to obtain the target object image.
[0131] In an optional embodiment, the first feature determination module 620 comprises:
[0132] a text feature extraction configured to perform text feature extraction processing on the target image description text input into a text feature extraction network, to obtain an initial description feature;
[0133] a prior feature conversion unit configured to perform processing on the initial description feature input into a prior feature conversion network, to obtain the first description feature; the first description feature is a prior feature in the process of generating an image by the preset image generation model.
[0134] In an optional embodiment, the information acquisition module 610 comprises:
[0135] a preset object image acquisition unit configured to perform acquisition of a preset object image containing the target object;
[0136] an object region image cropping unit configured to perform cropping of an object region image from the preset object image;
[0137] an enhancement processing unit configured to perform enhancement processing on the object region image, to obtain the original object image.
[0138] In an optional embodiment, the target posture image is acquired by using the following modules:
[0139] a target sample object image acquisition module configured to perform acquisition of a target sample object image containing an object with a posture being the target posture;
[0140] a posture image generation processing module configured to perform posture image generation processing on the target sample object image input into a preset posture generation model, to obtain the target posture image.
[0141] In an optional embodiment, the preset image generation model is obtained by training using the following modules:
[0142] a data acquisition module configured to perform acquisition of a plurality of original sample object images and training data corresponding to each original sample object image, the training data corresponding to each original sample object image including: a preset style clothing image corresponding to each original sample object image, a preset posture image corresponding to each original sample object image, a preset image description text corresponding to each original sample object image, and a preset object image corresponding to each original sample object image; a plurality of preset style clothing images, a plurality of preset posture images, and a plurality of sample image description texts corresponding to the plurality of original sample object images;
[0143] a current sample object image determination module configured to perform determination of a current sample object image from the plurality of original sample object images;
[0144] a second feature determination module configured to perform determination of a first current object feature corresponding to the current sample object image, a first current style feature corresponding to the current sample object image, and a first current description feature corresponding to the current sample object image;
[0145] a third fusion processing module configured to perform fusion processing of the first current object feature and the first current style feature to obtain a second current style feature;
[0146] a fourth fusion processing module configured to perform fusion processing of the first current object feature and the first current description feature to obtain a second current description feature;
[0147] a second image generation processing module configured to perform image generation processing of the second current style feature, the second current description feature, and a preset posture image corresponding to the current sample object image on a to-be-trained model to obtain a predicted object image corresponding to the current sample object image;
[0148] a model training module configured to perform training of the to-be-trained model based on the predicted object image corresponding to the current sample object image and a preset object image corresponding to the current sample object image to obtain the preset image generation model.
[0149] As to the apparatus in the above-described embodiments, the specific manners in which various modules perform operations have been described in details in the embodiments of the method, and will not be described in details here.
[0150] Figure 7 is a block diagram of an electronic device for image generation provided by an embodiment of the present application. The electronic device can be a terminal, and its internal structure diagram can be as shown in Figure 7As shown in the figure. The electronic device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capability. The memory of the electronic device includes non-volatile storage medium, internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network interface of the electronic device is used to communicate with the external terminal through the network connection. The computer program is executed by the processor to implement an image generation method. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the electronic device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the electronic device, or an external keyboard, touchpad or mouse, etc.
[0151] Figure 8 The figure shows another block diagram of an electronic device for image generation provided by the embodiments of the present application. The electronic device can be a server, and its internal structure diagram can be as shown in the figure. Figure 8 The electronic device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capability. The memory of the electronic device includes non-volatile storage medium, internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network interface of the electronic device is used to communicate with the external terminal through the network connection. The computer program is executed by the processor to implement an image generation method.
[0152] Those skilled in the art can understand that Figure 7 or Figure 8 The structure shown in the figure is only a block diagram of part of the structure related to the present disclosure scheme, and does not constitute a limitation on the electronic device to which the present disclosure scheme is applied. The specific electronic device can include more or less components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0153] In an exemplary embodiment, an electronic device is also provided, including: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement an image generation method as in the embodiments of the present disclosure.
[0154] In an exemplary embodiment, a computer readable storage medium is also provided, when the instructions in the storage medium are executed by the processor of the electronic device, the electronic device can execute the image generation method in the embodiments of the present disclosure.
[0155] In an example embodiment, a computer program product or computer program is also provided, which includes computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions, so that the computer device performs the image generation method provided in the various optional implementations described above.
[0156] It can be understood that, in the specific embodiments of the present application, data related to users is involved, and when the above embodiments of the present application are applied to specific products or technologies, the permission or consent of the users needs to be obtained, and the collection, use and processing of the related data need to comply with the relevant laws, regulations and standards of the countries and regions.
[0157] It can be understood by those skilled in the art that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the computer program can include the processes of the above-mentioned embodiments of the methods. In the embodiments of the present application, any reference to memory, storage, database or other medium can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0158] Other embodiments of the present disclosure will be apparent to those skilled in the art with the consideration of the specification and practice of the disclosure disclosed herein. The present application is intended to cover any variations, uses or adaptive changes of the present disclosure that follow the general principles of the present disclosure and include common knowledge or conventional technical means in the technical field of the present disclosure not disclosed by the present disclosure. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present disclosure are indicated by the claims.
[0159] It should be understood that the present disclosure is not limited to the precise construction that has been described above and shown in the accompanying drawings, and that various modifications and changes can be made by those skilled in the art without departing from the scope of the present disclosure. The scope of the present disclosure is limited only by the appended claims.
Claims
1. An image generation method, characterized in that, The method includes: Acquire an original object image, a target style clothing image, a target pose image, and a target image description text. The target image description text is used to indicate the generation of a target object corresponding to the original object image, and an image of a target style clothing corresponding to the target style clothing image in the target pose corresponding to the target pose image. Determine the first object feature corresponding to the original object image, the first style feature corresponding to the target style clothing image, and the first description feature corresponding to the target image description text; The first object feature and the first style feature are fused together to obtain the second style feature; The first object feature and the first description feature are fused together to obtain the second description feature; The second style feature, the second descriptive feature, and the target pose image are input into a preset image generation model for image generation processing to obtain the target object image corresponding to the target object.
2. The method according to claim 1, characterized in that, The preset image generation model includes: a pose feature extraction network, a first feature fusion network, a linear projection network, and an image generation network; The step of inputting the second style feature, the second descriptive feature, and the target pose image into a preset image generation model for image generation processing to obtain the target object image corresponding to the target object includes: The target pose image is input into the pose feature extraction network for pose feature extraction processing to obtain the first pose feature; The first pose feature, the second style feature, and the second descriptive feature are input into the first feature fusion network for fusion processing to obtain the first image fusion feature; The first image fusion feature is input into a linear projection network for linear projection processing to obtain the second image fusion feature; The second image fusion feature is input into the image generation network for image generation processing to obtain the target object image.
3. The method according to claim 2, characterized in that, The preset image generation model further includes: a cross-attention network; the method further includes: The second image fusion feature is input into the cross-attention network for cross-attention learning to obtain attention learning features; The step of inputting the second image fusion feature into the image generation network for image generation processing to obtain the target object image includes: The attention learning features are input into the image generation network for image generation processing to obtain the target object image.
4. The method according to claim 1, characterized in that, The first descriptive feature is determined in the following manner: The target image description text is input into a text feature extraction network for text feature extraction processing to obtain initial description features; The initial descriptive features are input into a prior feature transformation network for processing to obtain the first descriptive features; the first descriptive features are prior features in the process of generating images by the preset image generation model.
5. The method according to claim 1, characterized in that, The process of obtaining the original object image includes: Obtain a preset object image containing the target object; The object region image is cropped from the preset object image; The object region image is enhanced to obtain the original object image.
6. The method according to claim 1, characterized in that, The target pose image is obtained using the following method: Obtain a target sample object image, wherein the target sample object image contains an object whose pose is the target pose; The target sample object image is input into a preset pose generation model for pose image generation processing to obtain the target pose image.
7. The method according to any one of claims 1 to 6, characterized in that, The preset image generation model is trained in the following manner: Acquire multiple original sample object images and training data corresponding to each original sample object image. The training data corresponding to each original sample object image includes: a preset style clothing image, a preset pose image, a preset image description text, and a preset object image; and multiple preset style clothing images, multiple preset pose images, and multiple sample image description texts corresponding to the multiple original sample object images. The current sample object image is determined from the plurality of original sample object images; Determine the first current object feature, the first current style feature, and the first current description feature corresponding to the current sample object image; The first current object feature and the first current style feature are fused together to obtain the second current style feature; The first current object feature and the first current description feature are fused together to obtain the second current description feature; The second current style feature, the second current description feature, and the preset pose image corresponding to the current sample object image are input into the model to be trained for image generation processing to obtain the predicted object image corresponding to the current sample object image. The model to be trained is trained based on the predicted object image corresponding to the current sample object image and the preset object image corresponding to the current sample object image to obtain the preset image generation model.
8. An image generation apparatus, characterized in that, The device includes: The information acquisition module is configured to acquire an original object image, a target style clothing image, a target pose image, and a target image description text. The target image description text is used to indicate the generation of a target object corresponding to the original object image, and an image of a target object wearing the target style clothing corresponding to the target style clothing image in the target pose corresponding to the target pose image. The first feature determination module is configured to determine the first object feature corresponding to the original object image, the first style feature corresponding to the target style clothing image, and the first description feature corresponding to the target image description text. The first fusion processing module is configured to perform a fusion processing of the first object feature and the first style feature to obtain a second style feature; The second fusion processing module is configured to perform a fusion processing of the first object feature and the first descriptive feature to obtain the second descriptive feature; The first image generation and processing module is configured to perform image generation processing by inputting the second style feature, the second descriptive feature and the target pose image into a preset image generation model to obtain the target object image corresponding to the target object.
9. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the image generation method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the image generation method as described in any one of claims 1 to 7.