Image generation method and apparatus, image generation model training method and apparatus, and device, medium and program product

By fusing source images, noise images and preset posture images, using feature extractors and denoisers in the image generation model, combined with cross attention and diffusion models, the consistency and stability problems in image generation technology are solved, and higher quality image generation is achieved.

WO2025152951A1PCT designated stage expired Publication Date: 2025-07-24TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Application Number
PCT/CN2025/072426
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-17
Filing Date
2025-01-15
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

The existing image generation technology is unstable during the adversarial training process, poor image generation consistency, difficult to maintain real textures, and difficult to deal with complex deformation and occlusion, which limits the further development of image generation technology.

Method used

By acquiring the source image, noise image and preset posture image, the image feature extractor and denoiser in the image generation model are fused, and combined with the cross attention and diffusion model, the target image is generated to improve image consistency and generation accuracy.

Benefits of technology

It improves the stability and consistency of the image generation model, enhances the accuracy and nature of image generation, and improves the robustness and generalization ability of image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025072426_24072025_PF_FP_ABST
    Figure CN2025072426_24072025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the embodiments of the present application are an image generation method and apparatus, an image generation model training method and apparatus, and a device, a medium and a program product. The image generation method comprises: acquiring a source image, a noise image and a preset pose image, wherein the noise image is different from the source image; extracting a source pose image from the source image; at least fusing the source image, the noise image, the preset pose image and the source pose image to obtain an image to be processed; inputting the source image into an image feature extractor in an image generation model to extract a comprehensive feature of the source image; and inputting the image to be processed and the comprehensive feature of the source image into an image denoiser in the image generation model to obtain a target image, wherein the target image represents an image of an object in the source image in a preset pose.
Need to check novelty before this filing date? Find Prior Art

Description

Image generation method, image generation model training method, device, equipment, medium and program product

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on January 17, 2024, with application number 202410066677.9 and application name “Image Generation Method and Device”. Technical Field

[0002] The present application relates to the field of computer technology, and more specifically to an image generation method, an image generation model training method, an apparatus, a device, a medium, and a program product.

[0003] Background of the Invention

[0004] In the era of rapid development of Internet technology, with the development of deep learning, image generation technology has been widely used in many fields, such as e-commerce advertising, artistic creation, game design, virtual reality, etc.

[0005] However, related technologies often employ adversarial training, which can lead to unstable training processes, poor image generation consistency, and inconsistency. These approaches may not always preserve realistic textures, require dense correspondences, and struggle to handle complex deformations and severe occlusions. These issues and shortcomings have limited the further development of image generation technology. Summary of the Invention

[0006] In view of this, the present application provides an image generation method, an image generation model training method, an apparatus, a device, a medium and a program product, so as to alleviate, mitigate or even eliminate some or all of the above-mentioned problems and other possible problems.

[0007] In one aspect, an embodiment of the present application provides an image generation method, which is performed by an electronic device. The method includes:

[0008] Acquire a source image, a noise image, and a preset posture image, wherein the noise image is different from the source image;

[0009] extracting a source pose image from the source image;

[0010] fusing at least the source image, the noise image, the preset posture image, and the source posture image to obtain an image to be processed;

[0011] Inputting the source image into the image feature extractor in the image generation model to extract the comprehensive features of the source image; and

[0012] The comprehensive features of the image to be processed and the source image are input into the image denoiser in the image generation model to obtain a target image, where the target image represents an image of the object in the source image in a preset posture.

[0013] On the other hand, an embodiment of the present application provides an image generation model training method, which is performed by an electronic device, and the method includes:

[0014] Obtaining a source image sample, a target image sample, a preset pose image sample, a noise image sample, and a source pose image sample, wherein the noise image sample is different from the source image sample;

[0015] fusing at least the source image sample, the noise image sample, the preset posture image sample, the source posture image sample, and the target image sample to obtain an image sample to be processed;

[0016] Fusing the source image sample and the target image sample to obtain a target image label;

[0017] Inputting the source image sample into the image feature extractor in the image generation model to extract the comprehensive features of the source image sample;

[0018] Inputting the comprehensive features of the image sample to be processed and the source image sample into the image denoiser in the image generation model to obtain a denoised image;

[0019] Calculating a target loss based on the target image label and the denoised image; and

[0020] Based on the target loss, the parameters of the image generation model are iteratively updated until the target loss meets a preset condition.

[0021] On the other hand, an embodiment of the present application provides an image generating device, including:

[0022] A first acquisition module is configured to acquire a source image, a noise image, and a preset posture image, wherein the noise image is different from the source image;

[0023] A first extraction module, configured to extract a source posture image from the source image;

[0024] A first fusion module is configured to fuse at least the source image, the noise image, the preset posture image, and the source posture image to obtain an image to be processed;

[0025] A second extraction module is configured to input the source image into an image feature extractor in an image generation model to extract comprehensive features of the source image; and

[0026] The output module is used to input the comprehensive features of the image to be processed and the source image into the image denoiser in the image generation model to obtain a target image, where the target image represents the image of the object in the source image in a preset posture.

[0027] On the other hand, an embodiment of the present application proposes an image generation model training device, comprising:

[0028] A second acquisition module is configured to acquire a source image sample, a target image sample, a preset posture image sample, a noise image sample, and a source posture image sample, wherein the noise image sample is different from the source image sample;

[0029] A second fusion module is configured to fuse at least the source image sample, the noise image sample, the preset posture image sample, the source posture image sample, and the target image sample to obtain an image sample to be processed;

[0030] A third fusion module, configured to fuse the source image sample with the target image sample to obtain a target image label;

[0031] A third extraction module is used to input the source image sample into the image feature extractor in the image generation model to extract the comprehensive features of the source image sample;

[0032] a denoising module, configured to input the comprehensive features of the image sample to be processed and the source image sample into an image denoiser in the image generation model to obtain a denoised image;

[0033] a calculation module, configured to calculate a target loss based on the target image label and the denoised image; and

[0034] An iterative module is used to iteratively update the parameters of the image generation model based on the target loss until the target loss meets a preset condition.

[0035] On the other hand, an embodiment of the present application proposes an electronic device comprising: a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor is prompted to execute the image generation method and the image generation model training method according to some embodiments of the present application.

[0036] On the other hand, an embodiment of the present application proposes a computer-readable storage medium on which computer-readable instructions are stored. When the computer-readable instructions are executed, they implement the image generation method and image generation model training method according to some embodiments of the present application.

[0037] On the other hand, an embodiment of the present application proposes a computer program product, including a computer program, which, when executed by a processor, implements the image generation method and image generation model training method according to some embodiments of the present application.

[0038] BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The various aspects, features and advantages of the present application will be readily understood from the following detailed description and accompanying drawings, in which:

[0040] FIG1 schematically shows an example implementation environment of an image generation method according to some embodiments of the present application;

[0041] FIG2 schematically shows a flow chart of an image generation method according to some embodiments of the present application;

[0042] 3A-3C schematically illustrate examples of a source image, a preset pose image, and a target image according to some embodiments of the present application;

[0043] FIG4 schematically shows a principle diagram of an image generation method according to some embodiments of the present application;

[0044] FIG5 schematically shows a flow chart of an image generation method according to some embodiments of the present application;

[0045] FIG6 schematically shows a flow chart of an image generation method according to some embodiments of the present application;

[0046] FIG7 schematically shows a principle diagram of an image generation method according to some embodiments of the present application;

[0047] FIG8 schematically shows a flow chart of an image generation model training method according to some embodiments of the present application;

[0048] FIG9A schematically shows an example block diagram of an image generating apparatus according to some embodiments of the present application;

[0049] FIG9B schematically shows an example block diagram of an image generation model training apparatus according to some embodiments of the present application; and

[0050] FIG10 schematically shows an example block diagram of an electronic device according to some embodiments of the present application.

[0051] It should be noted that the above-mentioned drawings are merely schematic and illustrative and are not necessarily drawn to scale.

[0052] Implementation Method

[0053] Several embodiments of the present application will be described in more detail below with reference to the accompanying drawings so that those skilled in the art can implement the present application. The present application can be embodied in many different forms and for many different purposes and should not be limited to the embodiments described in the various embodiments of the present application. These embodiments are provided to make the present application comprehensive and complete and to fully convey the scope of the present application to those skilled in the art. The embodiments do not limit the present application.

[0054] It will be understood that, although the terms first, second, third, etc. may be used to describe various elements, components, and / or parts in various embodiments of the present application, these elements, components, and / or parts should not be limited by these terms. These terms are only used to distinguish one element, component, or part from another element, component, or part. Therefore, the first element, component, or part discussed below may be referred to as the second element, component, or part without departing from the teachings of the present application.

[0055] The terms used in the various embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the various embodiments of the present application, the singular forms "one", "one" and "the" are intended to also include plural forms, unless the context clearly indicates otherwise. It will be further understood that the terms "include" and / or "comprise" specify the existence of the features, integral bodies, steps, operations, elements and / or parts when used in this specification, but do not exclude the existence of one or more other features, integral bodies, steps, operations, elements, parts and / or their groups or add one or more other features, integral bodies, steps, operations, elements, parts and / or their groups. As used in the various embodiments of the present application, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0056] Unless otherwise defined, all terms (including technical and scientific terms) used in the embodiments of the present application have the same meaning as commonly understood by those skilled in the art to which the present application belongs. It will be further understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the relevant field and / or the context of this specification, and will not be interpreted in an idealized or overly formal sense unless explicitly defined in the embodiments of the present application.

[0057] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically separate entities. That is, these functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0058] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all promotional information and operations / steps, nor do they necessarily require execution in the order described. For example, some operations / steps may be broken down, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.

[0059] Those skilled in the art will understand that the drawings are merely example interfaces of example embodiments, and the modules or processes in the drawings are not necessarily necessary for implementing the present application, and therefore cannot be used to limit the scope of protection of the present application.

[0060] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to achieve the best results.

[0061] Machine Learning (ML) specializes in studying how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance.

[0062] Computer Vision (CV) technology refers to the use of cameras and computers to replace the human eye to identify and measure targets, and further perform graphic processing to make the computer processing into images more suitable for human observation or transmission to instrument detection.

[0063] Before introducing the embodiments of the present application in detail, for the sake of clarity, the concepts used in all the embodiments of the present application are first explained.

[0064] 1. Diffusion Model: The diffusion model is a probabilistic generative model for modeling and generating data, which can be used for image generation and transformation tasks. The diffusion model is based on a stochastic process and generates samples by iteratively propagating noise into the data space. The diffusion model is motivated by non-equilibrium thermodynamics and defines a Markov chain that slowly adds noise to the input sample (forward propagation) and then reconstructs the desired sample from the noise (backward propagation). Through a series of forward and backward diffusion steps, the diffusion model can learn reasonable transfer trajectories, rather than simulating complex feature transfers in a single process.

[0065] 2. Posture-guided Image Generation: Posture-guided image generation aims to render images of people or objects with a desired posture and appearance. Specifically, the appearance is defined by a given source image, and the posture is defined by a set of key points. In the embodiments of this application, posture generally refers to the posture of a person or object, including its position, posture, and angle.

[0066] 3. Cross-Attention: Cross-Attention is an attention mechanism used to establish correlations between two different input sequences. It compares each element in one sequence with all elements in the other sequence, thereby calculating the correlation between the two sequences. It can be used to handle multimodal tasks, such as image description generation.

[0067] 4. Self-Attention: Self-attention is an attention mechanism that establishes associations between different positions in an input sequence. It determines the importance of each element by calculating its similarity with other elements in the sequence. Self-attention can effectively capture long-range dependencies in a sequence and is not limited by sequence length.

[0068] 5. A feedforward neural network (FNN) is an artificial neural network model composed of multiple neural network layers. Each neural network layer contains multiple neurons. Each neuron receives the output of the previous layer, adds a weighted sum to it, performs a nonlinear transformation using an activation function, and then passes the result to the next layer. A feedforward neural network layer typically includes fully connected layers, convolutional layers, pooling layers, normalization layers, and recurrent layers.

[0069] 6. Generalization: This refers to a machine learning algorithm's ability to adapt to new samples. Simply put, it's the ability to train a machine learning algorithm to produce a reasonable output by adding a new dataset to the original one. The goal of learning is to learn the patterns underlying the data. This ability is called generalization, as it allows the trained network to produce appropriate outputs for data outside of the training set that exhibits the same patterns.

[0070] In related technologies, adversarial training is used. For example, the training process of generative adversarial networks (GANs) is relatively unstable. If it is based on a diffusion model, the consistency between the source image and the final generated image cannot be guaranteed, the generalization ability is weak, and the image generation accuracy is low.

[0071] An embodiment of the present application provides an image generation method that not only introduces source image features into an image generation model through cross-attention, but also introduces the source image and a preset posture image into the image generation model through splicing and masking, allowing the image generation model to more easily synthesize a consistent image, thereby improving the consistency between the source image and the final generated image.

[0072] Figure 1 schematically illustrates an example implementation environment 100 for image generation methods according to some embodiments of the present application. As shown in Figure 1, implementation environment 100 may include a terminal device 110, a server 120, and a network 130 for connecting terminal device 110 and server 120. In some embodiments, terminal device 110 may be used to implement the image generation methods according to the present application. For example, terminal device 110 may be deployed with corresponding programs or instructions for executing the various methods provided herein. Optionally, server 120 may also be used to implement the various methods according to the present application.

[0073] The terminal device 110 and the third-party terminal device 140 can be any type of mobile electronic device, including a mobile computer (e.g., a personal digital assistant (PDA), a laptop computer, a notebook computer, a tablet computer, a netbook computer, etc.), a mobile phone (e.g., a cellular phone, a smartphone, etc.) as shown in FIG1 , a wearable electronic device (e.g., a smart watch, a head-mounted device, including smart glasses), or other types of mobile devices. In some embodiments, the terminal device 110 can also be a stationary electronic device, such as a desktop computer, a game console, a smart TV, etc.

[0074] The server 120 may be a single server or a server cluster, or may be a cloud server or cloud server cluster that can provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. It should be understood that the servers mentioned in the various embodiments of the present application are typically server computers with large amounts of memory and processor resources, but other embodiments are also possible. Alternatively, the server 120 may also be an ordinary desktop computer, which includes a host, a display, and the like.

[0075] Examples of the network 130 include a local area network (LAN), a wide area network (WAN), a personal area network (PAN), and / or a combination of communication networks such as the Internet. The server 120 and the terminal device 110 may include at least one communication interface (not shown) capable of communicating via the network 130. Such a communication interface may be one or more of the following: any type of network interface (e.g., a network interface card (NIC)), a wired or wireless (such as an IEEE 802.11 wireless LAN (WLAN)) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth™ interface, a Near Field Communication (NFC) interface, etc.

[0076] As shown in Figure 1, the terminal device 110 may include a display screen 111 and the terminal user may interact with the terminal application 112 via the display screen 111. The terminal device 110 may, for example, interact with the server 120 via the network 130, for example, by sending data to it or receiving data from it. The terminal application 112 may be a local application, a web application, or a mini-program (LiteApp, such as a mobile app or WeChat app) as a lightweight application. In the case where the terminal application is a local application that needs to be installed, the terminal application 112 may be installed in the terminal device 110. In the case where the terminal application 112 is a web application, the terminal application 112 may be accessed through a browser. In the case where the terminal application 112 is a mini-program, the terminal application 112 may be opened directly on the terminal device 110 by searching for relevant information of the terminal application (such as the name of the terminal application, etc.), scanning a graphic code (such as a barcode, a QR code, etc.) of the terminal application, etc., without having to install the terminal application 112.

[0077] The example implementation environment of Figure 1 is merely illustrative, and the image generation method according to the embodiments of the present application is not limited to the example implementation environment and server shown. It should be understood that although the server 120 and the terminal device 110 are shown and described as separate structures in the various embodiments of the present application, they can be different components of the same electronic device. Optionally, all steps of the image generation method according to some embodiments of the present application can also be implemented on the server 120 side, or can also be implemented on both the terminal device 110 side and the server 120 side.

[0078] FIG2 schematically illustrates a flow chart of an image generation method according to some embodiments of the present application. The image generation method is performed by an electronic device. In some embodiments, as shown in FIG1 , the image generation method according to an embodiment of the present application can be performed on the terminal device 110. In other embodiments, the image generation method according to an embodiment of the present application can also be performed in combination by the server 120 and the terminal device 110.

[0079] As shown in FIG2 , the image generation method according to some embodiments of the present application may include steps S210 to S250:

[0080] S210, acquiring a source image, a noise image, and a preset posture image, wherein the noise image is different from the source image;

[0081] S220, extracting a source posture image from the source image;

[0082] S230, fusing at least the source image, the noise image, the preset posture image, and the source posture image to obtain an image to be processed;

[0083] S240, inputting the source image into the image feature extractor of the image generation model to extract comprehensive features of the source image;

[0084] S250, inputting the comprehensive features of the image to be processed and the source image into the image denoiser of the image generation model to obtain a target image, which represents the image of the object in the source image in a preset posture.

[0085] Steps S210-S250 are described in detail below in conjunction with Figures 3A-3C and Figure 4, where Figures 3A-3C respectively schematically illustrate examples of a source image, a preset posture image, and a target image according to some embodiments of the present application, and Figure 4 schematically illustrates a principle diagram of an image generation method according to some embodiments of the present application.

[0086] At step S210, a source image, a noise image, and a preset pose image may be obtained from an image library, wherein the noise image is different from the source image. In various embodiments of the present application, the source image may refer to an original reference image to be processed that includes various image features, an example of which is shown in FIG3A . These source images may be obtained from a relevant platform or system and stored in an image library.

[0087] A noise image may refer to an image that is different from the source image and is mainly used for stitching with the source image because the target image is unknown at this time. These noise images can be generated in advance and saved in an image library.

[0088] A preset posture image may refer to an image used to display a preset posture (e.g., information such as predetermined joint positions, skeletal structure, and posture angles), an example of which is shown in FIG3B . In the embodiment of the present application, each posture is pre-set, and a corresponding posture image is generated for each posture and stored in an image library.

[0089] The target image refers to an image of the object in the source image in a preset pose, or can also be understood as a composite image of the object in the source image and the image in the preset pose. The object here can include a recognizable subject in the image (such as a person or animal, and optionally other objects).

[0090] The image generation method according to the embodiments of the present application can be applied to fields such as e-commerce product display and game character design. A specific object, such as a model, is present in the acquired source image. Multiple source images are acquired based on the active authorization of the object in the source image after being aware of the platform or system's relevant intent.

[0091] In some embodiments, the types of images acquired may include, for example (but not limited to) pictures, photographs, drawings, color illustrations, cartoons, and the like. That is, the images described in this application may be images generally understood by those skilled in the art, and this application does not limit the types of images. Similarly, these types of images may include at least one object. For the purpose of convenience of description and clear illustration, people are used to represent the corresponding objects hereinafter, but this does not mean that the objects described in this application are limited to the field of people.

[0092] At step S220, a source pose image is extracted from the source image. In various embodiments of the present application, the source pose image may refer to an image having the same pose (such as joint positions, skeletal structure, and pose angle) as the object in the source image. In various embodiments of the present application, extracting a source pose image from an image may include the following steps:

[0093] Use a preset pose estimation algorithm to detect human pose information in the source image, including identifying and estimating joint positions, skeleton structure, and pose angles;

[0094] Extract the joint information of various parts of the human body from the human posture information, including the position and coordinates of the joints in the head, shoulders, arms, legs, etc.

[0095] Use the extracted joint point information to draw the source pose image.

[0096] Specifically, according to the positions and coordinates of the joints, skeleton lines including each joint can be drawn to represent the posture, thereby visualizing the posture of the human body as an image.

[0097] In some embodiments, both the source pose image and the preset pose image can be obtained from a pose image library, that is, these pose images can be pre-set. To this end, it is necessary to build a pose image library containing a certain number of pose images, and to retrieve an appropriate and corresponding pose image from it when needed. In this way, the process of obtaining the source pose image can be simplified and processing efficiency can be improved.

[0098] At step S230, at least the source image, the noise image, the preset pose image, and the source pose image are fused to obtain a processed image. In some embodiments, image fusion refers to a technique for combining multiple images into a single image. This technique can be used to enhance image quality, compensate for defects between different images, improve visual effects, and increase the amount of image information.

[0099] In some embodiments, image fusion may be achieved simply by stitching, for example, multiple images are stitched together to form a single image.

[0100] In some embodiments, image fusion refers to the process of synthesizing multiple images according to certain rules or algorithms to generate a complete image. In some embodiments, image fusion may include the following steps:

[0101] For multiple images to be fused, align the images in the same coordinate system;

[0102] The aligned images are fused in a certain direction to generate a complete fused image. Certain directions include horizontal, vertical, and channel directions.

[0103] Optionally, image fusion may also be achieved through other methods, such as spatial domain fusion and transform domain fusion.

[0104] The following further describes step S230 in detail with reference to FIG5 , where FIG5 schematically shows a flow chart of an image generating method according to some embodiments of the present application.

[0105] In some embodiments, as shown in FIG5 , step S230 may include:

[0106] S510, stitching the source posture image and the preset posture image along the horizontal direction to obtain a first stitched image;

[0107] S520, stitching the source image and the noise image together in a horizontal direction to obtain a second stitched image;

[0108] S530: stitching at least the first stitched image and the second stitched image along the channel direction to obtain an image to be processed, wherein the spatial positions of the source pose image and the preset pose image in the first stitched image correspond to the spatial positions of the source image and the noise image in the second stitched image, respectively.

[0109] In some embodiments, stitching images horizontally can be intuitively understood as shown in FIG4 . On the left side of FIG4 , there are four images, from top to bottom, namely, a first stitched image 401 (i.e., the “source pose image + preset pose image” shown in FIG4 ), a mask image 402, a second stitched image 403 (i.e., the “source image + noise image” shown in FIG4 ), and a noisy image 404. Mask image 402 and noisy image 404 are further described below.

[0110] Horizontal stitching means that, in a rectangular coordinate system, the images have a dimension along the x-direction (i.e., width) and a dimension along the y-direction (i.e., length), and the images to be stitched should have the same length in order to stitch them together horizontally (i.e., in the x-direction). This allows the first stitched image in step S510 and the second stitched image in step S520 to be achieved.

[0111] Likewise, in some embodiments, images may be stitched together in a vertical direction. That is, the images to be stitched together should have the same width so as to be stitched together in the vertical direction (ie, the y direction).

[0112] In some other embodiments, other stitching directions may exist as long as the images can be stitched together along a certain spatial direction.

[0113] In image processing, a channel refers to the grayscale image obtained by decomposing an image. A color channel is a grayscale image that records color information. A channel is composed of one or more color channels, each of which represents the brightness value of a specific color in the image. For example, a common three-channel image is created by decomposing a color image into three color channels: red, green, and blue. Each channel contains the brightness information of that color in the original image, and these three channels together form a complete color image.

[0114] Therefore, in some embodiments, in step S530, for example, when both the first stitched image and the second stitched image are RGB three-channel images, stitching along the channel direction means: stitching the R channel of the first stitched image with the R channel of the second stitched image, stitching the G channel of the first stitched image with the G channel of the second stitched image, and stitching the B channel of the first stitched image with the B channel of the second stitched image. In this way, the image to be processed stitched along the channel direction can be intuitively understood as "stacking" the first stitched image and the second stitched image together along the z direction (i.e., the direction perpendicular to the paper). In this way, the source image and the preset pose image can be "bound" together, so that the consistency of the generation can be guaranteed through the inherent feature interaction of the image generation model.

[0115] As shown in Figure 4, the spatial positions of the source pose image and the preset pose image in the first stitched image 401 correspond to the spatial positions of the source image and the noise image in the second stitched image 403, respectively. This means that the noise image acts as a placeholder and mask, ensuring the integrity and effectiveness of the stitching along the channel direction. An image different from the source image can be selected as the noise image in order to generate the final target image based on this noise image.

[0116] In some embodiments, step S530 may include:

[0117] Acquire a mask image, wherein the mask image includes a first mask portion and a second mask portion different from the first mask portion;

[0118] The mask image, the first stitched image, and the second stitched image are stitched together along a channel direction to obtain the image to be processed, wherein the spatial positions of the first mask portion and the second mask portion in the mask image correspond to the spatial positions of the source image and the noise image in the second stitched image, respectively.

[0119] In each embodiment of the present application, the function of the mask image is to selectively apply a certain filter and perform the corresponding operation only on a certain area in the image. The mask image needs to have the same size as the first stitched image and the second stitched image. At the same time, in order to correspond to the spatial positions of the first stitched image and the second stitched image, the mask image is divided into a corresponding first mask part and a second mask part. Then, the mask image, the first stitched image, and the second stitched image can be "stacked" together along the z direction (i.e., the direction perpendicular to the paper surface) to splice out the image to be processed. Among them, the mask image is "sandwiched" between the first stitched image and the second stitched image. In this way, the correspondence of the spatial position can be used to more accurately distinguish the two image parts in the second stitched image 403 (i.e., the source image and the noise image before stitching), thereby improving the accuracy of the generated image.

[0120] In some embodiments, the noise image may include a solid color image. While it's possible to generate a target image using the image generation model if the noise image is only slightly different from the source image, the computational effort will be high. Therefore, compared to noise images that are only slightly different from the source image, using a simple solid color image as the noise image can reduce computational effort and improve the computational efficiency of the image generation model.

[0121] In some embodiments, in the mask image, the first mask portion is a pure white image, and the second mask portion is a pure black image. The pure color image includes a pure black image. That is, the mask image can be single-channel. For example, as shown in FIG4 , the mask image 402 of FIG4 includes a pure white image (i.e., the first mask portion, the corresponding RGB pixel value of which is [255, 255, 255]) and a pure black image (i.e., the second mask portion, the corresponding RGB pixel value of which is [0, 0, 0]), while the noise image can include a pure black image (RGB pixel value is [0, 0, 0]). In this way, the source pose image and the source image correspond to the pure white image in the mask image, and the preset pose image and the noise image correspond to the pure black image in the mask image, thereby making it easier to distinguish between the source image and the noise image, further reducing the computational complexity of the image generation model and improving computational efficiency.

[0122] As shown in FIG4 , after the spliced ​​image to be processed 405 is input into the image generation model 406 to obtain a denoised image 407 , the denoised image 407 can be further processed to obtain a target image 408 (ie, the right half of the denoised image 407 ).

[0123] In some embodiments, this process may include: dividing the denoised image along the horizontal direction to obtain an image portion corresponding to the spatial position of the noise image in the second stitched image; and using the image portion as the target image.

[0124] Alternatively, the denoised image can be divided horizontally based on the spatial correspondence between the denoised image and a mask image consisting of a pure white image and a pure black image. The target image corresponds to the pure black image. This target image incorporates the features of the source image and has a preset pose, thus enabling a view of the person in the preset pose.

[0125] In step S240, the source image is input into the image feature extractor of the image generation model to extract comprehensive features of the source image.

[0126] In an embodiment of the present application, texture features can describe the properties of details and structure of local areas in an image, where texture can refer to the regularity and irregularity of the spatial relationship and grayscale distribution between pixels in the image, which includes the color, style, etc. of the image.

[0127] As shown in FIG4 , a source image 409 is input into an image feature extractor 410 to extract comprehensive features 413 of the source image.

[0128] The image feature extractor of the present application is not limited to a general image feature extractor, but also includes a deep learning model. The image feature extractor may include, for example (but not limited to) RNN (Recurrent Neural Network), CNN (Convolutional Neural Networks), SIFT (Scale-Invariant Feature Transform), SURF (Speeded Up Robust Features), HOG (Histogram of Oriented Gradients), LBP (Local Binary Patterns), etc. In other words, the image feature extractor described in the present application may be an image feature extractor generally understood by those skilled in the art, and the present application does not limit the type of image feature extractor.

[0129] Next, step S240 will be described in detail in conjunction with FIG6 and FIG7 , where FIG6 schematically shows a flow chart of the image generation method according to some embodiments of the present application, and FIG7 schematically shows a principle diagram of the image generation method according to some embodiments of the present application.

[0130] In some embodiments, as shown in FIG6 , the image feature extractor includes a first feature extractor and a second feature extractor, and step S240 may include:

[0131] S610, inputting the source image into a first feature extractor to extract texture features;

[0132] S620, inputting the texture feature into the second feature extractor to extract semantic features;

[0133] S630: Determine the comprehensive features of the source image according to the texture features and the semantic features.

[0134] At step S610, in some embodiments, the first feature extractor may be an image encoder that may employ a CLIP (Contrastive Language-Image Pretraining) model. The CLIP model is trained on a large number of images and associated text to learn a shared representation space, enabling images and text to correspond to each other in this space. The image encoder can extract rich texture features of images and has strong generalization capabilities.

[0135] The first feature extractor may also use a deep residual network (Residual Network, ResNet) trained on the Imagenet database, or an unsupervised trained DINO (Emerging Properties in Self-Supervised Vision Transformers) model, etc. This application does not limit the type of the first image feature extractor.

[0136] As described above, since the source image generally contains a large number of features, in some embodiments, the first feature extractor includes an image compressor and a texture feature extractor, and step S610 may include: inputting the source image into the image compressor to obtain a compressed source image; inputting the compressed source image into the texture feature extractor to extract texture features.

[0137] An image compressor is a tool or algorithm used to reduce the size of digital image files, which can include a pre-trained autoencoder model. In this way, the computational effort of the image feature extractor when extracting features can be reduced without degrading image quality.

[0138] At step S620 , semantic features can be used to describe the attributes of high-level information such as objects, scenes, and concepts in the image, that is, the meaning of the image content. Semantic features are more abstract and advanced than texture features.

[0139] In some embodiments, the second feature extractor may be a model capable of extracting semantic features of an image, which may include, for example, convolutional neural networks (CNNs), visual transformer models, feature pyramid networks (FPNs), and convolutional autoencoders. This application does not limit the type of the second feature extractor.

[0140] In some embodiments, step S620 may include: inputting predefined learnable features into a second feature extractor, the second feature extractor including a multi-layer encoder, each layer of the encoder including a self-attention layer, a cross-attention layer, and a feedforward neural network layer; inputting texture features into the second feature extractor to obtain encoding features; determining semantic features based on the learnable features and the encoding features.

[0141] As shown in Figure 7, first, learnable features 601 are defined (for example, here are 32 learnable features), which represent the most essential features of the image, and these learnable features 601 are input into the second feature extractor 605, which includes 8 layers of encoders, each layer of encoder includes a self-attention layer 602, a cross-attention layer 603, and a feedforward neural network layer 604.

[0142] The self-attention layer 602 is used to capture the dependencies between elements in the input sequence. It divides the input sequence into several subsequences and performs self-attention calculations on each subsequence to obtain the correlation weights between different positions.

[0143] The feedforward neural network layer 604 is responsible for performing nonlinear transformation and mapping on the features at each position, which can introduce more nonlinearity and help the second feature extractor 605 learn more complex feature representations.

[0144] After the source image 606 is input into the first feature extractor 607 and texture features 608 are extracted, the texture features 608 are input into the cross-attention layer 603 as a key-value sequence to obtain encoding features 609. Semantic features 610 are then determined based on the learnable features 601 and encoding features 609. In this way, the extracted features can contain rich semantic features, improving the accuracy and completeness of the semantic features.

[0145] An example of step S630 is shown in the upper right portion of FIG4 . It can be seen that the image feature extractor may include a first feature extractor and a second feature extractor. After the source image is input into the image feature extractor, comprehensive features of the source image can be obtained. Thereafter, step S250 can be performed.

[0146] At step S250, the combined features of the image to be processed and the source image are input into an image denoiser of the image generation model to obtain a target image, which represents an image of the object in the source image at a predetermined pose. The image denoiser can be used to denoise the image. In an embodiment of the present application, the image denoiser can be based on a diffusion model, in which the denoising process is achieved through iterative deconvolution operations.

[0147] In some embodiments, the diffusion model is implemented as a UNet architecture, consisting of four stages (encoder, decoder, skip connection, and output layer), with downsampling performed after each stage. Skip connections are a key component of UNet, connecting feature maps from different encoder layers with those from corresponding decoder layers, enabling the fusion of low-level and high-level features, thereby helping the decoder better locate and recover detailed information.

[0148] The target image represents an image of the object in the source image in a preset posture, that is, the posture of the object is transformed into the preset posture, an example of which is shown in FIG3C .

[0149] As shown in Figure 4, the module of each stage may include a convolution part and an attention part, wherein the attention includes self-attention and cross-attention. The image denoiser 414 includes a convolution layer, a self-attention layer and a cross-attention layer.

[0150] In some embodiments, step S250 may include:

[0151] Inputting the comprehensive features of the image to be processed and the source image into the image denoiser to obtain a cross attention weight;

[0152] Using the cross attention weights, performing weighted summation on the comprehensive features of the source image to obtain correlation information;

[0153] The target image is obtained according to the associated information and the image to be processed.

[0154] The cross-attention layer in the image denoiser compares each element in the source image's comprehensive features with all elements in the image to be processed, calculating the relevance weight for each element in the source image's comprehensive features. The relevance information represents information within the source image's comprehensive features that is relevant to the image to be processed.

[0155] By introducing the cross-attention layer into the image denoiser, the correlation between different input sequences can be effectively captured, thereby extracting more valuable information.

[0156] According to embodiments of the present application, given a front view of a character from a certain angle, the image generation method according to some embodiments of the present application can be used to generate side and back views of the character, thereby assisting in the 3D modeling of the character. Therefore, the image generation method according to some embodiments of the present application can generate new views of a character in a specific pose. This technology has enormous application potential and can significantly change the way work is done in multiple industries.

[0157] First, consider the scenario of displaying items in the e-commerce field. For example, if the item is a piece of clothing, by simply taking a photo of a person wearing the clothing, the methods described in the embodiments of this application can be used to synthesize photos of the person wearing the clothing in different poses. For example, given a front view of a person wearing the clothing, a new side view can be synthesized. This innovation not only significantly saves time for models and photographers and reduces shooting costs, but also provides richer and more diverse visual effects, enhances consumers' shopping experience, and thus increases product sales.

[0158] For another example, when designing a character in the gaming field, given a front view at an angle, the method described in this embodiment can be used to generate a side view and a back view of the character, thereby assisting in the 3D modeling of the character.

[0159] Therefore, the image generation methods of some embodiments of the present application can provide more efficient and flexible solutions to meet various specific needs. For example, if you need to show different poses of a person in a specific environment or at a specific time, you can generate a variety of different views using just a single base photo, without the need for reshooting or tedious post-processing. This is particularly valuable in fields such as filmmaking, game design, and virtual reality, greatly improving work efficiency while providing more realistic visual effects.

[0160] FIG8 schematically shows a flow chart of an image generation model training method according to some embodiments of the present application. The image generation method is performed by an electronic device. In some embodiments, as shown in FIG1 , the image generation model training method can be performed on the terminal device 110 side. In other embodiments, the image generation model training method can also be performed in combination by the server 120 and the terminal device 110. As shown in FIG8 , the image generation model training method according to some embodiments of the present application may include steps S810-S870:

[0161] S810 , obtaining source image samples, target image samples, preset posture image samples, noise image samples, and source posture image samples, wherein the noise image samples are different from the source image samples.

[0162] S820 , fusing at least the source image sample, the noise image sample, the preset posture image sample, the source posture image sample, and the target image sample to obtain an image sample to be processed.

[0163] S830: Fuse the source image sample and the target image sample to obtain a target image label.

[0164] S840: Input the source image sample into the image feature extractor of the image generation model to extract the comprehensive features of the source image sample.

[0165] S850: Input the comprehensive features of the image sample to be processed and the source image sample into the image denoiser of the image generation model to obtain a denoised image.

[0166] S860: Calculate target loss based on the target image label and the denoised image.

[0167] S870, iteratively updating the parameters of the image generation model based on the target loss until the target loss meets a preset condition.

[0168] With reference to the description of steps S210-S250 above, in step S810, the various samples used in the training method can be obtained from the image library. The target image sample is an image sample of the object in the source image sample in a preset posture (for example, an example of the source image sample is FIG3A, and an example of the target image sample is FIG3C). The target image label refers to the training label constructed during the training process and is used to guide the learning of the image generation model.

[0169] In some embodiments, step S820 may include:

[0170] S820a, horizontally splicing the source image sample and the target image sample and performing noise processing to obtain a first spliced ​​image sample;

[0171] S820b, splicing the source posture image sample and the preset posture image sample along the horizontal direction to obtain a second spliced ​​image sample;

[0172] S820c, splicing the source image sample and the noise image sample along the horizontal direction to obtain a third spliced ​​image sample;

[0173] S820d: stitch together at least the first stitched image sample, the second stitched image sample, and the third stitched image sample along a channel direction to obtain the image sample to be processed.

[0174] The first spliced ​​image sample is a noisy image as shown in FIG4 , and the noisy processing refers to the process of iteratively adding noise to the image sample during the training process.

[0175] In some embodiments, step S820d may include:

[0176] Acquire a mask image sample, where the mask image sample includes a first mask portion and a second mask portion different from the first mask portion;

[0177] The mask image sample, the first stitched image sample, the second stitched image sample, and the third stitched image sample are stitched together along a channel direction to obtain the image sample to be processed, wherein the spatial positions of the first mask portion and the second mask portion in the mask image sample correspond to the spatial positions of the source image sample and the noise image sample in the third stitched image sample, respectively.

[0178] The acquisition and stitching process here can refer to the description of the aforementioned step S530, except that an additional stitching image sample that has been subjected to noise processing is stitched together.

[0179] In some embodiments, step S830 may include: horizontally splicing the source image sample and the target image sample to obtain the target image label. The operation process is as described above.

[0180] At step S860, the target loss refers to an overall loss in the process of training using the image generation model. The target loss can be a mean square error loss (which facilitates the use of a gradient descent algorithm, thereby simplifying the amount of computation). For example, the target image label and the pixel value of each pixel in the denoised image can be obtained respectively, and then the mean square error of the target image label and the denoised image can be calculated based on the obtained pixel value of each pixel. In other examples, the target loss can also include other types of losses.

[0181] In actual model training, the weights can be balanced according to different goals and different application scenarios. Therefore, the target loss can be determined by assigning a corresponding weight to each loss, and then taking the weighted sum of all losses as the target loss, where the weight represents the importance of the corresponding loss.

[0182] At step S870, in some embodiments, during the initial optimization phase, larger weights and iteration steps may be used to allow the parameters to quickly converge to the optimal solution. Subsequently, the weights are gradually reduced, and the iteration step size is also reduced, so that the parameters are continuously updated during the iteration process, thereby slowly and accurately converging to the optimal solution. When the maximum number of iterations is reached or after several optimizations, the target loss no longer decreases, the optimization process can be stopped, thus completing the training of the image generation model.

[0183] It should be understood that some common functions (such as sigmoid function) can be used here to perform model iteration, or any other method can be used to perform model iteration, as long as the loss value of the loss function gradually decreases. In one example, when all the training images in the training data set are sampled, a training cycle ends; after completing a preset number of training cycles, the training process ends. This application does not list them all here. This process can be regarded as a supervised training process.

[0184] In some embodiments, classifier-free guidance can be used to enhance image generation. Classifier-free guidance refers to not using a classifier as a guidance or supervisory signal, but rather guiding the learning of the image generation model through other means, enabling learning in the absence of explicit label information. This process can be considered an unsupervised training process. In some embodiments, during training, in the current training round, the source image is input into the image feature extractor in the image generation model to extract the source image's comprehensive features. According to a preset probability, the source image's comprehensive features are set to 0. For example, if the preset probability is 10%, setting the source image's comprehensive features to 0 is equivalent to performing dropout. In this case, a source image with a zero comprehensive feature is essentially a pure black image. This pure black image and the image to be processed are then input into the image denoiser in the image generation model, which outputs a new image as the target image, which is used as the new source image for the next round of training. This can be used to mitigate or prevent overfitting of the image generation model, making the image generation model more generalizable.

[0185] In some embodiments, during the training process, the updated parameters include the weights of the first feature extractor (i.e., the image encoder), the weights of the second feature extractor, and the weights of the image denoiser (i.e., the diffusion model). In some examples, the pre-trained weights of the first feature extractor (i.e., the image encoder) can be frozen, and only the weights of the diffusion model and the second feature extractor can be trained. Alternatively, in some embodiments, during the training process, the weights of the first feature extractor, the diffusion model, and the second feature extractor can be trained simultaneously. In one example, the training process can use the AdamW optimizer (Adam Weight Decay Optimizer) with a fixed learning rate of 1e-4.

[0186] In this way, using the image generation method according to some embodiments of the present application, the image to be processed is obtained by splicing at least the acquired source image, the noise image, the preset posture image, and the source posture image extracted from the source image. The source image and the preset posture image can be directly spliced ​​together and input into the image generation model, so that the image generation model can more easily synthesize a consistent image, thereby improving the consistency and restoration of the source image and the final generated image, thereby enhancing the robustness and generalization ability of the image generation model, improving the authenticity, clarity and naturalness of the generated image, and improving the accuracy and recognition of image generation.

[0187] Fig. 9A schematically shows an example block diagram of an image generating apparatus 910 according to some embodiments of the present application. The image generating apparatus 910 shown in Fig. 9A, as an electronic device, may correspond to the terminal device 110 shown in Fig. 1 .

[0188] As shown in FIG9A , the image generating device 910 may include:

[0189] A first acquisition module 911 is configured to acquire a source image, a noise image, and a preset posture image, wherein the noise image is different from the source image;

[0190] A first extraction module 912 is configured to extract a source posture image from the source image;

[0191] A first fusion module 913 is configured to fuse at least the source image, the noise image, the preset posture image, and the source posture image to obtain an image to be processed;

[0192] The second extraction module 914 is configured to input the source image into the image feature extractor in the image generation model to extract comprehensive features of the source image; and

[0193] The output module 915 is used to input the comprehensive features of the image to be processed and the source image into the image denoiser in the image generation model to obtain a target image, where the target image represents the image of the object in the source image in a preset posture.

[0194] In some examples, the first fusion module 913 is used to stitch the source pose image and the preset pose image along the horizontal direction to obtain a first stitched image; stitch the source image and the noise image along the horizontal direction to obtain a second stitched image; and stitch at least the first stitched image and the second stitched image along the channel direction to obtain the image to be processed, wherein the spatial positions of the source pose image and the preset pose image in the first stitched image correspond to the spatial positions of the source image and the noise image in the second stitched image, respectively.

[0195] In some examples, the first fusion module 913 is used to obtain a mask image, wherein the mask image includes a first mask part and a second mask part different from the first mask part; the mask image, the first stitched image and the second stitched image are stitched along the channel direction to obtain the image to be processed, wherein the spatial positions of the first mask part and the second mask part in the mask image correspond to the spatial positions of the source image and the noise image in the second stitched image, respectively.

[0196] In some examples, the image feature extractor includes a first feature extractor and a second feature extractor, and the second extraction module 914 is used to input the source image into the first feature extractor to extract texture features; input the texture features into the second feature extractor to extract semantic features; and determine the comprehensive features of the source image based on the texture features and the semantic features.

[0197] In some examples, the second extraction module 914 is used to input the source image into the image compressor to obtain a compressed source image; and input the compressed source image into the texture feature extractor to extract the texture features.

[0198] In some examples, the output module 915 is used to input the comprehensive features of the image to be processed and the source image into the image denoiser to obtain cross-attention weights; use the cross-attention weights to perform weighted summation on the comprehensive features of the source image to obtain association information; and obtain the target image based on the association information and the image to be processed.

[0199] It should be noted that the various modules described above can be implemented in software or hardware or a combination of both. Multiple different modules can be implemented in the same software or hardware structure, or one module can be implemented by multiple different software or hardware structures.

[0200] Figure 9B schematically shows an example block diagram of an image generation model training apparatus 920 according to some embodiments of the present application. The image generation model training apparatus 920 shown in Figure 9B, as an electronic device, may correspond to the terminal device 110 shown in Figure 1 .

[0201] As shown in FIG9B , the image generation model training device 920 may include:

[0202] A second acquisition module 921 is configured to acquire a source image sample, a target image sample, a preset posture image sample, a noise image sample, and a source posture image sample, wherein the noise image sample is different from the source image sample;

[0203] A second fusion module 922 is configured to fuse at least the source image sample, the noise image sample, the preset posture image sample, the source posture image sample, and the target image sample to obtain an image sample to be processed;

[0204] A third fusion module 923 is configured to fuse the source image sample with the target image sample to obtain a target image label;

[0205] A third extraction module 924 is configured to input the source image sample into the image feature extractor in the image generation model to extract comprehensive features of the source image sample;

[0206] a denoising module 925 for inputting the comprehensive features of the image sample to be processed and the source image sample into an image denoiser in the image generation model to obtain a denoised image;

[0207] A calculation module 926 is configured to calculate a target loss based on the target image label and the denoised image; and

[0208] The iterative module 927 is used to iteratively update the parameters of the image generation model based on the target loss until the target loss meets a preset condition.

[0209] In some examples, the second fusion module 922 is configured to: stitch the source image sample and the target image sample in a horizontal direction and perform noise processing to obtain a first stitched image sample; stitch the source pose image sample and the preset pose image sample in a horizontal direction to obtain a second stitched image sample; stitch the source image sample and the noise image sample in a horizontal direction to obtain a third stitched image sample; and stitch at least the first stitched image sample, the second stitched image sample, and the third stitched image sample in a channel direction to obtain the image sample to be processed;

[0210] The third fusion module 923 is used to splice the source image sample and the target image sample along the horizontal direction to obtain the target image label.

[0211] In some examples, the second fusion module 922 is used to obtain a mask image sample, wherein the mask image sample includes a first mask part and a second mask part different from the first mask part; and stitching the mask image sample, the first stitched image sample, the second stitched image sample and the third stitched image sample along the channel direction to obtain the image sample to be processed, wherein the spatial positions of the first mask part and the second mask part in the mask image sample correspond to the spatial positions of the source image sample and the noise image sample in the third stitched image sample, respectively.

[0212] Figure 10 schematically shows an example block diagram of an electronic device 1000 according to some embodiments of the present application. The electronic device 1000 may represent a device for implementing the various devices or modules described in the various embodiments of the present application and / or performing the various methods described in the various embodiments of the present application. The electronic device 1000 may be, for example, a server, a desktop computer, a laptop computer, a tablet, a smart phone, a smart watch, a wearable device, or any other suitable electronic device or computing system, which may include various levels of devices ranging from full-resource devices with a large amount of storage and processing resources to low-resource devices with limited storage and / or processing resources. In some embodiments, the image generating device 910 described above with respect to Figure 9 may be implemented in one or more electronic devices 1000, respectively.

[0213] As shown in Figure 10, example electronic device 1000 includes a processing system 1001, one or more computer-readable media 1002, and one or more I / O interfaces 1003 that are communicatively coupled to each other. Although not shown, electronic device 1000 may also include a system bus or other data and command transmission system that couples various components to each other. The system bus may include any one or a combination of different bus structures, and the bus structure may be a processor or local bus such as a memory bus or memory controller, a peripheral bus, a universal serial bus, and / or any one of a variety of bus architectures. Alternatively, it may also include control and data lines.

[0214] Processing system 1001 represents the functionality of performing one or more operations using hardware. Thus, processing system 1001 is illustrated as including hardware elements 1004 that can be configured as processors, functional blocks, and the like. This can include implementation in hardware as application specific integrated circuits or other logic devices formed using one or more semiconductors. Hardware elements 1004 are not limited by the materials from which they are formed or the processing mechanisms employed therein. For example, a processor can be comprised of (a plurality of) semiconductors and / or transistors (e.g., electronic integrated circuits (ICs)). In such a context, processor-executable instructions can be electronically executable instructions.

[0215] The computer-readable medium 1002 is illustrated as including a memory / storage device 1005. The memory / storage device 1005 represents a memory / storage device associated with one or more computer-readable media. The memory / storage device 1005 may include volatile media (such as random access memory (RAM)) and / or non-volatile media (such as read-only memory (ROM), flash memory, optical disk, magnetic disk, etc.). The memory / storage device 1005 may include fixed media (e.g., RAM, ROM, fixed hard drive, etc.) and removable media (e.g., flash memory, removable hard drive, optical disk, etc.). Exemplarily, the memory / storage device 1005 may be used to store the first audio of the first category of users mentioned in the above embodiment, the queue list of requests, etc. The computer-readable medium 1002 may be configured in various other ways as further described below.

[0216] One or more I / O (input / output) interfaces 1003 represent functions that allow a user to enter commands and information into the electronic device 1000 and also allow the information to be displayed to the user and / or sent to other components or devices using various input / output devices. Examples of input devices include a keyboard, a cursor control device (e.g., a mouse), a microphone (e.g., for voice input), a scanner, a touch function (e.g., a capacitive or other sensor configured to detect physical touch), a camera (e.g., a device that can use visible or invisible wavelengths (such as infrared frequencies) to detect motion that does not involve touch as a gesture), a network card, a receiver, and the like. Examples of output devices include a display device (e.g., a monitor or projector), a speaker, a printer, a tactile response device, a network card, a transmitter, and the like. Exemplarily, in the embodiments described above, the first category of users and the second category of users can input to initiate requests and record audio and / or video, etc., through the input interfaces on their respective terminal devices, and can view various notifications and watch videos or listen to audio, etc., through the output interfaces.

[0217] Electronic device 1000 also includes a gesture-guided image generation strategy 1006. This gesture-guided image generation strategy 1006 can be stored as computer program instructions in memory / storage device 1005, or can be implemented as hardware or firmware. This gesture-guided image generation strategy 1006, together with processing system 1001 and other components, can implement all of the functions of the various modules of image generation device 910 described with respect to FIG. 9 .

[0218] The various embodiments of the present application can describe various technologies in the general context of software, hardware, elements or program modules. Generally, these modules include routines, programs, objects, elements, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The terms "module", "function", etc. used in the various embodiments of the present application generally represent software, firmware, hardware or a combination thereof. The characteristics of the technology described in the various embodiments of the present application are platform-independent, meaning that these technologies can be implemented on various computing platforms with various processors.

[0219] An implementation of the described modules and techniques may be stored on or transmitted across some form of computer-readable media. Computer-readable media may include various media accessible by the electronic device 1000. By way of example and not limitation, computer-readable media may include "computer-readable storage media" and "computer-readable signal media."

[0220] As opposed to a mere signal transmission, carrier wave, or signal itself, "computer-readable storage medium" refers to a medium and / or device, and / or tangible storage device, capable of persistently storing information. Thus, computer-readable storage media refers to non-signal-bearing media. Computer-readable storage media include hardware such as volatile and non-volatile, removable and non-removable media and / or storage devices implemented in a method or technology suitable for storing information (such as computer-readable instructions, data structures, program modules, logic elements / circuits, or other data). Examples of computer-readable storage media may include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage devices, hard disks, cassettes, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or other storage devices, tangible media, or articles of manufacture suitable for storing desired information and accessible by a computer.

[0221] "Computer-readable signal media" refers to signal-carrying media configured to transmit instructions to the hardware of the electronic device 1000, such as via a network. Signal media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave, data signal, or other transport mechanism. Signal media also includes any information transmission media. By way of example and not limitation, signal media include wired media such as a wired network or direct connection, and wireless media such as acoustic, RF, infrared, and other wireless media.

[0222] As previously mentioned, hardware element 1004 and computer-readable medium 1002 represent instructions, modules, programmable device logic and / or fixed device logic implemented in hardware form, which can be used to implement at least some aspects of the technology described in each embodiment of the present application in some embodiments. Hardware elements can include other implementations in integrated circuits or systems on a chip, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs) and silicon or components of other hardware devices. In this context, hardware elements can be used as processing equipment for executing program tasks defined by the instructions, modules and / or logic embodied by the hardware elements, and hardware devices for storing instructions for execution, such as the computer-readable storage media previously described.

[0223] The aforementioned combination may also be used to implement the various techniques and modules described in each embodiment of the present application. Therefore, software, hardware or program modules and other program modules may be implemented as one or more instructions and / or logic embodied on some form of computer-readable storage medium and / or by one or more hardware elements 1004. The electronic device 1000 may be configured to implement specific instructions and / or functions corresponding to the software and / or hardware modules. Therefore, for example, by using a computer-readable storage medium and / or hardware elements 1004 of a processing system, the module may be implemented as a module that can be executed by the electronic device 1000 as software, at least in part, in hardware. Instructions and / or functions may be executed / operable to implement the techniques, modules and examples described in each embodiment of the present application by, for example, one or more electronic devices 1000 and / or processing system 1001.

[0224] The technology described in the various embodiments of the present application can be supported by these various configurations of the electronic device 1000 and is not limited to the specific examples of the technology described in the various embodiments of the present application.

[0225] In particular, according to embodiments of the present application, the processes described above with reference to the flowcharts may be implemented as computer programs. For example, embodiments of the present application provide a computer program product comprising a computer program carried on a computer-readable medium, the computer program including program code for executing at least one step of the method embodiments of the present application.

[0226] In some embodiments of the present application, one or more computer-readable storage media are provided, on which computer-readable instructions are stored. When the computer-readable instructions are executed, the image generation method according to some embodiments of the present application is implemented. The various steps of the image generation method according to some embodiments of the present application can be converted into computer-readable instructions through programming and stored in a computer-readable storage medium. When such a computer-readable storage medium is read or accessed by an electronic device or computer, the computer-readable instructions therein are executed by a processor on the electronic device or computer to implement the method according to some embodiments of the present application.

[0227] In the description of this specification, the description of the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0228] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed in a sequence other than as shown or discussed (including in a substantially simultaneous manner or in reverse order depending on the functions involved), which should be understood by those skilled in the art to which the embodiments of the present application pertain.

[0229] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0230] It should be understood that various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, it can be implemented by any one of the following technologies well known in the art or a combination thereof: a discrete logic circuit having a logic gate circuit for implementing a logical function on a data signal, an application-specific integrated circuit having a suitable combinational logic gate circuit, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0231] A person skilled in the art of the methods described in the embodiments of the present application may understand that all or part of the steps of the methods of the above embodiments may be completed by hardware associated with program instructions, and the program may be stored in a computer-readable storage medium, which, when executed, includes executing one or a combination of the steps of the method embodiments.

[0232] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0233] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0234] It is understood that the specific embodiments of this application involve various data used for image generation and model training (e.g., source images, target images, image features, etc.). When the embodiments described in this application involving such data are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of the relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.

Claims

1. An image generation method, executed by an electronic device, the method comprising: Obtaining a source image, a noise image, and a preset pose image, wherein the noise image is different from the source image; Extracting a source pose image from the source image; Fusing at least the source image, the noise image, the preset pose image, and the source pose image to obtain a to-be-processed image; Inputting the source image into an image feature extractor in an image generation model to extract a comprehensive source image feature; and, Inputting the to-be-processed image and the comprehensive source image feature into an image denoiser in the image generation model to obtain a target image, the target image representing an image of an object in the source image in a preset pose.

2. The method according to claim 1, wherein The fusing at least the source image, the noise image, the preset pose image, and the source pose image to obtain a to-be-processed image includes: Stitching the source pose image and the preset pose image along the horizontal direction to obtain a first stitched image; Stitching the source image and the noise image along the horizontal direction to obtain a second stitched image; Stitching at least the first stitched image and the second stitched image along the channel direction to obtain the to-be-processed image, wherein the spatial positions of the source pose image and the preset pose image in the first stitched image respectively correspond to the spatial positions of the source image and the noise image in the second stitched image.

3. The method according to claim 2, wherein, The stitching at least the first stitched image and the second stitched image along the channel direction to obtain the to-be-processed image includes: Obtaining a mask image, the mask image including a first mask part and a second mask part different from the first mask part; Stitching the mask image, the first stitched image, and the second stitched image along the channel direction to obtain the to-be-processed image, wherein the spatial positions of the first mask part and the second mask part in the mask image respectively correspond to the spatial positions of the source image and the noise image in the second stitched image.

4. The method according to any one of claims 1-3, wherein The image feature extractor includes a first feature extractor and a second feature extractor, and the inputting the source image into the image feature extractor in the image generation model to extract a comprehensive source image feature includes: Inputting the source image into the first feature extractor to extract a texture feature; Inputting the texture feature into the second feature extractor to extract a semantic feature; Determining the comprehensive source image feature according to the texture feature and the semantic feature.

5. The method according to claim 4, wherein The first feature extractor includes an image compressor and a texture feature extractor, and the inputting the source image into the first feature extractor to extract a texture feature includes: Inputting the source image into the image compressor to obtain a compressed source image; Inputting the compressed source image into the texture feature extractor to extract the texture feature.

6. The method according to any one of claims 1-5, wherein, The inputting the to-be-processed image and the comprehensive source image feature into the image denoiser in the image generation model to obtain a target image includes: Inputting the to-be-processed image and the comprehensive source image feature into the image denoiser to obtain cross-attention weights; Using the cross-attention weights, perform weighted summation on the comprehensive features of the source image to obtain associated information; Obtain the target image according to the associated information and the image to be processed.

7. An image generation model training method, executed by an electronic device, the method comprising: Obtain a source image sample, a target image sample, a preset pose image sample, a noise image sample, and a source pose image sample, wherein the noise image sample is different from the source image sample; Fuse at least the source image sample, the noise image sample, the preset pose image sample, the source pose image sample, and the target image sample to obtain a sample of the image to be processed; Fuse the source image sample and the target image sample to obtain a target image label; Input the source image sample into an image feature extractor in the image generation model to extract the comprehensive features of the source image sample; Input the sample of the image to be processed and the comprehensive features of the source image sample into an image denoiser in the image generation model to obtain a denoised image; Calculate a target loss according to the target image label and the denoised image; and Based on the target loss, iteratively update the parameters of the image generation model until the target loss meets a preset condition.

8. The method according to claim 7, wherein, The step of fusing at least the source image sample, the noise image sample, the preset pose image sample, the source pose image sample, and the target image sample to obtain a sample of the image to be processed includes: Stitch the source image sample and the target image sample horizontally and perform noise addition processing to obtain a first stitched image sample; Stitch the source pose image sample and the preset pose image sample horizontally to obtain a second stitched image sample; Stitch the source image sample and the noise image sample horizontally to obtain a third stitched image sample; Stitch at least the first stitched image sample, the second stitched image sample, and the third stitched image sample along the channel direction to obtain the sample of the image to be processed; The step of fusing the source image sample and the target image sample to obtain a target image label includes: Stitch the source image sample and the target image sample horizontally to obtain the target image label.

9. The method according to claim 8, wherein, The step of stitching at least the first stitched image sample, the second stitched image sample, and the third stitched image sample along the channel direction to obtain the sample of the image to be processed includes: Obtain a mask image sample, the mask image sample including a first mask part and a second mask part different from the first mask part; Stitch the mask image sample, the first stitched image sample, the second stitched image sample, and the third stitched image sample along the channel direction to obtain the sample of the image to be processed, wherein the spatial positions of the first mask part and the second mask part in the mask image sample respectively correspond to the spatial positions of the source image sample and the noise image sample in the third stitched image sample.

10. An image generation device, comprising: A first acquisition module, configured to acquire a source image, a noise image, and a preset pose image, where the noise image is different from the source image; A first extraction module, configured to extract a source pose image from the source image; A first fusion module, configured to fuse at least the source image, the noise image, the preset pose image, and the source pose image to obtain a to-be-processed image; A second extraction module, configured to input the source image into an image feature extractor in an image generation model to extract a comprehensive source image feature; and An output module, configured to input the to-be-processed image and the comprehensive source image feature into an image denoiser in the image generation model to obtain a target image, where the target image represents an image of an object in the source image in a preset pose.

11. The apparatus according to claim 10, wherein, The first fusion module is configured to splice the source pose image and the preset pose image along the horizontal direction to obtain a first spliced image; splice the source image and the noise image along the horizontal direction to obtain a second spliced image; at least splice the first spliced image and the second spliced image along the channel direction to obtain the to-be-processed image, where the spatial positions of the source pose image and the preset pose image in the first spliced image respectively correspond to the spatial positions of the source image and the noise image in the second spliced image.

12. The apparatus according to claim 11, wherein, The first fusion module is configured to acquire a mask image, where the mask image includes a first mask part and a second mask part different from the first mask part; splice the mask image, the first spliced image, and the second spliced image along the channel direction to obtain the to-be-processed image, where the spatial positions of the first mask part and the second mask part in the mask image respectively correspond to the spatial positions of the source image and the noise image in the second spliced image.

13. The apparatus according to any one of claims 10-12, wherein, The image feature extractor includes a first feature extractor and a second feature extractor. The second extraction module is configured to input the source image into the first feature extractor to extract a texture feature; input the texture feature into the second feature extractor to extract a semantic feature; and determine the comprehensive source image feature according to the texture feature and the semantic feature.

14. The apparatus according to any one of claims 10-13, wherein, The output module is configured to input the to-be-processed image and the comprehensive source image feature into the image denoiser to obtain cross-attention weights; Use the cross-attention weights to perform weighted summation on the comprehensive source image feature to obtain associated information; and obtain the target image according to the associated information and the to-be-processed image.

15. An image generation model training device, including: A second acquisition module, configured to acquire a source image sample, a target image sample, a preset pose image sample, a noise image sample, and a source pose image sample, where the noise image sample is different from the source image sample; A second fusion module, configured to fuse at least the source image sample, the noise image sample, the preset pose image sample, the source pose image sample, and the target image sample to obtain a to-be-processed image sample; A third fusion module, configured to fuse the source image sample and the target image sample to obtain a target image label; A third extraction module, configured to input the source image sample into an image feature extractor in the image generation model to extract the comprehensive feature of the source image sample; A denoising module, configured to input the to-be-processed image sample and the comprehensive feature of the source image sample into an image denoiser in the image generation model to obtain a denoised image; A calculation module, configured to calculate a target loss according to the target image label and the denoised image; and An iteration module, configured to iteratively update the parameters of the image generation model based on the target loss until the target loss meets a preset condition.

16. The apparatus according to claim 15, wherein, The second fusion module is configured to splice the source image sample and the target image sample horizontally and perform noise addition processing to obtain a first spliced image sample; splice the source pose image sample and the preset pose image sample horizontally to obtain a second spliced image sample; splice the source image sample and the noise image sample horizontally to obtain a third spliced image sample; and splice at least the first spliced image sample, the second spliced image sample, and the third spliced image sample along the channel direction to obtain the to-be-processed image sample; The third fusion module is configured to splice the source image sample and the target image sample horizontally to obtain the target image label.

17. The apparatus according to claim 16, wherein, The second fusion module is configured to obtain a mask image sample, where the mask image sample includes a first mask part and a second mask part different from the first mask part; splice the mask image sample, the first spliced image sample, the second spliced image sample, and the third spliced image sample along the channel direction to obtain the to-be-processed image sample, where the spatial positions of the first mask part and the second mask part in the mask image sample respectively correspond to the spatial positions of the source image sample and the noise image sample in the third spliced image sample.

18. An electronic device, comprising: A memory and a processor, wherein, a computer program is stored in the memory, and when the computer program is executed by the processor, the method according to any one of claims 1-9 is implemented.

19. A computer-readable storage medium, storing computer-readable instructions, where the computer-readable instructions are loaded by a processor to execute the method according to any one of claims 1-9.

20. A computer program product, comprising a computer program, where the computer program is stored in a computer-readable storage medium, and when a processor of an electronic device reads the computer program from the computer-readable storage medium, the method according to any one of claims 1-9 is executed.

Citation Information

Patent Citations

  • Training method of multi-task model and virtual fitting method of clothes

    CN114639161A

  • Figure composite image generation method and device, computer equipment and storage medium

    CN114821811A

  • Virtual fitting method based on diffusion model

    CN117011207A

  • Image generation method and device based on attitude guidance

    CN117576248A

Cited By

  • Data set generation and annotation method, model training method, image generation method and equipment

    CN121811169A