Image generation method and related equipment

Through the methods of feature alignment and semantic spatial alignment, the generated image has improved similarity with the target subject and the editability is enhanced, which solves the problem that it is difficult to take into account both the similarity and editability of subjects in image generation.

CN119991852APending Publication Date: 2025-05-13BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510142215.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In image generation technology, subject similarity and editability are difficult to take into account. In the generated image, the subject and given subject are relatively low in similarity, and the editability is insufficient.

Method used

By acquiring the prompt text and the first image, a first generated image is generated, and the subject in it is characterized by aligning the target subject to obtain a second generated image. The second generated image is then semantically and spatially aligned with the first generated image to generate the final target image.

Benefits of technology

It effectively improves the similarity between the generated image and the target subject, and at the same time improves the editability and image quality of the generated image, solving the problem that it is difficult to take into account both the similarity and editability of the subject.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991852A_ABST
    Figure CN119991852A_ABST
Patent Text Reader

Abstract

The invention provides an image generation method and related equipment. The method comprises the steps of obtaining a prompt text and a first image used for indicating a target body; generating a first generated image based on the prompt text and the first image; performing feature alignment on a first main body in the first generated image and the target main body to generate a second generated image; and aligning the second generated image with the first generated image semantically and spatially to obtain a target image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing, and in particular to an image generation method and related equipment. Background Art

[0002] Image generation technology can generate an image containing a given subject under the text indication based on the input text and the image of the given subject. However, in actual applications, sometimes the subject in the generated image is inconsistent with the given subject, and the similarity is low; sometimes, even if the subject similarity meets the user's expectations, the subject, background, style and other contents in the generated image are less affected by the input text, resulting in low editability of the generated image. Therefore, there is a problem in image generation that subject similarity and editability cannot be taken into account at the same time. Summary of the invention

[0003] The present disclosure proposes an image generation method and related equipment, which at least to some extent solve the technical problems in the related art that the image generation results cannot take into account both the subject similarity and the editability.

[0004] In a first aspect, the present disclosure provides an image generation method, comprising:

[0005] Acquire a prompt text and a first image for indicating a target subject;

[0006] generating a first generated image based on the prompt text and the first image;

[0007] Aligning features of the first subject in the first generated image with the target subject to generate a second generated image;

[0008] The second generated image is semantically and spatially aligned with the first generated image to obtain a target image.

[0009] In a second aspect of the present disclosure, there is provided an image generating device, comprising:

[0010] An acquisition module, used for acquiring a prompt text and a first image indicating a target subject;

[0011] An image generation module is used to generate a first generated image based on the prompt text and the first image; align the features of the first subject in the first generated image with the target subject to generate a second generated image; and align the second generated image with the first generated image semantically and spatially to obtain a target image.

[0012] According to a third aspect of the present disclosure, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method described in the first aspect is implemented.

[0013] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in the first aspect.

[0014] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising computer program instructions, which, when executed on a computer, cause the computer to execute the method according to the first aspect.

[0015] From the above, it can be seen that the image generation method and related devices provided by the present disclosure generate a preliminary first generated image by combining the prompt text and the first image, and then align the subject in the first generated image with the target subject to obtain a second generated image. Then, the second generated image is semantically and spatially aligned with the first generated image to finally generate the target image. The similarity between the generated image and the target subject is effectively improved, while also improving the editability and image quality of the generated image. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the drawings required for use in the embodiments or related technical descriptions are briefly introduced below. Obviously, the drawings described below are only embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0017] Figure 1 A schematic diagram of an image generation architecture according to an embodiment of the present disclosure.

[0018] Figure 2 The figure is a schematic diagram of the hardware structure of an exemplary electronic device according to an embodiment of the present disclosure.

[0019] Figure 3 The figure is a flowchart of the image generating method according to the embodiment of the present disclosure.

[0020] Figure 4 is a schematic diagram of a first image according to an embodiment of the present disclosure.

[0021] Figure 5 This is a schematic diagram of a first generated image according to an embodiment of the present disclosure.

[0022] Figure 6 It is a schematic diagram of a second generated image according to an embodiment of the present disclosure.

[0023] Figure 7 This is a schematic diagram of a third generated image according to an embodiment of the present disclosure.

[0024] Figure 8Schematic diagram of an image generating device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0025] In order to make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.

[0026] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should be understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. "Including" or "comprising" and similar words mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0027] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the type, scope of use, and usage scenarios of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations. For example, in response to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require the acquisition and use of the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.

[0028] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet the relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0029] Figure 1 FIG. 1 is a schematic diagram showing an image generation architecture of an embodiment of the present disclosure. Figure 1, the image generation architecture 100 may include a server 110, a terminal 120, and a network 130 providing a communication link. The server 110 and the terminal 120 may be connected via a wired or wireless network 130. The server 110 may be an independent physical server, or a server cluster or distributed system consisting of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, security services, and CDN.

[0030] The terminal 120 may be implemented in hardware or software. For example, when the terminal 120 is implemented in hardware, it may be various electronic devices having a display screen and supporting page display, including but not limited to smart phones, tablet computers, e-book readers, laptop portable computers, and desktop computers, etc. When the terminal 120 device is implemented in software, it may be installed in the electronic devices listed above; it may be implemented as multiple software or software modules (such as software or software modules used to provide distributed services), or it may be implemented as a single software or software module, which is not specifically limited here.

[0031] It should be noted that the image generation method provided in the embodiment of the present application can be executed by the terminal 120 or by the server 110. It should be understood that Figure 1 The number of terminals, networks and servers in the embodiment is only for illustration and is not intended to limit the number of terminals, networks and servers. Any number of terminals, networks and servers may be provided as required.

[0032] Figure 2 FIG. 2 shows a schematic diagram of the hardware structure of an exemplary electronic device 200 provided in an embodiment of the present disclosure. Figure 2 As shown, the electronic device 200 may include: a processor 202, a memory 204, a network module 206, a peripheral interface 208 and a bus 210. The processor 202, the memory 204, the network module 206 and the peripheral interface 208 are connected to each other in communication within the electronic device 200 through the bus 210.

[0033] Processor 202 may be a central processing unit (CPU), a neural network processor (NPU), a microcontroller (MCU), a programmable logic device, a digital signal processor (DSP), an application specific integrated circuit (ASIC), or one or more integrated circuits. Processor 202 may be used to perform functions related to the technology described in this disclosure. In some embodiments, processor 202 may also include multiple processors integrated into a single logical component. For example, Figure 2 As shown, processor 202 may include a plurality of processors 202a, 202b, and 202c.

[0034] The memory 204 may be configured to store data (eg, instructions, computer code, etc.). Figure 2 As shown, the data stored in the memory 204 may include program instructions (e.g., program instructions for implementing the image generation method of the embodiment of the present disclosure) and data to be processed (e.g., the memory may store configuration files of other modules, etc.). The processor 202 may also access the program instructions and data stored in the memory 204, and execute the program instructions to operate on the data to be processed. The memory 204 may include a volatile storage device or a non-volatile storage device. In some embodiments, the memory 204 may include a random access memory (RAM), a read-only memory (ROM), an optical disk, a magnetic disk, a hard disk, a solid-state drive (SSD), a flash memory, a memory stick, etc.

[0035] The network module 206 can be configured to provide communication with other external devices to the electronic device 200 via a network. The network can be any wired or wireless network capable of transmitting and receiving data. For example, the network can be a wired network, a local wireless network (e.g., Bluetooth, WiFi, near field communication (NFC) etc.), a cellular network, the Internet or a combination thereof. It is understood that the type of network is not limited to the above specific examples. In some embodiments, the network module 206 can include any number of network interface controllers (NICs), radio frequency modules, transceivers, modems, routers, gateways, adapters, cellular network chips, etc., in any combination.

[0036] The peripheral interface 208 can be configured to connect the electronic device 200 to one or more peripheral devices to achieve information input and output. For example, the peripheral devices can include input devices such as a keyboard, a mouse, a touch pad, a touch screen, a microphone, and various sensors, and output devices such as a display, a speaker, a vibrator, and an indicator light.

[0037] The bus 210 can be configured to transmit information between various components of the electronic device 200 (e.g., the processor 202, the memory 204, the network module 206, and the peripheral interface 208), such as an internal bus (e.g., a processor-memory bus), an external bus (USB port, PCI-E bus), etc.

[0038] It should be noted that, although the architecture of the electronic device 200 only shows the processor 202, the memory 204, the network module 206, the peripheral interface 208 and the bus 210, in the specific implementation process, the architecture of the electronic device 200 may also include other components necessary for normal execution. In addition, it can be understood by those skilled in the art that the architecture of the electronic device 200 may also only include the components necessary for implementing the embodiments of the present disclosure, and does not necessarily include all the components shown in the figure.

[0039] Personalized image generation can be achieved based on pre-trained image models. For example, the user inputs one or several images representing a given subject and text, and generates an image containing the subject under text control. This enables the pre-trained image model to generate imaginative images for the subject specified by the user (for example, the user himself or the user's favorite toy, etc.). In the related art, one type is a technology based on model fine-tuning, which uses several images of a given subject to fine-tune the pre-trained image model, allowing the image model to learn the concept of the subject, and then use the fine-tuned image model to generate personalized images. The other type does not require fine-tuning, and often injects the subject concept in the image into the pre-trained image model by adding a new visual branch to achieve the generation of personalized images. However, personalized image generation technology that does not require fine-tuning usually faces the problem of balancing similarity and editability in actual use. If the injection strength of a given image is too strong, the similarity is high, but the text control of the subject, background, style, etc. in the generated image cannot be performed, and the editability is weak; if the injection strength of a given image is too weak, the editability is high, but the subject in the generated image is inconsistent with the subject expected by the user, and the similarity is weak. Both of these situations are difficult to meet the needs of most users, limiting the use scenarios of personalized image generation technology. Therefore, how to simultaneously improve the similarity between the subject of the generated image and the given subject, and the editability of the image, has become a technical problem that needs to be solved urgently.

[0040] In view of this, the embodiments of the present disclosure provide an image generation method and related devices. A preliminary first generated image is generated by combining the prompt text and the first image, and then the subject in the first generated image is feature-aligned with the target subject to obtain a second generated image. Then, the second generated image is semantically and spatially aligned with the first generated image to finally generate the target image. The similarity between the generated image and the target subject is effectively improved, and the editability and image quality of the generated image are also improved.

[0041] See also Figure 3 , Figure 3 The schematic flow chart of the image generation method according to the embodiment of the present disclosure is shown. The image generation method according to the embodiment of the present disclosure can be deployed on a server or a terminal. Figure 3In the method 300 for generating an image, the method 300 for generating an image may further include the following steps.

[0042] In step S310, a prompt text and a first image for indicating a target subject are acquired.

[0043] The prompt text may be used to provide information about the content, style, color or other visual features of the desired image to be generated, so as to instruct the image generation algorithm to generate or edit an image that meets these descriptions. For example, the prompt text may be "an old man playing chess in the park". It should be understood that the prompt text may be in a variety of languages, which are not limited here. The target subject may refer to the object that the user wants to generate during image generation. The first image may be an input image, which may be any image containing the target subject, such as Figure 4 As shown, Figure 4 A schematic diagram of a first image according to an embodiment of the present disclosure is shown. Figure 4 The elderly person shown in the video may be the target subject.

[0044] In step S320, a first generated image is generated based on the prompt text and the first image.

[0045] The first generated image may be an image initially generated based on the prompt text and the first image. The first generated image can retain key features of the target subject in the first image, such as facial features, body proportions, etc., so that the generated image is more visually coherent and consistent. By combining the descriptive information in the prompt text, a more vivid and expressive image can be generated. Specifically, see Figure 5 , Figure 5 A schematic diagram of a first generated image according to an embodiment of the present disclosure is shown.

[0046] Figure 5 In , the prompt text can be "an old man playing chess in the park", and a first image containing a child and an old man is provided, such as Figure 4 As shown. A new first generated image can be generated based on the prompt text and the first image, in which the old man in the first image is playing chess in the park. Specifically, the prompt text can be parsed to extract key information, such as "old man", "park", "chess", etc. The first image can also be subjected to object recognition to extract key information such as facial features of the object. The key information in the image is combined with the key information in the prompt text to generate a first generated image. As shown Figure 5 As shown, a portion of the first generated image has a low similarity with the elderly subject in the first image, and a portion of the image contains a child's image without being controlled by the prompt text, which fails to take into account both the subject similarity and the editability of the image.

[0047] Specifically, the first generated image can be generated based on the trained image generation model. A diffusion model can be used in the image generation model to generate the first generated image, and the corresponding first loss function can be the loss function of the diffusion model. For example, the loss function of the diffusion model can be obtained based on the prediction error of the denoising process, that is, the difference between the denoising result predicted by the model and the actual noise. Since a certain amount of noise is added to the initial image at each step in the forward diffusion process, the reverse diffusion process needs to learn how to remove this noise. Therefore, the diffusion model is trained to predict the noise that needs to be removed at each step, and the image is gradually restored based on this prediction. The first loss function of the diffusion model can be expressed as:

[0048] L simple =E x0,ε~N(0,I) [∥ε-εθ(xt,t)∥ 2 ].

[0049] Among them, x0 represents the original image, ε represents the added noise, and εθ(xt,t) represents the model at a given time and the current noisy image x t The goal of this first loss function is to minimize the difference between the noise predicted by the model and the noise actually added.

[0050] In step S330, feature alignment is performed between the first subject in the first generated image and the target subject to generate a second generated image.

[0051] Among them, by aligning the features of the first subject in the first generated image with the target subject (the user hopes to retain or achieve), the subject features in the first generated image can be adjusted to be more consistent with the target subject, so that the generated image better meets the user's expectations and requirements.

[0052] In some embodiments, aligning features of the first subject in the first generated image with the target subject to generate a second generated image includes:

[0053] Extracting features of a first subject in a first generated image to obtain first subject features, and extracting features of a target subject in the first image to obtain target subject features;

[0054] The first subject feature is aligned with the target subject feature to obtain the second generated image.

[0055] Among them, the first subject features obtained by extracting features from the first subject in the first generated image may include shape, texture, color, edge information, etc., which together describe the appearance and attributes of the first subject. At the same time, feature extraction is also performed on the target subject in the first image to obtain target subject features. These target subject features also describe the appearance and attributes of the target subject, but may be different from the first subject. The subject features in the generated image can be fine-tuned according to the target subject provided by the user, so as to generate an image that better meets the user's expectations. At the same time, since the feature alignment process maintains the coherence and naturalness of the image, the generated image is more visually realistic and credible.

[0056] Specifically, the second generated image can be generated based on the trained image generation model. The second loss function IP loss can be introduced to improve the similarity. In addition, considering that the diffusion model has poor generation effect at time steps with large noise intensity, the λ of the second loss function IP loss weight can be adaptively adjusted according to the signal-to-noise ratio at different time steps. IP . The key to improving the similarity of the subject is to let the model focus on the subject visual elements in the user input image and ignore other interfering elements, such as the background. Therefore, an IP encoder can be trained to extract the features of the subject visual elements in the image, and the IP encoder can be used to extract features from the first image and the first generated image generated by the model, align their subject visual representations, and improve the subject similarity of personalized image generation. See Figure 6 , Figure 6 A schematic diagram of a second generated image according to an embodiment of the present disclosure is shown. Figure 6 The second generated image in Figure 5 Compared with the first generated image in , the subject consistency with the first image is improved.

[0057] The second loss function IP loss can calculate the similarity between the first image I and the first generated image ψ θ (x t ,t,C txt ,C i ) The main features are aligned, and the calculation formula is as follows:

[0058]

[0059] Among them, Sim represents similarity calculation, represents the encoder (which can adaptively extract the subject in the image and extract features), ψ θ represents the image generation model, x t represents the intermediate result of diffusion at time step t, C txt Represents the text features of the prompt text, Ci Represents the visual embedding of the first image I, i.e., the target subject features.

[0060] In step S340, the second generated image is semantically and spatially aligned with the first generated image to obtain a target image.

[0061] Among them, the fundamental reason for the balance problem between the subject similarity and editability of image generation is the insufficiency of the training loss function of the relevant technology. The relevant technology only takes the injection of the subject information of the image as the optimization goal. Under the existing data conditions and model capabilities, it is not enough to simultaneously obtain high similarity and editability in the scenario of zero-sample personalized text images, so it faces a trade-off between similarity and editability in use. For editability, the model can retain the text instruction following ability of the first generated image to reduce the damage to the model's text instruction following ability after the reference image information is injected. Therefore, semantic and structural representation alignment can be performed to retain the model's text following ability. See Figure 7 , Figure 7 A schematic diagram of a target image according to an embodiment of the present disclosure is shown. Figure 7 In the example, the target image does not have a child, which means that the target image has enhanced its ability to follow the prompt text and excludes the display of the child according to the prompt text. In addition, the similarity between the old man and the old man in the first image is improved, ensuring the consistency of the subject. It can be seen that the target image can improve similarity and editability at the same time, greatly alleviating the balance problem between similarity and editability in image generation technology, making personalized image generation technology without fine-tuning more practical.

[0062] In some embodiments, semantically and spatially aligning the second generated image with the first generated image to obtain a target image includes:

[0063] The second generated image is spatially aligned with the first generated image, and the second generated image is semantically aligned with the first generated image based on the prompt text to obtain a target image.

[0064] Among them, spatial feature alignment can refer to the alignment of spatial attributes such as the position, shape, and size of objects in an image. This process usually involves techniques such as image registration and image deformation to ensure that the objects in the second generated image are consistent in position with the first generated image. For example, the correspondence between the feature points or feature areas in the second generated image and the first generated image can be compared to find the corresponding relationship between them, thereby determining the transformation relationship between the images. According to the transformation relationship, the second generated image is transformed to align it with the first generated image in space. Semantic feature alignment can refer to the alignment of semantic attributes such as the meaning, context, and theme of the image content to ensure that the content in the second generated image is semantically consistent with the first generated image. For example, the prompt text can be parsed to extract key information such as objects, actions, scenes, etc. to guide the alignment of semantic features. A deep learning model (such as a convolutional neural network CNN, a recurrent neural network RNN, or a transformer Transformer, etc.) is used to understand the image semantically and extract the semantic features in the image. The semantic features of the second generated image are matched with the semantic features of the first generated image to ensure that they are semantically consistent. Spatial feature alignment and semantic feature alignment are usually interdependent. On the one hand, spatial feature alignment provides accurate location information for semantic feature alignment; on the other hand, semantic feature alignment can guide deformation and transformation in the spatial feature alignment process to ensure that the generated image is consistent with the first generated image both semantically and spatially. The target image obtained after spatial feature alignment and semantic feature alignment not only maintains the spatial structure and semantic content of the first generated image, but also ensures the subject similarity in the second generated image, making the target image more visually coherent and natural, while meeting the user's expectations and requirements.

[0065] In some embodiments, aligning the second generated image with the first generated image using spatial features includes:

[0066] Performing spatial feature extraction on the first generated image to obtain a first spatial feature, and performing spatial feature extraction on the second generated image to obtain a second spatial feature;

[0067] The second spatial feature is aligned to the first spatial feature.

[0068] Among them, a feature matching algorithm, such as nearest neighbor search, RANSAC (random sampling consensus algorithm), etc., can be used to determine the correspondence between the spatial features in the first generated image and the second generated image, and the correspondence describes the relative positions of the two in space. According to the correspondence obtained by feature matching, the second generated image is transformed (e.g., translated, rotated, scaled, affine transformed) to align its spatial features with the spatial features of the first generated image.

[0069] In some embodiments, aligning semantic features of the second generated image with the first generated image based on the prompt text includes:

[0070] Extracting features from the prompt text to obtain text features;

[0071] A first matching feature is obtained by matching the text feature in the first spatial feature, and a second matching feature is obtained by matching the text feature in the second spatial feature;

[0072] Determining a first weight of the first matching feature based on the similarity between the text feature and the first matching feature, and determining a second weight of the second matching feature based on the similarity between the text feature and the second matching feature;

[0073] Obtaining a first semantic feature of the first generated image based on the first weight and the first matching feature, and obtaining a second semantic feature of the second generated image based on the second weight and the second matching feature;

[0074] The second semantic feature is aligned with the first semantic feature.

[0075] Among the spatial features of the first generated image and the second generated image, image features matching the prompt text features are queried respectively to determine which image features are semantically closest to the text features. Among the spatial features of the first generated image, the features that are most matched with the prompt text features can constitute the first matching features. Similarly, among the spatial features of the second generated image, the features that are matched with the prompt text features can form the second matching features. In order to measure the similarity between the first matching features and the second matching features and the prompt text features, a weight can be assigned to each matching feature, and the weight will reflect the degree of semantic proximity between the features. For example, the first weight is determined based on the similarity between the first matching feature and the prompt text feature; and the second weight is determined based on the similarity between the second matching feature and the prompt text feature. The semantic features of the image can be further extracted using the above weights and matching features. According to the first weight and the first matching feature, the first semantic feature of the first generated image is calculated, and this feature will reflect the part of the image that is most relevant to the prompt text. Similarly, the second semantic feature of the second generated image is calculated according to the second weight and the second matching feature. The second semantic feature of the second generated image is aligned with the first semantic feature of the first generated image to ensure that the second generated image is semantically consistent with the first generated image and the prompt text. This not only improves the semantic quality of the image, but also ensures that the generated image is consistent with the prompt text in content and theme, improving the followability of the prompt text.

[0076] In some embodiments, method 300 may further include:

[0077] The trained image generation model generates the target image based on the prompt text and the first image;

[0078] Among them, the model parameters of the image generation model are obtained based on minimizing the loss function, and the loss function includes a first loss function for generating the first generated image, a second loss function for generating the second generated image, and a third loss function for generating the target image.

[0079] In some embodiments, the second loss function is obtained based on a similarity between a first subject feature of the first subject and a target subject feature of the target subject.

[0080] In some embodiments, the third loss function is obtained based on spatial similarity and semantic similarity; wherein the spatial similarity includes the similarity between the first spatial features of the first generated image and the second spatial features of the second generated image; the semantic similarity includes the similarity between the first semantic features of the first generated image and the second semantic features of the second generated image.

[0081] Specifically, the third loss function can be divided into two parts, which are used to align the representation with the original text graph model in terms of semantics and structure. The first part of the third loss function can be used for semantic representation comparison, for example, the key text feature K of the attention module of the diffusion model can be used. txt The query feature Q1 of the image generation model and the query feature Q2 of the original text image model are queried to obtain the attention map, and a representation F is aggregated with Q1 and Q2 respectively. The representation represents the response of Q1 and Q2 to the text and has semantic information. Aligning the representation F can avoid the loss of some responses to the text in Q1 after the visual information of the reference image is injected. The second part of the third loss function can be aligned in structure to supplement the constraints of the spatial dimension missing in the first part, so as to avoid the influence of the reference image visual information injection on the structure and layout of the generated graph. For example, the third loss function can be obtained by weighted summing the first part and the second part. The complete model loss function can be the weighted sum of the first loss function, the second loss function and the third loss function. Among them, the first loss function can be a diffusion model loss function, and its corresponding weight can be a, such as 1. It can be seen that the entire model uses the third loss function to improve editability at the same time, and retains the second loss function IP loss as a regularization term to avoid the similarity from decreasing too quickly.

[0082] It can be seen that according to the method of the embodiment of the present disclosure, the main object in the generated image can be consistent with the given main object, while taking into account the high-quality editability of the generated image, that is, the part of the generated image that is not related to the main object should be subject to the control of the given text.

[0083] It should be noted that the method of the embodiment of the present disclosure can be performed by a single device, such as a computer or a server. The method of the present embodiment can also be applied in a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only perform one or more steps in the method of the embodiment of the present disclosure, and the multiple devices will generate images with each other to complete the described method.

[0084] It should be noted that the above describes some embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the above embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0085] Based on the same technical concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides an image generating device, see Figure 8 , the image generating device, the device comprising:

[0086] An acquisition module, used to acquire text resources and visual resources about the target object;

[0087] A matching module, used for generating a target text and a target visual material that match in space based on the text resource and the visual resource;

[0088] A generation module is used to generate a target video based on the target text and the target visual material.

[0089] For the convenience of description, the above device is described by dividing it into various modules according to its functions. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0090] The device of the above embodiment is used to implement the corresponding image generation method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.

[0091] Based on the same technical concept, corresponding to any of the above-mentioned embodiment methods, the present disclosure also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the image generation method described in any of the above embodiments.

[0092] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0093] The computer instructions stored in the storage medium of the above embodiments are used to enable the computer to execute the image generation method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0094] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples. Based on the concept of the present disclosure, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of simplicity.

[0095] In addition, to simplify the description and discussion, and in order not to make the embodiments of the present disclosure difficult to understand, the known power / ground connections to the integrated circuit (IC) chips and other components may or may not be shown in the provided figures. In addition, the device can be shown in the form of a block diagram to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure will be implemented (that is, these details should be fully within the scope of understanding of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present disclosure, it is apparent to those skilled in the art that the embodiments of the present disclosure can be implemented without these specific details or with changes in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0096] Although the present disclosure has been described in conjunction with specific embodiments of the present disclosure, many replacements, modifications and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.

[0097] The embodiments of the present disclosure are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present disclosure should be included in the scope of protection of the present disclosure.

Claims

1. A method for generating an image, comprising: Acquire a prompt text and a first image for indicating a target subject; generating a first generated image based on the prompt text and the first image; Aligning features of the first subject in the first generated image with the target subject to generate a second generated image; The second generated image is semantically and spatially aligned with the first generated image to obtain a target image.

2. The method according to claim 1, wherein: Aligning features of the first subject in the first generated image with the target subject to generate a second generated image includes: Extracting features of a first subject in a first generated image to obtain first subject features, and extracting features of a target subject in the first image to obtain target subject features; The first subject feature is aligned with the target subject feature to obtain the second generated image.

3. The method according to claim 1, wherein: Aligning the second generated image semantically and spatially with the first generated image to obtain a target image, comprising: The second generated image is spatially aligned with the first generated image, and the second generated image is semantically aligned with the first generated image based on the prompt text to obtain a target image.

4. The method according to claim 3, wherein: Aligning the second generated image with the first generated image using spatial features, comprising: Performing spatial feature extraction on the first generated image to obtain a first spatial feature, and performing spatial feature extraction on the second generated image to obtain a second spatial feature; The second spatial feature is aligned to the first spatial feature.

5. The method according to claim 4, wherein: Aligning semantic features of the second generated image with the first generated image based on the prompt text includes: Extracting features from the prompt text to obtain text features; A first matching feature is obtained by matching the text feature in the first spatial feature, and a second matching feature is obtained by matching the text feature in the second spatial feature; Determining a first weight of the first matching feature based on the similarity between the text feature and the first matching feature, and determining a second weight of the second matching feature based on the similarity between the text feature and the second matching feature; Obtaining a first semantic feature of the first generated image based on the first weight and the first matching feature, and obtaining a second semantic feature of the second generated image based on the second weight and the second matching feature; The second semantic feature is aligned with the first semantic feature.

6. The method according to claim 1, further comprising: The trained image generation model generates the target image based on the prompt text and the first image; Among them, the model parameters of the image generation model are obtained based on minimizing the loss function, and the loss function includes a first loss function for generating the first generated image, a second loss function for generating the second generated image, and a third loss function for generating the target image.

7. The method according to claim 6, wherein: The second loss function is obtained based on the similarity between the first subject feature of the first subject and the target subject feature of the target subject; The third loss function is obtained based on spatial similarity and semantic similarity; wherein the spatial similarity includes the similarity between the first spatial feature of the first generated image and the second spatial feature of the second generated image; the semantic similarity includes the similarity between the first semantic feature of the first generated image and the second semantic feature of the second generated image.

8. An image generating device, comprising: An acquisition module, used for acquiring a prompt text and a first image indicating a target subject; An image generating module, configured to generate a first generated image based on the prompt text and the first image; Aligning features of the first subject in the first generated image with the target subject to generate a second generated image; And aligning the second generated image with the first generated image semantically and spatially to obtain a target image.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the image generating method according to any one of claims 1 to 7 when executing the program. 10 . A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the image generation method according to any one of claims 1 to 7.