Image generation method and device, equipment, medium and product
By using different levels of text prompt information in stages during the image generation process for composition and details supplementation, the problem that image generation in the prior art does not meet expectations is solved, and high-quality and stable image generation effect is achieved.
Patent Information
- Application Number
- CN202510638520.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-15
AI Technical Summary
The existing technology of image generation based on text has failed to meet the object's expectations, resulting in insufficient accuracy and stability of image generation.
By obtaining the object input text, the first text prompt information is generated and the initial image composition is obtained by denoising, and the first text prompt information is enhanced based on the object input text to obtain the second text prompt information, which is used to supplement the image details, and finally denoising the initial image composition is performed to generate the target image.
The gradual guidance and robust control of image generation are achieved, the quality, efficiency and stability of image generation are improved, and the accuracy and consistency of the image generation process are ensured.
Smart Images

Figure CN120495475A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image generation, and more specifically, to an image generation method, an image generation device, an electronic device, a computer-readable storage medium, and a computer product. Background Art
[0002] In image generation tasks based on text-to-image (T2I) technology, image content is usually constructed and controlled based on user-input text descriptions to generate images that meet semantic requirements. However, current image generation based solely on input text has a certain degree of failure that does not meet the object's expectations. Therefore, how to improve the accuracy and stability of image generation is an urgent problem that needs to be solved. Summary of the Invention
[0003] The embodiments of the present application provide an image generation method, an image generation device, an electronic device, a computer-readable storage medium, and a computer program product, which can gradually guide and robustly control image generation, thereby improving efficiency and stability while ensuring the quality of image generation.
[0004] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by practice of the present application.
[0005] According to one aspect of an embodiment of the present application, an image generation method is provided, including: obtaining object input text for describing image content, and generating first text prompt information based on the object input text, the first text prompt information being used to describe image composition; denoising a random noise image based on the first text prompt information to obtain an initial image composition; performing text enhancement processing on the first text prompt information based on the object input text to obtain second text prompt information, the second text prompt being used to supplement image details; and denoising the initial image composition based on the second text prompt information to obtain a target image corresponding to the object input text.
[0006] According to one aspect of an embodiment of the present application, an image generation device is provided, including: an acquisition module, configured to acquire object input text used to describe image content, and generate first text prompt information based on the object input text, wherein the first text prompt information is used to describe the image composition; a denoising module, configured to perform denoising processing on a random noise image based on the first text prompt information to obtain an initial image composition; an enhancement module, configured to perform text enhancement processing on the first text prompt information based on the object input text to obtain second text prompt information, wherein the second text prompt is used to supplement image details; and the denoising module, further configured to perform denoising processing on the initial image composition based on the second text prompt information to obtain a target image corresponding to the object input text.
[0007] In one embodiment of the present application, the acquisition module is further used to perform intent recognition on the object input text to obtain key elements in the object input text and the context corresponding to the object input text; select a target text template corresponding to the object input text from multiple text templates based on the key elements and the context; and generate the first text prompt information based on the target text template.
[0008] In one embodiment of the present application, the context includes a selfie context, and the key elements include people; the acquisition module is further used to generate the first text prompt information based on a preset selfie perspective text template, and the selfie perspective text template is used to describe the selfie scene of a specified perspective; the key elements include people and selfie devices, and the acquisition module is further used to generate person prompt information based on a preset person text template, and generate device prompt information based on a preset device text template, wherein the person text template is used to describe the selfie image of the person, and the device text template is used to describe the device information of the selfie device; the person prompt information and the device prompt information are used as the first text prompt information.
[0009] In one embodiment of the present application, the acquisition module is further used to obtain the text length of the object input text; if the text length is less than a preset multiple of the template length of the target text template, the object input text is spliced onto the target text template to obtain the first text prompt information; if the text length is greater than or equal to a preset multiple of the template length, the target text template is used as the first text prompt information.
[0010] In one embodiment of the present application, the first text prompt information includes person prompt information and device prompt information in the selfie context; the denoising module is further used to denoise the randomly noisy image based on the person prompt information to obtain a person composition result, and to denoise the randomly noisy image based on the device prompt information to obtain a device composition result; the device composition result is scaled to obtain a target device composition result; the position of the selfie device is determined according to the object input text, and the target device composition result and the person composition result are merged according to the position of the selfie device to obtain the initial image composition; if the object input text does not include a position field for describing the selfie position, the position of the selfie device is determined to be the bottom of the image; the target device composition result is attached to the bottom of the person composition result to obtain the initial image composition.
[0011] In one embodiment of the present application, the denoising module is further used to fit the target device composition result to the character composition result according to the position of the selfie device, and adjust the rotation angle of the selfie device in the fitted character composition result to obtain a basic image composition; obtain the interaction area between the selfie device and the character in the basic image composition; divide the interaction area into multiple contact areas according to the contact boundary between the selfie device and the character, and set blur strength for the multiple contact areas according to the degree of contact; blur the multiple contact areas according to the blur strength to obtain the initial image composition.
[0012] In one embodiment of the present application, the denoising module is further used to determine the position of the selfie device based on the position field if the object input text includes the position field; obtain the main area of the character in the character composition result, and fuse the target device composition result and the character composition result according to the position of the selfie device and the main area of the character to obtain the initial image composition.
[0013] In one embodiment of the present application, the denoising module is further used to scale the device composition result according to a specified scaling ratio to obtain a target device composition result; or, determine a basic scaling ratio according to the device type of the selfie device, and adjust the basic scaling ratio according to the blank area in the character composition result to obtain a target scaling ratio; scale the device composition result according to the target scaling ratio to obtain a target device composition result.
[0014] In one embodiment of the present application, the enhancement module is further used to splice the object input text and the first text prompt information to obtain the second text prompt information; or, perform text amplification processing on the object input text according to the semantic information of the object input text to obtain the target object input text, and splice the target object input text and the first text prompt information to obtain the second text prompt information.
[0015] In one embodiment of the present application, the first text prompt information includes character prompt information and device prompt information in the selfie context; the enhancement module is further used to semantically enhance the character posture and selfie background of the character prompt information based on the object input text to generate visual point text; based on the character prompt information and the device prompt information, a corresponding description text of the interaction between the character and the device is generated, and based on the interaction description text, an interaction negative prompt is generated, and the interaction negative prompt is used for interaction detail exclusion and error control; the visual point text, interaction description text, interaction negative prompt and the first text prompt information are spliced to obtain the second text prompt information.
[0016] In one embodiment of the present application, the denoising module is further used to perform step-by-step denoising on the initial image composition according to the second text prompt information to obtain an intermediate image; perform detail missing detection on the intermediate image, and generate a supplementary text prompt based on the detection result, wherein the supplementary text prompt is used to describe supplementary local details; and perform step-by-step denoising on the intermediate image according to the supplementary text prompt to obtain the target image.
[0017] In one embodiment of the present application, the denoising module is further used to, in a first denoising sub-stage, denoise the initial image composition according to the second text prompt information to obtain a first noisy image; in a second denoising sub-stage, denoise the key area of the first noisy image according to the second text prompt information to obtain a second noisy image; in a third denoising sub-stage, denoise the non-key area of the second noisy image according to the second text prompt information to obtain the target image.
[0018] According to one aspect of an embodiment of the present application, an embodiment of the present application provides an electronic device, comprising one or more processors; a storage device for storing one or more computer programs, which, when executed by the one or more processors, enables the electronic device to implement the image generation method described above.
[0019] According to one aspect of an embodiment of the present application, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor of an electronic device, the electronic device executes the image generation method as described above.
[0020] According to one aspect of an embodiment of the present application, an embodiment of the present application provides a computer program product, including a computer program, wherein the computer program is stored in a computer-readable storage medium, and a processor of an electronic device reads and executes the computer program from the computer-readable storage medium, so that the electronic device performs the image generation method as described above.
[0021] In the technical solution provided in the embodiment of the present application, image text for describing the image content is obtained, and first text prompt information is generated based on the image text, and the first text prompt information is used to describe the image composition; the random noise image is denoised according to the first text prompt information to obtain an initial image composition, that is, the input text is converted into composition prompt information, and denoising is performed on the composition prompt information to form an initial image composition with reasonable composition but still fuzzy details; then, text enhancement processing is performed on the first text prompt information based on the image text to obtain richer and more specific second text prompt information for supplementing the image details, and then the initial image composition is denoised according to the second text prompt information. On the basis of maintaining the original composition, more refined information is introduced to obtain a high-quality target image corresponding to the image text. The solution provided in the embodiment of the present application divides the image generation process into a composition generation stage and a detail supplementation stage, and uses different levels of text prompt information in the two stages, thereby realizing step-by-step guidance and robust control of image generation, which not only ensures the image generation quality but also improves efficiency and stability.
[0022] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present application, and together with the specification, are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and it is clear that a person of ordinary skill in the art can derive other drawings based on these drawings without inventive effort. In the accompanying drawings.
[0024] Figure 1 It is a schematic diagram of an implementation environment involved in this application;
[0025] Figure 2 is a flowchart of an image generation method shown in an exemplary embodiment of the present application;
[0026] Figure 3 is a schematic diagram of another process of generating a selfie image according to an exemplary embodiment of the present application;
[0027] Figure 4 is a flowchart of another image generation method shown in an exemplary embodiment of the present application;
[0028] Figure 5 This is a schematic diagram showing the influence of the length of the object input text shown in an exemplary embodiment of the present application;
[0029] Figure 6 is a flowchart of another image generation method shown in an exemplary embodiment of the present application;
[0030] Figure 7 is a schematic diagram of an initial composition result shown in an exemplary embodiment of the present application;
[0031] Figure 8 is a flowchart of another image generation method shown in an exemplary embodiment of the present application;
[0032] Figure 9 is a flowchart of another image generation method shown in an exemplary embodiment of the present application;
[0033] Figure 10 FIG. 4 is a flowchart of an image generating method shown in another exemplary embodiment of the present application.
[0034] Figure 11 is a schematic diagram of an image without a selfie device, shown in another exemplary embodiment of the present application;
[0035] Figure 12 is a schematic diagram showing an image of a selfie device according to another exemplary embodiment of the present application;
[0036] Figure 13 FIG. 4 is a structural block diagram of an image generating device shown as an exemplary embodiment of the present application.
[0037] Figure 14 A schematic diagram of the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0038] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numerals in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0039] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically separate entities. That is, these functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0040] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations, nor must they be executed in the order described. For example, some operations may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.
[0041] It should also be noted that the term "plurality" used in this application refers to two or more. "And / or" describes the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.
[0042] The technical solutions of the embodiments of the present application are introduced in detail below.
[0043] See also Figure 1 , Figure 1 1 is a schematic diagram of an implementation environment involved in this application, which includes a terminal 10 and a server 20.
[0044] The terminal 10 is used to obtain the object input text and send the object input text to the server.
[0045] The server 20 is used to generate first text prompt information based on the object input text, where the first text prompt information is used to describe the image composition, and the random noise image is denoised according to the first text prompt information to obtain an initial image composition. Thereafter, the first text prompt information is enhanced based on the object input text to obtain second text prompt information, where the second text prompt is used to supplement the image details, and then the initial image composition is denoised according to the second text prompt information to obtain a target image corresponding to the object input text.
[0046] The server may send the target image to the terminal so that the terminal displays the target image to the object, or the terminal performs subsequent tasks based on the target image, such as video production.
[0047] In some embodiments, the server 20 and the terminal 10 may also independently implement the image generation process, that is, obtain the object input text themselves, and then generate a first text prompt information to denoise the random noise image to obtain an initial image composition, generate a second text prompt information, and denoise the initial image composition to obtain a target image.
[0048] Among them, the aforementioned terminal 10 can be an electronic device such as a smart phone, a tablet, a laptop, a computer, an intelligent voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, etc., and the server 20 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network) and big data and artificial intelligence platforms. There is no restriction on this here.
[0049] The terminal 10 and the server 20 establish a communication connection in advance through a network, so that the terminal 10 and the server 20 can communicate with each other through the network. The network can be a wired network or a wireless network, and there is no limitation here.
[0050] It should be noted that in the specific implementation of this application, at least one of the input text and text prompt information involves object-related information. When the embodiment of this application is applied to a specific product or technology, it is necessary to obtain the object's permission or consent, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0051] The various implementation details of the technical solutions of the embodiments of the present application are described in detail below.
[0052] like Figure 2 As shown, Figure 2 This is a flowchart of an image generation method shown in an embodiment of the present application, which can be applied to Figure 1 In the implementation environment shown, the method can be executed by the terminal or the server, or by the terminal and the server together. In the embodiment of the present application, the method is illustrated by taking the execution of the method by the server as an example. The image generation method may include S210 to S240, which are described in detail as follows.
[0053] S210: Acquire object input text for describing image content, and generate first text prompt information according to the object input text, where the first text prompt information is used to describe image composition.
[0054] In an embodiment of the present application, the object input text is used to describe the image content, which refers to the image content of the image that is to be generated; the object input text corresponds to an image content theme, and therefore the information used to describe the basic composition of the image can be determined based on the image content theme of the object input text to obtain a first text prompt information (Prompt 1), wherein the information describing the basic composition of the image includes character position, background space, etc., and then the first text prompt information is used to describe the image composition, such as spatial layout and subject position. The composition direction in the image generation process can be clarified through the first text prompt information; it should be noted that at this stage, the object input text is more to supplement the image generation background and is not used as detailed guidance.
[0055] S220: De-noising the random noise image according to the first text prompt information to obtain an initial image composition.
[0056] It can be understood that the overall generation process of the Wensheng map adopts an iterative denoising process, that is, starting from a random noise image as the starting point of generation, step-by-step denoising is performed, and each iteration uses the denoising result of the previous step, wherein the random noise image refers to an image using random Gaussian pure noise. When the random noise image is gradually denoised, the semantics of the first text prompt information is injected step by step. After the denoising iteration reaches a preset number of times, an initial image composition with a reasonable composition but still fuzzy details is obtained. The initial image composition is a noisy image. Since the first text prompt information describes the image composition, the initial image composition represents the prototype of the overall composition.
[0057] In one example, the denoising process includes: inputting the first text prompt information and the random noise image into a denoising module in a text graph model, the denoising module iteratively denoising the random noise image while taking the first text prompt vector corresponding to the first text prompt information as a condition, wherein the denoising module can predict the noise based on the random noise image and the first text prompt vector, and then subtract the predicted noise from the random noise image to obtain a predicted denoised image representation, and then predict the noise again based on the predicted denoised image representation and the first text prompt vector, and subtract the predicted noise again from the predicted denoised image representation, and iterate multiple times; wherein in the denoising process, the noise can be predicted by the injected first text prompt vector, and then the predicted noise is gradually subtracted to obtain the initial image composition.
[0058] S230: Perform text enhancement processing on the first text prompt information based on the object input text to obtain second text prompt information, where the second text prompt is used to supplement image details.
[0059] In an embodiment of the present application, text enhancement processing is performed on the first text prompt information based on the object input text. That is, at this stage, the object input text is used as a detailed guide. On the basis of the first text prompt information, text enhancement processing is performed to add detailed description to obtain a second text prompt information (Prompt 2). Compared with the first text prompt information, the details of the image content described in the second text prompt information are more specific.
[0060] Optionally, performing text enhancement processing on the first text hint information includes concatenating the object input text and the first text hint information to obtain a second text hint information. The second text hint information is obtained by prefixing the first text hint information with the object input text. Placing the first text hint information in front forces the image structure to be remembered during denoising, rather than being distracted by the object description. This ensures that details are consistently consistent with the overall composition, regardless of the object input text. Placing the object input text at the back ensures that it will eventually be seen, thereby enriching the details.
[0061] Optionally, performing text enhancement processing on the first text prompt information includes: performing text amplification processing on the object input text according to semantic information of the object input text to obtain target object input text, and splicing the target object input text and the first text prompt information to obtain second text prompt information.
[0062] Perform semantic analysis on the object input text to extract the core elements, attributes and scene factors. Based on these semantics, call a large language model or rule template to "amplify" the original input, such as automatically completing missed modifications, adding more scenes, materials, light and shadow details, to obtain a richer target object input text; after obtaining the target object input text, splice the target object input text to the first text prompt information to obtain a second text prompt information, so that the image details described by the second text prompt information are richer.
[0063] S240: Perform denoising processing on the initial image composition according to the second text prompt information to obtain a target image corresponding to the object input text.
[0064] Using the second text prompt information as a guide, the initial image composition is further subjected to iterative denoising processing to generate image details, such as image lighting, character details (clothing, expressions), etc.; it should be understood that when the second text prompt vector corresponding to the injected second text prompt is used in each detail generation, it is necessary to ensure that the core elements of the image remain consistent, and any additional details can serve as embellishments without destroying the overall composition, so that the final image is both realistic and aesthetically pleasing.
[0065] like Figure 3As shown, assuming that the object input text is "Please generate a person selfie", a first text prompt information is generated based on the object input text to describe the image structure as "person selfie", and the random noise image is denoised based on the first text prompt information to obtain an initial image composition. The initial image composition only has a person selfie outline, and then the first text prompt information is enhanced to obtain a second text prompt information for describing the "person selfie details", such as skin, clothes, expressions, etc., and on the basis of the initial layout generation, the image details are enhanced and fused through a further denoising process of the second text prompt information, so that the final target image is a person selfie, which has both a clear overall composition and rich details.
[0066] Optionally, a text-image model is pre-trained, and the first text prompt information and the random noise image are input into the text-image model to obtain an initial image composition; and the second text prompt information and the initial image composition are input into the text-image model to obtain a target image.
[0067] It is understandable that the image generation method provided in the embodiment of the present application can be applied to various scenes that require fine control of image composition and details, and is not only suitable for Figure 3 The selfie image generation in
[15] can also be applied to other image generation tasks with specific composition requirements, such as advertising images, customized images, etc.
[0068] In an embodiment of the present application, image text for describing image content is obtained, and first text prompt information is generated based on the image text, the first text prompt information is used to describe the image composition; the random noise image is denoised based on the first text prompt information to obtain an initial image composition, that is, the input text is converted into composition prompt information, and denoising is performed on the composition prompt information to form an initial image composition with reasonable composition but still fuzzy details; then, text enhancement is performed on the first text prompt information based on the image text to obtain richer and more specific second text prompt information for supplementing image details, and then the initial image composition is denoised based on the second text prompt information. While maintaining the original composition, more refined information is introduced to obtain a high-quality target image corresponding to the image text. The solution provided in the embodiment of the present application divides the image generation process into a composition generation stage and a detail supplementation stage, and uses different levels of text prompt information in the two stages, thereby achieving step-by-step guidance and robust control of image generation, improving efficiency and stability while ensuring image generation quality.
[0069] In one embodiment of the present application, another image generation method is provided, which can be applied to Figure 1 The implementation environment shown is described by taking the method executed by the server as an example. Figure 4As shown in Figure 2 On the basis of S210 to S240 shown in FIG, the process of generating the first text prompt information in S210 is expanded to S410 to S430; S410 to S430 are described in detail as follows.
[0070] S410: Perform intent recognition on the object input text to obtain key elements in the object input text and the context corresponding to the object input text.
[0071] In an embodiment of the present application, intent recognition is performed on the object input text to determine the text subject of the object input text, and the type of context is determined based on the text subject. Natural language processing (NLP) technology can be used to extract key elements representing the subject object or action in the context from the object input text.
[0072] Optionally, the intent of the object input text can be identified by matching topic keywords of different text topics, or by using a large language model to understand the meaning of the input text.
[0073] Optionally, in addition to providing object input text, the object can also provide data in other modalities, such as a reference image. When performing intent recognition, multimodal feature extraction is performed on the object input text and the reference image to obtain multimodal features, and intent recognition is performed based on the multimodal features.
[0074] S420: Select a target text template corresponding to the object input text from multiple text templates according to key elements and context.
[0075] In an embodiment of the present application, there are multiple predefined text templates, namely Prompt templates, and the template contents of the multiple text templates are used to describe text prompts of different image structures; wherein the templates are preset according to a combination of different contexts and key elements, as shown in Table 1 below.
[0076]
[0077]
[0078] Table 1
[0079] In an embodiment of the present application, based on key elements and context, the most matching one is dynamically selected from multiple Prompt templates to construct a basic composition for generating an image.
[0080] Optionally, when selecting a target text template, the key elements and context corresponding to the object input text can be compared with the context and element types in Table 1 for key phrase hit rate or scored by a large language model (LLM) to select the text template with the most similar semantics.
[0081] Optionally, a style hint template customized for the subject can be generated based on the subject's behavior data. For example, if the subject mentions "fashion style selfie" many times, a template that is closer to the subject's preferences can be generated, and then the customized style hint model can be added to the existing template library.
[0082] S430: Generate first text prompt information according to the target text template.
[0083] Optionally, only the template may be used, that is, the template content of the target text template is used as the first text prompt information, to ensure that the generated composition is highly stable.
[0084] Alternatively, for example, a context involving a selfie and a key element being a person is used as the context. In this case, the first text prompt is generated based on a preset selfie perspective text template. This template is used to describe a selfie scene from a specified perspective. For example, based on Table 1, T2 is selected as the target text template. This template is a preset selfie perspective text template used to describe a selfie scene from a specified perspective. For example, if the selfie perspective is level, the template does not need to include too much detail, but rather focuses on "composition information" such as spatial structure, person position, and shooting angle.
[0085] Optionally, in the context of a selfie, the key elements are the person and the selfie device. Person prompt information is generated based on a preset person text template, and device prompt information is generated based on a preset device text template. The person text template is used to describe the person's selfie image, and the device text template is used to describe the device information of the selfie device. The person prompt information and the device prompt information are used as the first text prompt information. For example, based on Table 1, T1 is selected as the target text template. The target text template includes a person generation template a for describing the person's selfie image and a device generation template b for describing the device information of the selfie device. The person generation template is used as the person prompt information, and the device generation template is used as the device prompt information to obtain the first text prompt information.
[0086] Among them, the template content of the device-generated template is such as "a small top portion of a{xxx}", where {xxx} is a selfie device, which can be determined from the object input text. If the object input text does not indicate the device description of the selfie device, it can be determined by keyword library matching or large language model analysis.
[0087] Through the above scheme, the generation of the templated first text prompt information makes the entire generation process more instructive and predictable, which not only clarifies the core character composition, but also ensures that the device part does not excessively interfere with the overall selfie effect.
[0088] It is understandable that if only the template is used to generate the first text prompt information, the customization intention of the object input is completely ignored. In order to maintain the personalization and composition stability of the generated target image, in an embodiment of the present application, after generating the first text prompt information, it can be decided whether to splice the object input text based on the length of the object input text.
[0089] Optionally, generating a first text prompt information based on a target text template includes: obtaining the text length of the object input text; if the text length is less than a preset multiple of the template length of the target text template, splicing the object input text onto the target text template to obtain the first text prompt information; if the text length is greater than or equal to a preset multiple of the template length, using the target text template as the first text prompt information.
[0090] Among them, if the text length is less than the preset multiple of the template length, it means that the object input text is short, and the object input text will be spliced to the back of the template to form a richer first text prompt. Because the text-based graph model is more sensitive to the front content when processing the prompt (i.e., "front position bias"), when the object input is short, the entire prompt is still dominated by the "selfie template" after splicing. The object content is just a supplementary detail and will not destroy the basic composition setting, making the initially generated picture closer to the object's expectations. Optionally, the preset multiple can be flexibly adjusted according to actual conditions, such as the preset multiple is 2 times.
[0091] like Figure 5 As shown, when the object input text is short, the template content in the first text prompt information is dominant. When the object input text is long, the object input text has an increasingly greater influence, and it is easy to misjudge the composition direction. Therefore, when the text length is greater than or equal to a preset multiple of the template length, only the template content of the target text template is used as the first text prompt information to avoid the text being too long, which causes the focus to shift when the image is generated. That is, it is better to sacrifice some details of the object description to ensure that the composition is highly stable and meets the expectations of selfies.
[0092] Optionally, whether to splice the object input text can be decided based on the length of the object input text and the semantic sparsity of the object input text. For example, semantic sparsity includes descriptiveness and narrative. If the text length is less than a preset multiple of the template length of the target text template, and the object input text is descriptive, that is, the content is mainly modifiers (such as "smile, soft light"), splicing will not disrupt the composition, so the object input text is spliced after the target text template; but if the object input text is narrative, that is, "sitting on the sofa flipping through a book, a car passing by outside the window", it is highly disruptive, and the template content of the target text template is used as the first text prompt information.
[0093] It can be understood that in the selfie context, when generating the first text prompt information based on the preset selfie perspective text template, and generating the character prompt information based on the preset character text template, it is possible to decide whether to splice the object input text based on the length of the object input text, or to decide whether to splice the object input text based on the length of the object input text and the semantic sparsity of the object input text.
[0094] It should be noted that when the object input text is short, it will be spliced with the template in the initial stage to obtain the first text prompt information, and the object input text will be spliced after the first text prompt information in the detail stage. That is to say, the object input text is spliced twice. In the initial stage, the composition is emphasized, and the response to the object input text is very limited. The detail stage truly controls the richness and accuracy of the image content. Therefore, the first object input text affects the skeleton, and the second affects the skin, clothes, light and shadow, background and other details; the object input text appears in both stages in order to control the composition (rough) and details (rich) respectively, and actually undertakes different generation purposes. It is an intentional reinforcement rather than redundancy.
[0095] Optionally, in order to avoid repetition causing transitional response to the object input text or deviation from the composition, a structural separator (such as #) can be used to clearly divide the prompt word into partitions, that is, the second text prompt information is: first text prompt information # object input text.
[0096] Optionally, in the selfie image generation scenario, when the selfie image includes a selfie device, the first text prompt information includes character prompt information and device prompt information in the selfie context. In order to further ensure the accuracy and authenticity of image generation in the detail stage, text enhancement processing is performed on the first text prompt information, including: semantically enhancing the character posture and selfie background of the character prompt information based on the object input text to generate visual point text; generating corresponding interaction description text between the character and the device based on the character prompt information and the device prompt information, and generating interaction negative prompts based on the interaction description text; splicing the visual point text, interaction description text, interaction negative prompts and the first text prompt information to obtain the second text prompt information.
[0097] As described above, the character prompt information is used to describe the character's selfie image, which is a selfie composition that does not include the character's posture and selfie background. Therefore, based on the core elements and attribute elements of the object input text, a semantic description of the character's posture and image background can be added to the character prompt information to obtain a perspective point text. This visual point text specifically describes the character's actions, postures and scene details.
[0098] Since the selfie image includes a selfie device, there is an interactive behavior between the person and the selfie device. Therefore, an interactive description text can be generated based on the person prompt information and the device prompt information. The interactive description text is used to describe the person's hand holding method and interactive behavior with the device. For example, if the person prompt information is "young woman" and the device prompt information is "xx mobile phone", the interactive description text is "holding xx mobile phone in the right hand and pressing the photo button". The negative interactive prompt refers to the interactive errors or inharmonious details that should be avoided when generating the image. It is used for interaction detail elimination and error control to constrain the generation model and improve the generation quality. For example, according to the interactive actions of the interactive description text, some situations that need to be avoided are listed, such as not twisting fingers, hands should not be obscured or overly blurred, and the generated negative interactive prompts are "do not twist fingers" and "avoid the person holding the wrong direction". Among them, the interactive description text is used as a positive constraint, and the negative interactive prompt is used as a negative constraint of the interactive description text, which helps the model understand which generation problems should be avoided and improves the accuracy of composition and semantics.
[0099] The visual point text, interactive description text, and interactive negative prompt are spliced together to obtain an enhanced text prompt, and then the enhanced text prompt is spliced after the first text prompt information to obtain the second text prompt information, such as the second text prompt information is: first text prompt information#enhanced text prompt.
[0100] Through the above scheme, when generating a selfie image based on the second text prompt information, the generated details will be forcibly guided by the first text prompt information, while enhancing the action, background, and posture details, clarifying the interaction between the person and the device, and providing negative constraints on incorrect behavior, making the generated selfie image more realistic and coordinated.
[0101] In an embodiment of the present application, the intention of the object input text is understood, the correct direction of the image composition is ensured, and a suitable template is selected to provide a structured expression framework through the template mechanism, thereby constructing the first text prompt information to ensure the accuracy of the first text prompt information, which is more instructive and expected.
[0102] The present invention provides another method for generating an image. The method can be applied to Figure 1 In the implementation environment shown, the method can be executed by the terminal or the server, or by the terminal and the server together. In the embodiment of the present application, the method is described by taking the server as an example. Figure 6 As shown in Figure 2 Based on the basis shown in Figure 2 The S220 shown in FIG is expanded to S610 to S630, wherein the first text prompt information includes person prompt information and device prompt information in the selfie context. S610 to S630 are described in detail as follows.
[0103] S610: De-noising the randomly noisy image based on the person prompt information to obtain a person composition result, and de-noising the randomly noisy image based on the device prompt information to obtain a device composition result.
[0104] In an embodiment of the present application, the character composition result and the device composition result are generated independently and in parallel based on the character prompt information and the device prompt information to avoid interference of the device elements on the character composition and improve the generation quality of each part. The character composition result and the device composition result are both noisy images, and the denoising process of the randomly noisy image to obtain the character composition result and the device composition result is the same.
[0105] Optionally, the randomly noisy images in the denoising process of the randomly noisy images based on the person prompt information and the denoising process of the randomly noisy images based on the device prompt information may be the same or different, which is not limited here.
[0106] S620: Perform scaling processing on the device composition result to obtain a target device composition result.
[0107] Since the device composition is generated separately, its size usually does not match the entire person image. Especially considering that the device generally only occupies a small part of the lower part of the picture in a selfie scene, it is necessary to scale the device composition result to obtain the target device composition result to ensure that the device does not visually dominate the image or take up too much space.
[0108] Optionally, the device composition result may be scaled according to a specified scaling ratio to obtain a target device composition result, such as scaling the device composition result to 1 / 16 to 1 / 9 of the person composition result.
[0109] Optionally, the scaling ratio of the device part can be dynamically adjusted to make the fusion of the device and the person image more natural and reasonable; for example, a basic scaling ratio is determined according to the device type of the selfie device, and the basic scaling ratio is adjusted according to the blank area in the person composition result to obtain a target scaling ratio; the device composition result is scaled according to the target scaling ratio to obtain a target device composition result.
[0110] It is understandable that the basic zoom ratios (corresponding to the character composition results) of different device types and sizes of selfie devices are also different. The basic zoom ratios corresponding to different device types are preset. For example, if the device type of the selfie device is a mobile phone, most of them are "slightly visible" devices, and only the top needs to be exposed, so the basic zoom ratio is 1 / 12 to 1 / 8. If the device type of the selfie device is a selfie stick, the basic zoom ratio is 1 / 16 to 1 / 10; if the device type of the selfie device is a camera (such as a micro single camera), the visible area is large, and the basic zoom ratio is 1 / 9 to 1 / 6.
[0111] After determining the basic scaling ratio, the space occupied by the blank area in the character composition result can be determined based on the character area, and the basic scaling ratio can be adjusted according to the space occupied by the blank area. If the space occupied by the blank area is greater than the preset space threshold, the middle value of the ratio range of the basic scaling ratio is increased; if the space occupied by the blank area is less than or equal to the preset space threshold, the middle value of the basic scaling ratio is reduced to avoid it blocking the character area.
[0112] Optionally, in other embodiments of the present application, the scaling ratio can be adjusted according to the resolution of the final target image, such as the preset target resolution of the target image, and the basic scaling ratio can be adjusted based on the quotient of the target resolution and the reference resolution, that is, the product of the resolution quotient and the basic scaling ratio is used as the target scaling ratio.
[0113] S630: Determine the location of the selfie device according to the object input text, and fuse the target device composition result and the character composition result according to the location of the selfie device to obtain an initial image composition.
[0114] In an embodiment of the present application, it is necessary to fuse the target device composition result and the character composition result to obtain an initial image composition in a selfie scenario containing a selfie device. When performing the fusion, it is necessary to determine the position of the selfie device. The position of the selfie device refers to the orientation of the device in the image, and then the fusion is performed based on the position of the selfie device.
[0115] The location of the selfie device may be determined based on whether the object input text mentions the device location, as shown in Table 2 below.
[0116] Condition Location of the selfie device default Bottom of image The text mentions "left hand" Left side of the image The text mentions "right-hand selfie" Right side of the image Text containing "close to face" Front of image (center and slightly above)
[0117] Table 2
[0118] As shown in Table 2, if the object input text does not contain a location field for describing the selfie location, the location of the selfie device is determined to be the bottom of the image; the target device composition result is attached to the bottom of the person composition result to obtain the initial image composition; that is, if the object does not specify the device location, the location of the selfie device is determined to be the default location, that is, the bottom of the image, as shown in Figure 7 As shown in (a), the device composition result is scaled, and the scaled target device composition result is attached to the bottom center of the character composition result.
[0119] Optionally, if the object input text contains a location field, it means that the object specifies the device location. In this case, the location of the selfie device is determined based on the location field; the main area of the character in the character composition result is obtained, and the target device composition result and the character composition result are fused based on the location of the selfie device and the main area of the character to obtain the initial image composition.
[0120] During fusion, in order to avoid the target device composition result obscuring the character composition result, it is necessary to also identify the main areas of the character in the character composition result, such as the facial area (head, eyes, nose, etc.), the upper body contour (shoulders, neck, etc.), etc., and then detect whether the position of the selfie device will conflict with the main areas of the character. If there is a conflict, the position of the selfie device will be automatically fine-tuned, such as up and down or left and right offsets, and then based on the offset position of the selfie device, the target device composition result will be fitted to the corresponding position of the character composition result.
[0121] As shown in Table 2, it is determined that the position of the selfie device is in front of the image, but it blocks the facial area of the character composition result. Then, the position of the selfie device is shifted left and right, and then fused, as shown in Figure 2. Figure 7 As shown in (b), the device composition result is scaled, and the scaled target device composition result is attached to the left side of the character composition result.
[0122] Optionally, when performing image pasting and fusion, the pasting can be performed through Gaussian edge or transparency fusion.
[0123] Optionally, when fusing the target device composition result and the person composition result based on the position of the selfie device, the composition direction can be fine-tuned based on the position of the selfie device to ensure that the focus of the image is on the portrait subject and the device details are properly processed. For example, the initial image composition includes:
[0124] According to the position of the selfie device, the target device composition result is attached to the person composition result, and the rotation angle of the selfie device in the attached person composition result is adjusted to obtain the basic image composition; the interaction area between the selfie device and the person in the basic image composition is obtained; according to the contact boundary between the selfie device and the person, the interaction area is divided into multiple contact areas, and blur strength is set for the multiple contact areas according to the degree of contact; the multiple contact areas are blurred according to the blur strength to obtain the initial image composition.
[0125] Among them, the target device composition result is pasted to the corresponding human body position in the character composition result according to the position of the selfie device. For example, if the selfie device is at the bottom of the image, the target device composition is pasted to the bottom of the character composition result. At this time, in order to make the angle, position and posture of the device image consistent in the character image, the rotation angle of the selfie device in the pasted character composition result can be adjusted to make it consistent with the character's perspective, such as slightly tilting the selfie device to conform to the upward shooting posture, and increasing the proportion of the character's front composition to obtain the basic image composition; among them, the natural angle that the device should hold can be inferred based on the line connecting the shoulder and the hand to rotate the selfie device.
[0126] If the selfie device is located on the left side of the image, fit the target device composition to the left hand side of the person in the person composition result, adjust the rotation angle of the selfie device, increase the proportion of the person's front face composition, make the person's face occupy more screen space, such as 60%-70% of the screen, reduce the prominence of the hand area, and obtain the basic image composition.
[0127] Since the target device composition result is fitted to the character composition result, there is an interaction area between the selfie device and the character in the basic image composition. The interaction area refers to the part where there is visual contact or spatial occlusion between the device and the character. The interaction area is larger than the device boundary area. The device boundary area is the area formed based on the contact boundary between the selfie device and the character. For example, the device boundary extends 5-15 pixels. Therefore, the interaction area can be divided into multiple contact areas based on the contact boundary. The interaction area is divided into three contact areas. The first contact area is the device body area, including the camera; the second contact area is the outer edge of the device boundary. The third contact area is the area where the device contacts the hand, and the blur intensity is set for each contact area. For example, if the contact between the main area of the device and the person is 0, the blur intensity of the first contact area is 0% blur to ensure a clear device structure. In the second contact area, the contact between the main area of the device and the person is 50%, and the corresponding blur intensity is 20%-40% blur to achieve a natural transition. In the third contact area, the contact between the main area of the device and the person is 100%, and the corresponding blur intensity is 40%-60% blur to use blur to achieve boundary transition, so that the device is naturally "embedded" into the person's hand or fits the body. Optionally, the blur intensity can also be adjusted according to the material of the selfie device. For example, for smooth metal devices, the sharpness of the reflective part can be appropriately enhanced, but the finger contact point area can be slightly blurred. For matte or non-glossy devices, the soft boundary blur is enhanced to make the transition softer.
[0128] Then, Gaussian blur or edge fusion operations of different intensities are applied to each contact area to obtain a preliminary completed and naturally integrated initial image composition, which serves as the input basis for subsequent detail generation.
[0129] In other embodiments of the present application, when the first text prompt information is generated based on the preset selfie perspective text template, the first text prompt information is directly used to denoise the randomly noisy image to obtain an initial composition result.
[0130] In the embodiment of the present application, a clear division of labor is ensured in the composition of each part (people and equipment). By independently generating and then stitching together, detailed local information is retained and the overall and coordinated integration into the final image is achieved. In particular, the scaling and fitting of the equipment part cleverly avoids common problems when generating hands.
[0131] It should be noted that the present application can realize text-generated images through two-stage denoising, that is, in the first stage, the initial image composition is generated by the first text prompt information, and in the second stage, the target image is generated by supplementing the details based on the initial image composition through the second text prompt information.
[0132] Optionally, the present application can also implement the denoising of the image in at least three stages, which can significantly improve the detail quality of the generated image. Figure 8 As shown, in Figure 2 Based on the basis shown in Figure 2 The S240 shown in FIG is expanded to S810 to S830, wherein S810 to S830 are described in detail as follows.
[0133] S810: Perform a step-by-step denoising process on the initial image composition according to the second text prompt information to obtain an intermediate image.
[0134] S820, performing detail missing detection on the intermediate image, and generating a supplementary text prompt based on the detection result, where the supplementary text prompt is used to describe and supplement local details;
[0135] S830 , performing step-by-step denoising processing on the intermediate image according to the supplementary text prompt to obtain a target image.
[0136] In an embodiment of the present application, in the first stage, an initial image composition is generated through the first text prompt information; in the second stage, the initial image composition is denoised through the second text prompt information to obtain an intermediate image, which is a noisy image and contains less noise than the initial image composition; in the third stage, the intermediate image is further denoised through the supplementary text prompt to obtain the final target image.
[0137] Among them, the initial image composition is denoised according to the second text prompt information until the number of denoising iterations reaches a preset number to obtain an intermediate image, and detail loss detection is performed on the intermediate image to identify areas or elements in the intermediate image that are incomplete, unclear, or do not match the object input text. For example, detail loss detection can be based on an image quality model to score each area in the intermediate image to detect low-scoring areas, such as blurred, artifact, and missing areas, and semantic matching is performed on the intermediate image and the object input text based on a graphic consistency model to detect whether the image truly reflects the "keyword" or semantic entity to obtain a detection result.
[0138] It is understandable that the detection results include areas with missing details, such as the hand area, and the regional problems of the areas with missing details are mapped into descriptive language. The descriptive language is used to supplement local details. For example, when a blurred finger area is detected, the descriptive language is "Enhance the fingers holding the phone".
[0139] Optionally, the descriptive language may be spliced after the second text prompt information to obtain a supplementary text prompt, and then denoising may be continued on the basis of the intermediate image using the supplementary text prompt until a target image free of noise is obtained.
[0140] Optionally, after obtaining the descriptive language, a preset refinement template is used to combine the key area name and the descriptive language to obtain the supplementary text prompt of the third stage. According to the supplementary text prompt, a corresponding area mask is generated, and then the masked area in the intermediate image is locally denoised according to the supplementary text prompt, and the area not selected by the mask is retained. The refined area and the original intermediate image are fused to output the final high-quality target image.
[0141] In an embodiment of the present application, a third stage of denoising is added on the basis of the two-stage denoising (initial composition denoising and fusion detail denoising), which can further refine the local areas in the picture or correct potential defects, thereby improving the final image quality and semantic alignment.
[0142] The present invention also provides another method for generating an image. The method can be applied to Figure 1 In the implementation environment shown, the method can be executed by the terminal or the server, or by the terminal and the server together. In the embodiment of the present application, the method is described by taking the server as an example. Figure 9 As shown, in Figure 2 Based on the basis shown in Figure 2 The S240 shown in FIG is expanded to S910 to S930, wherein S910 to S930 are described in detail as follows.
[0143] S910 : In a first denoising sub-stage, denoising is performed on the initial image composition according to the second text prompt information to obtain a first noisy image.
[0144] S920 . In a second denoising sub-stage, denoising is performed on a key area of the first noisy image according to the second text prompt information to obtain a second noisy image.
[0145] S930 : In the third denoising sub-stage, denoising is performed on the non-critical area of the second noisy image according to the second text prompt information to obtain a target image.
[0146] In an embodiment of the present application, during the second stage (fusion detail denoising), the denoising iterative process can be divided into three sub-stages. In the first denoising sub-stage (i.e., the early stage), the initial image composition is denoised using the second text prompt information, and a blurred form of the main composition elements (e.g., the position of the person, the general outline of the background) is generated globally in the image to obtain a first noisy image. The first noisy image is a preliminary composition sketch with low resolution or high blur. In the second denoising sub-stage (i.e., the middle stage), key areas of the first noisy image (e.g., the expression of the person, close-up of the device) are focused on denoising and optimizing to improve the quality of local details, thereby obtaining a second noisy image. The second noisy image already has high-quality performance in terms of local details, but the background and overall lighting may still have a certain degree of ambiguity. In the third denoising sub-stage (i.e., the late stage), based on the second noisy image, the details of non-critical areas are further denoised and optimized, and the overall lighting and color are adjusted to make the image harmonious and unified, thereby obtaining the final target image, where the non-critical areas include background, lighting, color, etc.
[0147] Optionally, in the middle and late stages of denoising, accurate optimization of key areas and non-key areas can be achieved through regional perception technology (such as attention mechanism) and dynamic weight allocation technology; for example, in the middle stage of denoising, the key areas in the first noisy image are identified through the attention mechanism, and more computing resources (such as higher resolution or finer diffusion steps) can be allocated to these key areas during the denoising iteration process, such as in the middle stage, 90% of resources are used for key areas and 10% of resources are used globally; and through dynamic weight adjustment, a larger weight is assigned to the key areas relative to other areas, and then when denoising is performed through the second text prompt information, the noise distribution and detail generation of the key areas are prioritized. Optionally, in the middle stage of denoising, when denoising the key areas, a mask can also be generated for the key areas, and then the resolution enhancement, texture reconstruction or contour completion of the masked key areas can be performed.
[0148] In the later stage of denoising, 60% of the resources are used for the global background and 40% of the resources are used to maintain the details of the key areas. When the resources are used for the global background, the background in the non-key areas of the second noisy image is identified, and the background and foreground are separated by bounding box division to ensure the independence of the optimized background. The texture in the background area is refined according to the second text prompt information, so that the background and global details are coordinated and unified with the key areas. In addition, the global color and light are set by the second text prompt information to ensure the unity of the global visual effect.
[0149] Through the above scheme, the denoising process of the second stage is divided into three cycles. In the early stage, the overall composition direction of the image and the preliminary form of the main elements are laid, in the middle stage, local details are enhanced, and in the later stage, global details are integrated. Through such a subdivision strategy, local detail enhancement and global detail fusion can effectively cooperate in the denoising process, gradually generating high-quality images with rich details and overall coordination.
[0150] In other embodiments, Figure 8 and Figure 9 The embodiments shown in the figure are combined. For example, when the initial image composition is gradually denoised according to the second text prompt information to obtain an intermediate image, steps S910 to S930 can also be performed, that is, the initial image composition is denoised in the early stage of denoising to obtain a first noisy image, the key area of the first noisy image is denoised in the middle stage of denoising, and the non-key area is denoised in the late stage of denoising to obtain an intermediate image, and then the intermediate image is further denoised after detail loss detection is performed on the intermediate image to obtain a target image.
[0151] For ease of understanding, the present application embodiment also provides an image generation method, which is described by taking the example of generating a selfie image through text. Figure 10 As shown, the image generation method proposed in the embodiment of the present application includes selfie intention recognition, initial prompt (i.e., the aforementioned first text prompt information) generation, initial composition generation, and fusion detail generation, and finally obtains the output result X0, i.e., the selfie image.
[0152] During the selfie intention recognition phase, the selfie type is determined based on the drawn text P input by the subject. The two main cases are:
[0153] Including selfie devices: A mobile phone or other selfie device appears in the picture (for example, the subject description is "holding a mobile phone", "taking a selfie in the mirror", etc.).
[0154] No selfie device: There is no selfie device in the picture, usually only the self-subject and background are shown.
[0155] A simple and fast approach is to use keyword matching to directly scan the subject's input text for keywords such as "holding a mobile phone" and "taking a selfie in the mirror." This method is fast and easy to implement. If more time is available, mature large language models can be used to perform semantic analysis on the input text to determine the specific selfie context. This method is more flexible and can capture more complex or subtle intent.
[0156] The overall generation process of the Wensheng map adopts an iterative denoising process, that is, starting from pure random noise and gradually denoising (i.e. continuously extracting structure and details), and each iteration uses the denoising result of the previous step. During the overall generation process, the time step t will gradually change from 1 to 0, corresponding to the image noise level gradually decreasing from pure random noise to 0. Here, t is defined as the time that distinguishes the initial generation stage and the fusion detail generation stage. thres When t<=t thres This is the initial composition generation stage, which mainly determines the core composition of the image. When t>t thres In order to integrate the detail generation stage, the focus is on supplementing details and refining the picture.
[0157] In the initial prompt generation stage, the core elements in the picture are mainly drawn. If there is no selfie device in the picture: a fixed selfie perspective is used to generate the text template P selfie , that is "(A person is taking aselfie.The background is interior, ground and environment.)(selfie Angle)", change P selfie As the initial prompt; if the subject's text is less than twice the length of the template, the user's text will be spliced onto the template to ensure that the basic information is preserved. If the text is longer, it will not be directly spliced to prevent the subject text from being too long and disrupting the original selfie composition and positioning, causing the initial generated result to not match the selfie pose.
[0158] For the case where a selfie device appears in the picture: the text input by the subject is divided into the person and the shooting device. Among them, the person will use a fixed text template P selfie_person, namely "(A close-up selfie of a person's face, focusing on his face and the top of his head. The composition centers the face, close-up)", where the model is required to generate an image of the upper body of a person; the other part will describe the shooting equipment, using a fixed template P selfie_device 『a small top portion of a{xxx}』. Fill in the camera device here, such as 『a small top portion of a cell phone with three cameras』; the camera device can be obtained from the text input by the object, using the keyword library matching or large language model analysis method, and P selfie_person and P selfie_device As the initial prompt.
[0159] In the initial composition generation phase, the image will be generated based on the extracted text template. In the case where there is no selfie device: the unique text P for generating the composition has been determined. selfie , the text graph model is based on this text P selfie Perform the denoising process to obtain the noisy initial composition image X comp .
[0160] For the case of selfie devices: Since the text is decomposed into portrait-related templates and shooting device-related templates, these two texts will be used independently and in parallel to generate the noisy character composition result X comp_person and the noisy device composition result of the shooting device X comp_device ; Then the output result of the shooting device will be X comp_device Zoom to X comp_person 1 / 16 to 1 / 9 of the image and fit it to the character composition with noise results X comp_person At the bottom, get X comp In this way, in the subsequent process of detail fusion, the shooting equipment only draws part of it, and at the bottom of the picture, the description of the hand will be relatively less, thus preventing the hand from collapsing.
[0161] In the fusion detail generation stage, further denoising will be performed on the generated noisy initial composition image. The text used here will be spliced with the initial prompt text in front of the object input text according to the situation of whether there is a camera or not, so as to prevent the composition from being misplaced; that is, in the case of no camera, P is spliced in front of the object input text P. selfie, and obtain the target prompt (i.e., the aforementioned second text prompt information), such as the target prompt is: "(A person is taking a selfie. The background is interior, ground and environment.) (selfie Angle). P".
[0162] In the case of a camera, splice P before the subject enters the text P selfie_person and P selfie_device , we get the target prompt, such as "(A close-up selfie of a person's face, focusing on his face and the top of his head. The composition centers the face, close-up). a small topportion of {xxx} is visible at the bottom of the frame. P", where xxx refers to the shooting device, and here it will be emphasized that the device is located at the bottom of the image.
[0163] In the embodiment of the present application, if more time is allowed, it can be considered to expand the object input text P based on LLM after obtaining the target prompt to obtain P ape , thus in P ape The pre-stitching initial prompt makes the subsequently generated images richer in details and fuller in picture layers.
[0164] like Figure 11 As shown, the subject enters "generate a selfie of a person" in the Wenshengtu application on the terminal, and Figure 10 The image of the selfie device does not appear in the application interface display.
[0165] like Figure 12 As shown, the subject enters "generate a selfie of a person in the mirror" in the Wenshengtu application on the terminal, and Figure 10 The image of the selfie device appears on the application interface display.
[0166] In an embodiment of the present application, in the cultural image model prediction stage, it is achieved by adjusting the segmented text and constructing the segmented layout in the prediction stage; such a method does not require training of the base model, nor does it require additional conditional input and additional module addition; when the purpose is achieved, the deployment cost is low and the running speed is fast; specifically, on the basis of the cultural image prediction process, a segmented prompt and segmented layout construction scheme is adopted, that is, different prompts are used in the initial denoising composition generation stage and the later detail supplementation stage; and in the initial composition generation stage, key targets other than the hands in the picture are pre-generated, and when the details are subsequently supplemented, the hand and finger details are further added; under the premise of ensuring the effect and efficiency of the generation results, the success rate of generating character selfies is maximized.
[0167] Here, an embodiment of the device of the present application is introduced, which can be used to perform the image generation method in the above embodiment of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the image generation method in the above embodiment of the present application.
[0168] The embodiment of the present application provides an image generating device, such as Figure 13 As shown, the device includes.
[0169] An acquisition module 1310 is configured to acquire object input text for describing image content, and generate first text prompt information based on the object input text, wherein the first text prompt information is used to describe the image composition;
[0170] a denoising module 1320 configured to perform denoising on the random noise image according to the first text prompt information to obtain an initial image composition;
[0171] An enhancement module 1330 is configured to perform text enhancement processing on the first text prompt information based on the object input text to obtain second text prompt information, where the second text prompt is used to supplement image details;
[0172] The denoising module 1320 is further configured to perform denoising processing on the initial image composition according to the second text prompt information to obtain a target image corresponding to the object input text.
[0173] In one embodiment of the present application, based on the aforementioned scheme, the acquisition module is further used to perform intent recognition on the object input text to obtain key elements in the object input text and the context corresponding to the object input text; select a target text template corresponding to the object input text from multiple text templates based on the key elements and the context; and generate the first text prompt information based on the target text template.
[0174] In one embodiment of the present application, based on the aforementioned scheme, the context includes a selfie context, and the key elements include people; the acquisition module is further used to generate the first text prompt information based on a preset selfie perspective text template, and the selfie perspective text template is used to describe the selfie scene of a specified perspective; the key elements include people and selfie devices, and the acquisition module is further used to generate person prompt information based on a preset person text template, and generate device prompt information based on a preset device text template, wherein the person text template is used to describe the selfie image of the person, and the device text template is used to describe the device information of the selfie device; the person prompt information and the device prompt information are used as the first text prompt information.
[0175] In one embodiment of the present application, based on the aforementioned scheme, the acquisition module is further used to obtain the text length of the object input text; if the text length is less than a preset multiple of the template length of the target text template, the object input text is spliced onto the target text template to obtain the first text prompt information; if the text length is greater than or equal to a preset multiple of the template length, the target text template is used as the first text prompt information.
[0176] In one embodiment of the present application, the first text prompt information includes person prompt information and device prompt information in the selfie context; the denoising module is further used to denoise the randomly noisy image based on the person prompt information to obtain a person composition result, and to denoise the randomly noisy image based on the device prompt information to obtain a device composition result; the device composition result is scaled to obtain a target device composition result; the position of the selfie device is determined according to the object input text, and the target device composition result and the person composition result are merged according to the position of the selfie device to obtain the initial image composition; if the object input text does not include a position field for describing the selfie position, the position of the selfie device is determined to be the bottom of the image; the target device composition result is attached to the bottom of the person composition result to obtain the initial image composition.
[0177] In one embodiment of the present application, the denoising module is further used to fit the target device composition result to the character composition result according to the position of the selfie device, and adjust the rotation angle of the selfie device in the fitted character composition result to obtain a basic image composition; obtain the interaction area between the selfie device and the character in the basic image composition; divide the interaction area into multiple contact areas according to the contact boundary between the selfie device and the character, and set blur strength for the multiple contact areas according to the degree of contact; blur the multiple contact areas according to the blur strength to obtain the initial image composition.
[0178] In one embodiment of the present application, based on the aforementioned scheme, the denoising module is further used to determine the position of the selfie device according to the position field if the object input text includes the position field; obtain the main area of the character in the character composition result, and fuse the target device composition result and the character composition result according to the position of the selfie device and the main area of the character to obtain the initial image composition.
[0179] In one embodiment of the present application, based on the aforementioned scheme, the denoising module is further used to scale the device composition result according to a specified scaling ratio to obtain a target device composition result; or, determine a basic scaling ratio according to the device type of the selfie device, and adjust the basic scaling ratio according to the blank area in the character composition result to obtain a target scaling ratio; scale the device composition result according to the target scaling ratio to obtain a target device composition result.
[0180] In one embodiment of the present application, based on the aforementioned scheme, the enhancement module is further used to splice the object input text and the first text prompt information to obtain the second text prompt information; or, perform text amplification processing on the object input text according to the semantic information of the object input text to obtain the target object input text, and splice the target object input text and the first text prompt information to obtain the second text prompt information.
[0181] In one embodiment of the present application, the first text prompt information includes character prompt information and device prompt information in the selfie context; the enhancement module is further used to semantically enhance the character posture and selfie background of the character prompt information based on the object input text to generate visual point text; based on the character prompt information and the device prompt information, a corresponding description text of the interaction between the character and the device is generated, and based on the interaction description text, an interaction negative prompt is generated, and the interaction negative prompt is used for interaction detail exclusion and error control; the visual point text, interaction description text, interaction negative prompt and the first text prompt information are spliced to obtain the second text prompt information.
[0182] In one embodiment of the present application, based on the aforementioned scheme, the denoising module is further used to perform step-by-step denoising on the initial image composition according to the second text prompt information to obtain an intermediate image; perform detail missing detection on the intermediate image, and generate supplementary text prompts based on the detection results, wherein the supplementary text prompts are used to describe supplementary local details; and perform step-by-step denoising on the intermediate image according to the supplementary text prompts to obtain the target image.
[0183] In one embodiment of the present application, based on the aforementioned scheme, the denoising module is further used to, in a first denoising sub-stage, denoise the initial image composition according to the second text prompt information to obtain a first noisy image; in a second denoising sub-stage, denoise the key area of the first noisy image according to the second text prompt information to obtain a second noisy image; in a third denoising sub-stage, denoise the non-key area of the second noisy image according to the second text prompt information to obtain the target image.
[0184] It should be noted that the apparatus provided in the above embodiment and the method provided in the above embodiment belong to the same concept, wherein the specific manner in which each module and unit performs operations has been described in detail in the method embodiment and will not be repeated here.
[0185] The device provided in the above embodiment may be arranged in a terminal or in a server.
[0186] An embodiment of the present application also provides an electronic device, comprising one or more processors and a storage device, wherein the storage device is used to store one or more computer programs, and when the one or more computer programs are executed by one or more processors, the electronic device implements the above image generation method.
[0187] Figure 14 A schematic diagram of the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application is shown.
[0188] It should be noted that Figure 14 The computer system 1400 of the electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present application.
[0189] like Figure 14 As shown, computer system 1400 includes a processor (Central Processing Unit, CPU) 1401, which can perform various appropriate actions and processes according to the program stored in read-only memory (Read-Only Memory, ROM) 1402 or the program loaded from storage part 1408 into random access memory (Random Access Memory, RAM) 1403, such as executing the method in the above embodiment. Various programs and data required for system operation are also stored in RAM 1403. CPU 1401, ROM 1402 and RAM 1403 are connected to each other via bus 1404. Input / output (I / O) interface 1405 is also connected to bus 1404.
[0190] In some embodiments, the following components are connected to the I / O interface 1405: an input section 1406 including a keyboard, a mouse, and the like; an output section 1407 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 1408 including a hard disk; and a communication section 1409 including a network interface card such as a LAN (Local Area Network) card or a modem. The communication section 1409 performs communication processing via a network such as the Internet. A drive 1410 is also connected to the I / O interface 1405 as needed. Removable media 1414, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like, is installed in the drive 1410 as needed, so that computer programs read therefrom can be installed into the storage section 1408 as needed.
[0191] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 1409, and / or installed from the removable medium 1414. When the computer program is executed by the processor (CPU) 1401, the various functions defined in the system of the present application are executed.
[0192] It should be noted that the computer-readable medium shown in the embodiment of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (Erasable Programmable Read Only Memory), a flash memory, an optical fiber, a portable compact disk read-only memory (Compact Disc Read-Only Memory, CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, wherein a computer-readable computer program is carried. This propagated data signal can take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. A computer program embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0193] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions and operations of the devices, methods and computer program products according to various embodiments of the present application. Among them, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, and the above-mentioned module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and a computer program.
[0194] The units or modules described in the embodiments of the present application may be implemented in software or hardware, and the units or modules described may also be provided in a processor. The names of these units or modules do not, in certain circumstances, limit the units or modules themselves.
[0195] Another aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the image generation method described above. The computer-readable storage medium may be included in the electronic device described in the above embodiments, or may exist independently and not be incorporated into the electronic device.
[0196] Another aspect of the present application further provides a computer program product, comprising a computer program stored in a computer-readable storage medium. A processor of an electronic device reads the computer program from the computer-readable storage medium and executes the computer program, causing the electronic device to perform the image generation method described above in each of the above embodiments.
[0197] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiment of the application, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.
[0198] Other embodiments of the present invention will readily occur to those skilled in the art after considering the specification and practicing the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of this application and include common knowledge or customary techniques in the art that are not disclosed herein.
[0199] The above content is only a preferred exemplary embodiment of the present application and is not intended to limit the implementation scheme of the present application. Ordinary technicians in this field can easily make corresponding changes or modifications based on the main concept and spirit of the present application. Therefore, the scope of protection of the present application shall be based on the scope of protection required by the claims.
Claims
1. An image generation method, characterized in that: include: Acquire object input text for describing image content, and generate first text prompt information according to the object input text, wherein the first text prompt information is used to describe the image composition; De-noising the random noise image according to the first text prompt information to obtain an initial image composition; Performing text enhancement processing on the first text prompt information based on the object input text to obtain second text prompt information, where the second text prompt is used to supplement image details; The initial image composition is subjected to denoising processing according to the second text prompt information to obtain a target image corresponding to the object input text.
2. The method according to claim 1, characterized in that Generating first text prompt information according to the object input text includes: Performing intent recognition on the object input text to obtain key elements in the object input text and a context corresponding to the object input text; selecting a target text template corresponding to the object input text from a plurality of text templates according to the key elements and the context; The first text prompt information is generated according to the target text template.
3. The method according to claim 2, characterized in that The context includes a selfie context; the key elements include characters; and generating the first text prompt information according to the target text template includes: Generate the first text prompt information based on a preset selfie perspective text template, where the selfie perspective text template is used to describe a selfie scene at a specified perspective; The key elements include a person and a selfie device; and generating the first text prompt information according to the target text template includes: Generate character prompt information based on a preset character text template, and generate device prompt information based on a preset device text template, wherein the character text template is used to describe the character's selfie image, and the device text template is used to describe device information of the selfie device; The character prompt information and the device prompt information are used as the first text prompt information.
4. The method according to claim 2, characterized in that The generating the first text prompt information according to the target text template includes: Get the text length of the input text of the object; If the text length is less than a preset multiple of the template length of the target text template, the object input text is concatenated with the target text template to obtain the first text prompt information; If the text length is greater than or equal to a preset multiple of the template length, the target text template is used as the first text prompt information.
5. The method according to claim 1, wherein The first text prompt information includes person prompt information and device prompt information in the selfie context; and the denoising process is performed on the randomly noisy image according to the first text prompt information to obtain the initial image composition, including: De-noising the randomly noisy image based on the person prompt information to obtain a person composition result, and de-noising the randomly noisy image based on the device prompt information to obtain a device composition result; Scaling the device composition result to obtain a target device composition result; The position of the selfie device is determined according to the object input text, and the target device composition result and the character composition result are fused according to the position of the selfie device to obtain the initial image composition.
6. The method according to claim 5, characterized in that The determining of the position of the selfie device according to the object input text, and fusing the target device composition result and the character composition result according to the position of the selfie device to obtain the initial image composition, includes: If the object input text does not include a location field for describing a selfie location, determining that the location of the selfie device is the bottom of the image; Pasting the target device composition result to the bottom of the character composition result to obtain the initial image composition; If the object input text includes the location field, determining the location of the selfie device according to the location field; The main area of the person in the person composition result is obtained, and the target device composition result and the person composition result are fused according to the position of the selfie device and the main area of the person to obtain the initial image composition.
7. The method according to claim 5, characterized in that The step of fusing the target device composition result and the person composition result according to the position of the selfie device to obtain the initial image composition includes: Adhere the target device composition result to the person composition result according to the position of the selfie device, and adjust the rotation angle of the selfie device in the adhered person composition result to obtain a basic image composition; Obtaining an interaction area between the selfie device and the person in the basic image composition; Dividing the interaction area into a plurality of contact areas according to a contact boundary between the selfie device and the person, and setting blur strengths for the plurality of contact areas according to a degree of contact; Blurring is performed on the multiple contact areas according to the blurring strength to obtain the initial image composition.
8. The method according to claim 5, characterized in that The scaling process of the device composition result to obtain the target device composition result includes: Scaling the device composition result according to a specified scaling ratio to obtain a target device composition result; Alternatively, a basic zoom ratio is determined according to the device type of the selfie device, and the basic zoom ratio is adjusted according to the blank area in the character composition result to obtain a target zoom ratio; The device composition result is scaled according to the target scaling ratio to obtain a target device composition result.
9. The method according to claim 1, characterized in that Performing text enhancement processing on the first text prompt information based on the object input text to obtain second text prompt information includes: splicing the object input text and the first text prompt information to obtain the second text prompt information; Alternatively, text augmentation processing is performed on the object input text according to semantic information of the object input text to obtain target object input text, and the target object input text and the first text prompt information are concatenated to obtain the second text prompt information.
10. The method according to claim 1, characterized in that The first text prompt information includes person prompt information and device prompt information in the selfie context; Performing text enhancement processing on the first text prompt information based on the object input text to obtain second text prompt information includes: Performing semantic enhancement on the character posture and selfie background of the character prompt information based on the object input text to generate visual point-delineated text; Generate a corresponding interaction description text between the character and the device based on the character prompt information and the device prompt information, and generate a negative interaction prompt based on the interaction description text, wherein the negative interaction prompt is used for interaction detail elimination and error control; The second text prompt information is obtained by splicing the visual point text, the interaction description text, the interaction negative prompt and the first text prompt information.
11. The method according to any one of claims 1 to 10, characterized in that The performing denoising on the initial image composition according to the second text prompt information to obtain the target image includes: performing step-by-step denoising processing on the initial image composition according to the second text prompt information to obtain an intermediate image; Performing detail missing detection on the intermediate image, and generating a supplementary text prompt according to the detection result, wherein the supplementary text prompt is used to describe and supplement local details; The target image is obtained by performing step-by-step denoising processing on the intermediate image according to the supplementary text prompt.
12. The method according to any one of claims 1 to 10, characterized in that The performing denoising on the initial image composition according to the second text prompt information to obtain the target image includes: In a first denoising sub-stage, denoising is performed on the initial image composition according to the second text prompt information to obtain a first noisy image; In the second denoising sub-stage, denoising is performed on the key area of the first noisy image according to the second text prompt information to obtain a second noisy image; In the third denoising sub-stage, denoising is performed on the non-critical area of the second noisy image according to the second text prompt information to obtain the target image.
13. An image generating device, characterized in that: include: an acquisition module, configured to acquire object input text for describing image content, and generate first text prompt information based on the object input text, wherein the first text prompt information is used to describe the image composition; a denoising module, configured to perform denoising processing on the random noise image according to the first text prompt information to obtain an initial image composition; an enhancement module, configured to perform text enhancement processing on the first text prompt information based on the object input text to obtain second text prompt information, wherein the second text prompt is used to supplement image details; The denoising module is further configured to perform denoising processing on the initial image composition according to the second text prompt information to obtain a target image corresponding to the object input text.
14. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, causes the electronic device to execute the method according to any one of claims 1 to 12.
15. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed by a processor of an electronic device, the electronic device is caused to execute the method according to any one of claims 1 to 12.
16. A computer program product, characterized in that The computer program product includes a computer program stored in a computer-readable storage medium. A processor of an electronic device reads and executes the computer program from the computer-readable storage medium, causing the electronic device to perform the method according to any one of claims 1 to 12.