A method for generating a theme image and an agent framework based on a large language model

Through the combination of large language model and pre-trained text to image model, the attention mechanism and cross-entropy weighted fusion technology are used to solve the problem of difficult separation of main elements in traditional image generation methods. The generated theme images have high precision and high quality, and are suitable for a variety of application scenarios.

CN120163903BActive Publication Date: 2025-07-22SHENZHEN LINGTU SHINE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510636068.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-07-22
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

Traditional image generation methods are difficult to effectively separate and modify the main elements in the image, resulting in the generated theme images with backgrounds and difficult to apply to specific scenes.

Method used

The user input is extended as prompt information through a large language model, the candidate images are generated using pre-trained text to the image model, and the attention map of the main elements is extracted through the attention mechanism, cross-entropy is calculated for weighted fusion, and the GrabCut algorithm is used for foreground segmentation to generate the theme image with transparency channels.

Benefits of technology

It realizes high-precision separation of the subject image, removes unnecessary elements, and the generated subject image can be directly applied to other scenes such as posters, improving the quality and accuracy of image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163903B_ABST
    Figure CN120163903B_ABST
Patent Text Reader

Abstract

The present application discloses a method for generating a subject image and an agent framework based on a large language model. The method includes: in the large language model, expanding key information related to the subject into prompt information; in a pre-trained text-to-image model, generating a candidate image including three color channels based on the prompt information and the key information; extracting an attention map corresponding to the main elements in the candidate image through an attention mechanism; calculating the cross entropy of the attention map at time step t and attention layer l below; according to the cross entropy, performing weighted fusion on the attention maps of the total time step T and the total attention layers L of the pre-trained text-to-image model; using the fused attention map as guidance information to predict a mask of the subject image and performing foreground segmentation to separate the subject image with an alpha channel. The present application realizes the application of entropy-based weighted fusion technology in image generation, can effectively remove unnecessary elements, and the separated subject image has higher accuracy and quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method for generating a theme image and an agent framework based on a large language model. Background Art

[0002] Traditional image generation methods, such as traditional models (e.g., GAN), learn the distribution of data through the latent space. However, the feature representations are highly coupled. For example, the object attributes (color, shape, position) in an image are often encoded in the same latent vector, lacking a mechanism for explicit separation. As a result, the generated theme images often have backgrounds. When a user wants to apply a specific element to certain specific scenarios (such as a poster), it is difficult for the user to separately remove or modify a certain element, difficult to separate unwanted elements, and unable to distinguish the main elements. Summary of the Invention

[0003] The purpose of the present application is to provide a method for generating a theme image and an agent framework based on a large language model to at least solve the above technical problems. The many technical effects that can be produced by the alternative technical solutions among the many technical solutions provided by the present invention are described in detail below.

[0004] To achieve the above purpose, in the first aspect, the present application provides a method for generating a theme image, including:

[0005] In a large language model, expand the key information related to the theme into prompt information;

[0006] In a pre-trained text-to-image model, generate a candidate image including three color channels based on the prompt information and the key information;

[0007] In the pre-trained text-to-image model, extract the attention map corresponding to the main element in the candidate image through an attention mechanism;

[0008] Perform a sampling operation on the attention map, and based on the time step t, the text token, the length of the text token, and the attention layer of the pre-trained text-to-image model l , calculate the cross entropy of the attention map at the time step t and the attention layer l ;

[0009] According to the cross entropy, perform weighted fusion on the attention maps of the total time step T and the total attention layer L of the pre-trained text-to-image model;

[0010] Use the fused attention map as guidance information to predict the mask of the theme image, and perform foreground segmentation to separate the theme image with an alpha channel.

[0011] In some embodiments, in the large language model, expanding the key information related to the theme into prompt information includes:

[0012] The large language model receives the key information related to the theme input by the user;

[0013] Fine-tune the large language model to expand the key information into prompt information and extract the main elements associated with the key information from the prompt information.

[0014] In some embodiments, generating a candidate image including three color channels based on the prompt information and the key information includes:

[0015] In the pre-trained text-to-image model, generate the candidate image based on the prompt information and the main elements.

[0016] In some embodiments, based on the time step t, the text token, the length of the text token, and the attention layer of the pre-trained text-to-image model l , calculate the cross-entropy of the attention map at the time step t and the attention layer l , which is represented by the following formula:

[0017] ;

[0018] Where represents the cross-entropy, t represents the sampling time step, l represents the attention layer of the pre-trained text-to-image model, N represents the length of the text token, n represents the text token, represents the time step t, the attention layer l , the histogram probability of the attention map of the text token n , represents the logarithmic function with base 2.

[0019] In some embodiments, performing weighted fusion on the attention maps of the total time step T and the total attention layers L of the pre-trained text-to-image model according to the cross-entropy includes:

[0020] Calculate the weights of the attention maps according to the cross-entropy ;

[0021] Based on the weights, the total time step T, the total number of attention layers L of the pre-trained text-to-image model, and the attention map , perform weighted fusion to obtain the fused attention map.

[0022] In some embodiments, the theme image generation method further includes:

[0023] Normalize the fused attention map and calculate the probability values of each region of the fused attention map;

[0024] In the fused attention map, divide the regions with probability values greater than the first threshold into definite foreground regions, divide the regions with probability values between the second threshold and the first threshold into possible foreground regions, divide the regions with probability values between the second threshold and the third threshold into possible background regions, and divide the regions with probability values less than the third threshold into definite background regions to obtain a four-value map;

[0025] Wherein, the third threshold is less than the second threshold, and the second threshold is less than the first threshold.

[0026] In some embodiments, using the fused attention map as guidance information to predict the mask of the subject image and perform foreground segmentation to separate the subject image with an alpha channel includes:

[0027] Use the four-value map and the candidate image as the input of the GrabCut algorithm to perform foreground segmentation on the candidate image to generate a final mask;

[0028] Generate a subject image with an alpha channel using the final mask.

[0029] In a second aspect, the present application provides an agent framework based on a large language model, including:

[0030] An input module for receiving the key information related to the subject from the user;

[0031] An extended agent module for expanding the key information related to the subject into prompt information in the large language model;

[0032] An output module for generating a candidate image including three color channels based on the prompt information and the key information in a pre-trained text-to-image model;

[0033] An extraction agent module for extracting the attention map corresponding to the main elements in the candidate image through an attention mechanism in the pre-trained text-to-image model;

[0034] A sampling module for performing a sampling operation on the attention map and calculating the cross-entropy of the attention map at the time step t, based on the text token, the length of the text token, and the attention layer of the pre-trained text-to-image model l and l at the attention layer

[0035] A weighted fusion module for weighted fusion of the attention maps of the total time steps T and the total attention layers L of the pre-trained text-to-image model according to the cross entropy.

[0036] A prediction module for using the fused attention map as guidance information to predict a mask of the subject image and perform foreground segmentation to separate the subject image with an alpha channel.

[0037] In some embodiments, the large language model-based proxy framework further includes a filtering module for filtering out abstract words in the attention map.

[0038] In some embodiments, the large language model-based proxy framework further includes an optimization proxy module for iteratively optimizing the prompt information until the prompt information meets a preset condition.

[0039] In a third aspect, the present application provides a processing device, including: one or more processors; a memory for storing one or more computer programs, and one or more of the processors for executing the one or more computer programs stored in the memory, so that the one or more processors execute a method for generating a subject image according to any one of the first aspect.

[0040] In a fourth aspect, the present application provides an electronic device, and the electronic device includes the processing device described in the third aspect.

[0041] Implementing one of the above technical solutions of the present application has the following advantages or beneficial effects:

[0042] The method for generating a subject image and the large language model-based proxy framework of the present application expand the user's simple input into a detailed scene description through the large language model as prompt information, and generate candidate images through the pre-trained text-to-image model; in the pre-trained text-to-image model, the attention maps corresponding to the main elements in the candidate images are extracted through the attention mechanism to achieve a preliminary positioning of the main elements; sampling operations are performed on the attention maps, and the cross entropy of the attention maps at time step t and attention layer l is calculated, and the attention maps of the total time steps T and the total attention layers L are weighted and fused to realize the application of the entropy-based weighted fusion technology in image generation. Different fused attention maps focus on different main elements, can be decoupled from other main elements, and can effectively remove unnecessary elements when predicting the subject image, and the separated subject image has higher accuracy and quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. In the drawings:

[0044] Figure 1 is a schematic flowchart of the method for generating the subject image in the embodiments of the present application;

[0045] Figure 2 is a schematic block diagram of the agent framework based on the large language model in the embodiments of the present application;

[0046] Figure 3 is another schematic block diagram of the agent framework based on the large language model in the embodiments of the present application;

[0047] Figure 4 is a schematic diagram of the subject image in the embodiments of the present application. Detailed Embodiments

[0048] In order to make the objectives, technical solutions and advantages of the present application more clear, the various exemplary embodiments to be described below will refer to the corresponding drawings, which form a part of the exemplary embodiments and describe various exemplary embodiments that may be adopted to implement the present application. Unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. It should be understood that they are only examples of processes, methods, devices, etc. consistent with some aspects of the present application disclosed in detail in the appended claims. Other embodiments may also be used, or structural and functional modifications may be made to the embodiments listed herein without departing from the scope and essence of the present application.

[0049] In the description of the present application, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. The meaning of the term "plurality" is two or more. The terms "connected" and "coupled" should be understood in a broad sense. For example, they can be fixedly connected, detachably connected, integrally connected, mechanically connected, electrically connected, communicatively connected, directly connected, indirectly connected through an intermediate medium, and can be the internal communication of two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.

[0050] To illustrate the technical solutions described in the present application, the following will be described through specific embodiments, and only the parts related to the embodiments of the present application are shown.

[0051] Example 1:

[0052] As Figure 1 shown, the present application provides a method for generating a theme image. The method for generating a theme image includes the following steps S10 to S60.

[0053] S10: In a large language model, expand key information related to a theme into prompt information.

[0054] In an image, a theme refers to the core idea, emotion, or narrative focus that a work intends to convey, and is the core information conveyed through visual elements such as composition, color, symbols, etc. Themes can be, for example, Christmas, Christmas trees, Santa Claus, etc.

[0055] The key information may include text information related to main elements or scenes. For example, if the theme is Christmas, the main element can be "Christmas".

[0056] A large language model is currently a natural language processing technology based on deep learning. By training a large amount of text data, it learns the patterns and rules of language and can be applied to text generation, question - answering systems, machine translation, etc.

[0057] In some embodiments, step S10 may include:

[0058] The large language model receives key information related to the theme input by the user;

[0059] Fine - tune the large language model to expand the key information into prompt information, and extract main elements associated with the key information from the prompt information.

[0060] Specifically, the user inputs key information related to the theme into the large language model, which can be an informal and simple input. After the large language model receives the key information, it can be fine - tuned, and the key information can be expanded into a detailed scene description to obtain prompt information.

[0061] The prompt information is a more abundant expression of the key information. For example, the key information input by the user into the large language model is several simple prompt words: "Santa Claus", "bell", "star". Then, when the large language model is fine - tuned, it enables the large language model to expand the above - mentioned prompt words. For example, the expanded prompt information is "cartoon style, a cheerful cartoon Santa Claus wearing red clothes, standing beside a decorated Christmas tree, a golden bell hanging on the branch, and a bright star twinkling above".

[0062] Moreover, the large language model can also extract the main elements associated with the key information from the prompt information. The main elements refer to the nouns directly related to the generation target, usually the key entities describing the foreground objects. For example, in the description "Santa Claus stands beside the Christmas tree", the main elements are "Santa Claus" and "Christmas tree".

[0063] S20: In the pre-trained text-to-image model, based on the prompt information and the key information, generate a candidate image including three color channels.

[0064] The pre-trained text-to-image model refers to a trained image generation model that can generate the required image using the text content. Here, the text content is the main elements in the prompt information and the key information.

[0065] In some of the embodiments, step S20 may include:

[0066] In the pre-trained text-to-image model, based on the prompt information and the main elements, generate the candidate image.

[0067] Specifically, the candidate image is generated by the pre-trained text-to-image model based on the prompt information and the main elements. The candidate image is an image including three color channels, namely the R (red) channel, the G (Green) channel, and the B (Blue) channel. The R channel represents the intensity of the red component in the image, the G channel represents the intensity of the green component in the image, and the B channel represents the intensity of the blue component in the image.

[0068] For example, if the prompt information is "Cartoon style, a happy cartoon Santa Claus in red clothes stands beside the decorated Christmas tree, with golden bells hanging on the branches and bright stars twinkling above", and the main elements are "Santa Claus", "Christmas tree", "bell", "star", then on the candidate image generated by the pre-trained text-to-image model, there are these main elements of "Santa Claus", "Christmas tree", "bell", "star", and moreover, the picture is related to the content of the prompt information.

[0069] S30: In the pre-trained text-to-image model, extract the attention map corresponding to the main elements in the candidate image through the attention mechanism.

[0070] There are one or more main elements in the candidate image, such as the aforementioned main elements of "Santa Claus", "Christmas tree", "bell", "star". One main element corresponds to one attention map, that is, "Santa Claus" corresponds to one attention map, "Christmas tree" corresponds to another attention map, and "star" corresponds to another attention map.

[0071] Through the attention mechanism, the attention map corresponding to the main elements initially separated does not well reflect the theme. It can only be initially seen that there are main elements, but the edge lines of the main elements are relatively blurred. If only using these initial attention maps to generate the theme image, the resulting theme image will carry the background and cannot be directly applied to other scenarios, such as directly pasted on a poster, and the background of the theme image will exist on the poster.

[0072] Extract the attention map of one or more main elements in the candidate image through the attention mechanism to achieve the initial positioning of the main elements. For subsequent cross-attention to extract the feature information of the main elements to generate guiding information, use this guiding information to guide the background removal and retain the foreground, which is beneficial to the generation of high-quality theme images.

[0073] S40. Perform a sampling operation on the attention map, and based on the time step t, the text token, the length of the text token, and the attention layer of the pre-trained text-to-image model l , calculate the cross-entropy of the attention map at the time step t and the attention layer l .

[0074] In some embodiments, taking the example that one main element corresponds to one attention map, in step S40, based on the time step t, the text token, the length of the text token, and the attention layer of the pre-trained text-to-image model l , calculate the cross-entropy of the attention map at the time step t and the attention layer l , which can be expressed by the following formula:

[0075] ;

[0076] where represents the cross-entropy, t represents the time step, l represents the attention layer of the pre-trained text-to-image model, N represents the length of the text token, n represents the text token, represents the time step t, the l attention layer, the histogram probability of the attention map of the text token n , represents the logarithmic function with base 2.

[0077] After calculating the cross-entropy of the attention map corresponding to the main element, the cross-entropy map of the main element can well reflect the theme, and its edge lines are relatively clear.

[0078] By introducing the cross-entropy mechanism, the attention map can more accurately reflect the main element and achieve the capture of the user's intention.

[0079] S50. Weightedly fuse the attention maps of all time steps T and the total attention layer L of the pre-trained text-to-image model according to the cross-entropy.

[0080] In some embodiments, taking one main element corresponding to one attention map as an example, step S50 may include:

[0081] According to the cross-entropy , calculate the weights of the attention maps;

[0082] Based on the weights, the total time step T, the total attention layer L of the pre-trained text-to-image model, and the attention maps, perform weighted fusion to obtain the fused attention map.

[0083] Specifically, calculating the weights of the attention maps is achieved through the following formula:

[0084] ;

[0085] where, represents the weight, represents the cross-entropy. is a very small constant used to avoid a zero denominator.

[0086] For one attention map corresponding to one main element, after calculating the cross-entropy of one attention map, calculate the weights and perform weighted fusion to obtain the fused attention map of one main element.

[0087] Based on the weights, the total time step T, the total attention layer L of the pre-trained text-to-image model, and the attention maps, perform weighted fusion to obtain the fused attention map, which is calculated through the following formula:

[0088] ;

[0089] where, represents the attention map that fuses the total time step T and the total attention layer L, represents the cross-attention map at time step t and attention layer l , represents the corresponding weight, T represents the total number of time steps, and L represents the total number of attention layers.

[0090] The highlighted part of the fused attention map is more consistent with the theme image to be generated. Specifically, different fused attention maps focus on different main elements and can be decoupled from other main elements, enabling each main element to be used independently.

[0091] For example, in the attention map corresponding to the main element "star", only the main element "star" in the fused attention map is the highlighted part, which is more consistent with the theme image of "star".

[0092] S60. Use the fused attention map as guidance information to predict the mask of the theme image and perform foreground segmentation to separate the theme image with an alpha channel.

[0093] In some embodiments, the theme image generation method may further include:

[0094] Normalize the fused attention map and calculate the probability values of each region of the fused attention map;

[0095] In the fused attention map, divide the regions with probability values greater than the first threshold into definite foreground regions, divide the regions with probability values between the second threshold and the first threshold into probable foreground regions, divide the regions with probability values between the second threshold and the third threshold into probable background regions, and divide the regions with probability values less than the third threshold into definite background regions to obtain a four-value map;

[0096] Wherein, the third threshold is less than the second threshold, and the second threshold is less than the first threshold.

[0097] For example, if the first threshold is 0.8, the second threshold is 0.2, and the third threshold is 0.1, then after normalizing the fused attention map and calculating the probability values, in the fused attention map, the regions with probability values greater than 0.8 are divided into definite foreground regions, the regions with probability values between 0.2 and 0.8 are divided into probable foreground regions, the regions with probability values between 0.1 and 0.2 are divided into probable background regions, and the regions with probability values less than 0.1 are divided into definite background regions. After dividing the fused attention map into four types of regions, a four-value map is obtained.

[0098] The obtained four-value map is used for image segmentation, so as to be able to separate the theme image.

[0099] Therefore, in some embodiments, step S60 may include:

[0100] Use the four-value map and the candidate image as the input of the GrabCut algorithm to perform foreground segmentation on the candidate image to generate a final mask;

[0101] Use the final mask to generate a theme image with an alpha channel.

[0102] Specifically, GrabCut is an interactive image segmentation method that uses the fused attention map as guidance information to optimize the foreground and background regions, accurately extract the target object, combines mask generation and transparency synthesis to achieve background removal, while retaining the foreground theme, thereby generating a high-quality thematic image with an alpha channel. The background of this thematic image is transparent, that is, the main elements can be separated independently, and this thematic image can be effectively applied to scenarios such as posters without a background.

[0103] In the embodiments of the present application, through a large language model, the simple input of the user is expanded into a detailed scene description as prompt information, and a candidate image is generated through a pre-trained text-to-image model; in the pre-trained text-to-image model, the attention map corresponding to the main elements in the candidate image is extracted through an attention mechanism to achieve a preliminary positioning of the main elements; a sampling operation is performed on the attention map, and the cross-entropy of the attention map at time step t and attention layer l is calculated, and the attention maps of the total time steps T and the total attention layers L are weighted and fused to realize the application of the entropy-based weighted fusion technology in image generation. Different fused attention maps focus on different main elements, can be decoupled from other main elements, and can effectively remove unnecessary elements when predicting the thematic image, and the separated thematic image has higher accuracy and quality.

[0104] As Figure 2 shown, the present application also provides a large language model-based proxy framework 100, including: an input module 101, an extension proxy module 102, an output module 103, an extraction proxy module 104, a sampling module 105, a weighted fusion module 106, and a prediction module 107.

[0105] Among them, the input module 101 is used to receive the key information related to the theme from the user;

[0106] The extension proxy module 102 is used to expand the key information related to the theme into prompt information in the large language model; the extension proxy module 102 can adopt the extension proxy function EXTENSION AGENT.

[0107] The output module 103 is used to generate a candidate image including three color channels in the pre-trained text-to-image model based on the prompt information and the key information;

[0108] The extraction proxy module 104 is used to extract the attention map corresponding to the main elements in the candidate image through an attention mechanism in the pre-trained text-to-image model;

[0109] The sampling module 105 is used to perform a sampling operation on the attention map and calculate the cross-entropy of the attention map at the time step t, based on the text token, the length of the text token, and the attention layer of the pre-trained text-to-image model l and the attention layer l at the time step t;

[0110] The weighted fusion module 106 is used to perform weighted fusion on the attention maps of the total time steps T and the total attention layers L of the pre-trained text-to-image model according to the cross-entropy;

[0111] The prediction module 107 is used to use the fused attention map as guidance information to predict the mask of the subject image and perform foreground segmentation to separate the subject image with an alpha channel.

[0112] In some embodiments, the extended proxy module 102 is further used for:

[0113] Receiving key information related to the subject input by the user in the large language model;

[0114] Fine-tuning the large language model to expand the key information into prompt information and extract the main elements associated with the key information from the prompt information.

[0115] In some embodiments, the output module 103 is further used for:

[0116] Generating the candidate image in the pre-trained text-to-image model based on the prompt information and the main elements.

[0117] In some embodiments, the sampling module 105 is further used for:

[0118] It is represented by the following formula:

[0119] ;

[0120] where represents the cross-entropy, t represents the sampling time step, l represents the attention layer of the pre-trained text-to-image model, N represents the length of the text token, n represents the text token, represents the time step t, the attention layer l and the attention map of the text token n represents the histogram probability, and represents the logarithmic function with base 2.

[0121] In some embodiments, the weighted fusion module 106 is further used for:

[0122] According to the cross-entropy , calculate the weights of the attention map;

[0123] Based on the weights, the total number of time steps T, the total number of attention layers L of the pre-trained text-to-image model, and the attention map , perform weighted fusion to obtain a fused attention map.

[0124] In some embodiments, it further includes a four-value map determination module for:

[0125] Normalize the fused attention map and calculate the probability values of each region of the fused attention map;

[0126] In the fused attention map, divide the regions with probability values greater than the first threshold into determined foreground regions, divide the regions with probability values between the second threshold and the first threshold into possible foreground regions, divide the regions with probability values between the second threshold and the third threshold into possible background regions, and divide the regions with probability values less than the third threshold into determined background regions to obtain a four-value map;

[0127] Wherein, the third threshold is less than the second threshold, and the second threshold is less than the first threshold.

[0128] In some embodiments, the prediction module 107 is further configured to:

[0129] Use the four-value map and the candidate image as the input of the GrabCut algorithm to perform foreground segmentation on the candidate image to generate a final mask;

[0130] Generate a subject image with an alpha channel using the final mask.

[0131] It should be noted that each module in the large language model-based proxy framework 100 has been described in detail in the subject image generation method and will not be elaborated here.

[0132] In some of these embodiments, as Figure 3 shown, the large language model-based proxy framework 100 may further include a filtering proxy module 108 for filtering out abstract words in the attention map.

[0133] The filtering proxy module 108 employs a filtering proxy function FILTER.

[0134] Abstract words such as "color" and "detail", after filtering, can retain specific words related to the main elements, ensuring that the attention map only focuses on the main elements and avoiding interference from irrelevant information.

[0135] In some of these embodiments, the large language model-based agent framework 100 may further include an optimization agent module 109 for iteratively optimizing the prompt information until the prompt information meets a preset condition.

[0136] The optimization agent module 109 employs an optimization agent function OPTIMIZER.

[0137] The optimization agent module 109 optimizes the prompt information through self-reflection and multiple rounds of iteration to ensure the accuracy and richness of the prompt information.

[0138] The entire agent framework 100 can automatically process the user's informal input and improve the quality of the subject image. The large language model-based agent framework 100 of this application has been proven through experiments to have significant advantages in terms of accuracy and efficiency. It can effectively solve the challenges faced by existing text-to-image models in practical applications, is more efficient than existing autoregressive model-based methods, and has higher accuracy. At the same time, the subject image generation method and the agent framework 100 of this application can be applied to various application scenarios, such as textile pattern design and meme generation, etc. It can help users better express their needs, effectively remove unnecessary elements, generate high-definition main elements, and further improve the quality and accuracy of image generation.

[0139] As Figure 4 shown, Figure 4 is a schematic diagram of the subject image of this application. It can be seen from Figure 4 that Figure 4 is a subject image of generating flowers. It can be seen that the main elements of the theme "flowers" are very obvious, with high definition, no irrelevant background, and can be directly applied to scenarios such as posters.

[0140] For a large language model-based agent framework 100 of this application, the input module 101 provides informal input to the user, such as simple descriptions like keywords. The extension agent module 102 uses a large language model to expand the key information into a detailed scene description to obtain prompt information. The output module 103 uses a pre-trained text-to-image model to generate candidate images and guides entropy-based feature weighted fusion. The extraction agent module 104 can extract the attention map corresponding to the main elements to achieve preliminary localization of the main elements. The sampling module 105 can calculate the cross entropy, and the weighted fusion module 106 performs weighted fusion on all attention maps according to the cross entropy. The prediction module 107 performs high-precision prediction to realize the application of entropy-based weighted fusion technology in image generation. Different fused attention maps focus on different main elements, can be decoupled from other main elements, and can effectively remove unnecessary elements when predicting the subject image, and the separated subject image has higher precision and quality.

[0141] Those of ordinary skill in the art can understand that all or part of the features / steps of implementing the above method embodiments can be realized by a method, a data processing system, or a computer program. These features can be implemented without using hardware, entirely using software, or using a combination of hardware and software. The aforementioned computer program can be stored in one or more computer-readable storage media. When the computer program stored on the storage media is executed (such as by a processor), it performs the steps of the method embodiments for generating the subject image as described above.

[0142] The aforementioned storage media that can store program codes include: a static hard disk, a solid-state drive, a random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), an optical storage device, a magnetic storage device, a flash memory, a magnetic disk or an optical disc, and / or a combination of the above devices, that is, it can be implemented by any type of volatile or non-volatile storage device or a combination thereof.

[0143] Embodiment 2:

[0144] This application also provides a processing device, including: one or more processors; a memory for storing one or more computer programs, and one or more of the processors are used to execute the one or more computer programs stored in the memory, so that one or more of the processors execute a method for generating a subject image as described in any one of Embodiment 1.

[0145] The processing device of the present invention can be a chip. The processing device provided by the present invention can be, for example, a large model processing device. When the processing device is a chip, each module in the chip can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computing device in the form of hardware, or stored in the memory of the computing device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules. The present invention can be applied to fields with high-precision, high-reliability, low-power consumption, and low-area-cost computing power requirements such as artificial intelligence and the metaverse, artificial intelligence training or inference chip products, autonomous driving chips, VR chips, robot-built-in chips, and parallel computing deep learning applications, etc. It can be applied to fields with high-precision, high-reliability, low-power consumption, and low-area-cost computing power requirements such as artificial intelligence and the metaverse, such as artificial intelligence training or inference chips, autonomous driving chips, VR chips, robot-built-in chips, etc.

[0146] This application also provides an electronic device, and the electronic device includes a processing device. The electronic device can be a series of electronic devices such as a smart phone, a tablet computer, a wearable electronic device, a smart home electronic product, and industry.

[0147] Since the implementation of the electronic device has been described in detail in Embodiment 1, it will not be elaborated here.

[0148] The above are only the preferred embodiments of the present invention. Those skilled in the art know that without departing from the spirit and scope of the present invention, various changes or equivalent replacements can be made to these features and embodiments. Additionally, under the teachings of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the protection scope of the present invention.

Claims

1. A method for generating a subject image, characterized in that, Comprising: In a large language model, expanding key information related to a theme into prompt information; In a pre-trained text-to-image model, generating a candidate image including three color channels based on the prompt information and the key information; In the pre-trained text-to-image model, extracting an attention map corresponding to the main elements in the candidate image through an attention mechanism; Sample the attention map and calculate the cross-entropy of the attention map at the time step t, based on the time step t, the text token, the length of the text token, and the attention layer of the pre-trained text-to-image model l , at the attention layer l at the time step t; According to the cross-entropy, performing weighted fusion on the attention maps of the total time steps T and the total attention layers L of the pre-trained text-to-image model; Using the fused attention map as guidance information, predicting a mask of the theme image, and performing foreground segmentation to separate the theme image with an alpha channel; Wherein, in the large language model, expanding key information related to a theme into prompt information includes: The large language model receives key information related to the theme input by the user; Fine-tuning the large language model to expand the key information into prompt information, and extracting the main elements associated with the key information from the prompt information; The theme image generation method further includes: Normalizing the fused attention map and calculating the probability values of each region of the fused attention map; In the fused attention map, dividing the regions with probability values greater than a first threshold into determined foreground regions, dividing the regions with probability values between a second threshold and the first threshold into possible foreground regions, dividing the regions with probability values between the second threshold and a third threshold into possible background regions, and dividing the regions with probability values less than the third threshold into determined background regions to obtain a four-value map; Wherein, the third threshold is less than the second threshold, and the second threshold is less than the first threshold; Using the fused attention map as guidance information, predicting a mask of the theme image, and performing foreground segmentation to separate the theme image with an alpha channel includes: Taking the four-value map and the candidate image as inputs of the GrabCut algorithm, performing foreground segmentation on the candidate image to generate a final mask; Using the final mask to generate a theme image with an alpha channel.

2. The method for generating a subject image according to claim 1, wherein Generating a candidate image including three color channels based on the prompt information and the key information includes: In the pre-trained text-to-image model, generating the candidate image based on the prompt information and the main elements.

3. The method for generating a subject image according to claim 1, wherein Based on the time step t, the text token, the length of the text token, and the attention layer of the pre-trained text-to-image model l , calculate the cross-entropy of the attention map at the time step t and the attention layer l as follows: ; Among them, represents the cross entropy, t represents the sampling time step, l represents the attention layer of the pre-trained text-to-image model, N represents the length of the text tokens, n represents the text token, represents the time step t, and the attention layer l and the attention map of the text token n histogram probability of represents the logarithm function with base 2.

4. The method for generating a subject image according to claim 3, wherein According to the cross-entropy, performing weighted fusion on the attention maps of the total time steps T and the total attention layers L of the pre-trained text-to-image model includes: According to the cross entropy , calculate the weight of the attention map; Based on the weights, the total number of time steps T, the total number of attention layers L of the pre-trained text-to-image model, and the attention maps , perform weighted fusion to obtain the fused attention map.

5. A proxy framework based on large language models, characterized in that, Including: An input module for receiving key information related to the theme input by the user; An expansion agent module for expanding key information related to the theme into prompt information in a large language model; An output module for generating a candidate image including three color channels based on the prompt information and the key information in a pre-trained text-to-image model; An extraction agent module for extracting an attention map corresponding to the main elements in the candidate image through an attention mechanism in the pre-trained text-to-image model; A sampling module for sampling the attention map and calculating the cross-entropy of the attention map at the time step t, based on the time step t, the text token, the length of the text token, and the attention layer of the pre-trained text-to-image model l , at the attention layer l at the time step t; A weighted fusion module for weighted fusion of the attention maps of the total time steps \(T\) and the total attention layers \(L\) of the pre-trained text-to-image model according to the cross-entropy; A prediction module for using the fused attention map as guidance information to predict a mask of the subject image and performing foreground segmentation to separate the subject image with an alpha channel; Wherein, the extended proxy module is further configured to receive, by the large language model, key information related to the subject input by the user; fine-tune the large language model to expand the key information into prompt information, and extract the main elements associated with the key information from the prompt information; It further includes a four-value map determination module for normalizing the fused attention map and calculating the probability values of each region of the fused attention map; in the fused attention map, regions with probability values greater than a first threshold are classified as determined foreground regions, regions with probability values between a second threshold and the first threshold are classified as possible foreground regions, regions with probability values between the second threshold and a third threshold are classified as possible background regions, and regions with probability values less than the third threshold are classified as determined background regions to obtain a four-value map; wherein the third threshold is less than the second threshold, and the second threshold is less than the first threshold; The prediction module is further configured to use the four-value map and the candidate image as inputs to the GrabCut algorithm to perform foreground segmentation on the candidate image to generate a final mask; and use the final mask to generate a subject image with an alpha channel.

6. The agent framework based on a large language model according to claim 5, characterized in that The large language model-based proxy framework further includes a filtering module for filtering out abstract words in the attention map.

7. The agent framework based on a large language model according to claim 5, wherein The large language model-based proxy framework further includes an optimization proxy module for iteratively optimizing the prompt information until the prompt information meets a preset condition.

Citation Information

Patent Citations

  • Image processing method and device and storage medium

    CN117392276A

  • Language tracking image editing method based on text and graph generation model

    CN117934657A