Image generation method and device, electronic equipment, storage medium and program product
By obtaining the color information of the image description text and user configuration, the color embedding features of the text-generated image model are generated, which solves the problem of image generation color deviation in the prior art, and achieves precise control and quality improvement of the target image color.
Patent Information
- Application Number
- CN202510578413.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-01
AI Technical Summary
In the prior art, the target image generated based on literary and biographical technology has a color deviation problem, making it difficult to achieve precise control of image color, affecting the quality of image generation.
Obtain the image description text and the color information configured by the user, generate the color embedding features matching the literary image model, and process the color embedding features and image description text through the literary image model to generate the target image, and the target image is composed of the target color.
Accurate control of the generated image color is achieved, color deviation is reduced, and image generation quality is improved.
Smart Images

Figure CN120411292A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of artificial intelligence technology, and in particular, to an image generation method, apparatus, electronic device, storage medium, and program product. Background Art
[0002] Text-to-Image is a technology based on a pre-trained multi-modal large model that converts text described in natural language input by a user into an image with corresponding content, achieving the purpose of quickly generating the required pictures. Currently, it is widely used in content creation fields such as making advertisements, posters, and promotional pictures.
[0003] In the prior art, during the process of content creation based on the Text-to-Image technology, it is usually necessary for the user to input an image description text to define the image content of the target image to be generated, and then use the Text-to-Image model to combine the image description text to generate the target image required by the user.
[0004] However, in the prior art, it is difficult to precisely control the image color only through the image description text, resulting in a color deviation problem in the generated target image, affecting the quality of image generation and making it difficult to meet the design requirements of users. Summary of the Invention
[0005] Embodiments of the present disclosure provide an image generation method, apparatus, electronic device, storage medium, and program product to overcome the problem of color deviation in the generated target image.
[0006] In a first aspect, embodiments of the present disclosure provide an image generation method, including:
[0007] Obtain an image description text and color information configured by a user, where the image description text is used to represent the image content of the target image to be generated, and the color information is used to represent at least two target colors;
[0008] Generate a color embedding feature that matches the Text-to-Image model based on the color information, and process the color embedding feature and the image description text through the Text-to-Image model to generate the target image, where the target image is composed of the target colors.
[0009] In a second aspect, embodiments of the present disclosure provide an image generation apparatus, including:
[0010] An interaction unit, configured to obtain an image description text and color information configured by a user, where the image description text is used to represent the image content of the target image to be generated, and the color information is used to represent at least two target colors;
[0011] A processing unit, configured to generate a color embedding feature that matches a text-to-image model based on the color information, and process the color embedding feature and the image description text through the text-to-image model to generate the target image, where the target image is composed of the target colors.
[0012] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: a processor and a memory;
[0013] The memory stores computer-executable instructions;
[0014] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the image generation method described in the first aspect and various possible designs of the first aspect above.
[0015] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored. When the processor executes the computer-executable instructions, the image generation method described in the first aspect and various possible designs of the first aspect above is implemented.
[0016] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, where when the computer program is executed by a processor, the image generation method described in the first aspect and various possible designs of the first aspect above is implemented.
[0017] The image generation method, device, electronic device, storage medium, and program product provided in this embodiment obtain an image description text and color information configured by a user. The image description text is used to represent the image content of the target image to be generated, and the color information is used to represent at least two target colors. Based on the color information, a color embedding feature that matches a text-to-image model is generated, and the color embedding feature and the image description text are processed through the text-to-image model to generate the target image, where the target image is composed of the target colors. By additionally obtaining color information representing multiple target colors on the basis of obtaining the image description text, then converting the color information into a color embedding feature that matches the text-to-image model, and jointly controlling the generation process of the target image based on the color embedding feature and the image description text, a target image composed of the target colors indicated by the color information is obtained, realizing precise control of the constituent colors of the generated target image, making the image colors presented by the target image match the color information configured by the user, reducing color deviation, and improving the quality of image generation. Description of the Drawings
[0018] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0019] Figure 1 It is an application scenario diagram of an image generation method provided by an embodiment of the present disclosure;
[0020] Figure 2 It is a schematic flowchart of the image generation method provided by an embodiment of the present disclosure Figure 1 ;
[0021] Figure 3 For Figure 2 It is a schematic flowchart of the step of generating color information in the illustrated embodiment;
[0022] Figure 4 For Figure 2 It is a flowchart of the specific implementation manner of step S102 in the illustrated embodiment;
[0023] Figure 5 It is a schematic diagram of the process of generating a target image provided by an embodiment of the present disclosure;
[0024] Figure 6 It is a schematic flowchart of the image generation method provided by an embodiment of the present disclosure Figure 2 ;
[0025] Figure 7 For Figure 6 It is a flowchart of the specific implementation manner of step S204 in the illustrated embodiment;
[0026] Figure 8 For Figure 7 It is a flowchart of the specific implementation manner of step S2044 in the illustrated embodiment;
[0027] Figure 9 It is a schematic diagram of displaying color ratio information provided by an embodiment of the present disclosure;
[0028] Figure 10 It is a structural block diagram of an image generation device provided by an embodiment of the present disclosure;
[0029] Figure 11 It is a structural schematic diagram of an electronic device provided by an embodiment of the present disclosure;
[0030] Figure 12 It is a hardware structural schematic diagram of the electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.
[0032] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to select authorization or rejection.
[0033] The application scenarios of the embodiments of the present disclosure will be explained below:
[0034] The image generation method provided by the embodiments of the present disclosure can be applied to an application program (APP, Application) with the function of generating images from text, such as artificial intelligence (AI) assistant application programs, image editing and video editing application programs, etc. More specifically, it can be applied to the application scenario of AI image generation. The execution subject of this embodiment can be a terminal device running the above application program with the function of generating images from text, or a server deploying the server side corresponding to the above application program, or other electronic devices with similar functions. Among them, when the execution subject is a terminal device, the terminal device executes the method provided by this embodiment by running the above application program; when the execution subject is a server, the server side of the above application program with the function of generating images from text can run partially or entirely on the server, and execute the method provided by this embodiment on the server side, while the terminal device runs the client of the application program. Based on the server-client communication between the server and the terminal device, the terminal device can obtain the execution result of the method provided by this embodiment and display it as needed.
[0035] Among them, in some embodiments, the terminal device or the server can implement the image generation method provided in the embodiments of the present disclosure by running various computer-executable instructions or computer programs. For example, the computer-executable instructions can be program-level commands, machine instructions, or software instructions. The computer program can be a native program or software module in the operating system; it can be a local application, that is, a program that needs to be installed in the operating system to run, or it can be a small program embedded in any APP, that is, a program running based on the browser environment. In summary, the above computer-executable instructions can be in any form of instructions, and the above computer programs can be in any form of application programs, modules, or plugins, and the specific implementation form can be configured according to needs. Further, in the process of implementing the image generation method provided in the embodiments of the present disclosure, the terminal device can execute the method by running the computer-executable instructions or computer programs set locally, or can execute the method by calling the computer-executable instructions or computer programs set in an external server. In some embodiments, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud storage, cloud communication, cloud databases, cloud computing, cloud functions, network services, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. Among them, the cloud service can be an interactive processing service for the terminal device to call.
[0036] Figure 1 A schematic diagram of an application scenario of the image generation method provided in the embodiments of the present disclosure. Refer to Figure 1 As shown in the figure, taking the terminal device as an example, a target application program with the function of generating images from text runs in the terminal device, such as a smart assistant application program. The user inputs an image description text through the interaction interface of the target application program. As shown in the figure, the content of the image description text input by the user is "A winding path leads to the mountains in the distance, and a person stands on the road looking into the distance". After that, the terminal device calls the text-to-image model deployed locally or in the cloud to perform semantic understanding on the above image description text and generate an image that matches the semantics of the image description text, such as the image P1 shown in the figure. Thus, the process of generating an image corresponding to the image description text based on the text is completed.
[0037] In the prior art, in the above text-to-image process, it is usually necessary for the user to input an image description text to define the image content of the target image to be generated, and then use a text-to-image model to combine the image description text to generate the corresponding target image. Among them, the image description text is essentially a description of the user's needs. The text-to-image model will generate a picture that matches the user's needs described in the image description text. However, for the colors that make up the target image, usually the text-to-image model automatically matches the corresponding colors according to the semantics of the image description text. Therefore, color control based on user needs cannot be achieved. In some other possible implementation manners, although color control in the image can be achieved to a certain extent by adding color limitations to the image description text, however, in the actual application process, the "understanding" of colors by the text-to-image model needs to be obtained through model training in advance. Therefore, when new colors or custom colors appear, the text-to-image model needs to be trained to enable it to obtain the ability to understand and use the color, resulting in the problem of too high cost. At the same time, as a natural language, the image description text itself is difficult to accurately describe colors, further increasing the difficulty of accurately controlling the colors of pictures through the above solutions in the prior art.
[0038] Embodiments of the present disclosure provide an image generation method to solve the above problems.
[0039] Refer to Figure 2 , Figure 2 is a flowchart of the image generation method provided by the embodiments of the present disclosure Figure 1 The method of this embodiment can be applied to a terminal device or a server. Among them, for the case where the terminal device executes the method provided by this embodiment, in one possible implementation manner, the terminal device can implement the image generation method provided by this embodiment by executing the program code deployed locally. In another possible implementation manner, the server can be used to deploy a function service implemented based on the image generation method provided by this embodiment, and the terminal device can access the above server and call the corresponding function service to implement the image generation method provided by this embodiment. Exemplarily, the image generation method provided by this embodiment includes:
[0040] Step S101: Obtain an image description text and color information configured by the user, where the image description text is used to represent the image content of the target image to be generated, and the color information is used to represent at least two target colors.
[0041] Refer to Figure 1The schematic diagram of the application scenario shown. In this embodiment, the terminal device is used as the execution entity to introduce the provided image generation method. Exemplarily, the terminal device provides a text-to-image function to the user by running a target application. The user inputs an image description text and color information to the terminal device through the interaction interface of the target application, so that the terminal device obtains the image description text and the color information configured by the user. Among them, on the one hand, the image description text is the text that characterizes the image content of the target image to be generated. For example, "A winding path leads to the mountains in the distance, and a person stands on the road looking into the distance" is the image description text that describes the image content of the target image to be generated. The specific content of this image description text can be various and is written according to the specific needs of the user, which will not be elaborated here. On the other hand, the color information is the information used to indicate the image colors that make up the target image, and it is used to indicate at least two target colors. In one possible implementation, the color information includes N RGB values, where N is an integer greater than 1, and the RGB value is used to characterize the target color. For example, the color information includes RGB_1 = [255, 0, 0], RGB_2 = [0, 255, 0]. The color indicated by RGB_1, that is, red, and the color indicated by RGB_2, that is, green. Since the color information is represented based on RGB values, the color information can represent more abundant colors that are difficult to accurately describe in natural language, such as RGB_3 = [102, 202, 52], etc. This color is a color similar to that of grass. The above color information is generated based on user configuration. In one possible implementation, the user can configure the color information by directly inputting at least one RGB value. In the case where the user only inputs one RGB value, the preset RGB values can be combined to generate the color information. The solution of this embodiment can generate the color information by the user directly inputting RGB values, which can achieve the most accurate expression of the color information and improve the color accuracy of the subsequent generated target image.
[0042] In another possible implementation, the color information can be generated by deploying a color configuration control in the target application and using the color configuration control to generate the color information. Specifically, in this embodiment, before step S101, a step of generating color information can also be included, as Figure 3 shown, the step of generating color information includes:
[0043] Step S100A: Display the color configuration control, and the color configuration control is at least used to display at least two candidate colors;
[0044] Step S100B: Respond to the trigger operation on the color configuration control and generate color information.
[0045] Exemplarily, a color configuration control is an interactive control for determining color information. For example, it displays multiple selectable colors through a color palette, and these selectable colors can be understood as common or process colors. At the same time, the color information (such as RGB values) mapped by each selectable color can be obtained through the color configuration control. After the terminal device displays the color configuration control, the user can, based on the visual effects of the selectable colors shown by the color configuration control, select several of the selectable colors through a trigger operation, determine them as target colors, and obtain the RGB values corresponding to the target colors, thereby generating color information. The solution of the steps in this embodiment, by combining the color configuration control, enables the user to select the required colors based on the color palette, thereby, on the basis of meeting the user's aesthetic needs, without the user directly inputting RGB values, reducing the user operation difficulty and professional knowledge requirements, and improving the interaction efficiency of generating color information.
[0046] Further, in a possible implementation manner, the number of target colors represented by the color information can be a preset fixed value, such as 5, that is, the user fixedly configures 5 target colors to generate a target image. And in another possible implementation manner, the number of target colors represented by the color information can also be dynamically determined based on the image description text. Specifically, in this embodiment, before step S100B, it further includes:
[0047] Step S100C: Generate a recommended number of colors according to the image description text.
[0048] In the case where step S100C is included, the specific implementation manner of step S100B includes: in response to a trigger operation on the color configuration control, generate color information based on the recommended number of colors.
[0049] Exemplarily, after receiving the image description text input by the user, first, semantic features are extracted from the image description text to obtain the text semantics representing the content of the image description text. Then, based on this text semantics, the corresponding recommended number of colors is determined. Exemplarily, the mapping rule between the text semantics of the image description text and the recommended number of colors can be: the more complex the content represented by the text semantics, the larger the recommended number of colors; or, the more unassociated independent objects in the content represented by the text semantics, the larger the recommended number of colors. The specific implementation process of generating the recommended number of colors according to the image description text can be implemented based on a pre-trained recommendation model, which will not be specifically elaborated here. After that, the recommended number of colors can be displayed in the color configuration control or in the interaction interface of the target application to prompt the user to control the number of target colors configured, so as to obtain the target colors of the recommended number of colors. Then, based on the identification identifier (such as RGB value) of the target colors of the recommended number of colors, color information is generated. Or, based on the recommended number of colors, from the initially selected candidate colors with the initial number triggered by the user, further select the candidate colors with the recommended number of colors as the target colors, so as to obtain the target colors of the recommended number of colors. Then, based on the identification identifier of the target colors of the recommended number of colors, color information is generated. For example, if the user's triggered operation selects 6 colors, the terminal device determines 5 (the recommended number of colors) of them as the target colors according to the recommended number of colors, and obtains the identification identifiers corresponding to these 5 target colors to generate color information.
[0050] In the steps of this embodiment, by generating the recommended number of colors through the image description text and then generating color information based on the recommended number of colors, the composition of the number of colors in the target image generated based on the color information can be made more reasonable, improving the visual effect of the finally generated target image.
[0051] Further, in a possible implementation manner, the color information includes color type information and color proportion information. The color type information is used to represent at least two target colors, and the color proportion information is used to represent the proportion of each target color indicated by the color information in the target image. Among them, the color information determined in the previous steps is the color type information in the steps of this embodiment. The color proportion information is the information used to represent the proportion of each target color in the target image, which corresponds to the target colors represented by the color type information. For example, the color type information indicates three colors that make up the target image, namely color A, color B, and color C. The corresponding color proportion information represents the respective proportions of color A, color B, and color C in the generated target image. For example, color A is 20%, color B is 45%, and color C is 35%, that is, in the target image, 20% of the area is color A, 45% of the area is color B, and 35% of the area is color C.
[0052] In this embodiment, by further configuring the color ratio information, more precise color control of the target image to be generated can be achieved, thereby meeting the personalized needs of users and improving the quality of the image.
[0053] Step S102: Based on the color information, generate a color embedding feature that matches the text-to-image model, and process the color embedding feature and the image description text through the text-to-image model to generate a target image, where the target image is composed of target colors.
[0054] Exemplarily, after obtaining the color information, perform feature transformation on the color information. For example, as introduced in the previous embodiment, the color information includes multiple RGB values. Then, by encoding the multiple RGB values, a color embedding feature that can match the text-to-image model is generated. Here, the color embedding feature that can match the text-to-image model means that the color embedding feature can be input into the text-to-image model and recognized and utilized by the text-to-image model, so that the text-to-image model generates a target image based on the meaning represented by the color embedding feature. Specifically, even the generated target image is composed of the target colors indicated by the color information. For example, the target colors indicated by the color information include color RGB_1, color RGB_2, and color RGB_3. After converting the color information into a color embedding feature and inputting it into the text-to-image model for processing, the target image generated by the text-to-image model based on the image description text is also composed of the three colors RGB_1, RGB_2, and RGB_3.
[0055] After that, input the color embedding feature and the image description text into the text-to-image model for processing. The text-to-image model is, for example, a diffusion model. After receiving the color embedding feature and the image description text, the text-to-image model processes the two and generates corresponding constraint features respectively, and generates an image based on the above two constraint features to obtain a target image. Since the generation process of the target image is constrained and controlled by the color embedding feature and the image description text, the generated target image contains both the content of the image description text and the target colors specified by the color information, thus achieving precise control over both the image content and the image color dimensions in the text-to-image process.
[0056] Further, in a possible implementation manner, the text-to-image model is composed of a diffusion transformer model (DiffusionTransformer, Dit) and a color encoding module. The diffusion transformer model is used to receive the image description text and generate a target image based on the image description text. The color encoding module is used to receive the color information, convert the color information into a color embedding feature, and input the color embedding feature into the diffusion transformer model, so that the diffusion transformer model generates a target image based on the image description text and the color embedding feature.
[0057] In a possible implementation, the color information includes N RGB values. As Figure 4 shown, the specific implementation of step S102 includes:
[0058] Step S1021: Concatenate the N RGB values into a first color vector;
[0059] Step S1022: Encode the first color vector through a color encoding module to generate color embedding features matching the diffusion transformer model, and input the color embedding features into the diffusion transformer model;
[0060] Step S1023: Process the color embedding features and the image description text through the diffusion transformer model to generate a target image.
[0061] Exemplarily, first, after obtaining the color information, the N RGB values represented by the color information are concatenated to obtain a first color vector with a data format of [1, N, 3] (where 3 represents the three color channels of R, G, and B). Then, the first color vector is input into the color encoding module. The color encoding module includes a Multilayer Perceptron (MLP) layer. After being processed by at least one (for example, 2) multilayer perceptrons, color embedding features with a data format of [1, N, D] are generated, where D is the encoding length and is determined according to the MLP structure. Then, the color embedding features and the image description text are processed through the diffusion transformer model to generate a target image. Among them, the diffusion transformer model generates the target image by reverse denoising the noisy image. In this process, the color embedding features and the image description text can be introduced as constraints through the multi-head attention mechanism to control the denoising process of the noisy image, so as to generate a target image that meets the constraint features corresponding to the color embedding features and the image description text.
[0062] Figure 5 FIG. is a schematic diagram of a process for generating a target image provided by an embodiment of the present disclosure. The following will be combined with Figure 5 to introduce the above process. As Figure 5As shown, exemplarily, first, the terminal device responds to the trigger operation input by the user by displaying a color configuration control within the interaction interface, and then obtains color information. For example, as shown in the figure, the color configuration control contains a color palette for displaying different candidate colors. In response to the trigger operation by the user, five colors (such as five RGB values) A, B, C, D, and E are determined as the target colors, and then color information is generated. After that, the color information is sent to the color encoding module for encoding processing to generate a color embedding feature, and this color embedding feature is input into the diffusion transformer model. On the other hand, the terminal device receives the image description text input by the user within the interaction interface and also inputs the image description text into the diffusion transformer model for processing. The diffusion transformer model performs inverse denoising on the noisy image based on the received image description text and the color embedding feature, and generates a target image that contains the image content described by the image description text and is composed of the target colors indicated by the color information.
[0063] In this embodiment, by obtaining the image description text and the color information configured by the user, where the image description text is used to characterize the image content of the target image to be generated, and the color information is used to characterize at least two target colors; based on the color information, a color embedding feature matching the text-to-image generation model is generated, and the color embedding feature and the image description text are processed by the text-to-image generation model to generate a target image, and the target image is composed of the target colors. By additionally obtaining the color information characterizing multiple target colors on the basis of obtaining the image description text, and then converting the color information into a color embedding feature matching the text-to-image generation model, and jointly controlling the generation process of the target image based on this color embedding feature and the image description text, a target image composed of the target colors indicated by the color information is obtained, realizing precise control over the constituent colors of the generated target image, making the image colors presented by the target image match the color information configured by the user, reducing color deviation, and improving the quality of image generation.
[0064] Reference Figure 6 , Figure 6 is a schematic flow chart of the image generation method provided by the embodiments of the present disclosure Figure 2 。This embodiment further refines step S102 on the basis of the embodiment shown in Figure 2 , and this image generation method includes:
[0065] Step S201: Obtain an image description text and color information configured by the user, where the image description text is used to characterize the image content of the target image to be generated, and the color information is used to characterize at least two target colors.
[0066] Step S202: Generate a color embedding feature that matches the text-to-image generation model based on the color information.
[0067] Step S203: Generate corresponding text embedding features based on the image description text.
[0068] Step S204: Based on the multi-head attention mechanism, the color embedding features and the text embedding features are input into the text-based graph model to generate the target image.
[0069] For example, in this embodiment, after obtaining the image description text and the color information configured by the user, the terminal device converts the color information into color embedding features using the color coding module in the text graph model. Figure 2 The corresponding embodiments have been described in detail and will not be repeated here. On the other hand, the diffusion transformer model (embedding model layer) in the text graph model is used to process the image description text to generate corresponding text embedding features. The text embedding feature is a feature vector that represents the semantics of the image description text. The specific generation method is existing technology and will not be repeated here.
[0070] Afterwards, based on the multi-head attention mechanism of the diffusion transformer model, the color embedding features generated by the color coding module are connected to the diffusion transformer model, so that the diffusion transformer model combines the color embedding features and the text embedding features. That is, under the joint constraints of the color embedding features and the text embedding features, a target image is generated that matches the image content described by the image description text and the target color described by the color information.
[0071] In one possible implementation, Figure 7 As shown, the specific implementation of step S204 includes:
[0072] Step S2041: Input the color embedding feature into the first attention mapping layer of the Wensheng graph model to obtain a first attention vector.
[0073] Step S2042: Input the text embedding feature into the second attention mapping layer of the text graph model to obtain a second attention vector.
[0074] Step S2043: Concatenate the first attention vector and the second attention vector to obtain a mixed attention vector.
[0075] Step S2044: Process the mixed attention vector through the text-based graph model to generate a target image.
[0076] For example, first, the color information includes N RGB values, and the first color vector formed after splicing is [1, N, 3]. The feature embedding is based on the color embedding feature generated by the color information. color , the color embedding feature Embedding colorThe size is [1, N, D], where D is the characteristic length. Subsequently, the first attention mapping layer of the text-to-image model for the color-embedded features is configured within the color encoding module to obtain the first attention vector output by the first attention mapping layer. The first attention vector includes the query vector Q color , the key vector K color and the value vector V color , where the first attention mapping layer consists of at least two multi-layer perceptron (MLP) layers. On the other hand, similarly, after processing the image description text to obtain the text embedding feature Embedding text , the text embedding feature Embedding text is input into the second attention mapping layer of the text-to-image model configured within the diffusion transformer model to obtain the second attention vector output by the second attention mapping layer. The second attention vector consists of the query vector Q text , the key vector K text and the value vector V text . Finally, the first attention vector and the second attention vector are concatenated to obtain the mixed attention vector. Among them, the mixed attention vector also consists of the query vector Q mm , the key vector K mm and the value vector V mm . Taking the query vector Q mm as an example, after concatenating and merging the query vector in the first attention vector and the query vector in the second attention vector, the obtained query vector Q mm = [Q text; Q color . The merging methods for the key vector K mm and the value vector V mm are similar and will not be elaborated here.
[0077] Furthermore, in this embodiment, before step S2043, it further includes:
[0078] Step S2045: Obtain the noise latent representation vector corresponding to the target image.
[0079] Correspondingly, when step S2045 is executed, the specific implementation manner of step S2043 includes:
[0080] Step S2043A: Concatenate the first attention vector, the second attention vector, and the noise latent to obtain the mixed attention vector.
[0081] Exemplarily, the noise latent representation vector is also a vector representing noise and can be represented by the corresponding query vector Q noise , the key vector K noise and the value vector Vnoise Composition, correspondingly, concatenate the first attention vector, the second attention vector, and the noise latent to obtain the mixed attention vector as: {Q mm =[Q text; Q color; Q color , K mm =[K text; K color; K color , V mm =[V text; V color; V color}. Wherein, Q mm is the mixed query vector, K mm is the mixed key vector, and V mm is the mixed value vector.
[0082] Furthermore, in a possible implementation manner, the mixed attention vector includes a mixed query vector, a mixed key vector, and a mixed value vector. As Figure 8 shown, the specific implementation manner of step S2044 includes:
[0083] Step S2044-1: Based on the mixed query vector and the mixed key vector, obtain the attention weight matrix;
[0084] Step S2044-2: According to the product of the attention weight matrix and the mixed value vector, obtain the sequence feature;
[0085] Step S2044-3: Through the text-to-image model, perform reverse denoising on the initial noise map based on the sequence feature to generate the target image.
[0086] Exemplarily, first calculate the attention weight matrix based on the mixed query vector and the mixed key vector, which can be obtained according to the following formula (1):
[0087]
[0088] Wherein, A is the attention weight matrix, is the transpose of K mm , and d k is the scaling factor.
[0089] is the attention score after scaling processing, and then it is normalized through the normalization function softmax to obtain the attention weight matrix A.
[0090] After that, calculate the product of the attention weight matrix and the mixed value vector, i.e., the product of A and Vmm, to obtain the sequence feature Attention. Then, the text-to-image model processes the initial noise map based on the sequence feature Attention to generate the target image. Among them, this step is executed by the DiT model in the text-to-image model. The sequence of image latent representations processed by the DiT model may contain a large number of elements, and there may be long-range dependencies between different parts. By calculating Attention, the DiT model can directly focus on the relationship between any two elements in the sequence, thereby effectively capturing this long-range dependence and generating a more logical and coherent image. In the image generation task, the model can focus on the information related to the current generation task through the sequence feature Attention according to conditions such as text descriptions, and ignore irrelevant information, thereby improving the quality and accuracy of the generated image and making it more in line with the expected semantic and visual effects.
[0091] Further optionally, after generating the target image, it further includes:
[0092] Step S205: Display the color ratio information corresponding to the target image, where the color ratio information is used to characterize the proportion of each target color indicated by the color information in the target image.
[0093] Exemplarily, after generating the target image, the color ratio information corresponding to the target image will be calculated. The color ratio information is used to characterize the proportion of each target color indicated by the color information in the target image, so that the user can make more detailed adjustments to the color composition of the target image based on the color ratio information, thereby improving the visual effect of the finally generated image work.
[0094] Figure 9 FIG. is a schematic diagram showing color ratio information provided by an embodiment of the present disclosure. The following will be combined with Figure 9 to introduce the above process in more detail, such as Figure 9As shown, for example, after generating the target image, the target image is composed of three colors: RGB_1, RGB_2, and RGB_3. As shown in the interactive interface in the figure, the color of the "mountain" area in the target image is RGB_1. According to the color ratio information, the color ratio corresponding to RGB_1 is obtained as 56%; the color of the "sky" area in the target image is RGB_2. According to the color ratio information, the color ratio corresponding to RGB_2 is obtained as 34%; the color of the "road surface" area in the target image is RGB_3. According to the color ratio information, the color ratio corresponding to RGB_3 is obtained as 10%. Among them, the above color ratio information can be generated by the text-to-image model based on the content semantics of the image description text during the process of generating the target image by the model (so that the obtained color ratio information matches the image content of the target image), or can be obtained by counting the pixel values of each pixel point in the target image after generating the target image. There is no limitation here. After that, the user can further perform a return operation (click the "return" button) based on the above color ratio information to readjust the color information or regenerate the target image to achieve the adjustment of the color composition and ratio in the target image; or can perform an export operation (click the "export" button) to export the target image. By displaying the color ratio information, the user can have a more intuitive feeling of the performance effect of the set target color, thereby improving the human-computer interaction efficiency during the process of adjusting the image.
[0095] In this embodiment, the implementation manner of step S201 is the same as that of step S101 in the embodiment shown in the present disclosure Figure 2 and will not be elaborated here one by one.
[0096] Corresponding to the image generation method in the above embodiment, Figure 10 is a structural block diagram of an image generation device provided by an embodiment of the present disclosure. The method introduced in the above embodiment can be executed by this image generation device. The device can be implemented in a software and / or hardware manner, and the device can be integrated in an electronic device with certain data processing capabilities. Among them, the electronic device can include, but is not limited to, a mobile terminal with big data processing capabilities, and a fixed terminal with big data processing capabilities such as a desktop computer and a supercomputer.
[0097] For the sake of convenience of description, only parts related to the embodiments of the present disclosure are shown. Referring to Figure 10 , the image generation device 3 includes:
[0098] An interaction unit 31, configured to obtain an image description text and color information configured by a user, where the image description text is used to characterize the image content of a target image to be generated, and the color information is used to characterize at least two target colors;
[0099] The processing unit 32 is configured to generate a color embedding feature that matches the text-to-image model based on color information, and process the color embedding feature and the image description text through the text-to-image model to generate a target image, where the target image is composed of target colors.
[0100] According to one or more embodiments of the present disclosure, the color information includes N RGB values, where N is an integer greater than 1, and the RGB values are used to represent the target color. The text-to-image model is composed of a diffusion transformer model and a color encoding module. The processing unit 32 is specifically configured to: concatenate the N RGB values into a first color vector; encode the first color vector through the color encoding module to generate a color embedding feature that matches the diffusion transformer model, and input the color embedding feature into the diffusion transformer model; process the color embedding feature and the image description text through the diffusion transformer model to generate a target image.
[0101] According to one or more embodiments of the present disclosure, when the processing unit 32 processes the color embedding feature and the image description text through the text-to-image model to generate a target image, it is specifically configured to: generate a corresponding text embedding feature based on the image description text; based on the multi-head attention mechanism, input the color embedding feature and the text embedding feature into the text-to-image model to generate a target image.
[0102] According to one or more embodiments of the present disclosure, when the processing unit 32 inputs the color embedding feature and the text embedding feature into the text-to-image model based on the multi-head attention mechanism to generate a target image, it is specifically configured to: input the color embedding feature into the first attention mapping layer of the text-to-image model to obtain a first attention vector; input the text embedding feature into the second attention mapping layer of the text-to-image model to obtain a second attention vector; concatenate the first attention vector and the second attention vector to obtain a mixed attention vector; process the mixed attention vector through the text-to-image model to generate a target image.
[0103] According to one or more embodiments of the present disclosure, the processing unit 32 is further configured to: obtain a noise latent representation vector corresponding to the target image; when the processing unit 32 concatenates the first attention vector and the second attention vector to obtain a mixed attention vector, it is specifically configured to: concatenate the first attention vector, the second attention vector, and the noise latent to obtain a mixed attention vector.
[0104] According to one or more embodiments of the present disclosure, the mixed attention vector includes a mixed query vector, a mixed key vector, and a mixed value vector. When the processing unit 32 processes the mixed attention vector through the text-to-image model to generate a target image, it is specifically configured to: obtain an attention weight matrix based on the mixed query vector and the mixed key vector; obtain a sequence feature according to the product of the attention weight matrix and the mixed value vector; through the text-to-image model, perform reverse denoising on the initial noise map based on the sequence feature to generate a target image.
[0105] According to one or more embodiments of the present disclosure, before obtaining the color information configured by the user, the interaction unit 31 is further configured to: display a color configuration control, where the color configuration control is at least used to display at least two candidate colors; and generate color information in response to a trigger operation on the color configuration control.
[0106] According to one or more embodiments of the present disclosure, the interaction unit 31 is further configured to: generate a recommended number of colors according to the image description text; when the interaction unit 31 generates color information in response to a trigger operation on the color configuration control, it is specifically configured to: generate color information based on the recommended number of colors in response to the trigger operation on the color configuration control.
[0107] According to one or more embodiments of the present disclosure, it further includes at least one of the following: the color information further includes color ratio information, where the color ratio information is used to represent the proportion of each target color indicated by the color information in the target image; the interaction unit 31 displays the color ratio information after generating the target image.
[0108] Wherein, the interaction unit 31 and the processing unit 32 are connected in sequence. The image generation device 3 provided in this embodiment can execute the technical solutions of the above method embodiments, and the implementation principles and technical effects are similar, which will not be elaborated here.
[0109] Figure 11 The following is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure, as Figure 11 shown, the electronic device 4 includes:
[0110] a processor 41, and a memory 42 communicatively connected to the processor 41;
[0111] The memory 42 stores computer-executable instructions;
[0112] The processor 41 executes the computer-executable instructions stored in the memory 42 to implement the image generation method in the embodiment as Figures 2 - 9 shown.
[0113] Optionally, the processor 41 and the memory 42 are connected through a bus 43.
[0114] For relevant descriptions, reference can be made to the relevant descriptions and effects corresponding to the steps in the Figures 2 - 9 corresponding embodiments, and no further elaboration will be made here.
[0115] An embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the present disclosure Figures 2 - 9The image generation method provided by any one of the corresponding embodiments.
[0116] Embodiments of the present disclosure provide a computer program product, including a computer program, which when executed by a processor implements the Figures 2 - 9 The image generation method provided by any one of the corresponding embodiments.
[0117] To implement the above embodiments, embodiments of the present disclosure also provide an electronic device.
[0118] Referring to Figure 12 , which shows a schematic structural diagram of an electronic device 900 suitable for implementing embodiments of the present disclosure. The electronic device 900 can be a terminal device or a server. Among them, the terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers, portable media players (PMPs), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 12 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0119] As Figure 12 shown, the electronic device 900 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 902 or the program loaded from the storage device 908 into the random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 are also stored. The processing device 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. The input / output (I / O) interface 905 is also connected to the bus 904.
[0120] Typically, the following devices can be connected to the I / O interface 905: input devices 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and a communication device 909. The communication device 909 can allow the electronic device 900 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 12 the electronic device 900 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices can be alternatively implemented or had.
[0121] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication device 909, or installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above functions defined in the method of the embodiment of the present disclosure are executed.
[0122] It should be noted that the computer-readable medium described above in the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. And in the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0123] The above computer-readable medium may be included in the above electronic device; or it may exist separately without being assembled into the electronic device.
[0124] The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to execute the method shown in the above embodiments.
[0125] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any kind of network, including a Local Area Network (LAN) or a Wide Area Network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0126] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0127] The units or modules involved in the embodiments described in the present disclosure may be implemented in software or in hardware. Among them, the name of the unit or module does not, in some cases, constitute a limitation on the unit itself.
[0128] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), Systems on Chip (SOC), Complex Programmable Logic Devices (CPLD), and so on.
[0129] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0130] In a first aspect, according to one or more embodiments of the present disclosure, there is provided an image generation method, including:
[0131] Obtaining an image description text and color information configured by a user, wherein the image description text is used to characterize the image content of a target image to be generated, and the color information is used to characterize at least two target colors; generating a color embedding feature matching a text-to-image model based on the color information, and processing the color embedding feature and the image description text through the text-to-image model to generate the target image, which is composed of the target colors.
[0132] According to one or more embodiments of the present disclosure, the color information includes N RGB values, where N is an integer greater than 1, and the RGB values are used to characterize the target colors. The text-to-image model is composed of a diffusion transformer model and a color encoding module; the generating a color embedding feature matching the text-to-image model based on the color information, and processing the color embedding feature and the image description text through the text-to-image model to generate the target image includes: concatenating the N RGB values into a first color vector; encoding the first color vector through the color encoding module to generate a color embedding feature matching the diffusion transformer model, and inputting the color embedding feature into the diffusion transformer model; processing the color embedding feature and the image description text through the diffusion transformer model to generate the target image.
[0133] According to one or more embodiments of the present disclosure, the processing the color embedding feature and the image description text through the text-to-image model to generate the target image includes: generating a corresponding text embedding feature based on the image description text; inputting the color embedding feature and the text embedding feature into the text-to-image model based on a multi-head attention mechanism to generate the target image.
[0134] According to one or more embodiments of the present disclosure, based on the multi-head attention mechanism, inputting the color embedding feature and the text embedding feature into the text-to-image model to generate the target image includes: inputting the color embedding feature into the first attention mapping layer of the text-to-image model to obtain a first attention vector; inputting the text embedding feature into the second attention mapping layer of the text-to-image model to obtain a second attention vector; concatenating the first attention vector and the second attention vector to obtain a mixed attention vector; and processing the mixed attention vector through the text-to-image model to generate the target image.
[0135] According to one or more embodiments of the present disclosure, the method further includes: obtaining a noise latent representation vector corresponding to the target image; and the concatenating the first attention vector and the second attention vector to obtain a mixed attention vector includes: concatenating the first attention vector, the second attention vector, and the noise latent to obtain a mixed attention vector.
[0136] According to one or more embodiments of the present disclosure, the mixed attention vector includes a mixed query vector, a mixed key vector, and a mixed value vector; and the processing the mixed attention vector through the text-to-image model to generate the target image includes: obtaining an attention weight matrix based on the mixed query vector and the mixed key vector; obtaining a sequence feature according to the product of the attention weight matrix and the mixed value vector; and performing reverse denoising on an initial noise map based on the sequence feature through the text-to-image model to generate the target image.
[0137] According to one or more embodiments of the present disclosure, before obtaining the color information configured by the user, the method further includes: displaying a color configuration control, where the color configuration control is at least used to display at least two candidate colors; and in response to a trigger operation on the color configuration control, generating the color information.
[0138] According to one or more embodiments of the present disclosure, the method further includes: generating a recommended number of colors according to the image description text; and the responding to a trigger operation on the color configuration control to generate the color information includes: responding to a trigger operation on the color configuration control, and generating the color information based on the recommended number of colors.
[0139] According to one or more embodiments of the present disclosure, at least one of the following is further included: the color information further includes color ratio information, and the color ratio information is used to characterize the proportion of each target color indicated by the color information in the target image; and after generating the target image, displaying the color ratio information.
[0140] In a second aspect, according to one or more embodiments of the present disclosure, there is provided an image generation device, including:
[0141] An interaction unit, configured to obtain an image description text and color information configured by a user, where the image description text is used to characterize the image content of a target image to be generated, and the color information is used to characterize at least two target colors;
[0142] A processing unit, configured to generate a color embedding feature matching a text-to-image model based on the color information, and process the color embedding feature and the image description text through the text-to-image model to generate the target image, where the target image is composed of the target colors.
[0143] According to one or more embodiments of the present disclosure, the color information includes N RGB values, N is an integer greater than 1, the RGB values are used to characterize the target colors, and the text-to-image model is composed of a diffusion transformer model and a color encoding module; the processing unit is specifically configured to: splice the N RGB values into a first color vector; encode the first color vector through the color encoding module to generate a color embedding feature matching the diffusion transformer model, and input the color embedding feature into the diffusion transformer model; process the color embedding feature and the image description text through the diffusion transformer model to generate the target image.
[0144] According to one or more embodiments of the present disclosure, when the processing unit processes the color embedding feature and the image description text through the text-to-image model to generate the target image, it is specifically configured to: generate a corresponding text embedding feature based on the image description text; based on a multi-head attention mechanism, input the color embedding feature and the text embedding feature into the text-to-image model to generate the target image.
[0145] According to one or more embodiments of the present disclosure, when the processing unit inputs the color embedding feature and the text embedding feature into the text-to-image model based on the multi-head attention mechanism to generate the target image, it is specifically configured to: input the color embedding feature into a first attention mapping layer of the text-to-image model to obtain a first attention vector; input the text embedding feature into a second attention mapping layer of the text-to-image model to obtain a second attention vector; splice the first attention vector and the second attention vector to obtain a mixed attention vector; process the mixed attention vector through the text-to-image model to generate the target image.
[0146] According to one or more embodiments of the present disclosure, the processing unit is further configured to: obtain a noise latent representation vector corresponding to the target image; when the processing unit splices the first attention vector and the second attention vector to obtain a mixed attention vector, specifically: splice the first attention vector, the second attention vector, and the noise latent to obtain a mixed attention vector.
[0147] According to one or more embodiments of the present disclosure, the mixed attention vector includes a mixed query vector, a mixed key vector, and a mixed value vector; when the processing unit processes the mixed attention vector through the text-to-image model to generate the target image, specifically: based on the mixed query vector and the mixed key vector, obtain an attention weight matrix; according to the product of the attention weight matrix and the mixed value vector, obtain a sequence feature; through the text-to-image model, based on the sequence feature, perform reverse denoising on the initial noise map to generate the target image.
[0148] According to one or more embodiments of the present disclosure, before obtaining the color information configured by the user, the interaction unit is further configured to: display a color configuration control, and the color configuration control is at least used to display at least two candidate colors; in response to a trigger operation on the color configuration control, generate the color information.
[0149] According to one or more embodiments of the present disclosure, the interaction unit is further configured to: generate a recommended number of colors according to the image description text; when the interaction unit generates the color information in response to a trigger operation on the color configuration control, specifically: in response to a trigger operation on the color configuration control, generate the color information based on the recommended number of colors.
[0150] According to one or more embodiments of the present disclosure, at least one of the following is further included: the color information further includes color ratio information, and the color ratio information is used to characterize the proportion of each target color indicated by the color information in the target image; after generating the target image, the interaction unit displays the color ratio information.
[0151] In a third aspect, according to one or more embodiments of the present disclosure, an electronic device is provided, including: at least one processor and a memory;
[0152] The memory stores computer-executable instructions;
[0153] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the image generation method as described in the first aspect and various possible designs of the first aspect above.
[0154] Fourthly, according to one or more embodiments of the present disclosure, there is provided a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the image generation method as described in the first aspect above and various possible designs of the first aspect.
[0155] Fifthly, according to one or more embodiments of the present disclosure, there is provided a computer program product including a computer program, which, when executed by a processor, implements the image generation method as described in the first aspect above and various possible designs of the first aspect.
[0156] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.
[0157] In addition, although the operations are depicted in a specific order, this should not be construed as requiring that the operations be performed in the specific order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of a single embodiment can also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0158] Although the subject matter has been described in language specific to structural features and / or method logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims.
Claims
1. An image generation method, characterized in that, Including: Obtain an image description text and color information configured by a user, where the image description text is used to characterize the image content of a target image to be generated, and the color information is used to characterize at least two target colors; Based on the color information, generate a color embedding feature that matches the text-to-image model, and process the color embedding feature and the image description text through the text-to-image model to generate the target image, where the target image is composed of the target colors.
2. The method according to claim 1, characterized in that, The color information includes N RGB values, where N is an integer greater than 1, the RGB values are used to characterize the target colors, and the text-to-image model is composed of a diffusion transformer model and a color encoding module; The generating a color embedding feature that matches the text-to-image model based on the color information, and processing the color embedding feature and the image description text through the text-to-image model to generate the target image includes: Concatenate the N RGB values into a first color vector; Encode the first color vector through the color encoding module to generate a color embedding feature that matches the diffusion transformer model, and input the color embedding feature into the diffusion transformer model; Process the color embedding feature and the image description text through the diffusion transformer model to generate the target image.
3. The method according to claim 1, characterized in that The processing the color embedding feature and the image description text through the text-to-image model to generate the target image includes: Generate a corresponding text embedding feature based on the image description text; Based on the multi-head attention mechanism, input the color embedding feature and the text embedding feature into the text-to-image model to generate the target image.
4. The method according to claim 3, characterized in that, The inputting the color embedding feature and the text embedding feature into the text-to-image model based on the multi-head attention mechanism to generate the target image includes: Input the color embedding feature into the first attention mapping layer of the text-to-image model to obtain a first attention vector; Input the text embedding feature into the second attention mapping layer of the text-to-image model to obtain a second attention vector; Concatenate the first attention vector and the second attention vector to obtain a mixed attention vector; Process the mixed attention vector through the text-to-image model to generate the target image.
5. The method according to claim 4, characterized in that, The method further includes: Obtain a noise latent representation vector corresponding to the target image; The concatenating the first attention vector and the second attention vector to obtain a mixed attention vector includes: Concatenate the first attention vector, the second attention vector, and the noise latent to obtain a mixed attention vector.
6. The method according to claim 4, wherein The mixed attention vector includes a mixed query vector, a mixed key vector, and a mixed value vector; the processing the mixed attention vector through the text-to-image model to generate the target image includes: Based on the mixed query vector and the mixed key vector, obtain an attention weight matrix; Obtain a sequence feature according to the product of the attention weight matrix and the mixed value vector; Through the text-to-image model, perform reverse denoising on an initial noise map based on the sequence feature to generate the target image.
7. The method according to claim 1, wherein Before obtaining the color information configured by the user, the method further includes: Displaying a color configuration control, where the color configuration control is at least used to display at least two candidate colors; In response to a trigger operation on the color configuration control, generating the color information.
8. The method according to claim 7, wherein The method further includes: Generating a recommended number of colors according to the image description text; The generating the color information in response to a trigger operation on the color configuration control includes: In response to a trigger operation on the color configuration control, generating the color information based on the recommended number of colors.
9. The method according to claim 1, characterized in that It further includes at least one of the following: The color information further includes color ratio information, and the color ratio information is used to characterize the proportion of each target color indicated by the color information in the target image; After generating the target image, displaying the color ratio information.
10. An image generation device, characterized in that, It includes: An interaction unit, configured to obtain an image description text and color information configured by a user, where the image description text is used to characterize the image content of a target image to be generated, and the color information is used to characterize at least two target colors; A processing unit, configured to generate a color embedding feature matching the text-to-image model based on the color information, and process the color embedding feature and the image description text through the text-to-image model to generate the target image, where the target image is composed of the target colors.
11. An electronic device, characterized in that, It includes: A processor and a memory; The memory stores computer execution instructions; The processor executes the computer execution instructions stored in the memory, so that the processor executes the image generation method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, Computer execution instructions are stored in the computer-readable storage medium, and when the processor executes the computer execution instructions, the image generation method according to any one of claims 1 to 9 is implemented.
13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the image generation method according to any one of claims 1 to 9 is implemented.