Multi-layer image generation method, device and storage medium

By receiving natural language instructions to generate editable multi-layered images, this technology solves the problems of high limitations in image generation, poor text rendering effects, and high barriers to entry in existing technologies, and achieves automated, professional-grade rendering and highly flexible image generation.

CN120388087BActive Publication Date: 2026-02-06BEIJING YUANSHI TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510429210.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2026-02-06
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

Existing technologies for image generation suffer from limitations, poor text rendering effects, low degree of freedom, and high barriers to entry. In particular, in automated graphic design tasks, users need a lot of professional knowledge and human intervention, and the generated text and design drawings lack flexibility.

Method used

By receiving natural language-based instruction information, an editable multi-layered image is generated using a large language model. The layer content includes image layers and non-image layers. Image layers are generated based on instruction information and embedded images, while non-image layers are generated only based on layer configuration information, thus realizing the automatic generation and editing of images.

Benefits of technology

It lowers the barrier to entry, improves text legibility and rendering effects, and enhances the freedom and flexibility of images. Users can generate and edit beautiful multi-layered images with simple commands.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388087B_ABST
    Figure CN120388087B_ABST
Patent Text Reader

Abstract

The application discloses a multi-layer image generation method, device and storage medium. The multi-layer image generation method comprises the following steps: receiving instruction information based on a natural language, wherein the instruction information can be used for generating a corresponding image, and contains description text related to image content of an image to be generated; and displaying an editable multi-layer image, wherein the content of at least one layer in the multi-layer image is generated according to the description text.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to a multi-layer image generation method and device and a storage medium. BACKGROUND

[0002] The success of scaling law makes many difficult tasks feasible. Graphic design uses visual as a way of communication and expression, by creating and combining text, images and graphics, to make visual presentation to convey information and data. Automated graphic design task design needs to realize functions such as multi-layer sorting, layer content generation and layout arrangement, involving multi-modal content understanding and generation. For such a difficult task as graphic design, there is currently no complete automated solution.

[0003] One existing solution is to retrieve, manually select and artificially design all text, image and graphic materials, and then predict the positions and sizes of these materials in the form of images through a layout generation model, and then achieve an aesthetic visual effect by splicing these elements. However, this approach has some limitations due to the need to determine all materials in advance. On the one hand, it cannot be fully automated, and users need to collect or design specific material presentation forms, including the visual form of images and graphics, the font, color, effect and arrangement of text paragraphs, etc., which requires a lot of professional knowledge and human cost. On the other hand, when the elements of graphic design are limited in advance, the possibility of graphic design is reduced, resulting in a low upper limit of the visual presentation of the design. For example, if a beginner without professional knowledge provides inharmonious materials, such as chaotic layout in paragraphs and chaotic color selection, the visual presentation of the entire design will not be beautiful.

[0004] Another solution uses diffusion models to complete the generation of graphic design with text. This solution can follow instructions and complete the creation process of graphic design in a single stage. However, this approach also faces some challenges. 1) Poor text rendering. The generated text often presents incorrect, missing, and extra spelling conditions. Especially in the case of small text, large text and dense text, the text generated by such methods is almost indistinguishable. 2) Low degree of freedom. The design generated by the diffusion model is presented in the form of a picture, and does not have any editing flexibility. 3) High usage threshold. Diffusion models often require users to input as detailed and professional instructions as possible to help the model get satisfactory results, and different models may have different instruction preferences, which greatly reduces the feasibility of users using such methods.

[0005] The prior art has the technical problems of high limitation of image generation, poor text rendering effect, low degree of freedom, and high use threshold. SUMMARY

[0006] Embodiments of the present application provide a multi-layer image generation method, device and storage medium to at least solve the technical problems of high limitation of image generation, poor text rendering effect, low degree of freedom, and high use threshold in the prior art.

[0007] According to one aspect of the embodiments of the present application, a multi-layer image generation method is provided, comprising: receiving instruction information based on natural language, wherein the instruction information can be used to generate a corresponding image, and contains description text related to the image content of the image to be generated; and displaying an editable multi-layer image, wherein the content of at least part of the layers in the multi-layer image is generated according to the description text.

[0008] According to another aspect of the embodiments of the present application, a multi-layer image generation method is also provided, comprising: receiving instruction information based on natural language, wherein the instruction information can be used to generate a corresponding image, and contains description text related to the image content of the image to be generated; generating multi-layer image information corresponding to the multi-layer image according to the instruction information by a large language model, wherein the multi-layer image information includes layer configuration information and embedded images, the layer configuration information indicates the configuration information of each layer of the multi-layer image, and the embedded images are used to be inserted into the image layers of the multi-layer image, and wherein the layers of the multi-layer image include image layers and non-image layers, the image layers are generated according to the corresponding layer configuration information and the embedded images, and the non-image layers can be generated only according to the corresponding layer configuration information; and generating an editable multi-layer image according to the multi-layer image information.

[0009] According to another aspect of the embodiments of the present application, a multi-layer image generation device is also provided, comprising: a first receiving module for receiving instruction information based on natural language, wherein the instruction information can be used to generate a corresponding image, and contains description text related to the image content of the image to be generated; and an image display module for displaying an editable multi-layer image, wherein the content of at least part of the layers in the multi-layer image is generated according to the description text.

[0010] According to another aspect of the embodiments of the present application, a multi-layer image generation apparatus is also provided, comprising: a second receiving module configured to receive instruction information based on a natural language, wherein the instruction information can be used to generate a corresponding image, and contains description text related to image content of an image to be generated; an information generation module configured to generate multi-layer image information corresponding to a multi-layer image according to the instruction information by a large language model, wherein the multi-layer image information comprises layer configuration information and embedded images, the layer configuration information indicates configuration information of each layer of the multi-layer image, and the embedded images are used to be inserted into image layers of the multi-layer image, and wherein the layers of the multi-layer image comprise image layers and non-image layers, the image layers are generated according to the corresponding layer configuration information and the embedded images, and the non-image layers can be generated only according to the corresponding layer configuration information; and an image generation module configured to generate an editable multi-layer image according to the multi-layer image information.

[0011] According to another aspect of the embodiments of the present application, a multi-layer image generation apparatus is also provided, comprising: a first processor; and a first memory connected with the first processor, configured to provide the first processor with instructions to process the following processing steps: receiving instruction information based on a natural language, wherein the instruction information can be used to generate a corresponding image, and contains description text related to image content of an image to be generated; and displaying an editable multi-layer image, wherein the content of at least part of the layers in the multi-layer image is generated according to the description text.

[0012] According to another aspect of the embodiments of the present application, a multi-layer image generation apparatus is also provided, comprising: a second processor; and a second memory connected with the second processor, configured to provide the second processor with instructions to process the following processing steps: receiving instruction information based on a natural language, wherein the instruction information can be used to generate a corresponding image, and contains description text related to image content of an image to be generated; generating multi-layer image information corresponding to a multi-layer image according to the instruction information by a large language model, wherein the multi-layer image information comprises layer configuration information and embedded images, the layer configuration information indicates configuration information of each layer of the multi-layer image, and the embedded images are used to be inserted into image layers of the multi-layer image, and wherein the layers of the multi-layer image comprise image layers and non-image layers, the image layers are generated according to the corresponding layer configuration information and the embedded images, and the non-image layers can be generated only according to the corresponding layer configuration information; and generating an editable multi-layer image according to the multi-layer image information.

[0013] According to another aspect of the embodiments of the present application, a computer readable storage medium having a computer program stored thereon is also provided, wherein the computer program is executed by a processor to implement the steps of the above method.

[0014] In the embodiment of the present application, the terminal device receives instruction information of generating an image input by a user, wherein the instruction information can be simple and not highly professional. The user can describe at least part of the content of the image to be generated by inputting text, thereby reducing the use threshold. Therefore, the technical solution automatically generates a multi-layer image that is beautiful and meets the user's requirements according to the instruction information. Moreover, the multi-layer image in the technical solution is generated by uniformly generating each layer, so that the text layer is consistent with the professional-level rendering of other layers, thereby improving the legibility and rendering effect of the text. Moreover, the multi-layer image in the technical solution can be edited after generation, thereby improving the degree of freedom and flexibility of the multi-layer image. Thus, the technical problems of high limitation of image generation, poor text rendering effect, low degree of freedom and high use threshold in the prior art are solved. BRIEF DESCRIPTION OF DRAWINGS

[0015] The accompanying drawings, which are included to provide a further understanding of the present application and are incorporated in and constitute a part of this application, illustrate embodiments of the present application and serve to explain the present application. In the drawings:

[0016] Figure 1 is a hardware structure block diagram of a computing device for implementing the method according to Embodiment 1 of the present application;

[0017] Figure 2A is a schematic diagram of a multi-layer image generation system according to Embodiment 1 of the present application;

[0018] Figure 2B is another schematic diagram of a multi-layer image generation system according to Embodiment 1 of the present application;

[0019] Figure 3 is a flowchart of a multi-layer image generation method according to the first aspect of Embodiment 1 of the present application;

[0020] Figure 4A is an interface diagram of displaying an image layer according to Embodiment 1 of the present application;

[0021] Figure 4B is an interface diagram of displaying a non-image layer according to Embodiment 1 of the present application;

[0022] Figure 5A is a first schematic diagram of a graphic layer according to Embodiment 1 of the present application;

[0023] Figure 5B is a second schematic diagram of a graphic layer according to Embodiment 1 of the present application;

[0024] Figure 5C is a third schematic view of the graphical layer according to the embodiment 1 of the present application;

[0025] Figure 5D is a fourth schematic view of the graphical layer according to the embodiment 1 of the present application;

[0026] Figure 5E is a fifth schematic view of the graphical layer according to the embodiment 1 of the present application;

[0027] Figure 6A is a first schematic view of the text layer according to the embodiment 1 of the present application;

[0028] Figure 6B is a second schematic view of the text layer according to the embodiment 1 of the present application;

[0029] Figure 6C is a third schematic view of the text layer according to the embodiment 1 of the present application;

[0030] Figure 7 is a schematic view of the image layer according to the embodiment 1 of the present application;

[0031] Figure 8 is a schematic view of the multi-layer image according to the embodiment 1 of the present application;

[0032] Figure 9 is a flowchart of generating the first embedded image according to the embodiment 1 of the present application;

[0033] Figure 10A is a flowchart of generating the fourth layer configuration information according to the embodiment 1 of the present application;

[0034] Figure 10B is a flowchart of generating the third text feature information according to the embodiment 1 of the present application;

[0035] Figure 11 is a flowchart of generating the second embedded image according to the embodiment 1 of the present application;

[0036] Figure 12 is a flowchart of the multi-layer image generation method according to the second aspect of the embodiment 1 of the present application;

[0037] Figure 13 is a schematic view of the multi-layer image generation apparatus according to the first aspect of the embodiment 2 of the present application;

[0038] Figure 14 is a schematic view of the multi-layer image generation apparatus according to the second aspect of the embodiment 2 of the present application;

[0039] Figure 15 is a schematic diagram of a multi-layer image generation apparatus according to a first aspect of the present embodiment 3; and

[0040] Figure 16 is a schematic diagram of a multi-layer image generation apparatus according to a second aspect of the present embodiment 3. DETAILED DESCRIPTION

[0041] In order to make the technical personnel in the art better understand the technical solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor should be within the scope of protection of the present application.

[0042] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0043] Embodiment 1

[0044] According to the present embodiment, a method embodiment of a multi-layer image generation method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that herein.

[0045] The method embodiment provided by the present embodiment can be executed in a mobile terminal, a computer terminal, a server or a similar computing device. Figure 1 A hardware structure block diagram of a computing device for implementing a multi-layer image generation method is shown. As Figure 1As shown, the computing device can include one or more processors (which can include, but are not limited to, processing devices such as microprocessors, MCUs, or programmable logic devices, FPGAs, etc.), a memory for storing data, and a transmission device for communication functions. In addition, a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera can be included. Those of ordinary skill in the art can understand that Figure 1 The structure shown is only schematic, and does not limit the structure of the electronic device described above. For example, the computing device can further include more or fewer components than those shown in Figure 1 or have a different configuration than that shown in Figure 1 .

[0046] It should be noted that the one or more processors and / or other data processing circuits described above can be referred to herein generally as "data processing circuits". The data processing circuits can be embodied in whole or in part as software, hardware, firmware, or any combination thereof. In addition, the data processing circuits can be a single independent processing module, or be incorporated in whole or in part into any one of the other elements of the computing device. As referred to in embodiments of the present application, the data processing circuits serve as a processor to control, for example, the selection of the variable resistance terminal path connected to the interface.

[0047] The memory can be used to store software programs and modules of application software, such as program instructions / data storage means corresponding to the multi-layer image generation method of the application embodiments. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, i.e., implements the multi-layer image generation method of the application program described above. The memory can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory can further include a memory remotely disposed with respect to the processor, which can be connected to the computing device through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0048] The transmission device is used to receive or send data via a network. Specific examples of the network can include a wireless network provided by a communication provider of the computing device. In one example, the transmission device includes a network adapter (NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device can be a radio frequency (RF) module, which is used to communicate with the Internet in a wireless manner.

[0049] The display can be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with the user interface of the computing device.

[0050] It is noted that in some alternative embodiments, the above-described Figure 1 The computing device shown can include hardware elements (including circuitry), software elements (including computer code stored on a computer readable medium), or a combination of both hardware and software elements. It should be noted that in some embodiments, hardware elements can include software elements that Figure 1 is merely one example of a particular implementation and is intended to illustrate the types of components that can be present in the computing device described above.

[0051] Figure 2A is a schematic diagram of a terminal device for implementing a multi-layer image generation method according to the present embodiment, which describes one specific application scenario of the present application. Referring to Figure 2A As shown, the terminal device includes an image processing application and an AI module. The image processing application can be, for example, an image processing application such as PhotoShop, or a graphic design application.

[0052] Thus, the terminal device receives instruction information of a user through the image processing application, then generates multi-layer image information for generating a multi-layer image according to the instruction information of the user through the AI module, then generates a multi-layer image according to the multi-layer image information through the image processing application, and displays the multi-layer image through the image processing application.

[0053] In addition, in another embodiment, Figure 2B is a schematic diagram of a multi-layer image generation system according to the present embodiment, which describes another specific application scenario of the present application. Referring to Figure 2B As shown, the system includes a terminal device, an image processing platform, and a large model server.

[0054] The terminal device logs in to the image processing platform through a client, and obtains instruction information of a user, and sends the instruction information to the image processing platform. The client can be a client of an image processing application, or a browser. Then the image processing platform sends the instruction information to the large model server, and the large model server generates multi-layer image information for generating a multi-layer image according to the instruction information through an AI module, and then returns the multi-layer image information to the image processing platform. The image processing platform generates a multi-layer image according to the received multi-layer image information, and returns the multi-layer image to the terminal device for display through the client.

[0055] In addition, although in the aboveFigure 2B The AI module is independently arranged on the large model server relative to the image processing platform, but the AI module can also be deployed on the image processing platform, and a server deployed by the image processing platform is accessed (not shown in the figure).

[0056] Under the above operating environment, according to a first aspect of the embodiment, a multi-layer image generation method is provided, which is implemented by Figure 2A or Figure 2B the terminal device shown in Figure 3 The flowchart of the method is shown, and the method includes: Figure 3

[0057] S302: receiving natural language-based instruction information, wherein the instruction information can be used to generate a corresponding image, and contains description text related to the image content of the image to be generated; and

[0058] S304: displaying an editable multi-layer image, wherein the content of at least one layer in the multi-layer image is generated according to the description text.

[0059] Specifically, referring to Figure 2A The terminal device is provided with an image processing application program for generating images, so that the user inputs instruction information for generating images through the terminal device in the image processing application program. The instruction information may, for example, be description text indicating the generation of a poster or the generation of a slide. The instruction information contains description text related to the image content of the image to be generated, for example: Please generate a poster with the words of "CARD", "preview", "HAPPY BIRTHDAY". Then the image processing application program receives the instruction information and calls the AI module to generate multiple layers according to the instruction information, and combines and superimposes each layer to generate a multi-layer image that meets the user's requirements. The multi-layer image may, for example, be a poster, a slide, a time image, a commercial image, and an emoticon, etc. And the content of the layer is, for example, generated according to the description text related to the image content of the image to be generated in the instruction information input by the user. Or, it is generated according to the description text related to the image content in the instruction information input by the user and the picture input by the user.

[0060] Further, the terminal device displays the multi-layer image through the image processing application program. The user can display each layer of the multi-layer image through the image processing application program, so as to edit each layer.

[0061] In another embodiment, referring to Figure 2B ​As shown, the user logs in the image processing platform through the client of the terminal device, so that the user receives the instruction information of the user through the terminal device, and sends the instruction information to the image processing platform. The instruction information may be, for example, a description text indicating to generate a poster or a slide. The instruction information includes a description text related to the image content to be generated. Then the image processing platform sends the instruction information to the large model server. The large model server receives the instruction information and generates multi-layer image information for generating a multi-layer image according to the instruction information through an AI module. Then the large model server returns the multi-layer image information to the image processing platform, so that the image processing platform generates a plurality of layers according to the received multi-layer image information, and then combines and superimposes each layer to generate a multi-layer image meeting the user's requirements. The multi-layer image may be, for example, a poster, a slide, a time image, a commercial image, an emoticon, etc. The content of the layer is generated, for example, according to the description text related to the image content of the image to be generated in the instruction information input by the user. Or, it is generated according to the description text related to the image content in the instruction information input by the user and the picture input by the user. Then the image processing platform returns the multi-layer image to the terminal device and displays it through the client. Thus, the user can open each layer of the multi-layer image through the client of the terminal device, and edit each layer.

[0062] As described in the background, one existing solution is to obtain all text, image and graphic materials through retrieval, manual selection and manual design, and then predict the position and size of these materials in image form through a layout generation model, and then achieve an aesthetic visual effect by splicing these elements. However, this method has some limitations because all materials need to be determined in advance. On the one hand, it cannot be fully automated, and users need to collect or design specific material forms, including the visual form of images and graphics, the font, color, effect and arrangement of text paragraphs, etc., which requires a lot of professional knowledge and human cost. On the other hand, when the elements of the graphic design are limited in advance, the possibility of the graphic design is reduced, resulting in a low upper limit of the aesthetic of the design. For example, if a beginner without professional knowledge provides inharmonious materials, such as chaotic layout in paragraphs and chaotic color selection, the visual presentation of the entire design will not be beautiful. Another solution uses a diffusion model to complete the generation of graphic design with text. This solution can follow instructions and complete the creation process of graphic design in a single stage. However, this method also faces some challenges. 1) Poor text rendering. The generated text often presents incorrect, missing, and additional spelling conditions. Especially in the case of small text, large text, and dense text, the text generated by such methods is almost illegible. 2) Low degree of freedom. The design generated by the diffusion model is presented in the form of a picture, and does not have any editing flexibility. 3) High usage threshold. Diffusion models often require users to input as detailed and professional instructions as possible to help the model get satisfactory results, and different models may have different instruction preferences, which greatly reduces the feasibility of users using such methods.

[0063] To solve the above technical problems, the technical solution of the embodiments of the present application is that the terminal device receives instruction information for generating an image input by the user, wherein the instruction information can be simple and not highly professional. The user can describe at least part of the content of the image to be generated by inputting text, thereby reducing the usage threshold. Thus, the technical solution automatically generates a multi-layer image that is beautiful and meets the user's requirements according to the instruction information. Moreover, the multi-layer image in the technical solution is generated by uniformly generating each layer, so the text layer will be consistently and professionally rendered with other layers, thereby improving the legibility and rendering effect of the text. Furthermore, the multi-layer image in the technical solution can be edited after generation, thereby improving the degree of freedom and flexibility of the multi-layer image. Thus, the technical problems of high limitation of image generation, poor text rendering, low degree of freedom and high usage threshold in the prior art are solved.

[0064] Optionally, the operation of displaying the multi-layer image comprises any one of the following: displaying a file link corresponding to the multi-layer image; displaying a file identifier corresponding to the multi-layer image; displaying layer configuration information of each layer of the multi-layer image; and displaying image content of each layer of the multi-layer image.

[0065] In particular, the method for displaying the multi-layer image by the terminal device comprises any one of the following, but is not limited to:

[0066] (1) After the multi-layer image is generated, the terminal device can display the multi-layer image in the form of a file link through an image processing application or a client, so that the user can click the file link to download or view the multi-layer image.

[0067] (2) After the multi-layer image is generated, the multi-layer image is displayed in the form of a downloaded file identifier on the system interface of the terminal device, so that the user can click the file identifier to view the multi-layer image.

[0068] (3) After the multi-layer image is generated, the multi-layer image is displayed in the form of a file identifier on the interface of the image processing application or the browser, so that the user can click the file identifier to download or view the multi-layer image.

[0069] (4) After the multi-layer image is generated, the terminal device displays layer configuration information of each layer of the multi-layer image. The layer configuration information is configuration data for describing the position, size and shape of each layer. Thus, the terminal device can display the layer configuration information through an image processing application or a client, so that the user can view the layer configuration information of the multi-layer image.

[0070] (5) After the multi-layer image is generated, the terminal device displays layer configuration information of each layer of the multi-layer image through an image processing application or a client. The layer configuration information is configuration data for describing the position, size and shape of each layer. Thus, the user can directly view and modify the layer configuration information on the image editor through the terminal device.

[0071] (6) After the multi-layer image is generated, the terminal device displays image content of each layer of the multi-layer image through an image processing application or a client, so that the user can directly edit each layer through the image processing application or the client.

[0072] It should be noted that the above-mentioned display method of the multi-layer image can be displayed in at least one way, which is not limited here.

[0073] Therefore, the technical solution can display the multi-layer image in various ways to meet different viewing needs of users.

[0074] Optionally, the operation of displaying the image content of each layer of the multi-layer image includes: displaying editable pixels corresponding to the image content in the layer; and displaying editable controls corresponding to the image content in the layer.

[0075] Specifically, the multi-layer image includes an image layer and a non-image layer, the image layer is an image, and the non-image layer is a control. Referring to Figure 4A As shown, for the image layer, when the image processing application or the client displays the image layer, the image content therein is editable pixels, so that the image on the image layer can be painted by a painting tool in the toolbar or zoomed in or out.

[0076] Referring to Figure 4B As shown, for the non-image layer, when the image processing application or the client displays the non-image layer, the image content therein is editable controls, so that the parameters (such as layer configuration information) of the controls (such as a graphic or a text box) constituting the non-image layer can be directly modified, such as zooming in or out of the graphic in the non-image layer, modifying the color of the graphic, or modifying the text content in the text box.

[0077] In addition, other controls can also be added to the image layer and the non-graphic layer, such as adding a text control to add text, etc.

[0078] Therefore, the technical solution displays the multi-layer image by using controls, thereby flexibly modifying the multi-layer image, and improving the autonomy of the user in modifying the multi-layer image.

[0079] Optionally, the method further includes: generating multi-layer image information corresponding to the multi-layer image according to the instruction information, wherein the multi-layer image information includes layer configuration information and embedded images, the layer configuration information indicates the configuration information of each layer of the multi-layer image, and the embedded images are used to be inserted into the image layer of the multi-layer image, and wherein the layers of the multi-layer image include an image layer and a non-image layer, the image layer is generated according to the corresponding layer configuration information and the embedded images, and the non-image layer can be generated only according to the corresponding layer configuration information.

[0080] Specifically, referring to Figure 2AAs shown, the image processing application of the terminal device inputs the received instruction information into the AI module. For example, the instruction information indicates the user to generate a poster with the words of "CARD", "preview", "HAPPY BIRTHDAY" and the like. Thus, the instruction information can be: Please generate a poster with the words of "CARD", "preview", "HAPPY BIRTHDAY".

[0081] Then, the AI module generates multi-layer image information corresponding to the multi-layer image according to the instruction information, where the multi-layer image information includes layer configuration information and embedded images. The layers of the multi-layer image include frame layer (frame), grouping layer (group), graphic layer (graphic), text layer (text) and image layer (image). The frame layer (frame), the grouping layer (group), the graphic layer (graphic) and the text layer (text) are non-image layers. The layer configuration information of the image layer (image) is used to describe the position, size, color, content and name of the embedded image and the like. When generating the layer configuration information, the AI module can predict the parameters in the layer configuration information of each layer and the stacking order of the layers according to the instruction information. The instruction information can specifically indicate some parameters (such as position, size, color, content and name) of the layers, so that the layer configuration information conforming to the instruction information is generated in the case that the instruction information specifically indicates the parameters. In the case that the instruction information does not indicate the specific parameters, the AI module predicts according to the instruction information to generate the corresponding layer configuration information.

[0082] For example, the layer configuration information can be: <position> 0,0,1200,628< / position> <color> #000000,1.0< / color> <graphic> <position> 0,0,996,445< / position> <color>#610B1C,1.0< / color> <path><move_to>0,0< / move_to><line_to>996,0< / line_to><line_to>996,445< / line_to><line_to>0,445< / line_to><close_path / >< / path> < / graphic> <position> 186,99,827,448< / position> <image_des>foreground, people, RGB< / image_des><image_id>{image_1}< / image_id> <group> <position> ... <graphic> ...< / graphic> < / position> < / group> <position> 86,70,887,440< / position> <image_des>foreground, people, RGB< / image_des><image_id>{image_2}< / image_id> <text> <position> 216,63,770,77< / position> <color>#B9223E,1.0< / color> <font><font_weight>900< / font_weight><font_family>Cinzel Decorative Black< / font_family><font_size>57< / font_size>< / font> <text_align>center< / text_align><letter_spacing>0.0< / letter_spacing> <content>Happy Valentine's Day< / content> < / text> <group> <position>... <graphic> ...< / graphic> .

[0083] wherein:

[0084] " <position> 0,0,1200,628< / position> <color> #000000,1.0< / color> " is the layer configuration information of the frame layer, according to which the coordinate of the upper left corner of the frame is (0, 0), the width is "1200", and the height is "628"; the color of the frame is "#000000, 1.0", wherein "#000000" represents pure black, and "1.0" represents that the transparency is completely opaque.

[0085] " <graphic> <position> 0,0,996,445< / position> <color>#610B1C,1.0< / color> <path><move_to>0,0< / move_to><line_to>996,0 < / line_to><line_to>996,445< / line_to><line_to>0,445< / line_to><close_path / >< / path> < / graphic> " is the layer configuration information of the graphic layer, according to which a graphic is drawn, so that in the layer configuration information, the coordinate of the upper left corner of the graphic is (0, 0), the width is "996", and the height is "445"; the color of the graphic is "#610B1C, 1.0", wherein "#610B1C" represents dark red, and "1.0" represents that the transparency is completely opaque; the graphic is a closed graphic drawn by a path, so that the starting point coordinate is (0.0), and then drawn horizontally to the right to (996, 0), then drawn vertically downward to (996, 445), then drawn horizontally to the left to (0, 445), and finally the path is closed to connect the end point (996, 0) and the starting point (0.0), generating a closed rectangular graphic.

[0086] " <text> <position> 216,63,770,77< / position> <color>#B9223E,1.0< / color> <font><font_weight>900< / font_weight><font_family>Cinzel Decorative Black< / font_family><font_size>57< / font_size>< / font> <text_align>center< / text_align><letter_spacing>0.0< / letter_spacing> <content>Happy Valentine's Day< / content> < / text> " is the layer configuration information of the text layer, according to which the coordinate of the upper left corner of the text is (216, 63), the width is "770", and the height is "77"; the color is "#B9223E, 1.0", wherein "#B9223E" is dark red, and "1.0" represents that the transparency is completely opaque; the font is bold, and the bold level is the thickest; the font is the Cinzel Decorative Black variant; the font size is 57; the text position is horizontally centered; the text character spacing is 0; and the text content is "Happy Valentine's Day".

[0087] " <group> <position> ... <graphic> ...< / graphic> < / position> < / group> " is the layer configuration information of the grouping layer, which is used to combine multiple graphic layers (graphic), text layers (text), and image layers (image).

[0088] " <position> 186,99,827,448< / position> The image layer configuration information of the image layer is "image_des foreground, people, RGB" and "image_id {image_1}". According to the image layer configuration information, the coordinates of the top-left corner of the image are (186, 99), the width is "827", and the height is "448". The image description is "foreground, people, RGB", in which "foreground" indicates that the image is located in the foreground, "people" indicates that the image content contains people, and "RGB" indicates that the color space of the image is in the RGB format. The name of the image is "image_1".

[0089] " <position> 86,70,887,440< / position> The image layer configuration information of the image layer is "image_des foreground, people, RGB" and "image_id {image_2}". According to the image layer configuration information, the coordinates of the top-left corner of the image are (86, 70), the width is "887", and the height is "440". The image description is "foreground, people, RGB", in which "foreground" indicates that the image is located in the foreground, "people" indicates that the image content contains people, and "RGB" indicates that the color space of the image is in the RGB format. The name of the image is "image_2".

[0090] According to the above layer configuration information, the layer configuration information includes the layer configuration information of the frame layer, the group layer, the graphic layer, the text layer, and the image layer. The image layer is generated according to the corresponding layer configuration information and the embedded image. The non-image layer can be generated only according to the corresponding layer configuration information.

[0091] Therefore, the technical solution can predict and generate complex and diverse layers through simple instruction information, realize complete automation, reduce the workload of the user, and reduce the limitations of image generation.

[0092] Optionally, the operation of displaying the multi-layer image includes: displaying the multi-layer image according to the multi-layer image information.

[0093] Specifically, after the AI module of the terminal device generates the multi-layer image information, the multi-layer image information is processed to generate multi-layer image information in a format suitable for processing by the image processing platform. The image processing platform may be, for example, a penpot design platform. Then the terminal device sends the multi-layer image information in the processed format to the image processing platform through the image processing application.

[0094] Further, the image processing platform generates non-image layers (including frame layers, group layers, graphic layers, and text layers) according to the layer configuration information of the frame layers, group layers, graphic layers, and text layers in the multi-layer image information. Among them Figure 5A to Figure 5E Five graphic layers are shown; Figure 6A to Figure 6C Three text layers are shown. And the image processing platform generates an image layer (image) according to the layer configuration information of the image layer and the embedded image. Among them Figure 7 The image layer is shown. Thus, the image processing platform displays a multi-layer image generated by combining non-image layers and image layers (see Figure 8 ).

[0095] Thus, the technical solution generates multi-layer image information suitable for processing by the image processing platform through the AI module, and generates a multi-layer image composed of non-image layers and image layers through the image processing platform, thereby improving the applicability between the AI module and the image processing platform, and more accurately generating a multi-layer image that meets user requirements.

[0096] Optionally, in the case where the instruction information indicates that the multi-layer image is generated based on single-modal information based on natural language, the operation of generating the multi-layer image information according to the instruction information comprises: generating first layer configuration information based on first text feature information corresponding to the instruction information using a large language model, wherein the first layer configuration information includes second layer configuration information and third layer configuration information, the second layer configuration information is used to indicate the layer configuration information of the first part of the non-image layer, and the third layer configuration information is used to indicate the part of the layer configuration information corresponding to the first image layer, and the third layer configuration information is located after the second layer configuration information; and generating a first embedded image based on the first text feature information and second text feature information corresponding to the first layer configuration information using a large language model and a diffusion model, wherein the first embedded image is used to insert the first image layer.

[0097] Specifically, in the present embodiment, the instruction information indicates to generate a multi-layer image by using single-modal information based on natural language, wherein the single-modal information is used to indicate only text information based on natural language. For example, the instruction information can be: Please generate a poster with the words of "CARD", "preview", "HAPPY BIRTHDAY". Thus, the image processing application program of the terminal device inputs the instruction information into the AI module. Then, the AI module receives the instruction information and generates text feature information (i.e., first text feature information) from the instruction information. Wherein the first text feature information is used to indicate a feature vector corresponding to each character of the instruction information.

[0098] Wherein the AI module is pre-provided with a large language model. Then, the AI module inputs the first text feature information into the large language model, and processes the first text feature information by using a causal mask through the large language model to generate second text feature information. Then, the AI module can generate first layer configuration information according to the second text feature information through a preset text decoder. Wherein the text decoder is used to decode the feature information into the layer configuration information (not shown in the figure). Wherein the second text feature information is used to indicate a feature vector corresponding to each character of the first layer configuration information.

[0099] For example, the first layer configuration information is: <position> 0,0,1200,628< / position> <color> #000000,1.0< / color> <graphic> <position> 0,0,996,445< / position> <color>#610B1C,1.0

[0100] < / color> <path><move_to>0,0< / move_to><line_to>996,0< / line_to><line_to>996,445< / line_to><line_to>0,445< / line_to><close_path / >< / path> < / graphic> <position> 186,99,827,448< / position> <image_des>foreground, people, RGB< / image_des><image_id>{image_gen}.

[0101] More specifically, for the process of generating the first layer configuration information, the AI module first generates second layer configuration information of the first part of the non-image layer through the text decoder.

[0102] For example, the second layer configuration information is: <position> 0,0,1200,628< / position> <color> #000000,1.0< / color> <graphic> <position> 0,0,996,445< / position> <color>#610B1C,1.0

[0103] < / color> <path><move_to>0,0< / move_to><line_to>996,0< / line_to><line_to>996,445< / line_to><line_to>0,445< / line_to><close_path / >< / path> < / graphic> .

[0104] Then, the AI module generates part of the layer configuration information corresponding to the first image layer after the second layer configuration information through the text decoder, for example: <position> 186,99,827,448< / position> foreground, people, RGB< / image_des><image_id>{image_gen}”. Wherein the first image layer is the first image layer generated when generating the multi-layer image. Therefore, the AI module groups the second layer configuration information and the third configuration information into the first layer configuration information.

[0105] Wherein the first layer configuration information ends at the {image_gen} tag, the {image_gen} tag of the first layer configuration information is the first generated {image_gen} tag, and the position of the {image_gen} tag is used to write the embedded image name, and the layer configuration information of the image layer has only generated part of the layer configuration information ending at the {image_gen} tag, and another part of the image layer configuration information has not been generated.

[0106] Further, the AI module is also pre-set with a diffusion model, so that the image processing application generates a first embedded image according to the first text feature information and the second text feature information corresponding to the first layer configuration information, using the large language model and the diffusion model. Wherein the first embedded image is used to insert the first image layer.

[0107] Therefore, in the present technical solution, the AI module generates text form layer configuration information, so that when generating a multi-layer image later, the text layer can be directly generated using the layer configuration information, avoiding the problem that the text cannot be recognized due to poor text rendering effect. And because language information itself has sequentiality, the large language model uses causal masking to simulate the sequential process of the language information of the instruction information, preventing information leakage.

[0108] Optionally, according to the first text feature information and the second text feature information corresponding to the first layer configuration information, the operation of generating the first embedded image by using the large language model and the diffusion model includes: splicing the first text feature information, the second text feature information and the preset initial embedding feature information to generate first fusion feature information, wherein the initial embedding feature information is feature information used to generate an embedded image; generating second fusion feature information suitable for image processing according to the first fusion feature information through the large language model; generating first semantic feature information according to the second fusion feature information through the first projection layer; and generating the first embedded image according to the first semantic feature information through the diffusion model.

[0109] Specifically, refer to Figure 9 As shown, the AI module splices the first text feature information, the second text feature information, and the preset initial embedding feature information to generate first fusion feature information, wherein the initial embedding feature information is a feature vector used to generate an embedded image. Then, the AI module inputs the first fusion feature information into a large language model, processes the first fusion feature information by using full attention masking of the large language model, and generates second fusion feature information suitable for image processing, wherein the second fusion feature information is a text feature. Then, the AI module inputs the second fusion feature information into a preset first projection layer, performs projection processing on the second fusion information by the first projection layer, and generates corresponding semantic feature information (i.e., first semantic feature information).

[0110] Further, the AI module inputs the first semantic feature information into a diffusion model, wherein the diffusion model includes a first linear layer, a visual encoder, and a second linear layer. The visual encoder may, for example, be a pre-trained CLIP visual encoder ViT-L / 14.

[0111] The diffusion model thus inputs the received first semantic feature information into the first linear layer, processes the first semantic feature information by the first linear layer, and generates first output feature. Then, the diffusion model inputs the first output feature into the visual encoder, encodes the first output feature by the visual encoder, and generates second output feature. Then, the diffusion model inputs the second output feature into the second linear layer, processes the second output feature by the second linear layer, and generates a first embedded image.

[0112] Thus, the technical solution can highlight key features by performing semantic processing on text feature information and image feature information by a large language model, predict parameters of multiple layers, and generate a corresponding embedded image by a diffusion model. The technical solution combines the use of a large language model and a diffusion model, which not only avoids the use of complex instruction information by the user, but also generates image content that meets the requirements of the user according to the description of the user.

[0113] Optionally, the method further includes: generating, by the large language model, fourth layer configuration information according to the first image feature information corresponding to the first embedded image and the first fusion feature information, wherein the fourth layer configuration information includes fifth layer configuration information and sixth layer configuration information, and wherein the fifth layer configuration information is used to indicate layer configuration information of a second part of non-image layers, and the sixth layer configuration information is used to indicate part of layer configuration information corresponding to a second image layer after the fifth layer configuration information; and generating, by the large language model and the diffusion model, a second embedded image according to the first fusion feature information, the first image feature information, and third text feature information corresponding to the fourth layer configuration information, wherein the second embedded image is used to be inserted into the second image layer.

[0114] Specifically, after generating the first image embedding image and the first image layer configuration information, the AI module generates image feature information (i.e., first image feature information) corresponding to the first image embedding image. Then the AI module inputs the first image feature information and the first fusion feature information into the large language model, processes the first image feature information and the first fusion feature information through the large language model, outputs corresponding feature information, and inputs the feature information into the text encoder. Then the text decoder generates an image name label corresponding to the first image embedding image according to the feature information, for example, {image_1}. Then the AI module replaces {image_gen} in the first image layer configuration information with {image_1}, and generates the image layer configuration information of another part of the image layer after {image_1} "< / image_id>". The name of the image is generated according to the generation order of the image layer, for example, the image name in the first generated image layer is automatically named as "image_1", the image name in the second generated image layer is automatically named as "image_2", and the image name in the n generated image layers is automatically named as "image_n".

[0115] Further, the large language model further generates third text feature information. Then the AI module inputs the third text feature information into the text decoder, and the text decoder generates fourth image layer configuration information according to the third text feature information.

[0116] For example, the fourth image layer configuration information is: <group> <position> ... <graphic> ...< / graphic> < / position> < / group> <position> 86,70,887,440< / position> <image_des>...< / image_des><image_id>{image_gen}.

[0117] More specifically, for the process of generating the fourth image layer configuration information, the AI module first generates fifth image layer configuration information after the image layer configuration information through the text decoder. The fifth image layer configuration information is used to indicate the image layer configuration information of the second part of the non-image layer, for example, it can be <group> <position> ... <graphic> ...< / graphic> < / position> < / group> .

[0118] Further, the text decoder generates sixth image layer configuration information after the fifth image layer configuration information. The sixth image layer configuration information is used to indicate the image layer configuration information corresponding to the second image layer. For example, <position> 186,99,827,448< / position> foreground, people, RGB< / image_des><image_id>{image_gen}”. Wherein the second image layer is the second image layer generated when generating the multi-layer image. Thus the AI module groups the fifth layer configuration information and the sixth configuration information into the fourth layer configuration information.

[0119] Wherein the fourth layer configuration information ends at the {image_gen} tag, the {image_gen} tag in the fourth layer configuration information is the second generated {image_gen} tag, and the position of the {image_gen} tag is used to generate the image name tag, and the layer configuration information of the image layer has only generated part of the layer configuration information ending at the {image_gen} tag, and another part of the image layer configuration information has not been generated.

[0120] It should be noted that the fifth layer configuration information can also not be generated. That is, after generating another part of the layer configuration information corresponding to the first image layer, the layer configuration information (i.e., the sixth layer configuration information) of the next image layer (i.e., the second image layer) is directly generated, and no other non-image layer needs to be generated before the sixth layer configuration information. Here, generation can be made according to actual conditions, and no specific limitation is made.

[0121] Further, the AI module generates a second embedded image according to the first fusion feature information, the first image feature information, and the third text feature information corresponding to the fourth layer configuration information, using a large language model and a diffusion model, wherein the second embedded image is used to insert the second image layer.

[0122] Further, after the AI module generates the fourth layer configuration information and the second embedded image, it can continue to generate other layer configuration information and embedded images until it is no longer necessary to generate.

[0123] Thus, the technical solution can generate multi-image layer information without restriction until the user requirements are met, so that the generated multi-layer image can be more detailed and accurate.

[0124] Optionally, the operation of generating the fourth layer configuration information using the large language model according to the first image feature information corresponding to the first embedded image and the first fusion feature information includes: generating second image feature information according to the first embedded image through a first image encoder; generating first image feature information suitable for processing by the large language model according to the second image feature information through a second projection layer; and generating the fourth layer configuration information using the large language model according to the first image feature information and the first fusion feature information.

[0125] Specifically, reference is made to Figure 10A As shown, the AI ​​module inputs the first embedded image to a preset first image encoder, and encodes the first embedded image through the first image encoder to generate second image feature information.

[0126] Furthermore, the AI ​​module inputs the second image feature information into a preset second projection layer, and performs projection processing on the second image feature information through the second projection layer to generate first image feature information suitable for processing by a large language model.

[0127] Further, refer to Figure 10A as well as Figure 10B As shown, the AI ​​module inputs the first image feature information and the first fused feature information into the large language model, which then generates feature information based on these two information. The AI ​​module then inputs this feature information into the text encoder, which generates fourth-layer configuration information based on the received feature information. The first fused feature information is composed of the first text feature information, the second text feature information, and the initial embedded feature information.

[0128] Therefore, this technical solution generates image feature information suitable for processing by a large language model by first encoding and then mapping the embedded image. This allows the large language model to better process the image feature information and adapt to various image feature information, thus improving the adaptability of the large language model.

[0129] Optionally, the operation of generating a second embedded image using a large language model and a diffusion model based on the first fusion feature information, the first image feature information, and the third text feature information corresponding to the fourth layer configuration information includes: concatenating the first fusion feature information, the first image feature information, and the third text feature information to generate third fusion feature information; generating fourth fusion information suitable for image processing using the large language model based on the third fusion feature information; generating second semantic feature information using the first projection layer based on the fourth fusion information; and generating a second embedded image using the diffusion model based on the second semantic feature information.

[0130] Specifically, refer to Figure 11 As shown, after generating the third text feature information, the AI ​​module concatenates the first fused feature information, the first image feature information, the third text feature information, and the initial embedded feature information to generate the third fused feature information.

[0131] Furthermore, the AI ​​module inputs the third fusion feature information into the large language model, and the large language model processes the third fusion feature information using a full attention mask to generate a fourth fusion feature information suitable for image processing.

[0132] Further, the AI module inputs the fourth fusion feature information into a first projection layer, and performs projection processing on the fourth fusion feature information through the first projection layer to generate second semantic feature information.

[0133] Further, the AI module inputs the second semantic feature information into a diffusion model, and generates a second embedded image according to the second semantic feature information through the diffusion model.

[0134] Therefore, the third fusion feature information is processed by the full attention mask, so that the model can not only pay attention to the previous content when creating and understanding the embedded image, but also capture spatial information related to generating and understanding the current image.

[0135] Optionally, the method further comprises: generating seventh layer configuration information by using the large language model according to the first image feature information corresponding to the first embedded image and the first fusion feature information, wherein the seventh layer configuration information is used to indicate the layer configuration information of the remaining non-image layers except the first part of non-image layers.

[0136] Specifically, after the AI module generates the first layer configuration information and the first embedded image, the first image feature information corresponding to the first embedded image is input into the large language model, and the first image feature information and the first fusion feature information are processed by the large language model to generate corresponding feature information. Then the AI module inputs the feature information into the text decoder, so that the text decoder generates an image name label corresponding to the first embedded image, for example, {image_1}, so as to replace {image_gen} in the first layer configuration information with {image_1}, and generate the layer configuration information of another part of image layers after {image_1} "< / image_id>".

[0137] Further, without generating other embedded images and corresponding layer configuration information, the large language model is used to generate corresponding feature information according to the first image feature information corresponding to the first embedded image and the first fusion feature information. Then the AI module inputs the feature information into the text decoder, and the text decoder generates the seventh layer configuration information according to the feature information. The seventh layer configuration information is used to indicate the layer configuration information of the remaining non-image layers except the first part of non-image layers, for example, including the end label of the frame layer, the text layer and the grouping layer, etc.

[0138] For example, the seventh layer configuration information is: <text> <position> 216,63,770,77< / position> <color>#B9223E,1.0< / color> <font><font_weight>900< / font_weight><font_family>Cinzel Decorative Black< / font_family><font_size>57< / font_size>< / font> <text_align>center< / text_align><letter_spacing>0.0< / letter_spacing> <content>Happy Valentine's Day< / content> < / text> <group> <position>... <graphic> ...< / graphic> .

[0139] Therefore, when the image is not needed to be generated, the AI module can directly generate the layer configuration information of the remaining non-image layer, thereby improving the efficiency of generating the layer configuration information.

[0140] Optionally, in the case that the instruction information indicates that the multi-layer image is generated by using the existing embedded image, the operation of generating the multi-layer image information according to the instruction information comprises: generating fifth image feature information by the second image encoder according to the preset third embedded image; generating sixth image feature information suitable for processing of the large language model according to the fifth image feature information by the third projection layer; and generating seventh layer configuration information for describing the image layer and eighth layer configuration information for describing the non-image layer by the large language model according to fourth text feature information corresponding to the instruction information and the sixth image feature information.

[0141] Specifically, in the embodiment, the instruction information indicates that the multi-layer image is generated by using the existing embedded image. The existing embedded image is an image for inserting the image layer and does not need to be generated in real time by the diffusion model.

[0142] Therefore, the image processing application program of the terminal device receives the instruction information for input and the embedded image (i.e., the third embedded image), and then inputs the instruction information and the third embedded image into the AI module. Then, the AI module inputs the third embedded image into the second image encoder, and the second image encoder encodes the third embedded image to generate the fifth image feature information.

[0143] Further, the AI module inputs the fifth image feature information into the third projection layer, and the third projection layer projects the fifth image feature information to generate the sixth image feature information suitable for processing of the large language model.

[0144] Further, the AI module generates corresponding fourth text feature information according to the instruction information, and then inputs the fourth text feature information and the sixth image feature information into the large language model, and the large language model generates corresponding feature information according to the fourth text feature information and the sixth image feature information. Then, the AI module inputs the feature information into the text decoder, and the text decoder generates the seventh layer configuration information for describing the image layer and the eighth layer configuration information for describing the non-image layer according to the feature information.

[0145] For example, the generated all layer configuration information is as follows: <position> 0,0,1200,628< / position> <color> #000000,1.0< / color> <graphic> <position> 0,0,996,445< / position> <color>#610B1C,1.0

[0146] < / color> <path><move_to>0,0< / move_to><line_to>996,0< / line_to><line_to>996,445< / line_to><line_to>0,445< / line_to><close_path / >< / path> < / graphic> <position> 186,99,827,448< / position> <image_des>foreground,people,RGB< / image_des><image_id>{image_1}< / image_id> <text> <position> 216,63,770,77< / position> <color>#B9223E,1.0< / color> <font><font_weight>900< / font_weight><font_family>Cinzel Decorative Black< / font_family><font_size>57< / font_size>< / font> <text_align>center< / text_align><letter_spacing>0.0< / letter_spacing> <content>Happy Valentine's Day< / content> < / text> <group> <position>... <graphic> ...< / graphic> .

[0147] The seventh layer configuration information is: <position> 186,99,827,448< / position> <image_des>foreground, people, RGB< / image_des><image_id>{image_1}< / image_id>. The layer configuration information other than the fourth layer configuration information in the entire layer configuration information is eighth layer configuration information for describing a non-image layer.

[0148] Therefore, the technical solution processes the input information including both instruction information and multi-modal information embedded with an image through a large language model, so as to quickly generate layer configuration information according to the instruction information and the embedded image, thereby improving the efficiency of generating layer configuration information.

[0149] Optionally, the large language model is trained by the following steps: collecting sample instruction information, first sample fusion feature information, and first sample image feature information, wherein the first sample fusion feature information is generated by splicing feature information corresponding to the sample instruction information, sample layer configuration information, and initial embedded feature information, and the first sample image feature information is feature information of a sample embedded image; generating second sample fusion feature information according to the sample instruction information by using the large language model; generating second sample image feature information according to the first sample fusion feature information; determining a fusion feature loss according to the first sample fusion feature information and the second sample fusion feature information; determining an image feature loss according to the first sample image feature information and the second sample image feature information; and training the large language model according to the fusion feature loss and the image feature loss, wherein the operation of training the large language model further includes adjusting the initial embedded feature information.

[0150] Specifically, the image processing application program is further provided with a training module, so that the training module of the image processing application program collects sample data, and then trains the parameter values of the large language model, the diffusion model, the first projection layer, the second projection layer, and the initial embedded feature information by using the sample data.

[0151] More specifically, the sample data collected by the training module of the image processing application program includes: sample instruction information, first sample fusion feature information, and first sample image feature information. The first sample fusion feature information is generated by splicing feature information corresponding to the sample instruction information and sample layer configuration information, and the initial embedded feature information; and the first sample image feature information is feature information of a sample embedded image.

[0152] For sample instruction information: the training module performs data augmentation and semantic augmentation operations on the layer configuration information of the text layer in the sample instruction information. For example, the text content of the text layer is "Lichun" "Spring is the season of all things coming to life." The random data augmentation used is to replace each character with random text, such as "Zhang San" "Zhang San Li Si Wang Wu Yi San Si Si", the content is random, but the length of the string remains the same, and all other attributes such as font, font size and color remain unchanged. Semantic augmentation is to generate context-consistent text using Qwen2, which introduces unique content while preserving semantic relevance, such as changing it to "Lidong" "Winter is a season of peace and tranquility." These data augmentations are used to expand the sample size and diversity of the training data, resulting in better model generation.

[0153] For the case where the instruction information indicates the generation of a multi-layer image based on natural language-based single-modal information, the training steps for the large language model are as follows:

[0154] (1) The training module collects sample instruction information, first sample fusion feature information and first sample image feature information, where the sample instruction information x = {x1, x2, …, xN} represents the i-th character in the sample instruction information, i = 1 ~ N. The first sample image feature information v = {v1, v2, …, vM} represents the j-th feature item in the first sample image feature information, j = 1 ~ M. The first sample fusion feature information c = {c1, c2, …, cQ} represents the k-th feature item in the first sample fusion feature information, k = 1 ~ Q. N} represents the i-th character in the sample instruction information, i = 1 ~ N. The first sample image feature information v = {v1, v2, …, vM} represents the j-th feature item in the first sample image feature information, j = 1 ~ M. The first sample fusion feature information c = {c1, c2, …, cQ} represents the k-th feature item in the first sample fusion feature information, k = 1 ~ Q. i M j Q k

[0155] (2) The training module inputs the text feature information corresponding to the sample instruction information into the large language model, processes the sample instruction information through the large language model, and outputs the first sample layer configuration information. Then, the feature information corresponding to the sample instruction information and the first sample layer configuration information, as well as the sample initial embedding feature information, are spliced to obtain the second sample fusion feature information. The second sample fusion feature information is the predicted value.

[0156] (3) The training module outputs the sample image by sequentially passing through the large language model, the first projection layer and the diffusion model according to the second sample fusion feature information. And the training module outputs the second sample image feature information by passing through the first encoder and the second projection layer according to the sample image. The second sample image feature information is the predicted value.

[0157] ​​​​​(4) The training module inputs the first sample image feature information, the text feature information of the sample instruction information, and the second sample fusion feature information into the large language model to generate multi-layer image information.

[0158] (5) The training module calculates a text loss of the first sample fusion feature information and the second sample fusion feature information

[0159] (6) The training module calculates an image loss between the first sample image feature information and the second sample image feature information

[0160] (7) The training module determines a model loss according to the text loss and the image loss

[0161] (8) The training module updates the parameter values of the first projection layer, the second projection layer, the large language model, the diffusion model, and the initial embedding feature information through gradient backpropagation according to the model loss.

[0162] In the case where the instruction information indicates that the multi-layer image is generated by using an existing embedded image, the training steps of the large language model are as follows:

[0163] (1) The training module collects sample instruction information and sample embedded images and second sample layer configuration information;

[0164] (2) The training module processes the sample embedded images through the second image encoder and the third projection layer to generate sample image feature information;

[0165] (3) The training module outputs third sample layer configuration information according to the text feature information corresponding to the sample instruction information and the sample image feature information through the large language model. The third sample layer configuration information is a predicted value.

[0166] (4) The training module calculates a text loss between the second sample layer configuration information and the third sample layer configuration information, and trains the large language model, the second image encoder, and the third projection layer according to the text loss.

[0167] Thus, the technical solution trains the models in the AI module based on single-modal information and multi-modal information in different ways, so that the AI module can be suitable for both single-modal information and multi-modal information, and the diversity of multi-layer image generation is improved.

[0168] It should be noted that the large language models for processing single-modal information and multi-modal information can be two different large language models or the same large language model, which is not limited here.

[0169] In addition, if the embedded image is a solid color background image, it will affect the training effect of the diffusion model, so that when training the diffusion model and the large language model by using the embedded image with a solid color background, the training module will use the k-means clustering method to select top-k color tones from the image to capture the main color tone of the image. Based on the Potrace algorithm, the training module converts each main color region into a vector path, and the paths of each main color tone are combined together to create a unified SVG file to capture the color structure. The SVG file is a graphic layer. That is, the embedded image with a solid color background generates a graphic in the graphic layer, so that the large language model is trained by the graphic. This process realizes the extensible and compact expression of simple design elements. When the embedded image has not yet been converted into an SVG file, the training module uses Qwen2-VL to generate a short descriptive label for it, thereby providing context information about the design element and enhancing controllability.

[0170] According to a first aspect of the present embodiment, the terminal device receives user input instruction information for generating an image, where the instruction information can be simple and not highly professional. The user can describe at least part of the content of the image to be generated by inputting text, thereby reducing the use threshold. Therefore, the present technical solution automatically generates a multi-layer image that is aesthetically pleasing and meets the user's requirements according to the instruction information. Moreover, the multi-layer image in the present technical solution is generated by uniformly generating each layer, so the text layer will be consistently and professionally rendered with other layers, thereby improving the legibility and rendering effect of the text. Furthermore, the multi-layer image in the present technical solution can be edited after generation, thereby improving the freedom and flexibility of the multi-layer image. In this way, the technical problems of high limitation of image generation, poor text rendering effect, low freedom, and high use threshold in the prior art are solved.

[0171] In addition, according to a second aspect of the present embodiment, a multi-layer image generation method is provided. Figure 12 A flowchart of the method is shown, and with reference to Figure 12 The method includes:

[0172] S1202: receiving natural language-based instruction information, where the instruction information can be used to generate a corresponding image, and contains a description text related to the image content of the image to be generated;

[0173] S1204: Generate multi-layer image information corresponding to a multi-layer image based on instruction information using a large language model. This multi-layer image information includes layer configuration information and an embedded image. The layer configuration information indicates the configuration information of each layer in the multi-layer image, and the embedded image is used to insert into the image layers of the multi-layer image. The layers of the multi-layer image include image layers and non-image layers. Image layers are generated based on the corresponding layer configuration information and the embedded image, while non-image layers can be generated solely based on the corresponding layer configuration information.

[0174] S1206: Generate an editable multi-layer image based on multi-layer image information.

[0175] Specifically, refer to Figure 2B As shown, the user logs into the image processing platform through a client on their terminal device. The user receives and sends instructions to the platform. These instructions can be, for example, descriptive text instructing the generation of a poster or a slideshow. The instructions include descriptive text related to the content of the image to be generated. The image processing platform then sends this instruction to a large model server. Upon receiving the instruction, the large model server uses its AI module to generate multi-layered image information for creating a multi-layered image. The large model server then returns this multi-layered image information to the image processing platform, which generates multiple layers based on the received information. These layers are then combined and overlaid to create a multi-layered image that meets the user's requirements. This multi-layered image can be, for example, a poster, a slideshow, a time-based image, a commercial image, or an emoji. The content of each layer is generated, for example, based on the descriptive text related to the image content in the user's instructions, or based on the descriptive text related to the image content in the user's instructions and the user-input image. Finally, the image processing platform returns the multi-layered image to the terminal device for display via the client. Therefore, users can open and edit each layer of the multi-layer image through a client on their terminal device.

[0176] More specifically, for example, the instruction message might be to instruct the user to generate a poster with the words "CARD", "preview", and "HAPPY BIRTHDAY". Thus, the instruction message could be: "Please generate a poster with the words 'CARD', 'preview', and 'HAPPY BIRTHDAY'".

[0177] After that, the AI module generates multi-layer image information corresponding to the multi-layer image according to the instruction information, wherein the multi-layer image information includes layer configuration information and embedded images. The layers of the multi-layer image include frame layer (frame), grouping layer (group), graphic layer (graphic), text layer (text) and image layer (image). The frame layer (frame), grouping layer (group), graphic layer (graphic) and text layer (text) are non-image layers. The layer configuration information of the image layer (image) is used to describe the position, size, color, content and name of the embedded image. And when generating the layer configuration information, the AI module can predict the parameters in the layer configuration information of each layer and the stacking order of the layers according to the instruction information. Wherein the instruction information may specifically indicate some parameters of the layers (such as position, size, color, content and name, etc.), so as to generate the layer configuration information in accordance with the instruction information in the case of specific parameters indicated by the instruction information. In the case where the instruction information does not indicate specific parameters, the AI module will make predictions according to the instruction information to generate the corresponding layer configuration information.

[0178] For example, the layer configuration information can be: <position> 0,0,1200,628< / position> <color> #000000,1.0< / color> <graphic> <position> 0,0,996,445< / position> <color>#610B1C,1.0

[0179] < / color> <path><move_to>0,0< / move_to><line_to>996,0< / line_to><line_to>996,445< / line_to><line_to>0,445< / line_to><close_path / >< / path> < / graphic> <position> 186,99,827,448< / position> <image_des>foreground,people,RGB< / image_des><image_id>{image_1}< / image_id> <group> <position> ... <graphic> ...< / graphic> < / position> < / group> <position> 86,70,887,440< / position> <image_des>foreground,people,RGB< / image_des><image_id>{image_2}< / image_id> <text> <position> 216,63,770,77< / position> <color>#B9223E,1.0< / color> <font><font_weight>900< / font_weight><font_family>Cinzel Decorative Black< / font_family><font_size>57< / font_size>< / font> <text_align>center< / text_align><letter_spacing>0.0< / letter_spacing> <content>Happy Valentine's Day< / content> < / text> <group> <position>... <graphic> ...< / graphic> .

[0180] wherein:

[0181] " <position> 0,0,1200,628< / position> <color> #000000,1.0< / color> " is the layer configuration information of the frame layer, according to which the coordinate of the upper left corner of the frame is (0, 0) and the coordinate of the lower right corner of the frame is (1200, 628); the color of the frame is "#000000, 1.0", wherein "#000000" represents pure black and "1.0" represents that the transparency is completely opaque.

[0182] " <graphic> <position> 0,0,996,445< / position> <color>#610B1C,1.0< / color> <path><move_to>0,0< / move_to><line_to>996,0 < / line_to><line_to>996,445< / line_to><line_to>0,445< / line_to><close_path / >< / path> < / graphic> " is the layer configuration information of the graphic layer, according to which a graphic is drawn, so that in the layer configuration information, the coordinate of the upper left corner of the graphic is (0, 0) and the coordinate of the lower right corner of the graphic is (996, 445); the color of the graphic is "#610B1C, 1.0", wherein "#610B1C" represents dark red and "1.0" represents that the transparency is completely opaque; the graphic is a closed graphic drawn by a path, so that the starting point coordinate is (0.0), then horizontally drawn to the right to (996, 0), then vertically drawn downward to (996, 445), then horizontally drawn to the left to (0, 445), and finally the path is closed to connect the end point (996, 0) and the starting point (0.0), generating a closed rectangular graphic.

[0183] " <text> <position> 216,63,770,77< / position> <color>#B9223E,1.0< / color> <font><font_weight>900< / font_weight><font_family>Cinzel Decorative Black< / font_family><font_size>57< / font_size>< / font> <text_align>center< / text_align><letter_spacing>0.0< / letter_spacing> <content>Happy Valentine's Day< / content> < / text> " is the layer configuration information of the text layer, according to which the coordinate of the upper left corner of the text is (216, 63) and the coordinate of the lower right corner of the text is (770, 77); the color is "#B9223E, 1.0", wherein "#B9223E" is dark red and "1.0" represents that the transparency is completely opaque; the font is bold and the bold level is the thickest; the font is the Cinzel Decorative Black variant; the font size is 57; the text position is horizontally centered; the text character spacing is 0; and the text content is "Happy Valentine's Day".

[0184] " <group> <position> ... <graphic> ...< / graphic> < / position> < / group> " is the layer configuration information of the grouping layer, which is used to combine multiple graphic layers (graphic), text layers (text) and image layers (image).

[0185] " <position> 186,99,827,448< / position> The image description is "foreground, people, RGB", wherein "foreground" indicates that the image is located in the foreground, "people" indicates that the image content contains people, and "RGB" indicates that the color space of the image is in an RGB format. The name of the image is "image_1".

[0186] The image description is "foreground, people, RGB", wherein "foreground" indicates that the image is located in the foreground, "people" indicates that the image content contains people, and "RGB" indicates that the color space of the image is in an RGB format. The name of the image is "image_1". <position> 86,70,887,440< / position> The image description is "foreground, people, RGB", wherein "foreground" indicates that the image is located in the foreground, "people" indicates that the image content contains people, and "RGB" indicates that the color space of the image is in an RGB format. The name of the image is "image_1".

[0187] Referring to the above layer configuration information, the layer configuration information includes layer configuration information of a frame layer (frame), a group layer (group), a graphic layer (graphic), a text layer (text), and an image layer (image). The image layer is generated according to the corresponding layer configuration information and embedded images. The non-image layer can be generated only according to the corresponding layer configuration information.

[0188] As described in the background, one existing solution is to obtain all text, image and graphic materials through retrieval, manual selection and manual design, and then predict the position and size of these materials in image form through a layout generation model, and then achieve an aesthetic visual effect by splicing these elements. However, this method has some limitations because all materials need to be determined in advance. On the one hand, it cannot be fully automated, and users need to collect or design specific material forms, including the visual form of images and graphics, the font, color, effect and arrangement of text paragraphs, etc., which requires a lot of professional knowledge and human cost. On the other hand, when the elements of the graphic design are limited in advance, the possibility of the graphic design is reduced, resulting in a low upper limit of the aesthetic of the design. For example, if a beginner without professional knowledge provides inharmonious materials, such as chaotic layout in paragraphs and chaotic color selection, the visual presentation of the entire design will not be beautiful. Another solution uses a diffusion model to complete the generation of graphic design with text. This solution can follow instructions and complete the creation process of graphic design in a single stage. However, this method also faces some challenges. 1) Poor text rendering. The generated text often presents incorrect, missing, and additional spelling conditions. Especially in the case of small text, large text, and dense text, the text generated by such methods is almost illegible. 2) Low degree of freedom. The design generated by the diffusion model is presented in the form of a picture, and does not have any editing flexibility. 3) High usage threshold. Diffusion models often require users to input as detailed and professional instructions as possible to help the model get satisfactory results, and different models may have different instruction preferences, which greatly reduces the feasibility of users using such methods.

[0189] To solve the above technical problems, the technical solution of the embodiments of the present application is that the terminal device receives instruction information input by the user, wherein the instruction information can be simple and not highly professional, thereby reducing the usage threshold, so that the terminal device automatically generates a multi-layer image that is beautiful and meets the user's requirements according to the instruction information. And the multi-layer image in this technical solution is generated by uniformly generating each layer, so the text layer will be consistently professionally rendered with other layers, thereby improving the legibility and rendering effect of the text. And the multi-layer image in this technical solution can be edited after generation, thereby improving the degree of freedom and flexibility of the multi-layer image. Thus, the technical problems of high limitation of image generation, poor text rendering, low degree of freedom and high usage threshold in the prior art are solved.

[0190] Optionally, the operation of displaying the multi-layer image comprises any of the following: displaying a file link corresponding to the multi-layer image; displaying a file identifier corresponding to the multi-layer image; displaying layer configuration information of each layer of the multi-layer image; and displaying image content of each layer of the multi-layer image.

[0191] Optionally, the operation of displaying the image content of each layer of the multi-layer image comprises: displaying an editable pixel corresponding to the corresponding image content in the layer; and displaying an editable control corresponding to the corresponding image content in the layer.

[0192] Optionally, the operation of displaying the multi-layer image comprises: displaying the multi-layer image according to the multi-layer image information.

[0193] Optionally, in the case where the instruction information indicates that the multi-layer image is generated based on the natural language-based single-modal information, the operation of generating the multi-layer image information according to the instruction information comprises: generating first layer configuration information according to first text feature information corresponding to the instruction information by using a large language model, wherein the first layer configuration information comprises second layer configuration information and third layer configuration information, the second layer configuration information is used to indicate layer configuration information of a first part of non-image layers, the third layer configuration information is used to indicate part of layer configuration information corresponding to a first image layer, and the third layer configuration information is located after the second layer configuration information; and generating a first embedded image for inserting the first image layer by using the large language model and a diffusion model according to the first text feature information and second text feature information corresponding to the first layer configuration information.

[0194] Optionally, the operation of generating the first embedded image by using the large language model and the diffusion model according to the first text feature information and the second text feature information corresponding to the first layer configuration information comprises: splicing the first text feature information, the second text feature information and preset initial embedded feature information to generate first fusion feature information, wherein the initial embedded feature information is feature information used to generate an embedded image; generating second fusion feature information suitable for image processing according to the first fusion feature information by using the large language model; generating first semantic feature information according to the second fusion feature information by using a first projection layer; and generating the first embedded image according to the first semantic feature information by using the diffusion model.

[0195] Optionally, the method further comprises: generating, by the large language model, fourth layer configuration information according to the first image feature information corresponding to the first embedded image and the first fusion feature information, wherein the fourth layer configuration information comprises fifth layer configuration information and sixth layer configuration information, and wherein the fifth layer configuration information is used to indicate layer configuration information of the second part of non-image layers, and the sixth layer configuration information is used to indicate part of layer configuration information corresponding to the second image layer located after the fifth layer configuration information; and generating, by the large language model and the diffusion model, a second embedded image according to the first fusion feature information, the first image feature information, and third text feature information corresponding to the fourth layer configuration information, wherein the second embedded image is used to be inserted into the second image layer.

[0196] Optionally, the operation of generating, by the large language model, fourth layer configuration information according to the first image feature information corresponding to the first embedded image and the first fusion feature information comprises: generating, by a first image encoder, second image feature information according to the first embedded image; generating, by a second projection layer, first image feature information suitable for processing of the large language model according to the second image feature information; and generating, by the large language model, fourth layer configuration information according to the first image feature information and the first fusion feature information.

[0197] Optionally, the operation of generating, by the large language model and the diffusion model, a second embedded image according to the first fusion feature information, the first image feature information, and third text feature information corresponding to the fourth layer configuration information comprises: splicing the first fusion feature information, the first image feature information, and the third text feature information to generate third fusion feature information; generating, by the large language model, fourth fusion information suitable for image processing according to the third fusion feature information; generating, by a first projection layer, second semantic feature information according to the fourth fusion information; and generating, by the diffusion model, a second embedded image according to the second semantic feature information.

[0198] Optionally, the method further comprises: generating, by the large language model, seventh layer configuration information according to the first image feature information corresponding to the first embedded image and the first fusion feature information, wherein the seventh layer configuration information is used to indicate layer configuration information of remaining non-image layers except for the first part of non-image layers.

[0199] Optionally, in the case where the instruction information indicates generating the multi-layer image by using the existing embedded image, the operation of generating the multi-layer image information according to the instruction information comprises: generating fifth image feature information according to the preset third embedded image by the second image encoder; generating sixth image feature information suitable for processing by the large language model according to the fifth image feature information by the third projection layer; and generating seventh layer configuration information for describing the image layer and eighth layer configuration information for describing the non-image layer by the large language model according to the fourth text feature information corresponding to the instruction information and the sixth image feature information.

[0200] Optionally, the large language model is trained by: collecting sample instruction information, first sample fusion feature information and first sample image feature information, wherein the first sample fusion feature information is generated by splicing the feature information corresponding to the sample instruction information and the sample layer configuration information and the initial embedded feature information, and the first sample image feature information is the feature information of the sample embedded image; generating second sample fusion feature information according to the sample instruction information by the large language model; generating second sample image feature information according to the first sample fusion feature information; determining fusion feature loss according to the first sample fusion feature information and the second sample fusion feature information; determining image feature loss according to the first sample image feature information and the second sample image feature information; and training the large language model according to the fusion feature loss and the image feature loss, wherein the operation of training the large language model further comprises: adjusting the initial embedded feature information.

[0201] According to the second aspect of the embodiment, the terminal device receives user input instruction information for generating an image, wherein the instruction information can be simple and not highly professional. The user can describe at least part of the content of the image to be generated by inputting text, thereby reducing the use threshold. According to the instruction information, the technical solution automatically generates a multi-layer image that is beautiful and meets the user's requirements. Moreover, the multi-layer image in the technical solution is generated by uniformly generating each layer, so the text layer is consistent with other layers in professional level rendering, thereby improving the legibility and rendering effect of the text. Moreover, the multi-layer image in the technical solution can be edited after generation, thereby improving the freedom and flexibility of the multi-layer image. In this way, the technical problems of high limitation of image generation, poor text rendering effect, low freedom and high use threshold in the prior art are solved.

[0202] In addition, referring to FIG. 8, according to a third aspect of the embodiment, a storage medium is provided. The storage medium comprises a stored program, wherein when the program is executed by a processor, the method described in any one of the above embodiments is executed. Figure 1 In addition, referring to FIG. 8, according to a third aspect of the embodiment, a storage medium is provided. The storage medium comprises a stored program, wherein when the program is executed by a processor, the method described in any one of the above embodiments is executed.

[0203] Thus, according to the present embodiment, the terminal device receives instruction information of the generated image input by the user, wherein the instruction information can be simple and not highly professional. The user can describe at least part of the content of the image to be generated by inputting text, thereby reducing the use threshold. Thus, the present technical solution automatically generates a multi-layer image that is beautiful and meets the user's requirements according to the instruction information. Moreover, the multi-layer image in the present technical solution is generated by uniformly generating each layer, so that the text layer therein is consistent with the professional-level rendering of other layers, thereby improving the legibility and rendering effect of the text. Moreover, the multi-layer image in the present technical solution can be edited after generation, thereby improving the degree of freedom and flexibility of the multi-layer image. Thus, the technical problems of high limitation of image generation, poor text rendering effect, low degree of freedom, and high use threshold in the prior art are solved.

[0204] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited to the order of the actions described, because according to the present application, certain steps can be performed in other order or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0205] From the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and the necessary general hardware platform, of course, it can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes a number of instructions to make a terminal device (which can be a mobile phone, computer, server, or network device, etc.) execute the methods described in various embodiments of the present application.

[0206] Embodiment 2

[0207] Figure 13 A multi-layer image generation device 1300 according to the first aspect of the present embodiment is shown, which corresponds to the method according to the first aspect of Embodiment 1. Referring to Figure 13 As shown, the apparatus 1300 comprises: a first receiving module 1310 configured to receive natural language-based instruction information, wherein the instruction information can be used to generate a corresponding image, and contains description text related to the image content of the image to be generated; and an image display module 1320 configured to display an editable multi-layer image, wherein the content of at least part of the layers in the multi-layer image is generated according to the description text.

[0208] Optionally, the image display module 1320 comprises: a first display submodule configured to display a file link corresponding to the multi-layer image; a second display submodule configured to display a file identifier corresponding to the multi-layer image; a third display submodule configured to display layer configuration information of each layer of the multi-layer image; and a fourth display submodule configured to display image content of each layer of the multi-layer image.

[0209] Optionally, the fourth display submodule comprises: a first display unit configured to display an editable pixel corresponding to the corresponding image content in the layer; and a second display unit configured to display an editable control corresponding to the corresponding image content in the layer.

[0210] Optionally, the multi-layer image generation apparatus 1300 further comprises: a first information generation module configured to generate multi-layer image information corresponding to the multi-layer image according to the instruction information, wherein the multi-layer image information comprises layer configuration information and embedded images, the layer configuration information indicates the configuration information of each layer of the multi-layer image, and the embedded images are used to be inserted into the image layers of the multi-layer image, and wherein the layers of the multi-layer image comprise image layers and non-image layers, the image layers are generated according to the corresponding layer configuration information and the embedded images, and the non-image layers can be generated only according to the corresponding layer configuration information.

[0211] Optionally, the image display module 1320 comprises: a fifth display submodule configured to display the multi-layer image according to the multi-layer image information.

[0212] Optionally, the first information generation module comprises: a first generation submodule configured to generate first layer configuration information using a large language model according to first text feature information corresponding to the instruction information, wherein the first layer configuration information comprises second layer configuration information and third layer configuration information, wherein the second layer configuration information is used to indicate the layer configuration information of the first part of the non-image layers, and the third layer configuration information is used to indicate the part of the layer configuration information corresponding to the first image layer, and the third layer configuration information is located after the second layer configuration information; and a second generation submodule configured to generate a first embedded image using a large language model and a diffusion model according to the first text feature information and second text feature information corresponding to the first layer configuration information, wherein the first embedded image is used to insert the first image layer.

[0213] Optionally, the second generation submodule comprises: a first generation unit configured to concatenate the first text feature information, the second text feature information, and preset initial embedding feature information to generate first fusion feature information, wherein the initial embedding feature information is feature information used to generate an embedded image; a second generation unit configured to generate second fusion feature information suitable for image processing according to the first fusion feature information by using a large language model; a third generation unit configured to generate first semantic feature information according to the second fusion feature information by using a first projection layer; and a fourth generation unit configured to generate the first embedded image according to the first semantic feature information by using a diffusion model.

[0214] Optionally, the second generation submodule further comprises: a fifth generation unit configured to generate fourth layer configuration information according to the first image feature information corresponding to the first embedded image and the first fusion feature information by using the large language model, wherein the fourth layer configuration information comprises fifth layer configuration information and sixth layer configuration information, and wherein the fifth layer configuration information is used to indicate layer configuration information of the second part of non-image layers, and the sixth layer configuration information is used to indicate part of layer configuration information corresponding to the second image layer after the fifth layer configuration information; and a sixth generation unit configured to generate a second embedded image according to the first fusion feature information, the first image feature information, and third text feature information corresponding to the fourth layer configuration information by using the large language model and the diffusion model, wherein the second embedded image is used to be inserted into the second image layer.

[0215] Optionally, the fifth generation unit comprises: generating second image feature information according to the first embedded image by using a first image encoder; generating first image feature information suitable for processing of the large language model according to the second image feature information by using a second projection layer; and generating the fourth layer configuration information according to the first image feature information and the first fusion feature information by using the large language model.

[0216] Optionally, the sixth generation unit comprises: concatenating the first fusion feature information, the first image feature information, and the third text feature information to generate third fusion feature information; generating fourth fusion information suitable for image processing according to the third fusion feature information by using the large language model; generating second semantic feature information according to the fourth fusion information by using the first projection layer; and generating the second embedded image according to the second semantic feature information by using the diffusion model.

[0217] Optionally, the second generation submodule further comprises: a seventh generation unit configured to generate seventh layer configuration information according to the first image feature information corresponding to the first embedded image and the first fusion feature information by using the large language model, wherein the seventh layer configuration information is used to indicate layer configuration information of the remaining non-image layers except for the first part of non-image layers.

[0218] Optionally, the first information generation module further comprises: a third generation submodule for generating fifth image feature information according to the preset third embedded image by the second image encoder; a fourth generation submodule for generating sixth image feature information suitable for processing of the large language model according to the fifth image feature information by the third projection layer; and a fifth generation submodule for generating seventh layer configuration information for describing the image layer and eighth layer configuration information for describing the non-image layer by the large language model according to the fourth text feature information corresponding to the instruction information and the sixth image feature information.

[0219] Optionally, the training module comprises: a collecting submodule for collecting sample instruction information, first sample fusion feature information and first sample image feature information, wherein the first sample fusion feature information is generated by splicing feature information corresponding to the sample instruction information and the sample layer configuration information and the initial embedded feature information, and the first sample image feature information is feature information of a sample embedded image; a first sample generation submodule for generating second sample fusion feature information according to the sample instruction information by the large language model; a second sample generation submodule for generating second sample image feature information according to the first sample fusion feature information; a third sample generation submodule for determining fusion feature loss according to the first sample fusion feature information and the second sample fusion feature information; a fourth sample generation submodule for determining image feature loss according to the first sample image feature information and the second sample image feature information; and a fifth sample generation submodule for training the large language model according to the fusion feature loss and the image feature loss, wherein the training module further comprises: an adjusting submodule for adjusting the initial embedded feature information.

[0220] In addition, Figure 14 A multi-layer image generation apparatus 1400 according to the second aspect of the present embodiment is shown, which corresponds to the method according to the second aspect of the present embodiment. Reference is made to Figure 14 As shown, the apparatus 1400 includes: a second receiving module, configured to receive instruction information based on a natural language, wherein the instruction information can be used to generate a corresponding image, and contains description text related to image content of an image to be generated; an information generation module, configured to generate multi-layer image information corresponding to a multi-layer image according to the instruction information by a large language model, wherein the multi-layer image information includes layer configuration information and embedded images, the layer configuration information indicates configuration information of each layer of the multi-layer image, and the embedded images are used to be inserted into image layers of the multi-layer image, and wherein the layers of the multi-layer image include image layers and non-image layers, the image layers are generated according to the corresponding layer configuration information and the embedded images, and the non-image layers can be generated only according to the corresponding layer configuration information; and an image generation module, configured to generate an editable multi-layer image according to the multi-layer image information.

[0221] According to the present embodiment, the terminal device receives instruction information for generating an image input by a user, wherein the instruction information can be simple and not highly professional. The user can describe at least part of the content of the image to be generated by inputting text, thereby reducing the use threshold. According to the present technical solution, a multi-layer image that is beautiful and meets the user's requirements is automatically generated according to the instruction information. Moreover, the multi-layer image in the present technical solution is generated by uniformly generating each layer, so that the text layer is consistent with other layers in professional-level rendering, thereby improving the legibility and rendering effect of the text. Moreover, the multi-layer image in the present technical solution can be edited after being generated, thereby improving the freedom and flexibility of the multi-layer image. In this way, the technical problems of high limitation of image generation, poor text rendering effect, low freedom, and high use threshold in the prior art are solved.

[0222] Embodiment 3

[0223] Figure 15 A multi-layer image generation apparatus 1500 according to a first aspect of the present embodiment is shown, which corresponds to the method according to the first aspect of Embodiment 1. Reference is made to Figure 15 As shown, the apparatus 1500 includes: a first processor 1510; and a first memory 1520 connected with the first processor 1510, configured to provide the first processor 1510 with instructions for processing the following processing steps: receiving instruction information based on a natural language, wherein the instruction information can be used to generate a corresponding image, and contains description text related to image content of an image to be generated; and displaying an editable multi-layer image, wherein the content of at least part of the layers in the multi-layer image is generated according to the description text.

[0224] Optionally, the operation of displaying the multi-layer image comprises any of the following: displaying a file link corresponding to the multi-layer image; displaying a file identifier corresponding to the multi-layer image; displaying layer configuration information of each layer of the multi-layer image; and displaying image content of each layer of the multi-layer image.

[0225] Optionally, the operation of displaying the image content of each layer of the multi-layer image comprises: displaying an editable pixel corresponding to the corresponding image content in the layer; and displaying an editable control corresponding to the corresponding image content in the layer.

[0226] Optionally, the first memory 1520 is further configured to provide the first processor 1510 with instructions for processing the following processing steps: generating multi-layer image information corresponding to the multi-layer image according to the instruction information, wherein the multi-layer image information comprises layer configuration information and embedded images, the layer configuration information indicates the configuration information of each layer of the multi-layer image, and the embedded images are used to be inserted into the image layers of the multi-layer image, and wherein the layers of the multi-layer image include image layers and non-image layers, the image layers are generated according to the corresponding layer configuration information and the embedded images, and the non-image layers can be generated only according to the corresponding layer configuration information.

[0227] Optionally, the operation of displaying the multi-layer image comprises: displaying the multi-layer image according to the multi-layer image information.

[0228] Optionally, in the case where the instruction information indicates that the multi-layer image is generated by using single-modal information based on natural language, the operation of generating the multi-layer image information according to the instruction information comprises: generating first layer configuration information by using a large language model according to first text feature information corresponding to the instruction information, wherein the first layer configuration information comprises second layer configuration information and third layer configuration information, the second layer configuration information is used to indicate the layer configuration information of the first part of the non-image layers, the third layer configuration information is used to indicate the part of the layer configuration information corresponding to the first image layer, and the third layer configuration information is located after the second layer configuration information; and generating a first embedded image by using the large language model and the diffusion model according to the first text feature information and second text feature information corresponding to the first layer configuration information, wherein the first embedded image is used to be inserted into the first image layer.

[0229] Optionally, the operation of generating the first embedded image by using the large language model and the diffusion model according to the first text feature information and the second text feature information corresponding to the first layer configuration information comprises: splicing the first text feature information, the second text feature information and preset initial embedding feature information to generate first fusion feature information, wherein the initial embedding feature information is feature information used for generating the embedded image; generating second fusion feature information suitable for image processing by the large language model according to the first fusion feature information; generating first semantic feature information by the first projection layer according to the second fusion feature information; and generating the first embedded image by the diffusion model according to the first semantic feature information.

[0230] Optionally, the first memory 1520 is further configured to provide the first processor 1510 with instructions for processing the following processing steps: generating fourth layer configuration information by using the large language model according to the first image feature information corresponding to the first embedded image and the first fusion feature information, wherein the fourth layer configuration information comprises fifth layer configuration information and sixth layer configuration information, and wherein the fifth layer configuration information is used for indicating layer configuration information of the second part of non-image layers, and the sixth layer configuration information is used for indicating part of layer configuration information corresponding to the second image layer located after the fifth layer configuration information; and generating a second embedded image by using the large language model and the diffusion model according to the first fusion feature information, the first image feature information and third text feature information corresponding to the fourth layer configuration information, wherein the second embedded image is used for inserting the second image layer.

[0231] Optionally, the operation of generating the fourth layer configuration information by using the large language model according to the first image feature information corresponding to the first embedded image and the first fusion feature information comprises: generating second image feature information by the first image encoder according to the first embedded image; generating first image feature information suitable for processing of the large language model by the second projection layer according to the second image feature information; and generating the fourth layer configuration information by using the large language model according to the first image feature information and the first fusion feature information.

[0232] Optionally, the operation of generating the second embedded image by using the large language model and the diffusion model according to the first fusion feature information, the first image feature information and the third text feature information corresponding to the fourth layer configuration information comprises: splicing the first fusion feature information, the first image feature information and the third text feature information to generate third fusion feature information; generating fourth fusion information suitable for image processing by the large language model according to the third fusion feature information; generating second semantic feature information by the first projection layer according to the fourth fusion information; and generating the second embedded image by the diffusion model according to the second semantic feature information.

[0233] Optionally, the first memory 1520 is further configured to provide the first processor 1510 with instructions to process the following processing steps: generating, by the large language model, seventh layer configuration information according to the first image feature information corresponding to the first embedded image and the first fusion feature information, wherein the seventh layer configuration information is used to indicate the layer configuration information of the remaining non-image layers except for the first part of the non-image layers.

[0234] Optionally, in the case where the instruction information indicates that the multi-layer image is generated by using the existing embedded image, the operation of generating the multi-layer image information according to the instruction information comprises: generating, by the second image encoder, fifth image feature information according to the preset third embedded image; generating, by the third projection layer, sixth image feature information suitable for processing by the large language model according to the fifth image feature information; and generating, by the large language model, seventh layer configuration information for describing the image layers and eighth layer configuration information for describing the non-image layers according to fourth text feature information corresponding to the instruction information and the sixth image feature information.

[0235] Optionally, the large language model is trained by the following steps: collecting sample instruction information, first sample fusion feature information and first sample image feature information, wherein the first sample fusion feature information is generated by splicing the feature information corresponding to the sample instruction information and the sample layer configuration information and the initial embedded feature information, and the first sample image feature information is the feature information of the sample embedded image; generating, by the large language model, second sample fusion feature information according to the sample instruction information; generating second sample image feature information according to the first sample fusion feature information; determining fusion feature loss according to the first sample fusion feature information and the second sample fusion feature information; determining image feature loss according to the first sample image feature information and the second sample image feature information; and training the large language model according to the fusion feature loss and the image feature loss, wherein the operation of training the large language model further comprises adjusting the initial embedded feature information.

[0236] In addition, Figure 16 A multi-layer image generation apparatus 1600 according to the second aspect of the present embodiment is shown, which corresponds to the method according to the second aspect of Embodiment 1. Reference is made to Figure 16 As shown, the apparatus 1600 includes: a second processor 1610; and a second memory 1620, connected with the second processor 1610, configured to provide the second processor 1610 with instructions to process the following processing steps: receiving natural language-based instruction information, wherein the instruction information can be used to generate a corresponding image, and contains description text related to the image content of the image to be generated; generating multi-layer image information corresponding to a multi-layer image according to the instruction information through a large language model, wherein the multi-layer image information includes layer configuration information and embedded images, the layer configuration information indicates the configuration information of each layer of the multi-layer image, and the embedded images are used to be inserted into the image layer of the multi-layer image, and wherein the layers of the multi-layer image include image layers and non-image layers, the image layers are generated according to the corresponding layer configuration information and the embedded images, and the non-image layers can be generated only according to the corresponding layer configuration information; and generating an editable multi-layer image according to the multi-layer image information.

[0237] According to the present embodiment, the terminal device receives user input instruction information for generating an image, wherein the instruction information can be simple and not highly professional. The user can describe at least part of the content of the image to be generated by inputting text, thereby reducing the use threshold. According to the instruction information, the present technical solution automatically generates a multi-layer image that is beautiful and meets the user's requirements. And the multi-layer image in the present technical solution is generated by uniformly generating each layer, so the text layer therein will be consistently rendered with other layers at a professional level, thereby improving the legibility and rendering effect of the text. And the multi-layer image in the present technical solution can be edited after being generated, thereby improving the degree of freedom and flexibility of the multi-layer image. Further, the technical problems of high limitation of image generation, poor text rendering effect, low degree of freedom and high use threshold in the prior art are solved.

[0238] The above-mentioned serial numbers of the embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.

[0239] In the above-mentioned embodiments of the present application, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.

[0240] In several embodiments provided in the present application, it should be understood that the disclosed technology can be implemented by other ways. Among them, the above-described device embodiments are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed units can be indirect coupling or communication connection through some interfaces, units or modules, and can be electrical or other forms.

[0241] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0242] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0243] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.

[0244] The above is only the preferred embodiment of the present application, and it should be pointed out that for ordinary skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should be considered as the protection scope of the present application.< / position> < / group> < / position> < / group> < / position> < / group> < / position> < / group>

Claims

1. A method for generating multi-layer images, characterized in that, include: Receive natural language-based instruction information, wherein the instruction information can be used to generate a corresponding image and includes descriptive text related to the image content of the image to be generated; Multi-layer image information corresponding to the multi-layer image is generated according to the instruction information. This multi-layer image information includes layer configuration information and an embedded image. The layer configuration information indicates the textual configuration information of each layer of the multi-layer image, and the embedded image is used to insert into the image layer of the multi-layer image. The layer configuration information includes image layers and non-image layers. The image layers are generated based on the corresponding layer configuration information and the embedded image. The non-image layers can be generated solely based on the corresponding layer configuration information. The non-image layers include frame layers, grouped layers, text layers, and graphic layers. Displaying an editable multi-layer image, wherein the content of at least a portion of the layers in the multi-layer image is generated based on the descriptive text, and wherein the operation of displaying the multi-layer image includes: displaying the multi-layer image based on the multi-layer image information. Furthermore, the method further includes: when the instruction information indicates that the multi-layer image is generated using unimodal information based on natural language, generating multi-layer image information according to the instruction information, including: generating first layer configuration information using a large language model based on first text feature information corresponding to the instruction information, wherein the first layer configuration information includes second layer configuration information and third layer configuration information, wherein the second layer configuration information is used to indicate the layer configuration information of a first portion of non-image layers, and the third layer configuration information is used to indicate the layer configuration information of a portion of the first image layer, wherein the third layer configuration information is located after the second layer configuration information; and generating a first embedded image using a large language model and a diffusion model based on the first text feature information and the second text feature information corresponding to the first layer configuration information, wherein the first embedded image is used to insert into the first image layer, wherein... The operation of generating a first embedded image using a large language model and a diffusion model based on the first text feature information and the second text feature information corresponding to the first layer configuration information includes: concatenating the first text feature information, the second text feature information and preset initial embedded feature information to generate first fused feature information, wherein the initial embedded feature information is feature information used to generate the embedded image; generating second fused feature information suitable for image processing using the large language model based on the first fused feature information; generating first semantic feature information using the first projection layer based on the second fused feature information; and generating the first embedded image using the diffusion model based on the first semantic feature information.

2. The method according to claim 1, characterized in that, The operation of displaying the multi-layer image includes any of the following methods: Display the file links corresponding to the multi-layered image; Display the file identifier corresponding to the multi-layer image; Displays the layer configuration information of each layer of the multi-layer image; as well as Displays the image content of each layer of the multi-layer image.

3. The method according to claim 2, characterized in that, The operation of displaying the image content of each layer of the multi-layer image includes: The layer displays editable pixels corresponding to the corresponding image content; and The layer displays editable controls corresponding to the respective image content.

4. The method according to claim 1, characterized in that, Also includes: A fourth layer configuration information is generated using a large language model based on the first image feature information corresponding to the first embedded image and the first fusion feature information. The fourth layer configuration information includes a fifth layer configuration information and a sixth layer configuration information. The fifth layer configuration information is used to indicate the layer configuration information of the second part of the non-image layer, and the sixth layer configuration information is used to indicate the layer configuration information of the part of the second image layer that is located after the fifth layer configuration information. as well as Based on the first fusion feature information, the first image feature information, and the third text feature information corresponding to the fourth layer configuration information, a second embedded image is generated using the large language model and the diffusion model, wherein the second embedded image is used to insert into the second image layer. The operation of generating fourth layer configuration information using a large language model based on the first image feature information corresponding to the first embedded image and the first fused feature information includes: The first image encoder generates second image feature information based on the first embedded image; The second projection layer generates first image feature information suitable for processing by the large language model based on the second image feature information; and The large language model is used to generate fourth layer configuration information based on the first image feature information and the first fusion feature information.

5. The method according to claim 4, characterized in that, The operation of generating a second embedded image using the large language model and the diffusion model based on the first fusion feature information, the first image feature information, and the third text feature information corresponding to the fourth layer configuration information includes: The first fused feature information, the first image feature information, and the third text feature information are concatenated to generate the third fused feature information; The large language model generates fourth fusion information suitable for image processing based on the third fusion feature information; The second semantic feature information is generated by the first projection layer based on the fourth fusion information; and The second embedded image is generated based on the second semantic feature information using the diffusion model.

6. The method according to claim 4, characterized in that, Also includes: The large language model is used to generate seventh layer configuration information based on the first image feature information corresponding to the first embedded image and the first fusion feature information. The seventh layer configuration information is used to indicate the layer configuration information of the remaining non-image layers other than the first part of non-image layers.

7. The method according to claim 1, characterized in that, When the instruction information indicates that a multi-layer image should be generated using an existing embedded image, the operation of generating multi-layer image information according to the instruction information includes: The second image encoder generates the fifth image feature information based on the preset third embedded image; The third projection layer generates sixth image feature information suitable for large language model processing based on the fifth image feature information; and The large language model generates seventh layer configuration information for describing image layers and eighth layer configuration information for describing non-image layers based on the fourth text feature information corresponding to the instruction information and the sixth image feature information.

8. The method according to claim 4, characterized in that, The large language model is trained using the following steps: Collect sample instruction information, first sample fusion feature information and first sample image feature information, wherein the first sample fusion feature information is generated by splicing the feature information corresponding to the sample instruction information and sample layer configuration information and the initial embedding feature information, and the first sample image feature information is the feature information of the sample embedded image; The large language model is used to generate second sample fusion feature information based on the sample instruction information; Generate second sample image feature information based on the first sample fusion feature information; Based on the first sample fusion feature information and the second sample fusion feature information, determine the fusion feature loss; Based on the feature information of the first sample image and the feature information of the second sample image, determine the image feature loss; as well as The large language model is trained based on the fusion feature loss and the image feature loss, wherein the operation of training the large language model further includes adjusting the initial embedded feature information.

Citation Information

Patent Citations

  • Systems and methods for layered image generation

    US20250078346A1