Multi-layer image generation method and device and storage medium

By receiving natural language instructions to generate multi-layer images, the problems of high limitations in image generation, poor text rendering effect and high usage threshold in the prior art are solved, and the beautiful and editable multi-layer images are automatically generated, which improves text rendering effect and freedom.

CN120388087AActive Publication Date: 2025-07-29BEIJING YUANSHI TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510429210.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-29
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

In the prior art, image generation has problems such as high limitations, poor text rendering effect, low degree of freedom and high usage threshold. Especially in graphic design, users need a lot of professional knowledge and manpower to intervene, and the generated text is difficult to recognize and edit flexibility is insufficient.

Method used

By receiving instruction information based on natural language, a multi-layer image is generated using a large language model, wherein at least a part of the content of the layer is generated based on the description text, and the editing of the layer is supported, which lowers the threshold for use and improves text rendering effect and freedom.

Benefits of technology

It realizes automatic generation of beautiful multi-layer images, improves text recognition and rendering effects, and allows users to edit images, reducing the difficulty and limitations of use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388087A_ABST
    Figure CN120388087A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-layer image generation method and device and a storage medium. The multi-layer image generation method comprises the steps that instruction information based on a natural language is received, and the instruction information can be used for generating a corresponding image and comprises a description text related to the image content of an image to be generated; and displaying an editable multi-layer image, wherein the content of at least a part of layers in the multi-layer image is generated according to the description text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and particularly to a multi-layer image generation method, apparatus, and storage medium. Background Art

[0002] The success of the scaling law has made many difficult tasks feasible. Graphic design uses vision as a means of communication and expression. By creating and combining text, images, and graphics, visual representations are made to convey information and data. Automating graphic design tasks requires implementing functions such as multi-layer sorting, layer content generation, and layout. It involves multi-modal content understanding and generation. For a difficult task like graphic design, there is currently no complete automated solution.

[0003] One existing solution is to obtain all text, image, and graphic materials through retrieval, manual selection, and manual design. Subsequently, these materials are all predicted for position and size through a layout generation model in the form of images, and then the elements are pieced together to achieve an aesthetic visual effect. However, this approach has some limitations due to all materials needing to be determined in advance. On the one hand, full automation cannot be achieved. Users need to collect or design specific material presentation forms by themselves, including the visual forms of images and graphics, the fonts, colors, effects, and arrangement methods of text paragraphs, etc. This requires a large amount of professional knowledge and human cost intervention. On the other hand, when the elements of graphic design are limited in advance, the possibilities of the graphic design will decrease, resulting in a low upper limit of the aesthetic feeling of the design drawing. For example, if the materials provided by a beginner without professional knowledge are not harmonious, such as chaotic typesetting within a paragraph, messy color selection, etc., then the visual presentation effect of the entire design drawing will not be beautiful.

[0004] Another solution uses a diffusion model to complete the generation of graphic design drawings with text. This solution can follow instructions and complete the creation process of graphic design drawings in a single stage. However, this approach also faces some challenges. 1) Poor text rendering effect. The generated text usually shows incorrect, missing, or extra spelling conditions. Especially in the case of small text, large paragraphs of text, and dense text, the text generated by such methods is almost unrecognizable. 2) Low degree of freedom. The design drawings generated using the diffusion model are presented in the form of pictures and do not have any editing flexibility. 3) High usage threshold. Diffusion models often require users to input as detailed and professional instructions as possible to help the model obtain satisfactory results, and different models may have different instruction preferences, which also greatly reduces the feasibility of users using such methods.

[0005] In view of the technical problems of high limitations in image generation, poor text rendering effect, low freedom, and high usage threshold existing in the above-mentioned prior art, no effective solution has been proposed yet. Summary of the Invention

[0006] Embodiments of the present application provide a multi-layer image generation method, apparatus, and storage medium to at least solve the technical problems of high limitations in image generation, poor text rendering effect, low freedom, and high usage threshold existing in the prior art.

[0007] According to one aspect of the embodiments of the present application, a multi-layer image generation method is provided, including: receiving instruction information based on natural language, where the instruction information can be used to generate a corresponding image and includes descriptive text related to the image content of the image to be generated; and displaying an editable multi-layer image, where the content of at least a part of the layers in the multi-layer image is generated according to the descriptive text.

[0008] According to another aspect of the embodiments of the present application, a multi-layer image generation method is further provided, including: receiving instruction information based on natural language, where the instruction information can be used to generate a corresponding image and includes descriptive text related to the image content of the image to be generated; generating multi-layer image information corresponding to the multi-layer image according to the instruction information through a large language model, where the multi-layer image information includes layer configuration information and embedded images, the layer configuration information indicates the configuration information of each layer of the multi-layer image, and the embedded images are used to be inserted into the image layers of the multi-layer image, and where the layers of the multi-layer image include image layers and non-image layers, the image layers are generated according to the corresponding layer configuration information and embedded images, and the non-image layers can be generated only according to the corresponding layer configuration information; and generating an editable multi-layer image according to the multi-layer image information.

[0009] According to another aspect of the embodiments of the present application, a multi-layer image generation apparatus is further provided, including: a first receiving module, configured to receive instruction information based on natural language, where the instruction information can be used to generate a corresponding image and includes descriptive text related to the image content of the image to be generated; and an image display module, configured to display an editable multi-layer image, where the content of at least a part of the layers in the multi-layer image is generated according to the descriptive text.

[0010] According to another aspect of the embodiments of the present application, there is also provided a multi-layer image generation device, including: a second receiving module, configured to receive instruction information based on natural language, where the instruction information can be used to generate a corresponding image and includes descriptive text related to the image content of the image to be generated; an information generation module, configured to generate multi-layer image information corresponding to the multi-layer image according to the instruction information through a large language model, where the multi-layer image information includes layer configuration information and embedded images, the layer configuration information indicates the configuration information of each layer of the multi-layer image, and the embedded images are used to be inserted into the image layers of the multi-layer image, and where the layers of the multi-layer image include image layers and non-image layers, the image layers are generated according to the corresponding layer configuration information and the embedded images, and the non-image layers can be generated only according to the corresponding layer configuration information; and an image generation module, configured to generate an editable multi-layer image according to the multi-layer image information.

[0011] According to another aspect of the embodiments of the present application, there is also provided a multi-layer image generation device, including: a first processor; and a first memory, connected to the first processor, configured to provide instructions for the first processor to perform the following processing steps: receiving instruction information based on natural language, where the instruction information can be used to generate a corresponding image and includes descriptive text related to the image content of the image to be generated; and displaying an editable multi-layer image, where the content of at least a part of the layers in the multi-layer image is generated according to the descriptive text.

[0012] According to another aspect of the embodiments of the present application, there is also provided a multi-layer image generation device, including: a second processor; and a second memory, connected to the second processor, configured to provide instructions for the second processor to perform the following processing steps: receiving instruction information based on natural language, where the instruction information can be used to generate a corresponding image and includes descriptive text related to the image content of the image to be generated; generating multi-layer image information corresponding to the multi-layer image according to the instruction information through a large language model, where the multi-layer image information includes layer configuration information and embedded images, the layer configuration information indicates the configuration information of each layer of the multi-layer image, and the embedded images are used to be inserted into the image layers of the multi-layer image, and where the layers of the multi-layer image include image layers and non-image layers, the image layers are generated according to the corresponding layer configuration information and the embedded images, and the non-image layers can be generated only according to the corresponding layer configuration information; and generating an editable multi-layer image according to the multi-layer image information.

[0013] According to another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0014] In the embodiments of the present application, the terminal device receives instruction information input by the user for generating an image, where the instruction information can be simple and not highly professional. The user can describe at least a part of the content of the image to be generated by inputting text, thereby reducing the usage threshold. Thus, according to this instruction information, the technical solution automatically generates a beautiful multi-layer image that meets the user's requirements. And the multi-layer image in this technical solution is generated by uniformly generating each layer, so the text layer therein will be rendered at a professional level in consistency with other layers, thereby improving the legibility and rendering effect of the text. And the multi-layer image in this technical solution can be edited for each image of the multi-layer image after generation, thereby improving the freedom and flexibility of the multi-layer image. Furthermore, the technical problems of high limitations in image generation, poor text rendering effect, low freedom, and high usage threshold existing in the prior art are solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0016] Figure 1 is a hardware structure block diagram of a computing device for implementing the method according to Embodiment 1 of the present application;

[0017] Figure 2A is a schematic diagram of a multi-layer image generation system according to Embodiment 1 of the present application;

[0018] Figure 2B is another schematic diagram of a multi-layer image generation system according to Embodiment 1 of the present application;

[0019] Figure 3 is a schematic flowchart of a multi-layer image generation method according to the first aspect of Embodiment 1 of the present application;

[0020] Figure 4A is a schematic diagram of an interface for displaying an image layer according to Embodiment 1 of the present application;

[0021] Figure 4B is a schematic diagram of an interface for displaying a non-image layer according to Embodiment 1 of the present application;

[0022] Figure 5A is the first schematic diagram of a graphic layer according to Embodiment 1 of the present application;

[0023] Figure 5B is the second schematic diagram of a graphic layer according to Embodiment 1 of the present application;

[0024] Figure 5C It is the third schematic diagram of the graphic layer according to Embodiment 1 of the present application;

[0025] Figure 5D It is the fourth schematic diagram of the graphic layer according to Embodiment 1 of the present application;

[0026] Figure 5E It is the fifth schematic diagram of the graphic layer according to Embodiment 1 of the present application;

[0027] Figure 6A It is the first schematic diagram of the text layer according to Embodiment 1 of the present application;

[0028] Figure 6B It is the second schematic diagram of the text layer according to Embodiment 1 of the present application;

[0029] Figure 6C It is the third schematic diagram of the text layer according to Embodiment 1 of the present application;

[0030] Figure 7 It is the schematic diagram of the image layer according to Embodiment 1 of the present application;

[0031] Figure 8 It is the schematic diagram of the multi-layer image according to Embodiment 1 of the present application;

[0032] Figure 9 It is the schematic flow diagram of generating the first embedded image according to Embodiment 1 of the present application;

[0033] Figure 10A It is the schematic flow diagram of generating the fourth layer configuration information according to Embodiment 1 of the present application;

[0034] Figure 10B It is the schematic flow diagram of generating the third text feature information according to Embodiment 1 of the present application;

[0035] Figure 11 It is the schematic flow diagram of generating the second embedded image according to Embodiment 1 of the present application;

[0036] Figure 12 It is the schematic flow diagram of the multi-layer image generation method according to the second aspect of Embodiment 1 of the present application;

[0037] Figure 13 It is the schematic diagram of the multi-layer image generation device according to the first aspect of Embodiment 2 of the present application;

[0038] Figure 14 It is the schematic diagram of the multi-layer image generation device according to the second aspect of Embodiment 2 of the present application;

[0039] Figure 15 It is a schematic diagram of a multi-layer image generation device according to the first aspect of Embodiment 3 of the present application; and

[0040] Figure 16 It is a schematic diagram of a multi-layer image generation device according to the second aspect of Embodiment 3 of the present application. Detailed implementation manners

[0041] In order to enable those skilled in the art to better understand the technical solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work shall fall within the protection scope of the present application.

[0042] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0043] Embodiment 1

[0044] According to this embodiment, a method embodiment of a multi-layer image generation method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.

[0045] The method embodiment provided in this embodiment can be executed in a mobile terminal, a computer terminal, a server or a similar computing device. Figure 1 A hardware structure block diagram of a computing device for implementing a multi-layer image generation method is shown. As Figure 1As shown, the computing device may include one or more processors (the processors may include, but are not limited to, processing devices such as microprocessor MCUs or programmable logic devices FPGAs), a memory for storing data, and a transmission device for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a Universal Serial Bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computing device may further include more or fewer components than those Figure 1 shown, or have a different configuration from that Figure 1 shown.

[0046] It should be noted that the above one or more processors and / or other data processing circuits are generally referred to as "data processing circuits" herein. The data processing circuit may be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computing device. As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistance terminal path connected to an interface).

[0047] The memory may be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the multi-layer image generation method in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the multi-layer image generation method of the above application program. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely located relative to the processor, and these remote memories may be connected to the computing device through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0048] The transmission device is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the computing device. In one instance, the transmission device includes a Network Interface Controller (NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device may be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0049] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables a user to interact with the user interface of the computing device.

[0050] It should be noted here that in some alternative embodiments, the above Figure 1 illustrated computing device can include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware and software elements. It should be pointed out that Figure 1 is only an example of a specific concrete instance and is intended to illustrate the types of components that can exist in the above computing device.

[0051] Figure 2A is a schematic diagram of a terminal device for implementing the multi-layer image generation method according to the present embodiment, which describes a specific application scenario of the present application. Referring to Figure 2A as shown, the terminal device includes: an image processing application program and an AI module. Among them, the image processing application program can be, for example, an image processing application program such as Photoshop, or a graphic design drawing application program.

[0052] Thus, the terminal device receives the instruction information of the user through the image processing application program, and then the terminal device generates multi-layer image information for generating a multi-layer image according to the instruction information of the user through the AI module. Then, the terminal device generates a multi-layer image according to the multi-layer image information through the image processing application program, and displays it through the image processing application program.

[0053] In addition, in another embodiment, Figure 2B is a schematic diagram of a multi-layer image generation system according to the present embodiment, which describes another specific application scenario of the present application. Referring to Figure 2B as shown, the system includes: a terminal device, an image processing platform, and a large model server.

[0054] Among them, the terminal device logs in to the image processing platform using the client and obtains the instruction information of the user, and sends the instruction information to the image processing platform. The client can be the client of the image processing application program or a browser. Then, the image processing platform sends the instruction information to the large model server. Then, the large model server generates multi-layer image information for generating a multi-layer image according to the instruction information through the AI module. Then, the large model server returns the multi-layer image information to the image processing platform. The image processing platform generates a multi-layer image according to the received multi-layer image information and returns the multi-layer image to the terminal device for display through the client.

[0055] In addition, although inFigure 2B The AI module is independently set up on the large model server relative to the image processing platform. However, the AI module can also be deployed on the image processing platform and accessed by the server deployed on the image processing platform (not shown in the figure).

[0056] Under the above operating environment, according to the first aspect of this embodiment, a multi-layer image generation method is provided. This method is implemented by Figure 2A or Figure 2B the terminal device shown in Figure 3 shows a schematic flowchart of this method. Referring to Figure 3 shown, this method includes:

[0057] S302: Receive instruction information based on natural language, where the instruction information can be used to generate a corresponding image and includes descriptive text related to the image content of the image to be generated; and

[0058] S304: Display an editable multi-layer image, where the content of at least some of the layers in the multi-layer image is generated according to the descriptive text.

[0059] Specifically, referring to Figure 2A shown, the terminal device is provided with an image processing application program for generating images. Thus, the user inputs instruction information for generating an image through the terminal device in this image processing application program. The instruction information can be, for example, descriptive text indicating the generation of a poster or indicating the generation of a slide. The instruction information includes descriptive text related to the image content of the image to be generated. For example: Please generate a poster with the words of “CARD”, “preview”, “HAPPY BIRTHDAY”. Then the image processing application program receives the instruction information and calls the AI module to generate multiple layers according to the instruction information, and combines and superimposes the respective layers to generate a multi-layer image that meets the user's requirements. The multi-layer image can be, for example, a poster, a slide, a time image, a commercial image, and an emoji, etc. And the content of the layer is, for example, generated according to the descriptive text related to the image content of the image to be generated in the instruction information input by the user. Or, it is generated according to the descriptive text related to the image content in the instruction information input by the user and the picture input by the user.

[0060] Furthermore, the terminal device displays the multi-layer image through the image processing application program. The user can display each layer of the multi-layer image through the image processing application program, so as to edit each layer.

[0061] In another embodiment, referring to Figure 2BAs shown, the user logs in to the image processing platform through the client of the terminal device. Thus, the user receives the instruction information through the terminal device and sends the instruction information to the image processing platform. The instruction information can be, for example, a description text indicating to generate a poster or indicating to generate a slide. And the instruction information contains the description text related to the image content to be generated. After that, the image processing platform sends the instruction information to the large model server. The large model server receives the instruction information and generates multi-layer image information for generating a multi-layer image through the AI module according to the instruction information. After that, the large model server returns the multi-layer image information to the image processing platform. Thus, the image processing platform generates multiple layers according to the received multi-layer image information, and then combines and overlays the respective layers to generate a multi-layer image that meets the user's requirements. The multi-layer image can be, for example, a poster, a slide, a time image, a commercial image, and an emoji, etc. And the content of the layer is, for example, generated according to the description text related to the image content of the image to be generated in the instruction information input by the user. Or, it is generated according to the description text related to the image content in the instruction information input by the user and the picture input by the user. After that, the image processing platform returns the multi-layer image to the terminal device for display through the client. Thus, the user can open each layer of the multi-layer image through the client of the terminal device to edit each layer.

[0062] As described in the background art, an existing solution is to obtain all text, image, and graphic materials through retrieval, manual selection, and manual design. Subsequently, the positions and sizes of these materials are predicted through a layout generation model in the form of images, and then these elements are pieced together to achieve an aesthetic visual effect. However, this method has some limitations because all materials need to be determined in advance. On the one hand, full automation cannot be achieved. Users need to collect or design specific material presentation forms by themselves, including the visual forms of images and graphics, the fonts, colors, effects, and arrangement methods of text paragraphs, etc. This requires the intervention of a large amount of professional knowledge and labor costs. On the other hand, when the elements of graphic design are limited in advance, the possibilities of graphic design drawings will decrease, resulting in a low upper limit of the aesthetics of the design drawings. For example, if the materials provided by a beginner without professional knowledge are not harmonious, such as chaotic typesetting within a paragraph, messy color selection, etc., then the visual presentation effect of the entire design drawing will not be beautiful. Another solution uses a diffusion model to complete the generation of graphic design drawings with text. This solution can follow instructions and complete the creation process of graphic design drawings in a single-stage form. However, this method also faces some challenges. 1) Poor text rendering effect. The generated text usually shows incorrect, missing, or extra spelling conditions. Especially in the case of small text, large paragraphs of text, and dense text, the text generated by such methods is almost unrecognizable. 2) Low degree of freedom. The design drawings generated using the diffusion model are presented in the form of pictures and do not have any editing flexibility. 3) High usage threshold. Diffusion models often require users to input as detailed and professional instructions as possible to help the model obtain satisfactory results, and different models may have different instruction preferences, which also greatly reduces the feasibility of users using such methods.

[0063] To address the above-mentioned technical problems, through the technical solution of the embodiments of the present application, the terminal device receives the instruction information for generating an image input by the user, where the instruction information can be simple instruction information with low professionalism. The user can describe at least part of the content of the image to be generated by inputting text, thereby reducing the usage threshold. Thus, this technical solution automatically generates a beautiful multi-layer image that meets the user's requirements according to the instruction information. Moreover, the multi-layer image in this technical solution is generated by uniformly generating each layer. Therefore, the text layer in it will be rendered at a professional level consistent with other layers, thereby improving the legibility and rendering effect of the text. And the multi-layer image in this technical solution can be edited for each image of the multi-layer image after generation, thereby improving the degree of freedom and flexibility of the multi-layer image. Furthermore, the technical problems of high limitations in image generation, poor text rendering effect, low degree of freedom, and high usage threshold existing in the prior art are solved.

[0064] Optionally, the operation of displaying a multi-layer image includes any of the following methods: displaying a file link corresponding to the multi-layer image; displaying a file identifier corresponding to the multi-layer image; displaying the layer configuration information of each layer of the multi-layer image; and displaying the image content of each layer of the multi-layer image.

[0065] Specifically, the method for a terminal device to display a multi-layer image includes but is not limited to any of the following methods:

[0066] (1) After generating the multi-layer image, the terminal device can display the multi-layer image in the form of a file link through an image processing application or a client, so that the user can click on the file link to download or view the multi-layer image;

[0067] (2) After generating the multi-layer image, the multi-layer image is displayed on the system interface of the terminal device in the form of a downloaded file identifier, so that the user can click on the file identifier to view the multi-layer image;

[0068] (3) After generating the multi-layer image, the multi-layer image is displayed on the interface of an image processing application or a browser in the form of a file identifier, so that the user can click on the file identifier to download or view the multi-layer image;

[0069] (4) After generating the multi-layer image, the terminal device displays the layer configuration information of each layer of the multi-layer image. The layer configuration information is configuration data for describing the position, size, shape, etc. of each layer. Thus, the terminal device can display the layer configuration information through an image processing application or a client, so that the user can view the layer configuration information of the multi-layer image.

[0070] (5) After generating the multi-layer image, the terminal device displays the layer configuration information of each layer of the multi-layer image through an image processing application or a client. The layer configuration information is configuration data for describing the position, size, shape, etc. of each layer. Thus, the user can directly view and modify the layer configuration information on the image editor through the terminal device.

[0071] (6) After generating the multi-layer image, the terminal device displays the image content of each layer of the multi-layer image through an image processing application or a client, so that the user can directly edit each layer through the image processing application or the client.

[0072] It should be noted that the above methods of displaying a multi-layer image can be displayed in at least one way, and no specific limitation is made here.

[0073] Thus, this technical solution can display multi-layer images in various ways, meeting different viewing needs of users.

[0074] Optionally, the operation of displaying the image content of each layer of the multi-layer image includes: displaying editable pixels corresponding to the corresponding image content in the layer; and displaying editable controls corresponding to the corresponding image content in the layer.

[0075] Specifically, the multi-layer image includes an image layer and a non-image layer. The image layer is an image, and the non-image layer is a control. Refer to Figure 4A As shown, for the image layer, when the image processing application or client displays the image layer, the image content therein is editable pixels. Therefore, operations such as smearing on the image of the image layer or zooming in and out of the image can be performed through the smearing tool in the toolbar.

[0076] Refer to Figure 4B As shown, for the non-image layer, when the image processing application or client displays the non-image layer, the image content therein is an editable control. Thus, the parameters (such as layer configuration information) of the control (such as a graphic or a text box) that makes up the non-image layer can be directly modified. For example, zoom in and out of the graphic in the non-image layer, modify the graphic color, or modify the text content in the text box, etc.

[0077] In addition, other controls can also be added to the image layer and the non-graphic layer, such as adding a text control to add text, etc.

[0078] Thus, this technical solution uses controls to display multi-layer images, thereby flexibly modifying multi-layer images and improving the autonomy of users in modifying multi-layer images.

[0079] Optionally, the method further includes: generating multi-layer image information corresponding to the multi-layer image according to the instruction information, where the multi-layer image information includes layer configuration information and an embedded image. The layer configuration information indicates the configuration information of each layer of the multi-layer image, and the embedded image is used to be inserted into the image layer of the multi-layer image. And the layers of the multi-layer image include an image layer and a non-image layer. The image layer is generated according to the corresponding layer configuration information and the embedded image, and the non-image layer can be generated only according to the corresponding layer configuration information.

[0080] Specifically, refer to Figure 2AAs shown, the image processing application of the terminal device inputs the received instruction information into the AI module. For example, the instruction information is that the user instructs to generate a poster with the words "CARD", "preview", and "HAPPY BIRTHDAY" written on it. Thus, the instruction information can be: Please generate a poster with the words of "CARD", "preview", "HAPPY BIRTHDAY".

[0081] After that, the AI module generates multi-layer image information corresponding to the multi-layer image according to the instruction information, where the multi-layer image information includes layer configuration information and embedded images. The layers of the multi-layer image include a frame layer (frame), a group layer (group), a graphic layer (graphic), a text layer (text), and an image layer (image). Among them, the frame layer (frame), the group layer (group), the graphic layer (graphic), and the text layer (text) are non-image layers. The layer configuration information of the image layer (image) is used to describe information such as the position, size, color, content, and name of the embedded image. And when generating the layer configuration information, the AI module can predict the parameters in the layer configuration information of each layer and the stacking order of the layers according to the instruction information. Among them, the instruction information may specifically specify some parameters of the layer (such as position, size, color, content, and name, etc.), so in the case where the instruction information specifically specifies the parameters, layer configuration information that conforms to the instruction information is generated. In the case where the instruction information does not specify specific parameters, the AI module will make predictions according to the instruction information and generate the corresponding layer configuration information.

[0082] Among them, the layer configuration information can be, for example: <position> 0,0,1200,628< / position> <color> #000000,1.0< / color> <graphic> <position> 0,0,996,445< / position> <color>#610B1C,1.0< / color> <path><move_to>0,0< / move_to><line_to>996,0< / line_to><line_to>996,445< / line_to><line_to>0,445< / line_to><close_path / >< / path> < / graphic> <position> 186,99,827,448< / position> <image_des>foreground,people,RGB< / image_des><image_id>{image_1}< / image_id> <group> <position> ... <graphic> ...< / graphic> < / position> < / group> <position> 86,70,887,440< / position> <image_des>foreground,people,RGB< / image_des><image_id>{image_2}< / image_id> <text> <position> 216,63,770,77< / position> <color>#B9223E,1.0< / color> <font><font_weight>900< / font_weight><font_family>CinzelDecorative Black< / font_family><font_size>57< / font_size>< / font> <text_align>center< / text_align><letter_spacing>0.0< / letter_spacing> <content>HappyValentine's Day< / content> < / text> <group> <position>... <graphic> ...< / graphic> 。

[0083] Wherein:

[0084] " <position> 0,0,1200,628< / position> <color> #000000,1.0< / color> ” is the layer configuration information of the frame layer. According to this layer configuration information, the coordinates of the upper left corner of the frame are (0, 0), the width is "1200", and the height is "628"; the color of the frame is "#000000, 1.0", where "#000000" represents pure black and "1.0" represents complete opacity.

[0085] " <graphic> <position> 0,0,996,445< / position> <color>#610B1C,1.0< / color> <path><move_to>0,0< / move_to><line_to>996,0 < / line_to><line_to>996,445< / line_to><line_to>0,445< / line_to><close_path / >< / path> < / graphic> ” is the layer configuration information of the graphic layer. According to this layer configuration information, a graphic is drawn. Thus, in the layer configuration information, the coordinates of the upper left corner of the graphic are (0, 0), the width is "996", and the height is "445"; the color of the graphic is "#610B1C, 1.0", where "#610B1C" represents dark red and "1.0" represents complete opacity; the graphic is a closed graphic drawn by a path. Thus, the starting point coordinates are (0, 0), draw horizontally to the right to (996, 0), then draw vertically downward to (996, 445), then draw horizontally to the left to (0, 445), and finally close the path, connecting the end point (996, 0) and the starting point (0, 0) to generate a closed rectangular graphic.

[0086] " <text> <position> 216,63,770,77< / position> <color>#B9223E,1.0< / color> <font><font_weight>900< / font_weight><font_family>CinzelDecorative Black< / font_family><font_size>57< / font_size>< / font> <text_align>center< / text_align><letter_spacing>0.0< / letter_spacing> <content>HappyValentine's Day< / content> < / text> ” is the layer configuration information of the text layer. According to this layer configuration information, the coordinates of the upper left corner of the text are (216, 63), the width is "770", and the height is "77"; the color is "#B9223E, 1.0", where "#B9223E" is dark red and "1.0" represents complete opacity; the font is bold, and the bold level is the thickest; the font is the Cinzel Decorative Black variant; the font size is 57; the text position is horizontally centered; the text character spacing is 0; the text content is "Happy Valentine's Day".

[0087] " <group> <position> ... <graphic> ...< / graphic> < / position> < / group> ” is the layer configuration information of the group layer, which is used to combine multiple graphic layers, text layers, and image layers.

[0088] " <position> 186,99,827,448< / position> "<image_des>foreground,people,RGB< / image_des><image_id>{image_1}< / image_id>" is the layer configuration information of the image layer. According to this layer configuration information, the coordinates of the upper left corner of the image are (186, 99), the width is "827", and the height is "448"; the image description is "foreground,people,RGB", where "foreground" indicates that the image is in the foreground, "people" indicates that the image content contains people, and "RGB" indicates that the color space of the image is in RGB format; the name of the image is "image_1".

[0089] " <position> 86,70,887,440< / position> "<image_des>foreground,people,RGB< / image_des><image_id>{image_2}< / image_id>" is the layer configuration information of the image layer. According to this layer configuration information, the coordinates of the upper left corner of the image are (86, 70), the width is "887", and the height is "440"; the image description is "foreground,people,RGB", where "foreground" indicates that the image is in the foreground, "people" indicates that the image content contains people, and "RGB" indicates that the color space of the image is in RGB format; the name of the image is "image_2".

[0090] Referring to the above layer configuration information, the layer configuration information includes the layer configuration information of the frame layer (frame), the group layer (group), the graphic layer (graphic), the text layer (text), and the image layer (image) respectively. Among them, the image layer is generated according to the corresponding layer configuration information and the embedded image, and the non-image layer can be generated only according to the corresponding layer configuration information.

[0091] Therefore, through simple instruction information, this technical solution predicts and generates complex and diverse layers, achieving full automation, without the need for users to set details, reducing the workload of users, and reducing the limitations of image generation.

[0092] Optionally, the operation of displaying a multi-layer image includes: displaying a multi-layer image according to the multi-layer image information.

[0093] Specifically, after the AI module of the terminal device generates multi-layer image information, the multi-layer image information is processed to generate multi-layer image information in a format suitable for processing by the image processing platform. The image processing platform can be, for example, the penpot design platform. Then, the terminal device sends the multi-image layer information after format processing to the image processing platform through an image processing application.

[0094] Further, the image processing platform generates non-image layers (including frame layer (frame), group layer (group), graphic layer (graphic), and text layer (text)) according to the layer configuration information of the frame layer (frame), group layer (group), graphic layer (graphic), and text layer (text) in the multi-image layer information. Among them Figures 5A to 5E shows 5 graphic layers; Figures 6A to 6C shows 3 text layers. And the image processing platform generates an image layer (image) according to the layer configuration information of the image layer (image) and the embedded image. Among them Figure 7 shows the image layer. Thus, the image processing platform displays a multi-layer image generated by combining non-image layers and image layers (refer to Figure 8 shown).

[0095] Therefore, in this technical solution, the AI module generates multi-layer image information in a format suitable for processing by the image processing platform, and the image processing platform generates a multi-layer image composed of non-image layers and image layers, thereby improving the applicability between the AI module and the image processing platform, and thus can more accurately generate a multi-layer image that meets the user's requirements.

[0096] Optionally, in the case where the instruction information indicates generating a multi-layer image using single-modal information based on natural language, the operation of generating multi-layer image information according to the instruction information includes: using a large language model to generate first layer configuration information according to the first text feature information corresponding to the instruction information, where the first layer configuration information includes second layer configuration information and third layer configuration information, where the second layer configuration information is used to indicate the layer configuration information of the first part of non-image layers, the third layer configuration information is used to indicate the part of layer configuration information corresponding to the first image layer, and the third layer configuration information is located after the second layer configuration information; and generating a first embedded image using the large language model and the diffusion model according to the first text feature information and the second text feature information corresponding to the first layer configuration information, where the first embedded image is used to be inserted into the first image layer.

[0097] Specifically, in this embodiment, the instruction information indicates generating a multi-layer image using single-modal information based on natural language, where the single-modal information is used to indicate only text information based on natural language. For example, the instruction information can be: Please generate a poster with the words of "CARD", "preview", "HAPPY BIRTHDAY". Then the image processing application of the terminal device inputs the instruction information into the AI module. After that, the AI module receives the instruction information and generates text feature information (i.e., the first text feature information). The first text feature information is used to indicate the feature vectors corresponding to each character of the instruction information.

[0098] The AI module is pre-set with a large language model. Then the AI module inputs the first text feature information into the large language model, and the large language model processes the first text feature information using a causal mask to generate second text feature information. Then the AI module can generate the first layer configuration information according to the second text feature information through a preset text decoder. The text decoder is used to decode the feature information into layer configuration information (not shown in the figure). The second text feature information is used to indicate the feature vectors corresponding to each character of the first layer configuration information.

[0099] For example, the first layer configuration information is: <position> 0,0,1200,628< / position> <color> #000000,1.0< / color> <graphic> <position> 0,0,996,445< / position> <color>#610B1C,1.0

[0100] < / color> <path><move_to>0,0< / move_to><line_to>996,0< / line_to><line_to>996,445< / line_to><line_to>0,445< / line_to><close_path / >< / path> < / graphic> <position> 186,99,827,448< / position> <image_des>foreground,people,RGB< / image_des><image_id>{image_gen}。

[0101] More specifically, for the process of generating the first layer configuration information, the AI module first generates the second layer configuration information of the first part of non-image layers through the text decoder.

[0102] For example, the second layer configuration information is: <position> 0,0,1200,628< / position> <color> #000000,1.0< / color> <graphic> <position> 0,0,996,445< / position> <color>#610B1C,1.0

[0103] < / color> <path><move_to>0,0< / move_to><line_to>996,0< / line_to><line_to>996,445< / line_to><line_to>0,445< / line_to><close_path / >< / path> < / graphic> 。

[0104] Then, the AI module generates the partial layer configuration information corresponding to the first image layer after the second layer configuration information through the text decoder. For example, it is " <position> 186,99,827,448< / position> <image_des>foreground,people,RGB< / image_des><image_id>{image_gen}”. The first image layer is the first image layer generated when generating a multi-layer image. Thus, the AI module composes the second layer configuration information and the third configuration information into the first layer configuration information.

[0105] The first layer configuration information ends at the {image_gen} mark. The {image_gen} mark of this first layer configuration information is the first generated {image_gen} mark, and the position of the {image_gen} mark is used to write the embedded image name. And only the part of the layer configuration information up to the {image_gen} mark of this image layer has been generated, and there is still another part of the layer configuration information of the image layer that has not been generated.

[0106] Furthermore, the AI module is also pre-set with a diffusion model. Thus, the image processing application generates the first embedded image using the large language model and the diffusion model according to the first text feature information and the second text feature information corresponding to the first layer configuration information. The first embedded image is used to be inserted into the first image layer.

[0107] Thus, in this technical solution, the AI module generates layer configuration information in text form, so that when generating a multi-layer image subsequently, the text layer, etc. can be directly generated using the layer configuration information, avoiding the problem that the text cannot be recognized due to poor text rendering effect. And since the language information itself has an order, the large language model uses a causal mask to simulate the order process of the language information of the instruction information to prevent information leakage.

[0108] Optionally, the operation of generating the first embedded image using the large language model and the diffusion model according to the first text feature information and the second text feature information corresponding to the first layer configuration information includes: splicing the first text feature information, the second text feature information with the preset initial embedded feature information to generate the first fusion feature information, where the initial embedded feature information is the feature information for generating the embedded image; generating the second fusion feature information suitable for image processing by the large language model according to the first fusion feature information; generating the first semantic feature information by the first projection layer according to the second fusion feature information; and generating the first embedded image by the diffusion model according to the first semantic feature information.

[0109] Specifically, refer to Figure 9 As shown, the AI module concatenates the first text feature information, the second text feature information, and the preset initial embedding feature information to generate the first fusion feature information, where the initial embedding feature information is the feature vector for generating the embedded image. After that, the AI module inputs the first fusion feature information into the large language model, and the large language model processes the first fusion feature information using the full attention mask to generate the second fusion feature information suitable for image processing, where the second fusion feature information is the text feature. After that, the AI module inputs the second fusion feature information into the preset first projection layer, and the first projection layer performs projection processing on the second fusion information to generate the corresponding semantic feature information (i.e., the first semantic feature information).

[0110] Further, the AI module inputs the first semantic feature information into the diffusion model, where the diffusion model includes the first linear layer, the visual encoder, and the second linear layer. The visual encoder can be, for example, the pre-trained CLIP visual encoder ViT-L / 14.

[0111] Thus, the diffusion model inputs the received first semantic feature information into the first linear layer, and the diffusion model processes the first semantic feature information through the first linear layer to generate the first output feature. After that, the diffusion model inputs the first output feature into the visual encoder, and the visual encoder encodes the first output feature to generate the second output feature. After that, the diffusion model inputs the second output feature into the second linear layer, and the second linear layer processes the second output feature to generate the first embedded image.

[0112] Thus, this technical solution performs semantic processing on the text feature information and the image feature information through the large language model, so as to highlight the key features, predict the parameters of multiple layers, and generate the corresponding embedded image through the diffusion model. This technical solution combines the large language model and the diffusion model, which not only avoids the user using complex instruction information, but also can generate image content that meets the user's requirements according to the user's description.

[0113] Optionally, the method further includes: using the large language model to generate the fourth layer configuration information according to the first image feature information corresponding to the first embedded image and the first fusion feature information, where the fourth layer configuration information includes the fifth layer configuration information and the sixth layer configuration information, and the fifth layer configuration information is used to indicate the layer configuration information of the second part of the non-image layer, and the sixth layer configuration information is used to indicate the partial layer configuration information corresponding to the second image layer after the fifth layer configuration information; and generating the second embedded image according to the first fusion feature information, the first image feature information, and the third text feature information corresponding to the fourth layer configuration information, using the large language model and the diffusion model, where the second embedded image is used to be inserted into the second image layer.

[0114] Specifically, after the AI module generates the first layer configuration information and the first embedded image, it generates image feature information from the first embedded image (i.e., the first image feature information). Then the AI module inputs the first image feature information and the first fusion feature information into the large language model. The large language model processes the first image feature information and the first fusion feature information, outputs the corresponding feature information, and inputs this feature information into the text encoder. Then the text decoder generates an image name tag corresponding to the first embedded image based on this feature information, such as {image_1}. Then, in the first layer configuration information, {image_gen} is replaced with {image_1}, and the layer configuration information of another part of the image layer after {image_1} is generated "< / image_id>”. The name of the image is generated according to the generation order of the image layers. For example, the name of the image in the first generated image layer is automatically named "image_1", the name of the image in the second generated image layer is automatically named "image_2", and the name of the image in the n generated image layers is automatically named "image_n".

[0115] Furthermore, the large language model further generates third text feature information. Then the AI module inputs the third text feature information into the text decoder, and the text decoder generates the fourth layer configuration information based on the third text feature information.

[0116] For example, the fourth layer configuration information is as follows: <group> <position> ... <graphic> ...< / graphic> < / position> < / group> <position> 86,70,887,440< / position> <image_des>...< / image_des><image_id>{image_gen}.

[0117] More specifically, for the process of generating the fourth layer configuration information, the AI module first generates the fifth layer configuration information after the layer configuration information of the image layer through the text decoder. The fifth layer configuration information is used to indicate the layer configuration information of the second part of the non-image layer. For example, it can be <group> <position> ... <graphic> ...< / graphic> < / position> < / group> .

[0118] Furthermore, the text decoder generates the sixth layer configuration information after the fifth layer configuration information. The sixth layer configuration information is used to indicate the partial layer configuration information corresponding to the second image layer. For example, it is " <position> 186,99,827,448< / position> <image_des>foreground, people, RGB< / image_des><image_id>{image_gen}”. The second image layer is the second image layer generated when generating the multi-layer image. Thus, the AI module composes the fifth layer configuration information and the sixth configuration information into the fourth layer configuration information.

[0119] Among them, the fourth layer configuration information ends at the {image_gen} mark. The {image_gen} mark in the fourth layer configuration information is the {image_gen} mark generated for the second time, and the position of the {image_gen} mark is used as the mark for generating the image name. And only the partial layer configuration information up to the {image_gen} mark of the layer configuration information of this image layer is generated, and there is still another part of the layer configuration information of the image layer that is not generated.

[0120] It should be noted that the fifth layer configuration information may not be generated either. That is, after generating the other part of the layer configuration information corresponding to the first image layer, directly generate the layer configuration information of the next image layer (i.e., the second image layer) (i.e., the sixth layer configuration information), without generating other non-image layers before the sixth layer configuration information, and it can be generated according to the actual situation here, without specific limitation.

[0121] Furthermore, the AI module generates the second embedded image according to the first fusion feature information, the first image feature information, and the third text feature information corresponding to the fourth layer configuration information, using the large language model and the diffusion model, where the second embedded image is used to be inserted into the second image layer.

[0122] Furthermore, after the AI module generates the fourth layer configuration information and the second embedded image, it can continue to generate other layer configuration information and embedded images until there is no need to generate.

[0123] Thus, this technical solution can generate multi-image layer information without limitation until the user requirements are met, making the generated multi-layer image more detailed and accurate.

[0124] Optionally, the operation of generating the fourth layer configuration information using the large language model according to the first image feature information corresponding to the first embedded image and the first fusion feature information includes: generating the second image feature information from the first embedded image through the first image encoder; generating the first image feature information suitable for processing by the large language model from the second image feature information through the second projection layer; and generating the fourth layer configuration information using the large language model according to the first image feature information and the first fusion feature information.

[0125] Specifically, refer to Figure 10A As shown, the AI module inputs the first embedded image into a preset first image encoder, and the first image encoder encodes the first embedded image to generate second image feature information.

[0126] Further, the AI module inputs the second image feature information into a preset second projection layer, and the second projection layer projects the second image feature information to generate first image feature information suitable for processing by a large language model.

[0127] Further, referring to Figure 10A and Figure 10B As shown, the AI module inputs the first image feature information and the first fusion feature information into a large language model, and the large language model generates feature information based on the first image feature information and the first fusion feature information. Then the AI module inputs the feature information into a text encoder, and the text encoder generates fourth layer configuration information according to the received feature information. The first fusion feature information is formed by splicing the first text feature information, the second text feature information, and the initial embedded feature information.

[0128] Thus, through the operations of first encoding and then mapping the embedded image, this technical solution generates image feature information suitable for processing by a large language model from the embedded image, so that the large language model can better process the image feature information, and thus the large language model can adapt to various image feature information, improving the adaptability of the large language model.

[0129] Optionally, the operation of generating a second embedded image using the large language model and the diffusion model according to the first fusion feature information, the first image feature information, and the third text feature information corresponding to the fourth layer configuration information includes: splicing the first fusion feature information, the first image feature information, and the third text feature information to generate third fusion feature information; generating fourth fusion information suitable for image processing by the large language model according to the third fusion feature information; generating second semantic feature information by the first projection layer according to the fourth fusion information; and generating a second embedded image by the diffusion model according to the second semantic feature information.

[0130] Specifically, referring to Figure 11 As shown, after generating the third text feature information, the AI module splices the first fusion feature information, the first image feature information, the third text feature information, and the initial embedded feature information to generate third fusion feature information.

[0131] Further, the AI module inputs the third fusion feature information into the large language model, and the large language model processes the third fusion feature information using a full attention mask to generate fourth fusion feature information suitable for image processing.

[0132] Further, the AI module inputs the fourth fused feature information into the first projection layer, and the first projection layer performs projection processing on the fourth fused feature information to generate second semantic feature information.

[0133] Further, the AI module inputs the second semantic feature information into the diffusion model, and the diffusion model generates a second embedded image according to the second semantic feature information.

[0134] Thus, this technical solution processes the third fused feature information through the full attention mask, enabling the model to not only focus on the previous content but also capture spatial information related to generating and understanding the current image when creating and understanding the embedded image.

[0135] Optionally, the method further includes: using a large language model to generate seventh layer configuration information according to the first image feature information corresponding to the first embedded image and the first fused feature information, where the seventh layer configuration information is used to indicate the layer configuration information of the remaining non-image layers except the first part of non-image layers.

[0136] Specifically, after the AI module generates the first layer configuration information and the first embedded image, the first image feature information corresponding to the first embedded image is input into the large language model. The large language model processes the first image feature information and the first fused feature information to generate corresponding feature information. Then the AI module inputs this feature information into the text decoder, and thus the text decoder generates an image name tag corresponding to the first embedded image, such as {image_1}. Then, {image_gen} in the first layer configuration information is replaced with {image_1}, and the layer configuration information of another part of the image layers after {image_1} is generated "< / image_id>".

[0137] Further, in the case where no other embedded images and corresponding layer configuration information need to be generated, the large language model generates corresponding feature information according to the first image feature information corresponding to the first embedded image and the first fused feature information. Then the AI module inputs this feature information into the text decoder, and the text decoder generates the seventh layer configuration information according to this feature information. The seventh layer configuration information is used to indicate the layer configuration information of the remaining non-image layers except the first part of non-image layers, such as including the end tag of the frame layer, the text layer, the grouped layer, etc.

[0138] For example, the seventh layer configuration information is: <text> <position> 216,63,770,77< / position> <color>#B9223E,1.0< / color> <font><font_weight>900< / font_weight><font_family>CinzelDecorative Black< / font_family><font_size>57< / font_size>< / font> <text_align>center< / text_align><letter_spacing>0.0< / letter_spacing> <content>HappyValentine's Day< / content> < / text> <group> <position>... <graphic> ...< / graphic> 。

[0139] Therefore, when the present technical solution does not need to generate an image, the AI module can directly generate the layer configuration information of the remaining non-image layers, improving the efficiency of generating the layer configuration information.

[0140] Optionally, when the instruction information indicates generating a multi-layer image using an existing embedded image, the operation of generating multi-layer image information according to the instruction information includes: generating fifth image feature information by a second image encoder according to a preset third embedded image; generating sixth image feature information suitable for processing by a large language model by a third projection layer according to the fifth image feature information; and generating seventh layer configuration information for describing the image layer and eighth layer configuration information for describing the non-image layer by the large language model according to the fourth text feature information corresponding to the instruction information and the sixth image feature information.

[0141] Specifically, in this embodiment, the instruction information indicates generating a multi-layer image using an existing embedded image. The existing embedded image is an image preset for inserting an image layer and does not need to be generated in real time by a diffusion model.

[0142] Thus, the image processing application program of the terminal device receives the instruction information for input and the embedded image (i.e., the third embedded image), and then inputs the instruction information and the third embedded image into the AI module. Then the AI module inputs the third embedded image into the second image encoder, and the second image encoder encodes the third embedded image to generate fifth image feature information.

[0143] Furthermore, the AI module inputs the fifth image feature information into the third projection layer, and the third projection layer projects the fifth image feature information to generate sixth image feature information suitable for processing by the large language model.

[0144] Furthermore, the AI module generates corresponding fourth text feature information from the instruction information, and then inputs the fourth text feature information and the sixth image feature information into the large language model. The large language model generates corresponding feature information according to the fourth text feature information and the sixth image feature information. Then the AI module inputs the feature information into the text decoder, and the text decoder generates seventh layer configuration information for describing the image layer and eighth layer configuration information for describing the non-image layer according to the feature information.

[0145] For example, the generated all layer configuration information is: <position> 0,0,1200,628< / position> <color> #000000,1.0< / color> <graphic> <position> 0,0,996,445< / position> <color>#610B1C,1.0

[0146] < / color> <path><move_to>0,0< / move_to><line_to>996,0< / line_to><line_to>996,445< / line_to><line_to>0,445< / line_to><close_path / >< / path> < / graphic> <position> 186,99,827,448< / position> <image_des>foreground,people,RGB< / image_des><image_id>{image_1}< / image_id> <text> <position> 216,63,770,77< / position> <color>#B9223E,1.0< / color> <font><font_weight>900< / font_weight><font_family>CinzelDecorative Black< / font_family><font_size>57< / font_size>< / font> <text_align>center< / text_align><letter_spacing>0.0< / letter_spacing> <content>HappyValentine's Day< / content> < / text> <group> <position>... <graphic> ...< / graphic> 。

[0147] Among them, the configuration information of the seventh layer is as follows: <position> 186,99,827,448< / position> <image_des>foreground,people,RGB< / image_des><image_id>{image_1}< / image_id>. The configuration information of the layers other than the configuration information of the fourth layer in all the layer configuration information is the configuration information of the eighth layer used to describe the non-image layer.

[0148] Therefore, in this technical solution, the large language model processes the input information, which includes both instruction information and embedded images, as multi-modal information, so that the layer configuration information can be quickly generated according to the instruction information and the embedded images, improving the efficiency of generating the layer configuration information.

[0149] Optionally, the large language model is trained through the following steps: collecting sample instruction information, first sample fusion feature information, and first sample image feature information, where the first sample fusion feature information is generated by splicing the feature information corresponding to the sample instruction information, sample layer configuration information, and initial embedding feature information, and the first sample image feature information is the feature information of the sample embedded image; using the large language model to generate second sample fusion feature information according to the sample instruction information; generating second sample image feature information according to the first sample fusion feature information; determining the fusion feature loss according to the first sample fusion feature information and the second sample fusion feature information; determining the image feature loss according to the first sample image feature information and the second sample image feature information; and training the large language model according to the fusion feature loss and the image feature loss, where the operation of training the large language model further includes: adjusting the initial embedding feature information.

[0150] Specifically, the image processing application program is also provided with a training module, so that the training module of the image processing application program collects sample data, and then uses the sample data to train the parameter values of the large language model, diffusion model, first projection layer, second projection layer, and initial embedding feature information.

[0151] More specifically, the sample data collected by the training module of the image processing application program includes: sample instruction information, first sample fusion feature information, and first sample image feature information. Among them, the first sample fusion feature information is generated by splicing the feature information corresponding to the sample instruction information and the sample layer configuration information, and the initial embedding feature information; the first sample image feature information is the feature information of the sample embedded image.

[0152] For the sample instruction information: The training module performs data augmentation and semantic augmentation operations on the layer configuration information of the text layer in the sample instruction information. For example, the text content of the text layer is "Beginning of Spring" and "Spring is the season when all things come back to life". The random data augmentation used is to replace each character with random text, such as "Zhang San" and "Zhang San, Li Si, Wang Wu, 1234", with random content but the same string length, and all other attributes such as font, font size, and color remain unchanged. The semantic augmentation is to generate contextually consistent text through Qwen2, while introducing unique content and retaining semantic relevance. For example, it is changed to "Beginning of Winter" and "Winter is the season of tranquility and peace". Using these data augmentations is to expand the sample quantity and diversity of the training data and make the model generate better results.

[0153] In the case where the instruction information indicates the generation of a multi-layer image using natural language-based unimodal information, the training steps for the large language model are as follows:

[0154] (1) The training module collects sample instruction information, first sample fusion feature information, and first sample image feature information, where the sample instruction information x = {x1, x2, …, x N}, x i represents the i-th character in the sample instruction information, i = 1~N. The first sample image feature information v = {v1, v2, …, v M}, v j represents the j-th feature item in the first sample image feature information, j = 1~M. The first sample fusion feature information c = {c1, c2, …, c Q}, c k represents the k-th feature item in the first sample fusion feature information, k = 1~Q.

[0155] (2) The training module inputs the text feature information corresponding to the sample instruction information into the large language model, processes the sample instruction information through the large language model, outputs the first sample layer configuration information, and then concatenates the feature information corresponding to the sample instruction information and the first sample layer configuration information and the sample initial embedding feature information to obtain the second sample fusion feature information. The second sample fusion feature information is the predicted value.

[0156] (3) The training module, according to the second sample fusion feature information, sequentially passes through the large language model, the first projection layer, and the diffusion model to output the sample image. And the training module, according to the sample image, passes through the first encoder and the second projection layer to output the second sample image feature information. The second sample image feature information is the predicted value.

[0157] (4) The training module inputs the first sample image feature information, the text feature information of the sample instruction information, and the second sample fusion feature information into the large language model to generate multi-layer image information.

[0158] (5) The training module calculates the text loss between the first sample fusion feature information and the second sample fusion feature information

[0159] (6) The training module calculates the image loss between the first sample image feature information and the second sample image feature information

[0160] (7) The training module determines the model loss according to the text loss and the image loss

[0161] (8) The training module updates the parameter values of the first projection layer, the second projection layer, the large language model, the diffusion model, and the initial embedding feature information through gradient backpropagation according to the model loss.

[0162] In the case where the instruction information indicates generating multi-layer images using the existing embedded images, the training steps for the large language model are as follows:

[0163] (1) The training module collects the sample instruction information, the sample embedded image, and the second sample layer configuration information;

[0164] (2) The training module processes the sample embedded image through the second image encoder and the third projection layer to generate sample image feature information;

[0165] (3) The training module outputs the third sample layer configuration information through the large language model according to the text feature information corresponding to the sample instruction information and the sample image feature information. The third sample layer configuration information is a predicted value.

[0166] (4) The training module calculates the text loss between the second sample layer configuration information and the third sample layer configuration information, and trains the large language model, the second image encoder, and the third projection layer according to this text loss.

[0167] Thus, this technical solution trains the models in the AI module based on single-modal information and multi-modal information in different ways, so that the AI module can be applicable to both single-modal information and multi-modal information, improving the diversity of multi-layer image generation.

[0168] It should be noted that the large language models for processing single-modal information and multi-modal information can be two different large language models or the same large language model, and no specific limitation is made here.

[0169] In addition, if the embedded image is a solid-color background image, it will affect the training effect of the diffusion model. Therefore, when training the diffusion model and the large language model using the embedded image with a solid-color background, the training module will use the k-means clustering method to select the top-k hues from the image to capture the dominant hues of the image. Based on the Potrace algorithm, the training module converts each dominant color area into a vector path, and the paths of each dominant hue are combined together to create a unified SVG file to capture the color structure. The SVG file is a graphic layer. That is, the embedded image with a solid-color background is used to generate the graphics in the graphic layer, and then the large language model is trained through this graphic. This process realizes the scalable and compact expression of simple design elements. Before the embedded image is converted into an SVG file, the training module uses Qwen2-VL to generate short descriptive labels for it, so as to provide context information about the design elements and enhance controllability.

[0170] Thus, according to the first aspect of this embodiment, the terminal device receives the instruction information for generating an image input by the user, and the instruction information can be simple instruction information with low professionalism. The user can describe at least part of the content of the image to be generated by inputting text, thereby lowering the usage threshold. Therefore, this technical solution automatically generates a beautiful multi-layer image that meets the user's requirements according to the instruction information. Moreover, the multi-layer image in this technical solution is generated by uniformly generating each layer, so the text layer in it will be rendered at a professional level consistent with other layers, thereby improving the legibility and rendering effect of the text. And the multi-layer image in this technical solution can be edited for each image of the multi-layer image after generation, thereby improving the freedom and flexibility of the multi-layer image. Furthermore, it solves the technical problems of high limitations in image generation, poor text rendering effect, low freedom, and high usage threshold existing in the prior art.

[0171] In addition, according to the second aspect of this embodiment, a method for generating a multi-layer image is provided. Figure 12 The flowchart of the method is shown. Refer to Figure 12 As shown, the method includes:

[0172] S1202: Receive instruction information based on natural language, where the instruction information can be used to generate a corresponding image and includes descriptive text related to the image content of the image to be generated;

[0173] S1204: Generate multi-layer image information corresponding to a multi-layer image based on instruction information through a large language model, where the multi-layer image information includes layer configuration information and embedded images. The layer configuration information indicates the configuration information of each layer of the multi-layer image, and the embedded images are used to be inserted into the image layers of the multi-layer image. And where the layers of the multi-layer image include image layers and non-image layers. The image layers are generated based on the corresponding layer configuration information and the embedded images, and the non-image layers can be generated only based on the corresponding layer configuration information; and

[0174] S1206: Generate an editable multi-layer image based on the multi-layer image information.

[0175] Specifically, as shown in Figure 2B The user logs in to the image processing platform through the client of the terminal device. Thus, the user receives the instruction information through the terminal device and sends the instruction information to the image processing platform. The instruction information can be, for example, a description text indicating the generation of a poster or indicating the generation of a slide. And the instruction information contains the description text related to the image content to be generated. Then the image processing platform sends the instruction information to the large model server. The large model server receives the instruction information and generates multi-layer image information for generating a multi-layer image through the AI module according to the instruction information. Then the large model server returns the multi-layer image information to the image processing platform. Thus, the image processing platform generates multiple layers according to the received multi-layer image information, and then combines and overlays each layer to generate a multi-layer image that meets the user's requirements. The multi-layer image can be, for example, a poster, a slide, a time image, a commercial image, and an emoji, etc. And the content of the layer is, for example, generated according to the description text related to the image content of the image to be generated in the instruction information input by the user. Or, it is generated according to the description text related to the image content in the instruction information input by the user and the picture input by the user. Then the image processing platform returns the multi-layer image to the terminal device for display through the client. Thus, the user can open each layer of the multi-layer image through the client of the terminal device to edit each layer.

[0176] More specifically, for example, the instruction information is that the user instructs to generate a poster, and the words "CARD", "preview", and "HAPPY BIRTHDAY" are written on the poster. Thus, the instruction information can be: Please generate a poster with the words of "CARD", "preview", "HAPPY BIRTHDAY".

[0177] After that, the AI module generates multi-layer image information corresponding to the multi-layer image according to the instruction information, where the multi-layer image information includes layer configuration information and embedded images. The layers of the multi-layer image include a frame layer (frame), a group layer (group), a graphic layer (graphic), a text layer (text), and an image layer (image). Among them, the frame layer (frame), the group layer (group), the graphic layer (graphic), and the text layer (text) are non-image layers. The layer configuration information of the image layer (image) is used to describe information such as the position, size, color, content, and name of the embedded image. And when generating the layer configuration information, the AI module can predict the parameters in the layer configuration information of each layer and the stacking order of the layers according to the instruction information. The instruction information may specifically specify some parameters of the layer (such as position, size, color, content, and name, etc.), so that when the instruction information specifically specifies the parameters, the layer configuration information that conforms to the instruction information is generated. When the instruction information does not specify specific parameters, the AI module will make predictions according to the instruction information and generate the corresponding layer configuration information.

[0178] Among them, the layer configuration information can be, for example: <position> 0,0,1200,628< / position> <color> #000000,1.0< / color> <graphic> <position> 0,0,996,445< / position> <color>#610B1C,1.0

[0179] < / color> <path><move_to>0,0< / move_to><line_to>996,0< / line_to><line_to>996,445< / line_to><line_to>0,445< / line_to><close_path / >< / path> < / graphic> <position> 186,99,827,448< / position> <image_des>foreground,people,RGB< / image_des><image_id>{image_1}< / image_id> <group> <position> ... <graphic> ...< / graphic> < / position> < / group> <position> 86,70,887,440< / position> <image_des>foreground,people,RGB< / image_des><image_id>{image_2}< / image_id> <text> <position> 216,63,770,77< / position> <color>#B9223E,1.0< / color> <font><font_weight>900< / font_weight><font_family>CinzelDecorative Black< / font_family><font_size>57< / font_size>< / font> <text_align>center< / text_align><letter_spacing>0.0< / letter_spacing> <content>HappyValentine's Day< / content> < / text> <group> <position>... <graphic> ...< / graphic> 。

[0180] Wherein:

[0181] " <position> 0,0,1200,628< / position> <color> #000000,1.0< / color> " is the layer configuration information of the frame layer. According to this layer configuration information, the coordinates of the upper left corner of the frame are (0, 0), and the coordinates of the lower right corner are (1200, 628); the color of the frame is "#000000, 1.0", where "#000000" represents pure black and "1.0" represents complete opacity.

[0182] " <graphic> <position> 0,0,996,445< / position> <color>#610B1C,1.0< / color> <path><move_to>0,0< / move_to><line_to>996,0 < / line_to><line_to>996,445< / line_to><line_to>0,445< / line_to><close_path / >< / path> < / graphic> " is the layer configuration information of the graphic layer. According to this layer configuration information, a graphic is drawn. Thus, in the layer configuration information, the coordinates of the upper left corner of the graphic are (0, 0), and the coordinates of the lower right corner are (996, 445); the color of the graphic is "#610B1C, 1.0", where "#610B1C" represents dark red and "1.0" represents complete opacity; the graphic is a closed graphic drawn through a path. Thus, the starting point coordinates are (0, 0), horizontally drawn to the right to (996, 0), then vertically drawn down to (996, 445), then horizontally drawn to the left to (0, 445), and finally the path is closed by connecting the end point (996, 0) and the starting point (0, 0) to generate a closed rectangular graphic.

[0183] " <text> <position> 216,63,770,77< / position> <color>#B9223E,1.0< / color> <font><font_weight>900< / font_weight><font_family>CinzelDecorative Black< / font_family><font_size>57< / font_size>< / font> <text_align>center< / text_align><letter_spacing>0.0< / letter_spacing> <content>HappyValentine's Day< / content> < / text> " is the layer configuration information of the text layer. According to this layer configuration information, the coordinates of the upper left corner of the text are (216, 63), and the coordinates of the lower right corner are (770, 77); the color is "#B9223E, 1.0", where "#B9223E" is dark red and "1.0" represents complete opacity; the font is bold, and the bold level is the thickest; the font is the Cinzel Decorative Black variant; the font size is 57; the text is horizontally centered; the text character spacing is 0; the text content is "Happy Valentine's Day".

[0184] " <group> <position> ... <graphic> ...< / graphic> < / position> < / group> " is the layer configuration information of the group layer, which is used to combine multiple graphic layers, text layers, and image layers.

[0185] " <position> 186,99,827,448< / position> "<image_des>foreground,people,RGB< / image_des><image_id>{image_1}< / image_id>" is the layer configuration information of the image layer. According to this layer configuration information, the coordinates of the upper left corner of the image are (186, 99), and the coordinates of the lower right corner are (827, 448); the image description is "foreground,people,RGB", where "foreground" indicates that the image is in the foreground, "people" indicates that the image content contains people, and "RGB" indicates that the color space of the image is in RGB format; the name of the image is "image_1".

[0186] " <position> 86,70,887,440< / position> "<image_des>foreground,people,RGB< / image_des><image_id>{image_2}< / image_id>" is the layer configuration information of the image layer. According to this layer configuration information, the coordinates of the upper left corner of the image are (86, 70), and the coordinates of the lower right corner are (887, 440); the image description is "foreground,people,RGB", where "foreground" indicates that the image is in the foreground, "people" indicates that the image content contains people, and "RGB" indicates that the color space of the image is in RGB format; the name of the image is "image_2".

[0187] Referring to the above layer configuration information, the layer configuration information includes the layer configuration information of the frame layer (frame), group layer (group), graphic layer (graphic), text layer (text) and image layer (image) respectively. Among them, the image layer is generated according to the corresponding layer configuration information and the embedded image, and the non-image layer can be generated only according to the corresponding layer configuration information.

[0188] As described in the background art, an existing solution is to obtain all text, image, and graphic materials through retrieval, manual selection, and manual design. Subsequently, these materials are all predicted for their positions and sizes through a layout generation model in the form of images, and then these elements are pieced together to achieve an aesthetic visual effect. However, this method has some limitations due to the need for all materials to be determined in advance. On the one hand, full automation cannot be achieved. Users need to collect or design specific material presentation forms by themselves, including the visual forms of images and graphics, the fonts, colors, effects, and arrangement methods of text paragraphs, etc. This requires the intervention of a large amount of professional knowledge and labor costs. On the other hand, when the elements of graphic design are limited in advance, the possibilities of graphic design drawings will be reduced, resulting in a low upper limit of the aesthetic feeling of the design drawings. For example, if the materials provided by a beginner without professional knowledge are not harmonious, such as chaotic typesetting within a paragraph, messy color selection, etc., then the visual presentation effect of the entire design drawing will not be beautiful. Another solution uses a diffusion model to complete the generation of graphic design drawings with text. This solution can follow instructions and complete the creation process of graphic design drawings in a single-stage form. However, this method also faces some challenges. 1) Poor text rendering effect. The generated text usually shows incorrect, missing, or extra spelling conditions. Especially in the case of small text, large paragraphs of text, and dense text, the text generated by such methods is almost unrecognizable. 2) Low degree of freedom. The design drawings generated using the diffusion model are presented in the form of pictures and do not have any editing flexibility. 3) High usage threshold. Diffusion models often require users to input as detailed and professional instructions as possible to help the model obtain satisfactory results, and different models may have different instruction preferences, which also greatly reduces the feasibility of users using such methods.

[0189] To address the above-mentioned technical problems, through the technical solutions of the embodiments of the present application, the terminal device receives the instruction information input by the user, where the instruction information can be simple and not highly professional instruction information, reducing the usage threshold. Thus, the terminal device automatically generates a multi-layer image that is beautiful and meets the user's requirements according to the instruction information. Moreover, the multi-layer image in this technical solution is generated by uniformly generating each layer. Therefore, the text layer in it will be subjected to consistent professional-level rendering with other layers, thereby improving the legibility and rendering effect of the text. And the multi-layer image in this technical solution can be edited for each image of the multi-layer image after generation, thereby improving the degree of freedom and flexibility of the multi-layer image. Furthermore, the technical problems of high limitations in image generation, poor text rendering effect, low degree of freedom, and high usage threshold existing in the prior art are solved.

[0190] Optionally, the operation of displaying a multi-layer image includes any of the following methods: displaying a file link corresponding to the multi-layer image; displaying a file identifier corresponding to the multi-layer image; displaying the layer configuration information of each layer of the multi-layer image; and displaying the image content of each layer of the multi-layer image.

[0191] Optionally, the operation of displaying the image content of each layer of the multi-layer image includes: displaying editable pixels corresponding to the respective image content on the layer; and displaying editable controls corresponding to the respective image content on the layer.

[0192] Optionally, the operation of displaying a multi-layer image includes: displaying the multi-layer image according to the multi-layer image information.

[0193] Optionally, in the case where the instruction information indicates generating a multi-layer image using single-modal information based on natural language, the operation of generating multi-layer image information according to the instruction information includes: using a large language model to generate first layer configuration information according to first text feature information corresponding to the instruction information, where the first layer configuration information includes second layer configuration information and third layer configuration information, where the second layer configuration information is used to indicate the layer configuration information of the first part of non-image layers, the third layer configuration information is used to indicate the partial layer configuration information corresponding to the first image layer, and the third layer configuration information is located after the second layer configuration information; and generating a first embedded image using the large language model and a diffusion model according to the first text feature information and second text feature information corresponding to the first layer configuration information, where the first embedded image is used to be inserted into the first image layer.

[0194] Optionally, the operation of generating a first embedded image using the large language model and a diffusion model according to the first text feature information and second text feature information corresponding to the first layer configuration information includes: splicing the first text feature information, the second text feature information with preset initial embedded feature information to generate first fusion feature information, where the initial embedded feature information is feature information for generating an embedded image; generating second fusion feature information suitable for image processing according to the first fusion feature information through the large language model; generating first semantic feature information according to the second fusion feature information through a first projection layer; and generating a first embedded image according to the first semantic feature information through the diffusion model.

[0195] Optionally, the method further includes: using a large language model to generate fourth layer configuration information based on first image feature information corresponding to the first embedded image and first fusion feature information, where the fourth layer configuration information includes fifth layer configuration information and sixth layer configuration information, and where the fifth layer configuration information is used to indicate the layer configuration information of the second part of non-image layers, and the sixth layer configuration information is used to indicate partial layer configuration information corresponding to the second image layer located after the fifth layer configuration information; and generating a second embedded image using the large language model and a diffusion model based on the first fusion feature information, the first image feature information, and third text feature information corresponding to the fourth layer configuration information, where the second embedded image is used to be inserted into the second image layer.

[0196] Optionally, the operation of using a large language model to generate fourth layer configuration information based on first image feature information corresponding to the first embedded image and first fusion feature information includes: generating second image feature information from the first embedded image through a first image encoder; generating first image feature information suitable for processing by the large language model from the second image feature information through a second projection layer; and using the large language model to generate fourth layer configuration information based on the first image feature information and the first fusion feature information.

[0197] Optionally, the operation of generating a second embedded image using the large language model and a diffusion model based on the first fusion feature information, the first image feature information, and third text feature information corresponding to the fourth layer configuration information includes: concatenating the first fusion feature information, the first image feature information, and the third text feature information to generate third fusion feature information; generating fourth fusion information suitable for image processing from the third fusion feature information through the large language model; generating second semantic feature information from the fourth fusion information through a first projection layer; and generating a second embedded image from the second semantic feature information through the diffusion model.

[0198] Optionally, the method further includes: using a large language model to generate seventh layer configuration information based on first image feature information corresponding to the first embedded image and first fusion feature information, where the seventh layer configuration information is used to indicate the layer configuration information of the remaining non-image layers except the first part of non-image layers.

[0199] Optionally, when the instruction information indicates generating a multi-layer image by using an existing embedded image, the operation of generating multi-layer image information according to the instruction information includes: generating fifth image feature information by a second image encoder according to a preset third embedded image; generating sixth image feature information suitable for processing by a large language model by a third projection layer according to the fifth image feature information; and generating seventh layer configuration information for describing an image layer and eighth layer configuration information for describing a non-image layer by the large language model according to fourth text feature information corresponding to the instruction information and the sixth image feature information.

[0200] Optionally, the large language model is trained by the following steps: collecting sample instruction information, first sample fusion feature information, and first sample image feature information, where the first sample fusion feature information is generated by splicing feature information corresponding to the sample instruction information and sample layer configuration information, and initial embedded feature information, and the first sample image feature information is the feature information of a sample embedded image; generating second sample fusion feature information by the large language model according to the sample instruction information; generating second sample image feature information according to the first sample fusion feature information; determining a fusion feature loss according to the first sample fusion feature information and the second sample fusion feature information; determining an image feature loss according to the first sample image feature information and the second sample image feature information; and training the large language model according to the fusion feature loss and the image feature loss, where the operation of training the large language model further includes: adjusting the initial embedded feature information.

[0201] Thus, according to the second aspect of this embodiment, the terminal device receives the instruction information for generating an image input by the user, where the instruction information may be simple and not highly professional. The user can describe at least part of the content of the image to be generated by inputting text, thereby reducing the usage threshold. Thus, this technical solution automatically generates a beautiful multi-layer image that meets the user's requirements according to the instruction information. And the multi-layer image in this technical solution is generated by uniformly generating each layer, so the text layer in it will be rendered at a professional level in consistency with other layers, thereby improving the legibility and rendering effect of the text. And the multi-layer image in this technical solution can be edited for each image of the multi-layer image after generation, thereby improving the freedom and flexibility of the multi-layer image. Furthermore, it solves the technical problems of high limitations in image generation, poor text rendering effect, low freedom, and high usage threshold existing in the prior art.

[0202] In addition, as shown in Figure 1 According to the third aspect of this embodiment, a storage medium is provided. The storage medium includes a stored program, where, when the program runs, the method described in any one of the above is executed by a processor.

[0203] Thus, according to this embodiment, the terminal device receives the instruction information for generating an image input by the user, where the instruction information can be simple and not highly professional. The user can describe at least a part of the content of the image to be generated by inputting text, thereby reducing the usage threshold. Thus, this technical solution automatically generates a beautiful multi-layer image that meets the user's requirements according to the instruction information. Moreover, the multi-layer image in this technical solution is generated by uniformly generating each layer, so the text layer therein will be rendered at a professional level in consistency with other layers, thereby improving the legibility and rendering effect of the text. And the multi-layer image in this technical solution can be edited for each image of the multi-layer image after generation, thereby improving the freedom and flexibility of the multi-layer image. Furthermore, the technical problems of high limitations in image generation, poor text rendering effect, low freedom, and high usage threshold existing in the prior art are solved.

[0204] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0205] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present invention.

[0206] Embodiment 2

[0207] Figure 13 Shows a multi-layer image generation device 1300 according to the first aspect of this embodiment. The device 1300 corresponds to the method according to the first aspect of Embodiment 1. Refer to Figure 13 As shown, the device 1300 includes: a first receiving module 1310 for receiving instruction information based on natural language, where the instruction information can be used to generate a corresponding image and includes descriptive text related to the image content of the image to be generated; and an image display module 1320 for displaying an editable multi-layer image, where the content of at least some of the layers in the multi-layer image is generated according to the descriptive text.

[0208] Optionally, the image display module 1320 includes: a first display sub-module for displaying a file link corresponding to the multi-layer image; a second display sub-module for displaying a file identifier corresponding to the multi-layer image; a third display sub-module for displaying the layer configuration information of each layer of the multi-layer image; and a fourth display sub-module for displaying the image content of each layer of the multi-layer image.

[0209] Optionally, the fourth display sub-module includes: a first display unit for displaying editable pixels corresponding to the corresponding image content in the layer; and a second display unit for displaying editable controls corresponding to the corresponding image content in the layer.

[0210] Optionally, the multi-layer image generating device 1300 further includes: a first information generating module for generating multi-layer image information corresponding to the multi-layer image according to the instruction information, where the multi-layer image information includes layer configuration information and an embedded image, the layer configuration information indicates the configuration information of each layer of the multi-layer image, and the embedded image is used to be inserted into the image layer of the multi-layer image, and where the layers of the multi-layer image include an image layer and a non-image layer, the image layer is generated according to the corresponding layer configuration information and the embedded image, and the non-image layer can be generated only according to the corresponding layer configuration information.

[0211] Optionally, the image display module 1320 includes: a fifth display sub-module for displaying the multi-layer image according to the multi-layer image information.

[0212] Optionally, the first information generating module includes: a first generating sub-module for using a large language model to generate first layer configuration information according to first text feature information corresponding to the instruction information, where the first layer configuration information includes second layer configuration information and third layer configuration information, where the second layer configuration information is used to indicate the layer configuration information of the first part of the non-image layers, the third layer configuration information is used to indicate part of the layer configuration information corresponding to the first image layer, and the third layer configuration information is located after the second layer configuration information; and a second generating sub-module for generating a first embedded image according to the first text feature information and second text feature information corresponding to the first layer configuration information, using the large language model and a diffusion model, where the first embedded image is used to be inserted into the first image layer.

[0213] Optionally, the second generation sub-module includes: a first generation unit configured to splice the first text feature information, the second text feature information, and the preset initial embedding feature information to generate first fused feature information, where the initial embedding feature information is the feature information for generating an embedded image; a second generation unit configured to generate second fused feature information suitable for image processing based on the first fused feature information through a large language model; a third generation unit configured to generate first semantic feature information based on the second fused feature information through a first projection layer; and a fourth generation unit configured to generate a first embedded image based on the first semantic feature information through a diffusion model.

[0214] Optionally, the second generation sub-module further includes: a fifth generation unit configured to generate fourth layer configuration information based on the first image feature information corresponding to the first embedded image and the first fused feature information through a large language model, where the fourth layer configuration information includes fifth layer configuration information and sixth layer configuration information, and where the fifth layer configuration information is used to indicate the layer configuration information of the second part of non-image layers, and the sixth layer configuration information is used to indicate the partial layer configuration information corresponding to the second image layer after the fifth layer configuration information; and a sixth generation unit configured to generate a second embedded image based on the first fused feature information, the first image feature information, and the third text feature information corresponding to the fourth layer configuration information through a large language model and a diffusion model, where the second embedded image is used to be inserted into the second image layer.

[0215] Optionally, the fifth generation unit includes: generating second image feature information based on the first embedded image through a first image encoder; generating first image feature information suitable for large language model processing based on the second image feature information through a second projection layer; and generating fourth layer configuration information based on the first image feature information and the first fused feature information through a large language model.

[0216] Optionally, the sixth generation unit includes: splicing the first fused feature information, the first image feature information, and the third text feature information to generate third fused feature information; generating fourth fused information suitable for image processing based on the third fused feature information through a large language model; generating second semantic feature information based on the fourth fused information through a first projection layer; and generating a second embedded image based on the second semantic feature information through a diffusion model.

[0217] Optionally, the second generation sub-module further includes: a seventh generation unit configured to generate seventh layer configuration information based on the first image feature information corresponding to the first embedded image and the first fused feature information through a large language model, where the seventh layer configuration information is used to indicate the layer configuration information of the remaining non-image layers except the first part of non-image layers.

[0218] Optionally, the first information generation module further includes: a third generation sub-module, configured to generate fifth image feature information according to a preset third embedded image through a second image encoder; a fourth generation sub-module, configured to generate sixth image feature information suitable for processing by a large language model according to the fifth image feature information through a third projection layer; and a fifth generation sub-module, configured to generate seventh layer configuration information for describing an image layer and eighth layer configuration information for describing a non-image layer according to fourth text feature information corresponding to the instruction information and the sixth image feature information through the large language model.

[0219] Optionally, the training module includes: training the large language model through the following steps: an acquisition sub-module, configured to acquire sample instruction information, first sample fusion feature information, and first sample image feature information, where the first sample fusion feature information is generated by splicing feature information corresponding to the sample instruction information and the sample layer configuration information, and initial embedded feature information, and the first sample image feature information is the feature information of a sample embedded image; a first sample generation module, configured to generate second sample fusion feature information according to the sample instruction information by using the large language model; a second sample generation module, configured to generate second sample image feature information according to the first sample fusion feature information; a third sample generation module, configured to determine a fusion feature loss according to the first sample fusion feature information and the second sample fusion feature information; a fourth sample generation module, configured to determine an image feature loss according to the first sample image feature information and the second sample image feature information; and a fifth sample generation module, configured to train the large language model according to the fusion feature loss and the image feature loss, where the training module further includes: an adjustment sub-module, configured to adjust the initial embedded feature information.

[0220] In addition, Figure 14 Fig. shows a multi-layer image generation device 1400 according to the second aspect of the present embodiment, and the device 1400 corresponds to the method according to the second aspect of Embodiment 1. Refer to Figure 14 As shown, the device 1400 includes: a second receiving module, configured to receive instruction information based on natural language, where the instruction information can be used to generate a corresponding image and includes descriptive text related to the image content of the image to be generated; an information generation module, configured to generate multi-layer image information corresponding to a multi-layer image according to the instruction information through a large language model, where the multi-layer image information includes layer configuration information and embedded images, the layer configuration information indicates the configuration information of each layer of the multi-layer image, and the embedded images are used to be inserted into the image layers of the multi-layer image, and where the layers of the multi-layer image include image layers and non-image layers, the image layers are generated according to the corresponding layer configuration information and the embedded images, and the non-image layers can be generated only according to the corresponding layer configuration information; and an image generation module, configured to generate an editable multi-layer image according to the multi-layer image information.

[0221] Thus, according to this embodiment, the terminal device receives instruction information for generating an image input by the user, where the instruction information can be simple instruction information with low professionalism. The user can describe at least part of the content of the image to be generated by inputting text, thereby reducing the usage threshold. Thus, this technical solution automatically generates a beautiful multi-layer image that meets the user's requirements according to the instruction information. Moreover, the multi-layer image in this technical solution is generated by uniformly generating each layer. Therefore, the text layer therein will be rendered at a professional level in consistency with other layers, thereby improving the legibility and rendering effect of the text. And the multi-layer image in this technical solution can be edited for each image of the multi-layer image after generation, thereby improving the degree of freedom and flexibility of the multi-layer image. Furthermore, the technical problems of high limitations in image generation, poor text rendering effect, low degree of freedom, and high usage threshold existing in the prior art are solved.

[0222] Embodiment 3

[0223] Figure 15 Shows a multi-layer image generation device 1500 according to the first aspect of this embodiment, and the device 1500 corresponds to the method according to the first aspect of Embodiment 1. Refer to Figure 15 As shown, the device 1500 includes: a first processor 1510; and a first memory 1520, connected to the first processor 1510, for providing instructions for the first processor 1510 to perform the following processing steps: receiving instruction information based on natural language, where the instruction information can be used to generate a corresponding image and includes descriptive text related to the image content of the image to be generated; and displaying an editable multi-layer image, where the content of at least part of the layers in the multi-layer image is generated according to the descriptive text.

[0224] Optionally, the operation of displaying a multi-layer image includes any of the following methods: displaying a file link corresponding to the multi-layer image; displaying a file identifier corresponding to the multi-layer image; displaying the layer configuration information of each layer of the multi-layer image; and displaying the image content of each layer of the multi-layer image.

[0225] Optionally, the operation of displaying the image content of each layer of the multi-layer image includes: displaying editable pixels corresponding to the respective image content on the layer; and displaying editable controls corresponding to the respective image content on the layer.

[0226] Optionally, the first memory 1520 is further configured to provide instructions for the first processor 1510 to process the following processing steps: generating multi-layer image information corresponding to the multi-layer image according to the instruction information, where the multi-layer image information includes layer configuration information and an embedded image, the layer configuration information indicates the configuration information of each layer of the multi-layer image, and the embedded image is used to be inserted into the image layer of the multi-layer image, and where the layers of the multi-layer image include an image layer and a non-image layer, the image layer is generated according to the corresponding layer configuration information and the embedded image, and the non-image layer can be generated only according to the corresponding layer configuration information.

[0227] Optionally, the operation of displaying a multi-layer image includes: displaying the multi-layer image according to the multi-layer image information.

[0228] Optionally, in the case where the instruction information indicates generating a multi-layer image using natural language-based unimodal information, the operation of generating multi-layer image information according to the instruction information includes: using a large language model to generate first layer configuration information according to first text feature information corresponding to the instruction information, where the first layer configuration information includes second layer configuration information and third layer configuration information, where the second layer configuration information is used to indicate the layer configuration information of the first part of the non-image layers, the third layer configuration information is used to indicate the partial layer configuration information corresponding to the first image layer, and the third layer configuration information is located after the second layer configuration information; and generating a first embedded image using the large language model and a diffusion model according to the first text feature information and second text feature information corresponding to the first layer configuration information, where the first embedded image is used to be inserted into the first image layer.

[0229] Optionally, the operation of generating the first embedded image using the large language model and the diffusion model according to the first text feature information and the second text feature information corresponding to the first layer configuration information includes: splicing the first text feature information, the second text feature information, and the preset initial embedded feature information to generate first fused feature information, where the initial embedded feature information is the feature information for generating the embedded image; generating second fused feature information suitable for image processing by the large language model according to the first fused feature information; generating first semantic feature information by the first projection layer according to the second fused feature information; and generating the first embedded image by the diffusion model according to the first semantic feature information.

[0230] Optionally, the first memory 1520 is further configured to provide instructions for the first processor 1510 to process the following steps: generating fourth layer configuration information using the large language model according to the first image feature information corresponding to the first embedded image and the first fused feature information, where the fourth layer configuration information includes fifth layer configuration information and sixth layer configuration information, and where the fifth layer configuration information is used to indicate the layer configuration information of the second part of the non-image layer, and the sixth layer configuration information is used to indicate the partial layer configuration information corresponding to the second image layer after the fifth layer configuration information; and generating a second embedded image using the large language model and the diffusion model according to the first fused feature information, the first image feature information, and the third text feature information corresponding to the fourth layer configuration information, where the second embedded image is used to be inserted into the second image layer.

[0231] Optionally, the operation of generating the fourth layer configuration information using the large language model according to the first image feature information corresponding to the first embedded image and the first fused feature information includes: generating second image feature information by the first image encoder according to the first embedded image; generating first image feature information suitable for large language model processing by the second projection layer according to the second image feature information; and generating the fourth layer configuration information using the large language model according to the first image feature information and the first fused feature information.

[0232] Optionally, the operation of generating the second embedded image using the large language model and the diffusion model according to the first fused feature information, the first image feature information, and the third text feature information corresponding to the fourth layer configuration information includes: splicing the first fused feature information, the first image feature information, and the third text feature information to generate third fused feature information; generating fourth fused information suitable for image processing by the large language model according to the third fused feature information; generating second semantic feature information by the first projection layer according to the fourth fused information; and generating the second embedded image by the diffusion model according to the second semantic feature information.

[0233] Optionally, the first memory 1520 is further configured to provide instructions for the first processor 1510 to process the following processing steps: generating seventh layer configuration information by using a large language model according to the first image feature information corresponding to the first embedded image and the first fusion feature information, where the seventh layer configuration information is used to indicate the layer configuration information of the remaining non-image layers except the first part of non-image layers.

[0234] Optionally, in the case where the instruction information indicates generating a multi-layer image by using an existing embedded image, the operation of generating multi-layer image information according to the instruction information includes: generating fifth image feature information by a second image encoder according to a preset third embedded image; generating sixth image feature information suitable for processing by a large language model by a third projection layer according to the fifth image feature information; and generating seventh layer configuration information for describing image layers and eighth layer configuration information for describing non-image layers by the large language model according to the fourth text feature information corresponding to the instruction information and the sixth image feature information.

[0235] Optionally, the large language model is trained by the following steps: collecting sample instruction information, first sample fusion feature information, and first sample image feature information, where the first sample fusion feature information is generated by splicing the feature information corresponding to the sample instruction information and the sample layer configuration information, and the initial embedded feature information, and the first sample image feature information is the feature information of the sample embedded image; generating second sample fusion feature information by using the large language model according to the sample instruction information; generating second sample image feature information according to the first sample fusion feature information; determining a fusion feature loss according to the first sample fusion feature information and the second sample fusion feature information; determining an image feature loss according to the first sample image feature information and the second sample image feature information; and training the large language model according to the fusion feature loss and the image feature loss, where the operation of training the large language model further includes: adjusting the initial embedded feature information.

[0236] In addition, Figure 16 There is shown a multi-layer image generating device 1600 according to the second aspect of the present embodiment, and the device 1600 corresponds to the method according to the second aspect of Embodiment 1. Refer to Figure 16 As shown, the device 1600 includes: a second processor 1610; and a second memory 1620, connected to the second processor 1610, for providing instructions for the second processor 1610 to process the following processing steps: receiving instruction information based on natural language, where the instruction information can be used to generate a corresponding image and includes descriptive text related to the image content of the image to be generated; generating, by a large language model, multi-layer image information corresponding to a multi-layer image according to the instruction information, where the multi-layer image information includes layer configuration information and embedded images, the layer configuration information indicates the configuration information of each layer of the multi-layer image, and the embedded images are used to be inserted into the image layers of the multi-layer image, and where the layers of the multi-layer image include image layers and non-image layers, the image layers are generated according to the corresponding layer configuration information and embedded images, and the non-image layers can be generated only according to the corresponding layer configuration information; and generating an editable multi-layer image according to the multi-layer image information.

[0237] Accordingly, in this embodiment, the terminal device receives the instruction information for generating an image input by the user, where the instruction information can be simple instruction information with low professionalism. The user can describe at least part of the content of the image to be generated by inputting text, thus reducing the usage threshold. Accordingly, this technical solution automatically generates a beautiful multi-layer image that meets the user's requirements according to the instruction information. And the multi-layer image in this technical solution is generated by uniformly generating each layer, so the text layer in it will be rendered at a professional level in consistency with other layers, thus improving the legibility and rendering effect of the text. And the multi-layer image in this technical solution can be edited for each image of the multi-layer image after generation, thus improving the freedom and flexibility of the multi-layer image. Furthermore, it solves the technical problems of high limitations in image generation, poor text rendering effect, low freedom, and high usage threshold existing in the prior art.

[0238] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0239] In the above embodiments of the present invention, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0240] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.

[0241] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0242] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0243] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. And the aforementioned storage medium includes: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks or optical discs and other various media that can store program codes.

[0244] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.< / position> < / group> < / position> < / group> < / position> < / group> < / position> < / group>

Claims

1. A method for generating a multi-layer image, characterized in that, Including: Receiving natural language-based instruction information, where the instruction information can be used to generate corresponding images and includes descriptive text related to the image content of the image to be generated; And Displaying an editable multi-layer image, where the content of at least some of the layers in the multi-layer image is generated according to the descriptive text.

2. The method according to claim 1, wherein The operations of displaying the multi-layer image include any of the following methods: Displaying a file link corresponding to the multi-layer image; Displaying a file identifier corresponding to the multi-layer image; Displaying the layer configuration information of each layer of the multi-layer image; And Displaying the image content of each layer of the multi-layer image.

3. The method according to claim 2, wherein The operations of displaying the image content of each layer of the multi-layer image include: Displaying editable pixels corresponding to the corresponding image content in the layer; and Displaying editable controls corresponding to the corresponding image content in the layer.

4. The method according to claim 1, characterized in that, It further includes: Generating multi-layer image information corresponding to the multi-layer image according to the instruction information, where the multi-layer image information includes layer configuration information and embedded images, the layer configuration information indicates the configuration information of each layer of the multi-layer image, and the embedded images are used to be inserted into the image layers of the multi-layer image, and where The layers of the multi-layer image include image layers and non-image layers, the image layers are generated according to the corresponding layer configuration information and the embedded images, and the non-image layers can be generated only according to the corresponding layer configuration information, where The operations of displaying the multi-layer image include: displaying the multi-layer image according to the multi-layer image information, where In the case where the instruction information indicates generating the multi-layer image using natural language-based unimodal information, the operations of generating multi-layer image information according to the instruction information include: using a large language model to generate first layer configuration information according to first text feature information corresponding to the instruction information, where the first layer configuration information includes second layer configuration information and third layer configuration information, the second layer configuration information is used to indicate the layer configuration information of the first part of non-image layers, the third layer configuration information is used to indicate part of the layer configuration information corresponding to the first image layer, and the third layer configuration information is located after the second layer configuration information; and generating a first embedded image using the large language model and a diffusion model according to the first text feature information and second text feature information corresponding to the first layer configuration information, where the first embedded image is used to be inserted into the first image layer, where The operation of generating the first embedded image using the large language model and the diffusion model based on the first text feature information and the second text feature information corresponding to the first layer configuration information includes: splicing the first text feature information, the second text feature information with the preset initial embedded feature information to generate first fused feature information, where the initial embedded feature information is the feature information for generating the embedded image; generating second fused feature information suitable for image processing by the large language model according to the first fused feature information; generating first semantic feature information by the first projection layer according to the second fused feature information; and generating the first embedded image by the diffusion model according to the first semantic feature information.

5. The method according to claim 4, wherein It further includes: Using the large language model to generate fourth layer configuration information according to the first image feature information corresponding to the first embedded image and the first fused feature information, where the fourth layer configuration information includes fifth layer configuration information and sixth layer configuration information, and where the fifth layer configuration information is used to indicate the layer configuration information of the second part of non-image layers, and the sixth layer configuration information is used to indicate the partial layer configuration information corresponding to the second image layer located after the fifth layer configuration information; And According to the first fused feature information, the first image feature information, and the third text feature information corresponding to the fourth layer configuration information, using the large language model and the diffusion model to generate a second embedded image, where the second embedded image is used to be inserted into the second image layer, where The operation of using the large language model to generate fourth layer configuration information according to the first image feature information corresponding to the first embedded image and the first fused feature information includes: Generating second image feature information by the first image encoder according to the first embedded image; Generating first image feature information suitable for processing by the large language model by the second projection layer according to the second image feature information; and Using the large language model to generate fourth layer configuration information according to the first image feature information and the first fused feature information.

6. The method according to claim 5, characterized in that, The operation of generating the second embedded image using the large language model and the diffusion model according to the first fused feature information, the first image feature information, and the third text feature information corresponding to the fourth layer configuration information includes: Splicing the first fused feature information, the first image feature information, and the third text feature information to generate third fused feature information; Generating fourth fused information suitable for image processing by the large language model according to the third fused feature information; Generating second semantic feature information by the first projection layer according to the fourth fused information; and Generating the second embedded image by the diffusion model according to the second semantic feature information.

7. The method according to claim 5, characterized in that, It further includes: Using the large language model, generate seventh layer configuration information based on the first image feature information corresponding to the first embedded image and the first fusion feature information, where the seventh layer configuration information is used to indicate the layer configuration information of the remaining non-image layers except the first part of non-image layers.

8. The method according to claim 4, wherein In the case where the instruction information indicates generating a multi-layer image using an existing embedded image, the operation of generating multi-layer image information according to the instruction information includes: Generating fifth image feature information through a second image encoder based on a preset third embedded image; Generating sixth image feature information suitable for processing by the large language model through a third projection layer according to the fifth image feature information; and Generating seventh layer configuration information for describing image layers and eighth layer configuration information for describing non-image layers through the large language model according to the fourth text feature information corresponding to the instruction information and the sixth image feature information.

9. The method according to claim 5, wherein Training the large language model through the following steps: Collecting sample instruction information, first sample fusion feature information, and first sample image feature information, where the first sample fusion feature information is generated by splicing the feature information corresponding to the sample instruction information and sample layer configuration information, and initial embedded feature information, and the first sample image feature information is the feature information of the sample embedded image; Using the large language model to generate second sample fusion feature information according to the sample instruction information; Generating second sample image feature information according to the first sample fusion feature information; Determining a fusion feature loss according to the first sample fusion feature information and the second sample fusion feature information; Determining an image feature loss according to the first sample image feature information and the second sample image feature information; And Training the large language model according to the fusion feature loss and the image feature loss, where The operation of training the large language model further includes: adjusting the initial embedded feature information.

10. A method for generating a multi-layer image, characterized in that, Including: Receiving instruction information based on natural language, where the instruction information can be used to generate a corresponding image and includes descriptive text related to the image content of the image to be generated; Generating multi-layer image information corresponding to a multi-layer image through a large language model according to the instruction information, where the multi-layer image information includes layer configuration information and an embedded image, the layer configuration information indicates the configuration information of each layer of the multi-layer image, and the embedded image is used to be inserted into the image layer of the multi-layer image, and where The layers of the multi-layer image include image layers and non-image layers, the image layers are generated according to the corresponding layer configuration information and the embedded image, and the non-image layers can be generated only according to the corresponding layer configuration information; And Generating an editable multi-layer image according to the multi-layer image information.

Citation Information

Patent Citations

  • Image editing method and device, equipment, storage medium and program product

    CN117611709A

  • Apparatus and method for generating text from image and method of training model for generating text from image

    US20240177507A1

  • Systems and methods for layered image generation

    US20250078346A1

Cited By

  • Image generation method and device, electronic equipment and storage medium

    CN120953422A

  • Image generation method and device, electronic equipment and storage medium

    CN120953422B