Unified multi-task and multi-field layout generation method and device, equipment and medium

Through a method based on a large language model, unified multi-task and multi-domain layout generation is solved, and the problem of lack of universality and flexibility in layout generation in the existing technology is solved, and a high-performance layout generation engine is realized.

CN120068798AActive Publication Date: 2025-05-30SOUTH CHINA UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510026123.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-30
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

It is difficult for the existing technology to achieve unified multi-task and multi-field layout generation. The existing multi-task model only focuses on single-field layout generation and lacks universality and flexibility.

Method used

Using a method based on a large language model, the information of all elements of each layout is leveled into a sequence by obtaining layouts in multiple fields, and a random mask is used as input, and the complete sequence is trained as a label. Use interval quantization position encoding to avoid using placeholders and improve generation efficiency.

Benefits of technology

It realizes unified multi-layered generation tasks and multi-field layout generation, improves the universality and performance of the generation engine, and can surpass other single-field layout generation methods in complex multi-task and multi-field unified scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068798A_ABST
    Figure CN120068798A_ABST
Patent Text Reader

Abstract

The invention discloses a method, a device and equipment for unifying multi-task and multi-field layout generation and a medium. The method comprises the following steps: acquiring layout generation data of a plurality of fields; planishing information of all elements of each layout into a sequence, randomly masking the sequence as input, and taking the complete sequence as a label; the sequence and the tag are input into a large language model for training, and layout data in different fields are used in a mixed mode in training; and inputting partial layout information generated according to different task requirements and different field requirements into the trained model, and enabling the model to generate a complete layout sequence. According to the method, deep learning and a sequence generation technology based on a large language model are used, various layout generation tasks and layout generation in multiple fields are unified, and a universal layout generation engine with good performance is realized. The method can be widely applied to the field of deep learning and pattern recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning and pattern recognition, and in particular to a method, device, equipment and medium for unified multi-task and multi-domain layout generation. Background Art

[0002] Layout generation refers to automatically generating a structured document layout with a reasonable typesetting layout according to the requirements of users. Specifically, a complete layout consists of multiple layout elements, and each layout element contains 5 pieces of information: element category (text, table, picture, etc.), the x and y coordinates of the upper left corner of the box, and the width and height of the box. Users can give some information of the layout elements, such as only giving the element category or the width and height of the box, and the program generates a complete document layout according to these preset and incomplete information. This technology is widely used in a variety of actual scenarios, such as web and UI design (quickly generating web layouts or user interface designs according to natural language descriptions), advertising poster generation (automatically generating marketing posters, advertising copy layout, etc.), education and publishing (automatically generating test papers, teaching plans, book catalog layouts, etc.).

[0003] Most of the existing layout generation methods only focus on the layout generation of a single task. For example, inputting the categories of all elements and letting the model automatically infer the x, y coordinates and width and height of this element. In recent years, layout generation models that take into account multiple tasks have received more and more attention and shown better generality and flexibility. However, these multi-task models only focus on the layout generation of a single domain, such as only generating paper layouts or magazine layouts. Summary of the Invention

[0004] To at least to some extent solve one of the technical problems existing in the prior art, the purpose of the present invention is to provide a method, device, equipment and medium for unified multi-task and multi-domain layout generation based on a large language model.

[0005] The first technical solution adopted by the present invention is:

[0006] A method for unified multi-task and multi-domain layout generation, comprising the following steps:

[0007] Obtain layout generation data of multiple domains;

[0008] Flatten the information of all elements of each layout into a sequence, perform random masking on the sequence as input, and use the complete sequence as a label;

[0009] Input the sequence and the label into a large language model for training, and the layout data of different domains are mixed and used in the training;

[0010] Input partial layout information generated according to different task requirements and different domain requirements into the trained model, so that the model generates a complete layout sequence.

[0011] Furthermore, the layout generation data includes thesis layouts, mobile App UI layouts, magazine layouts, and slide layouts;

[0012] Each complete layout consists of N layout elements, and the format of each layout element is (c, x, y, w, h), where c is the category of the current element, such as text, table, image, etc., (x, y) is the upper left corner coordinate of the current element in the entire layout canvas, w is the width of the current element, and h is the height of the current element; the entire layout is represented as: {x 1 , y 1 , w 1 , h 1 , …, x N , y N , w N , h N}.

[0013] Furthermore, flattening the information of all elements of each layout into a sequence, performing random masking on the sequence as the input, and using the complete sequence as the label, includes:

[0014] Construct an arbitrary layout prompt template, and after masking or adding noise, use it as the input of the model;

[0015] Construct a unified layout answer template, and use the sequence composed of the complete and noise-free layout element information as the unified layout answer;

[0016] Use interval quantization position encoding to avoid using placeholders.

[0017] Furthermore, the arbitrary layout prompt template consists of two parts: a prefix part and a body part;

[0018] The prefix part consists of a "denoising flag", a "layout type", a "number of elements", and a "number of columns"; the body part consists of multiple layout elements and descriptions of the relationships between different elements, and the format of each element is (c, x, y, w, h);

[0019] Perform random masking or add noise to the information of each element to simulate arbitrary layout generation conditions; the masked information will be directly discarded, and adding noise means adding random noise to x, y, w, h, and requires the model to denoise the added noise; the unmasked layout element information is directly spliced together as the input of the model, and the model generates the complete layout element information in an autoregressive manner.

[0020] Further, the position encoding by interval quantization to avoid using placeholders includes:

[0021] Set a sufficiently large interval value l, which is greater than the length and width of all layout canvases; then encode the x, y, w, and h of each layout element according to the following rules:

[0022] x = x + 0 × h

[0023] y = y + 1 × h

[0024] w = w + 2 × h

[0025] h = h + 3 × h

[0026] This encoding rule makes x, y, w, and h each fall into a different numerical interval, that is, x ∈ [0, l), y ∈ [l, 2l), w ∈ [2l, 3l), h ∈ [3l, 4l).

[0027] Further, the inputting the sequence and label into the large language model for training, and mixing the layout data in different fields for training includes:

[0028] Adopt the GPT2-XL model as the large language model and train the model:

[0029] Take any layout prompt template as the input of the GPT2-XL model, and use the unified layout answer template as the label, and require the model to generate a complete layout in the specified field for any task and generation requirement in any field.

[0030] Further, it also includes the following steps:

[0031] Compare the generated layout with the real layout to calculate metrics, or render the generated layout sequence into a two-dimensional layout picture.

[0032] The second technical solution adopted by the present invention is:

[0033] An apparatus for unified multi-task and multi-field layout generation includes:

[0034] A data acquisition module for acquiring layout generation data in multiple fields;

[0035] An input-output construction module for flattening the information of all elements of each layout into a sequence, randomly masking the sequence as the input, and using the complete sequence as the label;

[0036] A model training module for inputting the sequence and label into the large language model for training, and mixing the layout data in different fields for training;

[0037] A model inference module, which is used to input partial layout information generated according to different task requirements and different domain requirements into the trained model, so that the model generates a complete layout sequence.

[0038] The third technical solution adopted by the present invention is:

[0039] An electronic device, the electronic device includes a processor and a memory, and at least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the above-mentioned method for generating a unified multi-task and multi-domain layout.

[0040] The fourth technical solution adopted by the present invention is:

[0041] A computer-readable storage medium, and at least one instruction, at least one program, a code set or an instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the above-mentioned method for generating a unified multi-task and multi-domain layout.

[0042] The fifth technical solution adopted by the present invention is:

[0043] A computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned method for generating a unified multi-task and multi-domain layout.

[0044] The beneficial effects of the present invention are: The present invention uses deep learning and sequence generation technology based on large language models to unify various layout generation tasks and layout generation in multiple domains, and realizes a general and high-performance layout generation engine. In addition, the performance of the model is enhanced by compressing the information in the prompt, and it can exceed the layout generation methods in other single scenarios even in more difficult unified scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following introduces the accompanying drawings related to the technical solutions in the embodiments of the present invention or the prior art. It should be understood that the accompanying drawings below are only for conveniently and clearly expressing some embodiments of the technical solutions in the present invention. For those skilled in the art, other drawings can be obtained according to these drawings without creative efforts.

[0046] Figure 1It is a flowchart of the steps of a method for unified multi-task and multi-domain layout generation in an embodiment of the present invention. Detailed implementation manners

[0047] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as a limitation of the present invention. For the step numbers in the following embodiments, they are only set for the convenience of description and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0048] In the description of the present invention, it should be understood that for the orientation description, such as the orientation or positional relationship indicated by up, down, front, back, left, right, etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.

[0049] In the description of the present invention, the meaning of several is one or more, the meaning of multiple is more than two, and understandings such as greater than, less than, exceeding, etc. do not include the present number, and understandings such as above, below, within, etc. include the present number. If there is a description of first and second, it is only for the purpose of distinguishing technical features and should not be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.

[0050] In the description of the present invention, unless otherwise clearly defined, words such as setting, installing, connecting, etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above words in the present invention in combination with the specific content of the technical solution.

[0051] In order to achieve more general and comprehensive layout generation, the present invention proposes a method for unified multi-task and multi-domain layout generation. Since the unification of multi-domains and multi-tasks brings greater difficulties, it is difficult to obtain good generation effects with existing small models. Therefore, the present invention proposes to implement this method based on large language models. This method proposes an arbitrary layout prompt template, which accommodates the conditions of any layout generation task and the requirements of multi-domain layout generation, and uses interval quantization position encoding to compress the length of the input to improve the generation efficiency. The method of the present invention uses a large language model as the generation engine and can achieve performance beyond existing methods in complex unified scenarios of multi-tasks and multi-domains.

[0052] Embodiment 1

[0053] As shown Figure 1 in the figure, this embodiment provides a method for unified multi-task and multi-domain layout generation based on a large language model, including the following steps:

[0054] S1. Obtain layout generation data for multiple domains.

[0055] As an implementation manner, the embodiments of the present invention consider layout data for four domains, including thesis layouts, mobile App UI layouts, magazine layouts, and slide layouts. It should be noted that the present invention is not limited to the data of these four layout types and can include layout data of any type. Each complete layout consists of N layout elements, and the format of each layout element is (c, x, y, w, h), where c is the category of the current element, such as text, table, image, etc., (x, y) is the upper left corner coordinate of the current element in the entire layout canvas, w is the width of the current element, and h is the height of the current element. The entire layout can be represented as {x 1 , y 1 , w 1 , h 1 , …, x N , y N , w N , h N}.

[0056] S2. Flatten the information of all elements of each layout into a sequence, randomly mask the sequence as the input, and use the complete sequence as the label.

[0057] In some embodiments, step S2 specifically includes the following steps S21 - S23:

[0058] S21. Construct an arbitrary layout prompt template.

[0059] This embodiment proposes an arbitrary layout prompt template as the input of the model, which consists of two parts: a prefix part and a main body part. The prefix part consists of a "denoising flag", a "layout type", a "number of elements", and a "number of columns", where the "layout type" is "thesis", "App UI", "magazine", and "slide". The "number of elements" is N in step S1. The main body part consists of multiple layout elements and descriptions of the relationships between different elements, and the format of each element is (c, x, y, w, h) as described above.

[0060] In this embodiment, the information of each element is randomly masked or noise-added to simulate arbitrary layout generation conditions. The masked information will be directly discarded. Adding noise means adding random noise to x, y, w, h, and the model is required to denoise the added noise. The unmasked layout element information will be directly spliced together as the input. For example, assuming that the sequence of a complete layout is "denoising; paper; N; 2; text x 1 ,y 1 ,w 1 ,h 1 ; …; table x N ,y N ,w N ,h N ", after masking and noise-adding, the possible input may be "denoising; paper; N; 2; text …; table y N ". By changing the "layout type", the model can learn to generate a specific type of layout; by inputting arbitrary layout generation conditions, the model can learn to generate layouts under arbitrary conditions. The combined effect of the two can train the model's ability to generate unified multi-domain and multi-task layouts.

[0061] S22. Construct a unified layout answer template.

[0062] As described in step S21, we use the sequence composed of complete and noise-free layout element information as the unified layout answer, that is, no matter what input is given, the model is required to output a complete layout, such as "text x 1 ,y 1 ,w 1 ,h 1 ; …; table x N ,y N ,w N ,h N ". Combining any layout prompt template with the unified layout answer template can enable the model to learn to output a complete layout in any task and any specified domain, achieving the ability of unified layout generation.

[0063] S23. Interval quantization position encoding.

[0064] In the description of step S21, the embodiments of the present invention directly discard the masked information, splice the remaining unmasked information together, and then let the model generate the complete layout element information in an autoregressive manner. However, this will face a problem, that is, the model may not be able to infer which of x, y, w, and h a certain unmasked value is. The conventional solution is to replace the position of a piece of information with a placeholder after masking it, which is equivalent to telling the model that there is a missing position here that needs to be predicted. However, this approach will introduce a large number of placeholders, reducing the layout information density in the input instruction and affecting the generation performance. Therefore, the embodiments of the present invention propose a method that can avoid both model confusion about unmasked values and the use of placeholders, namely interval quantization position encoding. Specifically, in this embodiment, a sufficiently large interval value l is set, and this interval is greater than the length and width of all layout canvases. Then, the x, y, w, and h of each layout element are encoded according to the following rules:

[0065] x = x + 0 × h

[0066] y = y + 1 × h

[0067] w = w + 2 × h

[0068] h = h + 3 × h

[0069] This encoding rule makes x, y, w, and h each in a different numerical interval, that is, x ∈ [0, l), y ∈ [l, 2l), w ∈ [2l, 3l), h ∈ [3l, 4l). The model can accurately determine what the unmasked element is and what the information to be predicted is based on the magnitude of the value, so there is no need to use placeholders to replace the masked information. This interval quantization position encoding can retain only the valid layout information in the input, greatly increasing the layout information density and improving the model performance; at the same time, it can shorten the length of the input sequence, improve the input and output efficiency, and thus improve the training and inference efficiency of the model.

[0070] S3. Input the sequence and labels into the large language model for training, and the layout data in different fields are mixed and used in the training.

[0071] Use existing large language models as sequence generation models. In recent years, large language models have mainly been developed based on pure decoder architectures, such as GPT2, LLaMA, etc., demonstrating quite amazing text generation capabilities. In the embodiments of the present invention, the GPT2-XL model is used as the generation engine. GPT2-XL is a pure decoder Transformer with 1.5 billion parameters, including 48 Transformer decoder layers, a hidden layer dimension of 1600, and 25 attention heads, and uses an autoregressive paradigm for text generation. In this embodiment, any layout prompt template is used as the input to GPT2-XL, and a unified layout answer template is used as a label, requiring the model to generate a complete layout in a specified field for any task and any field generation requirements.

[0072] Since generating layouts that unify multiple tasks and multiple fields simultaneously is more difficult than only unifying one of the scenarios or single-task, single-field layout generation, we hope to utilize the extremely strong reasoning ability of large language models themselves to solve this problem, and experiments have also proven that using large language models can obtain better performance than other methods compared to using small language models.

[0073] As an implementation method, the training parameters are specifically as follows:

[0074] 1) Number of iterations: 23000

[0075] 2) Optimizer: AdamW

[0076] 3) Learning rate: 0.0001, with a cosine annealing strategy for decay, and the learning rate drops to 0 at the last iteration step.

[0077] S4. Input the partial layout information generated according to different task requirements and different field requirements into the trained model, and let the model generate a complete layout sequence.

[0078] Model inference: Input the partial layout information generated according to different task requirements and different field requirements into the model, and let the model generate a complete layout sequence. The generated layout can be compared with the real layout to calculate metrics, or the layout sequence can be rendered into a two-dimensional layout image.

[0079] Specifically, according to a given layout generation task and a given layout data type, corresponding inputs can be generated according to any layout prompt template. For example, if the generation task is type c with only 10 layout elements given, generating the x, y, w, h of these elements, and the layout data type to be generated is a paper, then the input is "non-denoising; paper; 10; 2; text; table; …; image". After inputting into the model, a sequence should be generated, which should be "text x 1 ,y 1 ,w 1,h 1 ; Table x 2 ,y 2 ,w 2 ,h 2 ; …; Image x N ,y N ,w N ,h N ". The generated sequence and the real sequence can be compared to calculate metrics, or the generated sequence can be rendered on a blank canvas according to the coordinates and width and height to visualize each layout element of the layout.

[0080] In summary, the present invention discloses a layout generation method, which includes an arbitrary layout prompt template, a unified layout answer template, and interval quantization position encoding. The arbitrary layout prompt template encompasses any combination of layout generation requirements, can support any layout generation task, and at the same time specifies the type of layout data to be generated (such as magazines, papers, etc.) through specific type indications, supporting layout generation in multiple fields. Therefore, it can support layout generation for multiple tasks and multiple fields simultaneously. The interval quantization position encoding encodes the position information (x coordinate, y coordinate, width, and height) of layout elements into independent numerical intervals, maintaining the discriminability between different position information without the need to use placeholders, shortening the length of the input information and increasing the information density, enhancing the performance and input-output efficiency of the model. The present invention inputs the layout information into the large language model according to the arbitrary layout prompt template, generates a complete layout as the output, and realizes unified layout generation for multiple fields and multiple tasks. The present invention first proposes a method for unified multi-task and multi-field layout generation, and at the same time enhances the performance of the model through information compression of the prompt, and can surpass the layout generation methods in other single scenarios even in more difficult unified scenarios.

[0081] The method of the present invention has at least the following advantages and beneficial effects compared with the prior art:

[0082] (1) The method proposed by the present invention is the first method for unified multi-task and multi-field layout generation, while the previous methods only focused on the unification of multi-tasks or the unification of multi-fields, and did not unify the two scenarios. At the same time, even in the more difficult situation brought about by the two unifications, this method can still achieve performance that surpasses the existing methods for only multi-task unification / multi-field unification.

[0083] (2) The present invention proposes an arbitrary layout prompt template and a unified layout answer template. The arbitrary layout prompt template consists of a prefix part and a main body part. The prefix part can specify the type of layout data to be generated, thus accommodating the layout generation requirements in multiple fields; the main body part can simulate arbitrary layout generation requirements by randomly masking the information in the layout elements (i.e., c, x, y, w, h). Each requirement can be regarded as a task, so doing so can cover arbitrary layout generation tasks. The two parts work together as the input of the model. By requiring the model to output a complete layout (i.e., the unified layout answer template), the unified layout generation ability for multiple tasks and multiple fields can be trained.

[0084] (3) The present invention proposes an interval quantization coding method. By encoding the values of x, y, w, and h into non-overlapping intervals, the model can accurately determine what the unmasked elements in the input sequence are and what the information to be predicted is based on the numerical size, avoiding using placeholders to replace the masked information. On the one hand, this can retain only the valid layout information without using placeholders, increasing the layout information density in the input and enhancing the generation performance of the model. On the other hand, it can shorten the length of the input sequence, improving the training and inference efficiency of the model.

[0085] (4) The present invention uses a large language model as the generation engine, which is a rare exploration of using a large language model for layout generation. Since unifying layout generation for multiple tasks and multiple fields simultaneously is more difficult than unifying only one of the scenarios or for single-task and single-field layout generation, the present invention makes good use of the extremely strong reasoning ability of the large language model itself to solve this problem and obtains better performance than other methods.

[0086] Embodiment 2

[0087] This embodiment provides a device for unified multi-task and multi-field layout generation, including:

[0088] A data acquisition module for acquiring layout generation data in multiple fields;

[0089] An input-output construction module for flattening the information of all elements of each layout into a sequence, randomly masking the sequence as the input, and using the complete sequence as the label;

[0090] A model training module for inputting the sequence and the label into a large language model for training, and mixing layout data in different fields during training;

[0091] A model inference module for inputting partial layout information generated according to different task requirements and different field requirements into the trained model, and enabling the model to generate a complete layout sequence.

[0092] Since this device is a device for unified multi-task and multi-domain layout generation in an embodiment of the present invention, and the principle of how this device solves problems is similar to that of the method, the implementation of this device can refer to the implementation process of the above method embodiment, and repeated parts will not be elaborated again.

[0093] Embodiment 3

[0094] An embodiment of the present invention further provides an electronic device, the electronic device includes a processor and a memory, and at least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement a method for unified multi-task and multi-domain layout generation as Figure 1 shown.

[0095] It can be understood that the memory may include a random access memory (RAM), and may also include a read-only memory (ROM). Optionally, the memory includes a non-transitory computer-readable storage medium. The memory can be used to store instructions, programs, codes, code sets or instruction sets. The memory may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the above various method embodiments, etc.; the data storage area may store data created according to the use of the server, etc.

[0096] The processor may include one or more processing cores. The processor uses various interfaces and lines to connect various parts within the entire server, and by running or executing instructions, programs, code sets or instruction sets stored in the memory, and calling data stored in the memory, it executes various functions of the server and processes data. Optionally, the processor may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor may integrate one or a combination of several of a central processing unit (CPU) and a modem, etc. Among them, the CPU mainly processes the operating system and application programs, etc.; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor and may be implemented separately by a single chip.

[0097] Since the electronic device is an electronic device corresponding to a method for unified multitasking and multi-domain layout generation according to an embodiment of the present invention, and the principle of the electronic device for solving problems is similar to that of the method, the implementation of the electronic device can refer to the implementation process of the above method embodiment, and the repeated parts will not be described again.

[0098] Embodiment 4

[0099] An embodiment of the present invention further provides a computer-readable storage medium, in which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement a method for unified multitasking and multi-domain layout generation as Figure 1 shown.

[0100] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electrically-erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc memories, a magnetic disk memory, a tape memory, or any other computer-readable medium capable of carrying or storing data.

[0101] Since the storage medium is a storage medium corresponding to a method for unified multitasking and multi-domain layout generation according to an embodiment of the present invention, and the principle of the storage medium for solving problems is similar to that of the method, the implementation of the storage medium can refer to the implementation process of the above method embodiment, and the repeated parts will not be described again.

[0102] Embodiment 5

[0103] In some possible embodiments, aspects of the method of the embodiments of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is used to cause the computer device to execute the steps of a method for generating a unified multi-tasking and multi-domain layout according to various exemplary embodiments described above in this specification. Among them, the executable computer program code or "code" for executing each embodiment can be written in high-level programming languages such as C, C++, C#, Smalltalk, Java, JavaScript, Visual Basic, Structured Query Language (e.g., Transact-SQL), Perl, or in various other programming languages.

[0104] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0105] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0106] The above embodiments are only for illustrating the technical concept and characteristics of the present invention, and their purpose is to enable those of ordinary skill in the art to understand the content of the present invention and implement it accordingly, and should not be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the essence of the content of the present invention should be covered by the protection scope of the present invention.

Claims

1. A method for unified multi-task and multi-domain layout generation, characterized in that: The following steps are involved: Obtain layout generation data in multiple fields; Flatten the information of all elements of each layout into a sequence, randomly mask the sequence as input, and use the complete sequence as the label; The sequences and labels are input into a large language model for training, and layout data from different fields are mixed and used in training; Part of the layout information generated according to different task requirements and different field requirements is input into the trained model to allow the model to generate a complete layout sequence.

2. A method for unified multi-task and multi-domain layout generation according to claim 1, characterized in that: The layout generation data includes paper layout, mobile app UI layout, magazine layout and slide layout; Each complete layout is composed of N layout elements. The format of each layout element is (c, x, y, w, h), where c is the category of the current element, (x, y) is the coordinate of the upper left corner of the current element in the entire layout canvas, w is the width of the current element, and h is the height of the current element. The entire layout is represented by: {x1, y1, w1, h1, ..., x N ,y N ,w N ,h N }.

3. A method for unified multi-task and multi-domain layout generation according to claim 1, characterized in that: The information of all elements of each layout is flattened into a sequence, the sequence is randomly masked as input, and the complete sequence is used as a label, including: Construct any layout prompt template, and use it as the input of the model after masking or adding noise; Construct a unified layout answer template, and use a sequence of complete and noise-free layout element information as a unified layout answer; Positional encoding is done by interval quantization to avoid using placeholders.

4. A method for unified multi-task and multi-domain layout generation according to claim 3, characterized in that: The arbitrary layout prompt template consists of two parts: a prefix part and a main body part; The prefix part consists of "denoising flag", "layout type", "number of elements" and "number of columns". The main part consists of multiple layout elements and the relationship between different elements. The format of each element is (c,x,y,w,h). The information of each element is randomly masked or noised to simulate any layout generation condition; the masked information will be discarded directly, and noise addition refers to adding random noise to x, y, w, and h, requiring the model to denoise the noise after adding noise; the unmasked layout element information is directly spliced ​​together as the input of the model, allowing the model to generate complete layout element information through autoregression.

5. A method for unified multi-task and multi-domain layout generation according to claim 4, characterized in that: The positional encoding is done by interval quantization to avoid the use of placeholders, including: Set a sufficiently large interval value l that is larger than the length and width of all layout canvases; then encode the x, y, w, and h of each layout element according to the following rules: x=x+0×h y=y+1×h w=w+2×h h=h+3×h The encoding rule makes x, y, w, and h each in a different numerical interval, that is, x∈[0,l), y∈[l,2l), w∈[2l,3l), h∈[3l,4l).

6. A method for unified multi-task and multi-domain layout generation according to claim 1, characterized in that: The sequences and labels are input into the large language model for training. Layout data from different fields are mixed and used in the training, including: Use the GPT2-XL model as the large language model and train the model: Any layout prompt template is used as the input of the GPT2-XL model, and a unified layout answer template is used as the label. The model is required to generate a complete layout of a specified field given any task or generation requirement in any field.

7. A method for unified multi-task and multi-domain layout generation according to claim 1, characterized in that: The following steps are also included: The generated layout is compared with the real layout to calculate indicators, or the generated layout sequence is rendered into a two-dimensional layout image.

8. A device for unified multi-task and multi-domain layout generation, characterized in that: include: A data acquisition module, used to acquire layout generation data in multiple fields; The input-output construction module is used to flatten the information of all elements of each layout into a sequence, randomly mask the sequence as input, and use the complete sequence as the label; The model training module is used to input sequences and labels into the large language model for training. Layout data from different fields are mixed and used in training; The model reasoning module is used to input partial layout information generated according to different task requirements and different field requirements into the trained model, allowing the model to generate a complete layout sequence.

9. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Automatic label layout and typesetting method and device, electronic equipment and storage medium

    CN116629201A

  • Interface layout model training method, interface layout method, device and medium

    CN118069134A

  • Layout generation method and device, medium and equipment

    CN118607031A