Layout generation and page generation based on text
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- MICROSOFT TECHNOLOGY LICENSING LLC
- Filing Date
- 2024-06-11
- Publication Date
- 2026-04-22
AI Technical Summary
Layout design for page composition is a complex task that requires professional knowledge, making it difficult for ordinary users to complete independently and reducing efficiency as they often need to spend significant time communicating with professional designers.
A page generation solution that determines the layout of a target page based on input sequences corresponding to text descriptions, allowing for the generation of candidate layouts and improving efficiency and accuracy by constructing input sequences based on layout constraints indicated by text.
This solution enhances page generation efficiency and accuracy by enabling users to create pages independently, improving the layout design process through automated layout generation based on text descriptions.
Smart Images

Figure US2024033329_26122024_PF_FP_ABST
Abstract
Description
LAYOUT GENERATION AND PAGE GENERATION BASED ON TEXTBACKGROUND ART
[0001] In recent years, layout design is an important task in content creation. For example, people need to determine a size, a dimension and other information about each element to complete a desired page layout and such a process is a complex task that depends on a professional design capability.
[0002] This makes it difficult for an ordinary user to complete such layout design independently. Generally, the ordinary user needs to spend a large amount of time in communicating with a professional designer, which greatly reduces efficiency of layout design and page composition.SUMMARY OF THE INVENTION
[0003] According to implementations of the present disclosure, a page generation solution is provided. In this solution, a first portion of a target page to be generated may be provided, wherein a layout of the first portion is determined based on an input sequence corresponding to a first description text; and the input sequence is generated based on a type of at least one layout constraint indicated by the first description text. Further, a second description text for a second portion of the target page may be acquired. Accordingly, a first set of candidate layouts for the second portion may be provided, wherein the first set of candidate layouts are determined at least based on the second description text. Therefore, a process of generating a page by regions can be supported according to the implementations of the present disclosure, thereby improving page generation efficiency. In addition, according to the implementations of the present disclosure, layout generation accuracy can also be improved by constructing an input sequence based on a type of a constraint indicated by text description.
[0004] The SUMMARY section of the present disclosure is provided to introduce identification of concepts in a simplified form, which will be further described below in the DETAILED DESCRIPTION section. The SUMMARY section is not intended to identify key features or main features of the claimed subject matter, or to limit the scope of the claimed subject matter.BRIEF DESCRIPTION OF THE DRAWINGS
[0005] FIG. 1 is a schematic diagram of an example environment according to some implementations of the present disclosure;
[0006] FIG. 2 is a schematic diagram of an example layout generation system for generating a page layout according to some implementations of the present disclosure;
[0007] FIG. 3A and 3B show example page generation interfaces according to some implementations of the present disclosure;
[0008] FIG. 4 is a flowchart of an example page generation method according to someimplementations of the present disclosure; and
[0009] FIG. 5 is a block diagram of an example computing device according to some implementations of the present disclosure.
[0010] In these drawings, the same or similar reference numerals are used to denote the same or similar elements.DETAILED DESCRIPTION
[0011] The present disclosure is now discussed with reference to several example implementations. It should be understood that these implementations are discussed only to enable those of ordinary’ skill in the art to better understand and thus implement the present disclosure, rather than to imply any limitation on the scope of the subject matter.
[0012] As used herein, the term "include" and its variants should be construed as open terms meaning "including, but not limited to". The term "based on" should be construed as "at least partially based on". The terms "one implementation" and "an implementation" should be construed as "at least one implementation". The term "another implementation" should be construed as "at least one other implementation". The terms "first", "second", and the like may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0013] As discussed above, during page composition or design, layout design is a complex task that depends on professional knowledge. Therefore, how to improve layout design efficiency has become the focus of attention.
[0014] According to implementations of the present disclosure, a page generation solution is provided. In this solution, a first portion of a target page to be generated may be provided, wherein a layout of the first portion is determined based on an input sequence corresponding to a first description text; and the input sequence is generated based on a type of at least one layout constraint indicated by the first description text. Further, a second description text for a second portion of the target page may be acquired. Accordingly, a first set of candidate layouts for the second portion may be provided, wherein the first set of candidate layouts are determined at least based on the second description text. Therefore, a process of generating a page by regions can be supported according to the implementations of the present disclosure, thereby improving page generation efficiency. In addition, according to the implementations of the present disclosure, layout generation accuracy can also be improved by constructing an input sequence based on a type of a constraint indicated by text description.
[0015] The basic principle and several example implementations of the present disclosure are described below with reference to the accompanying drawings.
[0016] Example environment
[0017] FIG. 1 shows an example environment 100 according to some implementations of the present disclosure. As shown in FIG. 1, the environment 100 may include an electronic device 110. The electronic device 110 may be, for example, configured to provide an interface 130 for page creation or page generation.
[0018] As shown in FIG. 1, the electronic device 110 may use. for example, the interface 130 to acquire a description text 140 about a page to be generated, for example, "A header". It should be understood that a user may input the description text 140 into the electronic device 110 in an appropriate input manner. For example, the user may type a text with an input device, or input audio information and then convert the same to the description text 140. The present disclosure is not intended to limit an obtaining manner of the description text 140.
[0019] Further, the electronic device 110 may provide a corresponding page layout 150 based on the description text 140. In some implementations, the page layout 150 may be generated by the electronic device 110 based on processing of the description text 140. Alternatively, the electronic device 110 may further send, for example, the description text 140 to an appropriate remote device 120 (for example, a server) and receive, from the remote device 120, the page layout 150 determined by the remote device 120 based on the description text 140.
[0020] A specific generation process of the page layout is described in detail below with reference to FIG. 2.
[0021] Further, the electronic device 110 may use the interface 130 to provide the generated page layout 150 for the user. Such a page layout 150 may indicate, for example, information, such as positions and sizes, of one or more page elements in a page.
[0022] The user may also use, for example, the interface 130 to perform operations such as modification and replacement on the provided page layout 150, and use the interface 130 to generate a final page. A page generation process is described in detail below with reference to FIG. 3A and FIG. 3B.
[0023] Layout generation
[0024] FIG. 2 is a schematic diagram of an example layout generation system 200 for generating a page layout according to some implementations of the present disclosure. In some implementations, the generation system 200 may be implemented by the electronic device 110 shown in FIG. 1, the remote device 120 shown in FIG. 1, and / or a combination thereof.
[0025] As shown in FIG. 2, the layout generation system 200 may include a parsing module 210 and a disposing module 220. The parsing module 210 may be configured to acquire the description text 140 and generate an intermediate representation (IR) 215 for the description text 140.
[0026] In some implementations, the description text 140 may be configured to indicate at leastone layout constraint associated with the page to be generated. For example, the description text "A header" may indicate that the page includes a page element whose type is "header".
[0027] Considering that users' habits about the description text 140 may be different, the parsing module 210 may convert the description text 140 to the intermediate representation 215 expressed in a formal language. The formal language is intended to represent a language defined in a precise mathematical or machine-processable formula. In some implementations, the formal language for representing the intermediate representation may have, for example, a limited character set and a predetemiined grammatical rule.
[0028] The intermediate representation 215 may describe, in the formal language, the at least one layout constraint indicated by the text 140, thereby improving normativity of a representation of the layout constraint and improving accuracy and generation efficiency of a subsequent layout generation process. In some implementations, the parsing module 210 may generate corresponding intermediate representations for description texts 140 of different types.
[0029] In some examples, the description text 140 may indicate a function of a corresponding page element. For example, the description text may include "a news heading" and accordingly, the parsing module 210 may convert the same to an intermediate representation "[e:title]".
[0030] As another example, the description text 140 may further indicate an absolute position and / or a relative position of the corresponding page element. For example, the description text 140 may include "a news heading at the top" and accordingly, the parsing module 210 may convert the same to an intermediate representation " [e: title [prop:position"top"]]".
[0031] As still another example, the description text may further indicate a relative size and / or an absolute size of the corresponding page element. For example, the description text 140 may include "a brief news summary" and accordingly, the parsing module 210 may convert the same to an intermediate representation "[e: description [prop:size"small"]]".
[0032] As yet still another example, the description text may further indicate a hierarchy of the corresponding page element. For example, the description text 140 may include "3 news pieces, each has a title and summary" and accordingly, the parsing module 210 may convert the same to an intermediate representation "[group [prop:repeat"3"] [item[e:title][e:description]]]".
[0033] Based on the above examples, it can be seen that the parsing module 210 may convert different types of description texts into intermediate representations in a structured style. In this way, in one aspect, a degree of freedom of text description input by a user can be improved according to the implementations of the present disclosure, thereby helping the user input text description according to his / her habits. In another aspect, according to the implementations of the present disclosure, the accuracy of subsequent layout generation can also be improved by converting text description expressed in a natural language into structured description.
[0034] In some implementations, the parsing module 210 may be implemented by using an appropriate language model. The language model may include, for example, a transformer-based pre-trained language model (PLM), which may include, for example, an encoder and a decoder.
[0035] Specifically, the encoder of the model may take the description text as an input, which can be expressed asX~. wherein 'l''1denotes an ithtoken in theL > description text. Further, the encoder may convert an input token into an expression11~ f x . . . / J;11 1 ’ ’ n J containing context information. Accordingly, the decoder of the model may autoregressively predict a token of an intermediate representation (IR). This process may be expressed asin a specific example, the decoder of the model may alternatively acquire, for example, the token of the intermediate representation (IR) based on a generated IR tokenand an outputof the encoder.
[0036] In some examples, in a process of training the language model, a parameter 9 of the language model may be trained by minimizing a negative log-likelihood of a labeled IR z of a given text input x. This process may be expressed as:
[0038] It should be understood that other appropriate language models may also be used to implement the parsing module 210. The present disclosure is not intended to limit a specific architecture or specific training process of the language model.
[0039] Further, the disposing module 220 may be configured to determine the page layout 150 based on the received intermediate representation 215. In some implementations, a generation process from the intermediate representation 215 to the page layout 150 may be converted into a sequence-to-sequence conversion task.
[0040] To improve layout generation accuracy, the disposing module 220 may determine, for example, an input sequence based on the intermediate representation 215. In some implementations, the disposing module 220 may determine a corresponding input sequence based on a ty pe of a layout constraint indicated in the intermediate representation 215.
[0041] In some implementations, corresponding layout constraints may be processed differentially based on differences among page elements targeted by the description text 140.
[0042] In an example, the description text 140 may7indicate, for example, a first type of constraint, for example, a layout constraint for a single page element. The layout constraint is also referred to as an element-wise constraint or a point-wise constraint. For example, adescription text "a large text box" is intended to represent a type layout constraint, a position layout constraint, and a size layout constraint about a single page element.
[0043] In another example, the description text 140 may indicate, for example, a second ty pe of constraint, for example, a layout constraint between two page elements. The layout constraint is also referred to as a pair-wise constraint. For example, a description text "an image below the text box" is intended to represent a relative position layout constraint between two page elements.
[0044] In still another example, the description text 140 may indicate, for example, a third type of constraint, for example, a layout constraint about a set of page elements. The layout constraint is also referred to as a group-wise constraint. For example, a description text "a tool bar with 4 icons" is intended to represent a hierarchical layout constraint of a layout.
[0045] Additionally, the disposing module 220 may generate a corresponding input sequence based on a type of a layout constraint.
[0046] In an example, the disposing module 220 may use a specific token k to represent a first type of constraint, whereinX denotes a vocabulary constructed by the type of the layout constraint. For example, a vocabulary of a type layout constraint for a page element, , {image, ■ ■ - . text) „ may be expressed asL J. Further, a plurality ol constraints of the first type for a single page element may be expressed as a token sequence: ( P° _ . L 1I1‘ J. For example, "a large image on the left" may be expressed as "image large left".
[0047] Specifically, the disposing module 220 may identify the first type of constraint from the intermediate representation (IR), and generate an input sequence corresponding to the first type of constraint according to an input sequence representation method corresponding to the first type of constraint.
[0048] By taking an intermediate representation "[e:title [prop:position"top"]]" as an example, the disposing module 220 may determine that the intermediate representation (IR) is a constraint for a single element, which includes a type constraint "title" and a position constraint "top". Therefore, the disposing module 220 may extract corresponding constraint information from the intermediate representation, and compare the constraint information with a vocabulary, thereby generating a corresponding token sequence, for example, "title top".
[0049] In another example, the disposing module 220 may express the second ty pe of constraint as rn three tokens, name 1ly i < wlationship; and 81 andindicate indexes of a source page element and a target page element in a finalinput sequence. For example, "an image below a text box" may be expressed as "text-1 bottom image- 1".
[0050] Specifically, the disposing module 220 may identify the second type of constraint from the intermediate representation (IR), and generate an input sequence corresponding to the second type of constraint according to an input sequence representation method corresponding to the second type of constraint.
[0051] For example, the disposing module 220 may identify, from the intermediate representation (IR), a specific token (for example, bottom) representing a positional relationship, and further determine that the token corresponds to the second type of constraint. Further, the disposing module 220 may extract corresponding page elements and determine index information of the page elements, thereby converting the intermediate representation (IR) into a corresponding input sequence.
[0052] In still another example, the disposing module 220 may use a special token to indicate the third type of constraint. For example, the disposing module 220 may use square brackets to indicate an element group C9PFor example, "a toolbar with 4 icons" may be expressed as " [icon|icon|icon|icon] " .
[0053] Further, the disposing module 220 may use, for example, separators "|" to achieve cascade of subsequences of a plurality' of page elements corresponding to the intermediate representation. This process may be. for example, expressed as:
[0055] Specifically, the disposing module 220 may identify the third type of constraint from the intermediate representation (IR), and generate an input sequence corresponding to the third ty pe of constraint according to an input sequence representation method corresponding to the third fype of constraint.
[0056] By taking an intermediate representation "[group [prop:repeat"3"] [item[e:title][e:description]]]" as an example, the disposing module 220 may determine, by detecting a group-type token (for example, group), that the intermediate representation corresponds to the third type of constraint. Further, the disposing module 220 may extract corresponding page elements from the intermediate representation, and generate an input sequence as shoyvn in Formula (2).
[0057] Therefore, the disposing module 220 may convert the intermediate representation 215 into an input sequence 5 for a subsequent generation process.
[0058] In some implementations, the disposing module 220 may convert the input sequence 5 into an output sequence y to indicate a generated page layout.
[0059] In some implementations, the output sequence y may include an element representation1 1 J? corresponding to a set of page elements, wherein m denotes the quantity of page elements in a layout. In some implementations, an element representation corresponding to each page element may include a plurality’ of tags (also referred to as attributes). - p
[0060] Exemplarily, an element representation "lmay include a flag z to indicate a type of the page element. Additionally or alternatively, the element representation7may include a flag to indicate a left coordinate of the page element. Additionally or alternatively, the element representationmay include a flag ' to indicate a top coordinate of the page element. It should be understood thatmay be used to indicate a position of the page element in the page.
[0061] Additionally or alternatively, the element representation may include a flagto indicate the width of the page element. Additionally or alternatively, the element representation E ■ may include a flag h1to indicate the height of the page element. It should be understood that W1■ and h1• may be used to indicate the size of the page element in the page.
[0062] In some implementations, to help the disposing module 220 complete a page element that is not indicated in the description text 140, the element representationmay include a flagai to indicate whether the corresponding page element is indicated in the description text. For example, a value of the flagmay be {indicate, supplement} to indicate whether the page element is additionally added by the disposing module 220.
[0063] Further, the disposing module 220 may use, for example, separators "|" to achieve cascade of element representations of a plurality of elements. Therefore, the output sequence y for indicating the page layout may be expressed as:
[0065] In some implementations, the disposing module 220 may use an appropriate machine learning model to implement the conversion from the input sequence 5 to the output sequence y. For example, the disposing module 220 may be implemented by using a transformer-based model, which may include an encoder and a decoder.
[0066] In some implementations, as shown in FIG. 2, the machine learning model used for implementing the disposing module 220 may be trained with a training data set 230. In some implementations, the training data set 230 may include, for example, a labeled data portion, which may be. for example, generated by tagging a set of training pages. For example, a set oftraining pages may be tagged by a tagger to generate a corresponding set of description texts. An output sequence corresponding to this set of training pages may be, for example, acquired based on automatic analysis of this set of training pages.
[0067] In this tagging process, the tagger may omit description of some page elements in a page according to his / her personal habits. Therefore, labeled data provided by the tagger may include a description text for a training page. The description text may only correspond to, for example, a subset of all page elements in a corresponding training page. Exemplarily, if a page includes four icons and a top picture, the tagger may provide a description text of "Page including four icons" and omit description about a page element "Top picture".
[0068] The disposing module 220 can leam, by using the training data set 230 to train the machine learning model, a capability of expanding another appropriate page element based on description of limited page elements. For example, in a case that text description input by a user includes only four icons, the disposing module 220 may provide a layout that includes a top picture in addition to the four icons. In this tagging manner, tagging costs of the tagger can also be reduced according to the embodiments of the present disclosure, so that the tagger does not have to pay attention to all page elements in a page, thereby improving tagging efficiency.
[0069] In some implementations, the training data set 230 may further include, for example, an unlabeled data portion. The unlabeled data portion may be, for example, automatically generated based on a set of training pages.
[0070] Specifically, the unlabeled data portion may be generated by an appropriate electronic device. For example, the electronic device may acquire a set of training pages, and generate a training intermediate representation corresponding to the set of training pages based on a page layout of the set of training pages. Exemplarily, the electronic device may first extract page elements in the set of training pages, and further extract a third type of constraint (group-wise constraint), a second type of constraint (pair-wise constraint), and a first type of constraint (point-wise constraint) of the page elements at a time, thereby constructing an intermediate representation corresponding to the set of training pages. In another aspect, the electronic device may further determine a training output sequence corresponding to the set of training pages. Therefore, the electronic device may construct the unlabeled data portion in the training data set 230 based on the training intermediate representation and the training output sequence.
[0071] Based on such a process, according to the embodiments of the present disclosure, labelling costs can be greatly reduced, and labeled data can be generated automatically by using a large quantity of existing pages, thereby7reducing production costs of the labeled data and improving model training efficiency.
[0072] Based on the process discussed above, a page layout can be created automatically basedon the description text according to the embodiments of the present disclosure, thereby improving layout generation efficiency. In addition, according to the embodiments of the present disclosure, layout generation accuracy can be improved by constructing the input sequence based on the type of the layout constraint.
[0073] Page generation
[0074] In some implementations, as discussed with reference to FIG. 1, the electronic device 110 further supports a user in designing or composing a page by using an interface.
[0075] FIG. 3A shows an example interface 300A according to some example embodiments of the present disclosure. As shown in FIG. 3A, the interface 300A may include an input region 302 for a description text, a preview region 304 for a page layout, and a selection region 306 for a candidate layout.
[0076] Exemplarily, the input region 302 may include, for example, a text input box to receive a description text 320 input by a user, for example, "A header". As discussed with reference to FIG. 1. the user may use, for example, any appropriate manner to input the description text 320. The manner includes, but is not limited to, typing the text by using an input device, and transcribing audio based on voice.
[0077] Further, the interface 300A may include, for example, a generation control 325. In a case that triggering the generation control 325 is received, the electronic device 110 may trigger generation and provision of at least one page layout. For a generation process of the page layout, refer to content described above. Details are not described herein again.
[0078] Exemplarily, the electronic device 110 may provide a candidate layout 310 in the preview region. Additionally, the electronic device 110 may further provide one or more other selectable candidate layouts 335 in the selection region 306. The user may click, for example, on a candidate layout 335 to replace a selected candidate layout 310 in the preview region 304.
[0079] The candidate layout 310 and the one or more other candidate layouts 335 may be generated, for example, by using one or more of the layout generation processes discussed above, which may have, for example, certain generation randomness, so that different layout styles are presented.
[0080] Additionally, the user may also trigger, for example, generation of a new set of candidate layouts based on the description text 320. For example, in a case that the user is unsatisfied with the provided candidate layouts 310 and 335, the generation and provision of the new set of candidate layouts can also be tnggered.
[0081] Further, the user may modify the provided candidate layout 310. For example, the user may use the interface 300A to adjust information, such as ty pes, sizes, and positions, of one or more page elements in the candidate layout 310.
[0082] In some implementations, the electronic device 110 may support, for example, the user in generating a page by stages. For example, the description text 320 may be, for example, configured to generate a first portion of the page. After the user determines a page layout of the first portion, the user may continue to, for example, design another portion of the page.
[0083] Exemplarily, as shown in FIG. 3A, the interface 300A may include an add control 330 used for adding a second portion of the page. Upon receiving selection for the add control 330, the electronic device 110 may present, for example, an interface 300B as show n in FIG. 3B.
[0084] As shown in FIG. 3B, the electronic device 110 may receive a new description text 340 input by the user, for example, "A tool bar with four icons". For example, the user may use a text box in the input region 302 to input the new description text 340. An input process of the new description text may be similar to a generation process of the description text 320. The description text 340 may be, for example, a second portion for a page to be generated.
[0085] Further, the electronic device 110 may accordingly provide, in the preview region, a candidate layout 345 of the second portion. Additionally, the electronic device 110 may further provide one or more other candidate layouts 350 in the selection region 306. The user may click, for example, on a candidate layout 355 to replace a selected candidate layout 345 in the preview region 304.
[0086] The candidate layouts 345 and 350 may be generated, for example, by using one or more of the layout generation processes discussed above, which may have, for example, certain generation randomness, so that different layout styles are presented.
[0087] In some implementations, a generation process of the candidate layouts 345 and 350 may be different from that of the candidate layouts 310 and 335, and generation of the candidate layouts 345 and 350 may alternatively be. for example, determined based on layout information associated with the first portion.
[0088] As an example, in a case that the user modifies the candidate layout 310 of the first portion, the electronic device 110 may further acquire final layout information of the first portion. For example, the user may adjust a left coordinate of the leftmost page element and a right coordinate of the rightmost page element in the first portion.
[0089] Accordingly, such layout information may also be provided for a layout generation device to take layout information of an existing portion of the page into consideration. For example, in the generation processes of the candidate layouts 345 and 350 of the second portion, the left and right coordinates of the first portion may be taken into consideration, so that a left coordinate of the leftmost page element in a candidate layout generated for the second portion is aligned with the left coordinate of the first portion, and / or a right coordinate of the rightmost page element in the candidate layout generated for the second portion is aligned with the rightcoordinate of the first portion.
[0090] In some implementations, layout information associated with the first portion may further include another appropriate type of information, such as style information, size information, orientation information, and other positional information. Such layout information can be provided for a generation process to assist in the generation of a page layout of the second portion, thereby improving style consistency of a generated page.
[0091] Similarly, the user may also trigger, for example, generation of a new set of candidate layouts based on the description text 340. For example, in a case that the user is unsatisfied with the provided candidate layouts 345 and 350, the generation and provision of the new set of candidate layouts can also be triggered.
[0092] Further, the user may modify the provided candidate layouts 345 and 350. For example, the user may use the interface 300B to adjust information, such as types, sizes, and positions, of one or more page elements in the candidate layouts 345 and 350.
[0093] Additionally, the user may also use, for example, the add control 330 to add one or more other portions in the page to be generated, and may complete generation or publication of the page based on the added page portions.
[0094] Therefore, a process of generating a page by regions can be supported according to the embodiments of the present disclosure, thereby improving page generation efficiency.
[0095] Example process
[0096] FIG. 4 is a flowchart of an example page generation process 400 according to some implementations of the present disclosure. The process 400 may be implemented, for example, by the electronic device 110 in FIG. 1 or another appropriate device (such as a device 500 discussed with reference to FIG. 5).
[0097] As shown in FIG. 4, at Block 410, the electronic device 110 provides a first portion of a target page to be generated, wherein a layout of the first portion is determined based on an input sequence corresponding to a first description text; and the input sequence is generated based on a type of at least one layout constraint indicated by the first description text.
[0098] At Block 420, the electronic device 1 10 acquires a second description text for a second portion of the target page.
[0099] At Block 430, the electronic device 110 provides a first set of candidate layouts for the second portion, wherein the first set of candidate layouts are determined based on at least the second description text.
[0100] In some implementations, the input sequence is generated by converting an intermediate representation corresponding to the first description text, and the intermediate representation describes, in a formal language, the at least one layout constraint indicated by thefirst description text.
[0101] In some implementations, the type of the at least one constraint includes at least one of:
[0102] a first constraint type indicating an element-wise constraint for a single page element; a second constraint type indicating a pair-wise constraint between two page elements; and a third constraint type indicating a group-wise constraint on a set of page elements.
[0103] In some implementations, providing the first portion of the target page to be generated includes: obtaining an output sequence determined by a target model based on the input sequence, wherein the output sequence indicates distribution of a set of page elements in the first portion; and providing the first portion of the target page based on the output sequence.
[0104] In some implementations, the target model is trained by using a training data set, wherein the training data set includes at least a first data portion, the first data portion including unlabeled data generated based on a first set of training pages.
[0105] In some implementations, the first data portion includes: a first set of training input sequences generated based on a page layout of the first set of training pages; and a first set of training output sequences corresponding to the first set of training pages.
[0106] In some implementations, the training data set further includes a second data portion, the second data portion including labeled data corresponding to a second set of training pages.
[0107] In some implementations, the second data portion includes: a second set of training input sequences generated based on the labeled data, wherein the labeled data includes a set of training description texts corresponding to the second set of training pages; and a second set of training output sequences corresponding to the second set of training pages.
[0108] In some implementations, at least one training input sequence in the second set of training input sequences corresponds to a subset of all page elements in a corresponding training page.
[0109] In some implementations, the output sequence includes an element representation corresponding to a set of page elements, wherein the element representation at least includes a first tag, the first flag indicating whether a corresponding page element is indicated in the description text.
[0110] In some implementations, the element representation further includes at least one of the following items: a second flag indicating a type of the corresponding page element; a third flag indicating a position of the corresponding page element; and a fourth flag indicating a size of the corresponding page element.
[0111] In some implementations, the first set of candidate layouts are further determinedbased on layout information associated with the first portion.
[0112] In some implementations, providing the first portion of the target page to be generated includes: providing a second set of candidate layouts for the first portion, wherein the second set of candidate layouts are determined based on the first description text; and providing the first portion of the target page based on at least a selection for a target layout of the second set of candidate layouts.
[0113] In some implementations, providing the first portion of the target page based on at least the selection for the target layout of the second set of candidate layouts includes: generating the first portion of the target page based on a modification operation for the selected target layout.
[0114] In some implementations, providing the second set of candidate layouts for the first portion includes: providing a third set of candidate layouts determined based on the first description text; and providing the second set of candidate layouts for the first portion based on a regeneration request.
[0115] In some implementations, obtaining the second description text for the second portion of the target page includes: providing, in association with the first portion, an add control for adding the second portion; and obtaining the second description text for the second portion of the target page based on selection for the add control.
[0116] In some implementations, the process 400 further includes: determining the second portion of the target page based on the first set of candidate layouts; and generating the target page at least based on the first portion and the second portion.
[0117] Example device
[0118] FIG. 5 is a schematic block diagram of an example device 500 that can be configured to implement embodiments of the present disclosure. It should be understood that the device 500 shown in FIG. 5 is merely example, and should not constitute any limitation to functions and scopes of the embodiments described in the present disclosure. As shown in FIG. 5, components of the device 500 may include, but are not limited to, one or more processors or processing units 510, a memory 520, a storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560.
[0119] In some implementations, the device 500 may be implemented as various user terminals or service terminals. The service terminals may be servers, large-scale computing devices, and the like that are provided by various service providers. The user terminals may be, for example, any type of mobile terminals, fixed terminals, or portable terminals, including a mobile phone, a multimedia computer, a multimedia tablet, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tabletcomputer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is also foreseeable that the device 500 can support any type of user- oriented interface (such as a "wearable" circuit).
[0120] The processing unit 510 may be an actual or virtual processor and can perform various processing based on programs stored in the memory 520. In a multi-processor system, a plurality of processing units execute computer-executable instructions in parallel to improve a parallel processing capability of the device 500. The processing unit 510 may also be referred to a central processing unit (CPU), a micro-processor, a controller, and a micro-controller.
[0121] The device 500 generally includes a plurality of computer storage media. The media may be any available media accessible to the device 500, which include, but are not limited to, volatile media, non-volatile media, detachable media, and non-detachable media. The memory 520 may be a volatile memory (for example, a register, a cache, or a random access memory (RAM), a non-volatile memory7(for example, a read-only memory' (ROM), an electrically erasable programmable read-only memoiy (EEPROM), or a flash memoiy), or a combination thereof. The memory 520 may include one or more generation modules and / or a program product 525. These program modules are configured to perform functions of layout generation and / or page generation of various embodiments described herein. The program product 525 may be accessed and run by the processing unit 510 to implement a corresponding function. The storage device 530 may be a detachable or non-detachable medium, and may include a machine-readable medium that can be used to store information and / or data and maybe accessed in the device 500.
[0122] Functions of the components of the device 500 may be implemented by using a single computing cluster or a plurality- of computing machines. These computing machines can communicate through communication connections. Therefore, the device 500 may use a logical connection with one or more other servers, a personal computer (PC), or another general network node to operate in a networked environment. The device 500 may also communicate with one or more external devices (not shown) by using the communication unit 540 as required. The external device may be a database, another storage device, a server, a display device, or the like, and communicate with one or more devices that enable a user to interact with the device 500, or communicate with any- device (for example, a network adapter or a modem) that enables the device 500 to communicate with one or more other computing devices. The communication may be performed through an input / output (I / O) interface (not shown).
[0123] The input device 550 may be one or more various input devices, such as a mouse, a keyboard, a trackball, a voice input device, and a camera. The output device 560 may be one or more output devices, such as a display, a loudspeaker, and a printer.
[0124] Example embodiment
[0125] Some example embodiments of the present disclosure are described below.
[0126] According to a first aspect of the present disclosure, a page generation method is provided. The method includes: providing a first portion of a target page to be generated, wherein a layout of the first portion is determined based on an input sequence corresponding to a first description text: and the input sequence is generated based on a type of at least one layout constraint indicated by the first description text; obtaining a second description text for a second portion of the target page; and providing a first set of candidate layouts for the second portion, wherein the first set of candidate layouts are determined based on at least the second description text.
[0127] In some implementations, the input sequence is generated by converting an intermediate representation corresponding to the first description text, and the intermediate representation describes, in a formal language, the at least one layout constraint indicated by the first description text.
[0128] In some implementations, the type of the at least one constraint includes at least one of:
[0129] a first constraint type indicating an element-wise constraint for a single page element; a second constraint ty pe indicating a pair-wise constraint between two page elements; and a third constraint type indicating a group-wise constraint on a set of page elements.
[0130] In some implementations, providing the first portion of the target page to be generated includes: obtaining an output sequence determined by a target model based on the input sequence, wherein the output sequence indicates distribution of a set of page elements in the first portion; and providing the first portion of the target page based on the output sequence.
[0131] In some implementations, the target model is trained by using a training data set; the training data set includes at least a first data portion, the first data portion including unlabeled data generated based on a first set of training pages.
[0132] In some implementations, the first data portion includes: a first set of training input sequences generated based on a page layout of the first set of training pages; and a first set of training output sequences corresponding to the first set of training pages.
[0133] In some implementations, the training data set further includes a second data portion, the second data portion including labeled data corresponding to a second set of training pages.
[0134] In some implementations, the second data portion includes: a second set of training input sequences generated based on the labeled data, wherein the labeled data includes a set of training description texts corresponding to the second set of training pages; and a second set of training output sequences corresponding to the second set of training pages.
[0135] In some implementations, at least one training input sequence in the second set of training input sequences corresponds to a subset of all page elements in a corresponding training page.
[0136] In some implementations, the output sequence includes an element representation corresponding to a set of page elements, wherein the element representation at least includes a first tag, the first flag indicating whether a corresponding page element is indicated in the description text.
[0137] In some implementations, the element representation further includes at least one of the following items: a second flag indicating a type of the corresponding page element; a third flag indicating a position of the corresponding page element; and a fourth flag indicating a size of the corresponding page element.
[0138] In some implementations, the first set of candidate layouts are further determined based on layout information associated with the first portion.
[0139] In some implementations, providing the first portion of the target page to be generated includes: providing a second set of candidate layouts for the first portion, wherein the second set of candidate layouts are determined based on the first description text; and providing the first portion of the target page based on at least a selection for a target layout of the second set of candidate layouts.
[0140] In some implementations, providing the first portion of the target page based on at least the selection for the target layout of the second set of candidate layouts includes: generating the first portion of the target page based on a modification operation for the selected target layout.
[0141] In some implementations, providing the second set of candidate layouts for the first portion includes: providing a third set of candidate layouts determined based on the first description text; and providing the second set of candidate layouts for the first portion based on a regeneration request.
[0142] In some implementations, obtaining the second description text for the second portion of the target page includes: providing, in association with the first portion, an add control for adding the second portion; and obtaining the second description text for the second portion of the target page based on selection for the add control.
[0143] In some implementations, the method further includes: determining the secondportion of the target page based on the first set of candidate layouts; and generating the target page at least based on the first portion and the second portion.
[0144] According to a second aspect of the present disclosure, an electronic device is provided. The device includes: a processing unit; and a memory coupled to the processing unit and containing instructions stored thereon, wherein when executed by the processing unit, the instructions enable the device to perform the following actions: providing a first portion of a target page to be generated, wherein a layout of the first portion is determined based on an input sequence corresponding to a first description text; and the input sequence is generated based on a type of at least one layout constraint indicated by the first description text; obtaining a second description text for a second portion of the target page; and providing a first set of candidate layouts for the second portion, wherein the first set of candidate layouts are determined based on at least the second description text.
[0145] In some implementations, the input sequence is generated by converting an intermediate representation corresponding to the first description text, and the intermediate representation describes, in a formal language, the at least one layout constraint indicated by the first description text.
[0146] In some implementations, the type of the at least one constraint includes at least one of:
[0147] a first constraint type indicating an element-wise constraint for a single page element; a second constraint type indicating a pair- wise constraint between two page elements; and a third constraint type indicating a group-wise constraint on a set of page elements.
[0148] In some implementations, providing the first portion of the target page to be generated includes: obtaining an output sequence determined by a target model based on the input sequence, wherein the output sequence indicates distribution of a set of page elements in the first portion; and providing the first portion of the target page based on the output sequence.
[0149] In some implementations, the target model is trained by using a training data set; the training data set includes at least a first data portion, the first data portion including unlabeled data generated based on a first set of training pages.
[0150] In some implementations, the first data portion includes: a first set of training input sequences generated based on a page layout of the first set of training pages; and a first set of training output sequences corresponding to the first set of training pages.
[0151] In some implementations, the training data set further includes a second data portion, the second data portion including labeled data corresponding to a second set of training pages.
[0152] In some implementations, the second data portion includes: a second set oftraining input sequences generated based on the labeled data, wherein the labeled data includes a set of training description texts corresponding to the second set of training pages; and a second set of training output sequences corresponding to the second set of training pages.
[0153] In some implementations, at least one training input sequence in the second set of training input sequences corresponds to a subset of all page elements in a corresponding training page.
[0154] In some implementations, the output sequence includes an element representation corresponding to a set of page elements, wherein the element representation at least includes a first tag, the first flag indicating whether a corresponding page element is indicated in the description text.
[0155] In some implementations, the element representation further includes at least one of the following items: a second flag indicating a type of the corresponding page element; a third flag indicating a position of the corresponding page element; and a fourth flag indicating a size of the corresponding page element.
[0156] In some implementations, the first set of candidate layouts are further determined based on layout information associated with the first portion.
[0157] In some implementations, providing the first portion of the target page to be generated includes: providing a second set of candidate layouts for the first portion, wherein the second set of candidate layouts are determined based on the first description text; and providing the first portion of the target page based on at least a selection for a target layout of the second set of candidate lay outs.
[0158] In some implementations, providing the first portion of the target page based on at least the selection for the target layout of the second set of candidate layouts includes: generating the first portion of the target page based on a modification operation for the selected target layout.
[0159] In some implementations, providing the second set of candidate layouts for the first portion includes: providing a third set of candidate layouts determined based on the first description text; and providing the second set of candidate layouts for the first portion based on a regeneration request.
[0160] In some implementations, obtaining the second description text for the second portion of the target page includes: providing, in association with the first portion, an add control for adding the second portion; and obtaining the second description text for the second portion of the target page based on selection for the add control.
[0161] In some implementations, the actions further include: determining the second portion of the target page based on the first set of candidate layouts; and generating the targetpage at least based on the first portion and the second portion.
[0162] According to a third aspect, a computer program product is provided. The computer program product is tangibly stored in a non-transitory computer storage medium and contains machine-executable instructions, wherein when executed by a device, the machineexecutable instructions enable the device to perform the following actions: providing a first portion of a target page to be generated, wherein a layout of the first portion is determined based on an input sequence corresponding to a first description text; and the input sequence is generated based on a ty pe of at least one layout constraint indicated by the first description text; obtaining a second description text for a second portion of the target page; and providing a first set of candidate layouts for the second portion, wherein the first set of candidate layouts are determined based on at least the second description text.
[0163] In some implementations, the input sequence is generated by converting an intermediate representation corresponding to the first description text, and the intermediate representation describes, in a formal language, the at least one layout constraint indicated by the first description text.
[0164] In some implementations, the ty pe of the at least one constraint includes at least one of:
[0165] a first constraint type indicating an element-wise constraint for a single page element; a second constraint type indicating a pair-wise constraint between two page elements; and a third constraint type indicating a group-wise constraint on a set of page elements.
[0166] In some implementations, providing the first portion of the target page to be generated includes: obtaining an output sequence determined by a target model based on the input sequence, wherein the output sequence indicates distribution of a set of page elements in the first portion; and providing the first portion of the target page based on the output sequence.
[0167] In some implementations, the target model is trained by using a training data set; the training data set includes at least a first data portion, the first data portion including unlabeled data generated based on a first set of training pages.
[0168] In some implementations, the first data portion includes: a first set of training input sequences generated based on a page layout of the first set of training pages; and a first set of training output sequences corresponding to the first set of training pages.
[0169] In some implementations, the training data set further includes a second data portion, the second data portion including labeled data corresponding to a second set of training pages.
[0170] In some implementations, the second data portion includes: a second set of training input sequences generated based on the labeled data, wherein the labeled data includes aset of training description texts corresponding to the second set of training pages; and a second set of training output sequences corresponding to the second set of training pages.
[0171] In some implementations, at least one training input sequence in the second set of training input sequences corresponds to a subset of all page elements in a corresponding training page.
[0172] In some implementations, the output sequence includes an element representation corresponding to a set of page elements, wherein the element representation at least includes a first tag, the first flag indicating whether a corresponding page element is indicated in the description text.
[0173] In some implementations, the element representation further includes at least one of the following items: a second flag indicating a type of the corresponding page element; a third flag indicating a position of the corresponding page element; and a fourth flag indicating a size of the corresponding page element.
[0174] In some implementations, the first set of candidate layouts are further determined based on layout information associated with the first portion.
[0175] In some implementations, providing the first portion of the target page to be generated includes: providing a second set of candidate layouts for the first portion, wherein the second set of candidate layouts are determined based on the first description text; and providing the first portion of the target page based on at least a selection for a target layout of the second set of candidate layouts.
[0176] In some implementations, providing the first portion of the target page based on at least the selection for the target layout of the second set of candidate layouts includes: generating the first portion of the target page based on a modification operation for the selected target layout.
[0177] In some implementations, providing the second set of candidate layouts for the first portion includes: providing a third set of candidate layouts determined based on the first description text; and providing the second set of candidate layouts for the first portion based on a regeneration request.
[0178] In some implementations, obtaining the second description text for the second portion of the target page includes: providing, in association with the first portion, an add control for adding the second portion; and obtaining the second description text for the second portion of the target page based on selection for the add control.
[0179] In some implementations, the actions further include: determining the second portion of the target page based on the first set of candidate layouts; and generating the target page at least based on the first portion and the second portion.
[0180] The functions described above herein may be at least partially performed by one or more hardware logic units. For example, non-restrictively, example types of hardware logic components that may be used include: a field programmable gate array (FPGA), an applicationspecific integrated circuit (ASIC), an application-specific standard product (ASSP), a system- on-chip (SOC), a complex programmable logic device (CPLD). and the like.
[0181] Program codes for implementing the method of the present disclosure may be written in any combination of one or more programming languages. The program codes may be provided for a processor or controller of a general-purpose computer, a special-purpose computer, or another programmable data processing apparatus, so that functions / operations specified in the flowchart and / or block diagram are implemented when the program codes are executed by the processor or controller. The program codes may be completely executed on a machine, or partially executed on a machine, or may be, as an independent software package, partially executed on a machine and partially executed on a remote machine, or completely executed on a remote machine or server.
[0182] In the context of the present disclosure, the machine-readable medium may be a tangible medium that may contain or store a program used by or used in combination with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination thereof. More specific examples of the machine-readable storage medium may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or a flash memonj. an optical fiber, a portable compact disk read-only memory (CD- ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0183] In addition, although various operations are depicted in a specific order, it should be understood as requiring such operations to be performed in the specific order shown or in a sequential order, or requiring all illustrated operations to be performed to achieve desired results. Under given conditions, multi-task processing and parallel processing may be advantageous. Similarly, although details of several specific implementations are included in the foregoing discussion, these details should not be construed as a limitation on a scope of the present disclosure. Some features described in the context of individual implementations may also be implemented in a single implementations in combination. On the contrary, various features described in the context of a single implementation may also be implemented in a plurality of implementations separately or in any suitable sub-combination.
[0184] Although the subject matter has been described in a language that is specific to structural features and / or logical actions of the method, it should be understood that the subject matter defined in the appended claims is not limited to the specific features or actions described above. On the contrary, the specific features and actions described above are only forms of examples that implement the claims.
Claims
CLAIMS1. A computer-implemented method, comprising: providing a first portion of a target page to be generated, a layout of the first portion being determined based on an input sequence corresponding to a first description text, the input sequence being generated based on a type of at least one layout constraint indicated by the first description text; obtaining a second description text for a second portion of the target page; and providing a first set of candidate layouts for the second portion, the first set of candidate layouts being determined at least based on the second description text.
2. The method according to claim 1, wherein the input sequence is generated by converting an intermediate representation corresponding to the first description text, and the intermediate representation describes, in a formal language, the at least one layout constraint indicated by the first description text.
3. The method according to claim 1, wherein the type of the at least one constraint comprises at least one of: a first constraint type indicating an element-wise constraint for a single page element; a second constraint type indicating a pair-wise constraint between two page elements; and a third constraint type indicating a group- wise constraint on a set of page elements.
4. The method according to claim 1, wherein providing the first portion of the target page to be generated comprises: obtaining an output sequence determined by a target model based on the input sequence, the output sequence indicating distribution of a set of page elements in the first portion; and providing the first portion of the target page based on the output sequence.
5. The method according to claim 4, wherein the target model is trained with a training data set; the training data set comprises at least a first data portion, the first data portion comprising unlabeled data generated based on a first set of training pages.
6. The method according to claim 5, wherein the first data portion comprises: a first set of training input sequences generated based on a page layout of the first set of training pages; and a first set of training output sequences corresponding to the first set of training pages.
7. The method according to claim 5. wherein the training data set further comprises a second data portion, the second data portion comprising labeled data corresponding to a second set of training pages.
8. The method according to claim 7, wherein the second data portion comprises: a second set of training input sequences generated based on the labeled data, wherein thelabeled data comprises a set of training description texts corresponding to the second set of training pages; and a second set of training output sequences corresponding to the second set of training pages.
9. The method according to claim 8, wherein at least one training input sequence in the second set of training input sequences corresponds to a subset of all page elements in a corresponding training page.
10. The method according to claim 4, wherein the output sequence comprises an element representation corresponding to the set of page elements, the element representation at least comprising a first tag, the first flag indicating whether a corresponding page element is indicated in the description text.
11. The method according to claim 10, wherein the element representation further comprises at least one of: a second flag indicating a type of the corresponding page element; a third flag indicating a position of the corresponding page element; and a fourth flag indicating a size of the corresponding page element.
12. The method according to claim 1, wherein the first set of candidate layouts are further determined based on layout information associated with the first portion.
13. The method according to claim 1, wherein providing the first portion of the target page to be generated comprises: providing a second set of candidate layouts for the first portion, the second set of candidate layouts being determined based on the first description text; and providing the first portion of the target page based on at least a selection for a target layout of the second set of candidate layouts.
14. An electronic device, comprising: a processing unit; and a memory being coupled to the processing unit and containing instructions stored thereon, wherein when executed by the processing unit, the instructions enable the electronic device to perform actions comprising: providing a first portion of a target page to be generated, a layout of the first portion being determined based on an input sequence corresponding to a first description text, the input sequence being generated based on a type of at least one layout constraint indicated by the first description text; obtaining a second description text for a second portion of the target page; and providing a first set of candidate layouts for the second portion, the first set of candidate layouts being determined at least based on the second description text.
15. A computer program product being tangibly stored in a non-transitory computer storage medium and comprising machine-executable instructions, wherein when executed by a device, the machine-executable instructions enable the device to perform actions comprising: providing a first portion of a target page to be generated, a layout of the first portion being determined based on an input sequence corresponding to a first description text, the input sequence being generated based on a type of at least one layout constraint indicated by the first description text; obtaining a second description text for a second portion of the target page: and providing a first set of candidate layouts for the second portion, the first set of candidate layouts being determined based on at least the second description text.