Layout information generation method and device, medium, equipment and product
By generating a first code including a mask symbol and using a diffusion language model to predict layout information, the problem of low efficiency in generating element layout information in the prior art is solved, and more efficient layout information generation is achieved.
Patent Information
- Application Number
- CN202511064683.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-10-17
AI Technical Summary
The existing methods for generating element layout information are inefficient and cannot meet the complex and changing creative needs.
By obtaining the description information of the target content display interface, a first code including a mask symbol is generated, the layout information is predicted using a diffusion language model, and denoising is performed to obtain a second code, and finally the layout information is parsed.
The efficiency of layout information generation is improved, the number of prediction steps and the number of tokens that need to be predicted are reduced, and more efficient layout information generation is achieved.
Smart Images

Figure CN120803452A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a layout information generation method, apparatus, medium, device, and product. Background Art
[0002] Layout generation is widely used in scenarios such as page design and document typesetting. The layout of elements affects the display of content. A reasonable element layout can ensure that the rendered element arrangement better meets user needs. Currently, to meet complex and ever-changing creative needs, models are often used to predict element layout information. However, the methods used in related technologies for generating element layout information are inefficient. Summary of the Invention
[0003] This summary is provided to briefly introduce concepts that will be described in detail in the detailed description below. This summary is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0004] In a first aspect, the present disclosure provides a method for generating layout information, the method comprising: Acquire description information of a target content display interface, the description information including element information of elements to be displayed in the target content display interface; generating, according to the description information, a first code for determining layout information of the element, wherein the first code includes a code corresponding to the description information, and the layout information is a mask symbol in the first code; Obtaining a second code through a diffusion language model, wherein the diffusion language model is used to predict the layout information according to the first code, and denoising the mask symbol according to the predicted layout information to obtain the second code; The layout information is parsed from the second code.
[0005] In a second aspect, the present disclosure provides a layout information generating device, the device comprising: An acquisition module, configured to acquire description information of a target content display interface, wherein the description information includes element information of elements to be displayed in the target content display interface; a generating module, configured to generate, based on the description information, a first code for determining layout information of the element, wherein the first code includes a code corresponding to the description information, and the layout information is a mask symbol in the first code; obtaining a second code by a diffusion language model, the diffusion language model being configured to predict the layout information according to the first code, and to denoise the mask symbol according to the predicted layout information, to obtain the second code; parsing the layout information from the second code.
[0006] In a third aspect, the present disclosure provides a computer readable medium having stored thereon a computer program, which, when executed by a processing apparatus, implements the steps of the layout information generation method according to the first aspect of the present disclosure.
[0007] In a fourth aspect, the present disclosure provides an electronic device, comprising: a storage device having stored thereon a computer program; a processing apparatus configured to execute the computer program in the storage device to implement the steps of the layout information generation method according to the first aspect of the present disclosure.
[0008] In a fifth aspect, the present disclosure provides a computer program product comprising a computer program, which, when executed by a processor, implements the steps of the layout information generation method according to the first aspect of the present disclosure.
[0009] According to the above technical solution, the first code for determining the layout information of the element is generated according to the description information of the target content display interface, the first code contains the code corresponding to the description information, and the layout information is a mask symbol in the first code. Then, the second code is obtained by a diffusion language model, the diffusion language model is configured to predict the layout information according to the first code, and to denoise the mask symbol according to the predicted layout information, to obtain the second code. On the one hand, the diffusion language model is used to predict the layout information, the diffusion language model can predict multiple word pieces at each step, the number of required prediction steps is less, and the efficiency of generating the layout information is higher. On the other hand, the code corresponding to the description information is not masked and can be used as known information of the diffusion language model, the diffusion language model does not need to predict and denoise the code corresponding to the description information again, the number of word pieces that need to be predicted is greatly reduced, and the efficiency of generating the layout information is higher.
[0010] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF DRAWINGS
[0011] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail the following specific embodiments thereof with reference to the attached drawings. The same or similar elements are denoted by the same or similar reference numerals throughout the drawings. It is to be understood that the drawings are schematic, and the sizes of the components and elements are not necessarily to scale. In the drawings: Figure 1 is a flowchart of a layout information generation method according to an example embodiment.
[0012] Figure 2A is a schematic diagram of a document, shown by way of example.
[0013] Figure 2B is a schematic diagram of predicted layout information for individual elements, shown.
[0014] Figure 2C is a schematic diagram of annotated layout information for individual elements, shown.
[0015] Figure 3 (a) is a schematic diagram of an actual document image, shown by way of example.
[0016] Figure 3 (b) is a schematic diagram of a document image rendered according to predicted layout information, shown.
[0017] Figure 4 is a block diagram of a layout information generation apparatus according to an example embodiment.
[0018] Figure 5 shows a schematic diagram of the structure of an electronic device suitable for use in implementing embodiments of the present disclosure. DETAILED DESCRIPTION
[0019] Embodiments of the present disclosure will now be described in more detail below with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein; rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of the present disclosure.
[0020] It should be understood that the various steps in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. Additionally, the method embodiments can include additional steps and / or omit performing the steps shown. The scope of the present disclosure is not limited in this regard.
[0021] The term "comprising" and variations thereof as used herein are open-ended, and mean "including but not limited to". The term "based on" means "based, at least in part, on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related terms are defined as follows: "in one embodiment" means "in at least one embodiment"; "in another embodiment" means "in at least one additional embodiment"; "in some embodiments" means "in at least some embodiments".
[0022] It should be noted that the terms "first", "second", and the like in the present disclosure are merely used to distinguish different devices, modules or units, and do not limit the order or interdependence of the functions performed by these devices, modules or units.
[0023] It should be noted that the terms "one", "multiple" in the present disclosure are illustrative rather than limiting, and those skilled in the art should understand that "one or more" should be understood unless the context clearly indicates otherwise.
[0024] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0025] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type, scope of use, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained in a proper manner according to relevant laws and regulations.
[0026] For example, in response to receiving the active request of the user, the user is sent a prompt information to explicitly prompt the user that the operation requested to be performed will require obtaining and using the personal information of the user. Thus, the user can voluntarily choose whether to provide personal information to the electronic device, application program, server or storage medium, etc. software or hardware performing the operation of the technical solutions of the present disclosure according to the prompt information.
[0027] As an optional but non-limiting implementation, in response to receiving the active request of the user, the user can be sent a prompt information in the form of a pop-up window, for example, in which the prompt information can be presented in the form of text. In addition, the pop-up window can also carry a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0028] It can be understood that the above notification and user authorization process is only illustrative and does not limit the implementation of the present disclosure, and other ways that meet the relevant laws and regulations can also be applied to the implementation of the present disclosure.
[0029] At the same time, it can be understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and relevant provisions.
[0030] Figure 1 is a flowchart of a layout information generation method according to an exemplary embodiment. The layout information generation method can be applied in an electronic device with processing capability, such as a terminal or a server, as shown in Figure 1 The layout information generation method includes steps 11 to 14.
[0031] In step 11, description information of the target content display interface is acquired.
[0032] The target content display interface can be an interface for displaying content. For example, the target content display interface can be a document image, and the elements to be displayed in the document image are document elements, such as text elements, table elements, image elements, etc. The target content display interface can also be a target page, which can be a web page, an application page, etc. The elements to be displayed in the target page are page elements, which can be control elements, such as button controls and text input box controls, and component elements, such as card components and list components. That is, the layout information generation method in the present disclosure can be applied to document layout and page design scenarios, and the form of the target content display interface is not limited.
[0033] The description information can include element information of the elements to be displayed in the target content display interface. The element information can include element type information and element content information. The element type information is used to indicate the type of the element, and the element content information can be the specific text content in the element or the content used to describe the element.
[0034] Figure 2A FIG. 1 is a schematic diagram of a document, which is used as an example to illustrate the present disclosure. Figure 2A The document shown in FIG. 1 has text elements to be displayed in the document image, and the specific text content is replaced by X.
[0035] The element type information of the page number 2 is page number, and the element content information is 2. The element type information of the page header starting with A is page header, and the element content information is "AXXXX" (hereinafter referred to as page header A). The element type information of the paragraph starting with B is body, and the element content information is "BX...". (hereinafter referred to as body B). The element type information of the paragraph starting with C is body, and the element content information is "CX...". (hereinafter referred to as body C). The element type information of the paragraph starting with D is body, and the element content information is "DX...". (hereinafter referred to as body D). The element type information of the title starting with E is title, and the element content information is "EXXXX" (hereinafter referred to as title E). The element type information of the sub-title starting with F is sub-title, and the element content information is "FXXX" (hereinafter referred to as sub-title F). The element type information of the paragraph starting with G is body, and the element content information is "GX...". (hereinafter referred to as body G). The element type information of the sub-title starting with H is sub-title, and the element content information is "HX...". (hereinafter referred to as sub-title H). The element type information of the paragraph starting with J is body, and the element content information is "JX...". (hereinafter referred to as body J). The element type information of the sub-title starting with K is sub-title, and the element content information is "KX...". (hereinafter referred to as sub-title K). The element type information of the paragraph starting with L is body, and the element content information is "LX...". (hereinafter referred to as body L). The element type information of the page footer starting with M is page footer, and the element content information is "MX...". (hereinafter referred to as page footer M). The ellipsis in the element content information indicates that the text content therein is omitted.
[0036] Figure 2A Taking the text element as an example, in other embodiments, for example, a table is to be displayed in the document image, the element type information of the element is table, and the element content information can include the number of rows, the number of columns, and the content in each cell of the table. For example, an image is to be displayed in the document image, the element type information of the element is image, and the element content information can be the address of the image. Taking the target page to be generated as an example, a button with content "Confirm" is to be displayed in the target page, the element type information of the page element is button control, and the element content information is "Confirm".
[0037] In addition, the description information can also include size information of the target display interface, which can include the size of the target display interface in the horizontal direction and the size in the vertical direction. The description information can also include the positional relationship between the elements, such as the up-down and front-back relationship, for example, information indicating that body D is located after body C, i.e., the next paragraph, information indicating that sub-title F is the next level of title of title E, and the like.
[0038] In step 12, first code for determining the layout information of the elements is generated according to the description information.
[0039] The first code contains a code corresponding to the description information, and the layout information is a mask symbol in the first code. The layout information includes position information and size information of the element in the target content display interface. The position information can include horizontal and vertical coordinates of the upper left corner of a rectangular frame for displaying the content of the element, or coordinates of any other vertex of the rectangular frame. The size information includes the length of the rectangular frame in the horizontal direction and the length in the vertical direction.
[0040] For example, the element shown in the figure can include the following part of the code: Figure 2A For example, the element shown in the figure can include the following part of the code: <svg width="w1" height="h1"> <rect data-category="header",x= <m> ,y= <m>,width= <m>,height= <m>text=AXXXX> <rect data-category="para",x= <m> ,y= <m>,width= <m>,height= <m>text=DX……> <rect data-category="sec",x= <m> ,y= <m>,width= <m>,height= <m>text=EXXXX> wherein width="wl" indicates that the length of the document image in the horizontal direction is wl, height="hl" indicates that the length of the document image in the vertical direction is hl. "header" indicates that the element type information is a header, "para" indicates that the element type information is a body, and "sec" indicates that the element type information is a section title. x indicates the horizontal coordinate of the upper left corner of the rectangular frame, y indicates the vertical coordinate of the upper left corner of the rectangular frame, width indicates the length of the rectangular frame in the horizontal direction, and height indicates the length of the rectangular frame in the vertical direction, <m>represents a mask symbol, and text represents element content information. As shown in the above code, the layout information of each element is indicated in the first code by a mask symbol <m>The representation is performed, that is, the model needs to predict the layout information of the element.
[0041] For example, the first code can be generated according to a code generation function, and the code generation function is provided with rules for generating the first code according to the description information. For example, header represents a page header, the element type information is a page header, and the element content information is "AXXXX". Through the code generation function, the code corresponding to the page header A can be generated: <rect data-category="header",x= <m> ,y= <m>,width= <m>,height= <m>text=AXXXX> The codes corresponding to other elements are similar.
[0042] It should be noted that the above-mentioned codes are only examples of part of the first code, and the codes corresponding to other elements such as page 2, text B, text D, subheading F, footer M, etc. can refer to the above examples. Among them, the codes in the first code except the mask symbol are the codes corresponding to the description information, and the codes corresponding to the description information are used to embody the element type information and the element content information, and can also be used to embody other information in the description information such as the size information of the target content display interface, etc.
[0043] In addition, the first code can also include prompt information, which can be used to prompt the model to generate the layout information of the generated document style, and to generate the layout information according to the element type information and the element content information of the provided elements.
[0044] The type of the first code is not limited, and the first code is an HTML (Hyper Text Markup Language) code in the example, which is only an implementation manner.
[0045] Step 13, obtaining the second code by the diffusion language model, the diffusion language model is used to predict the layout information according to the first code, and to denoise the mask symbol according to the predicted layout information, so as to obtain the second code.
[0046] Step 14, parsing the layout information from the second code.
[0047] In the present disclosure, the diffusion language model (Diffusion Language Models, DLM) is used for prediction of the layout information.
[0048] The autoregressive modeling (ARM) based large language model has some problems. The ARM predicts the next token according to all the previously output tokens, and the inference direction is a one-way generation process from left to right, and only one token can be predicted at a time. In the related art, the autoregressive modeling based large language model is used to predict the layout information, and only one token can be predicted and output at each step. The token-by-token output method is low in efficiency. Moreover, for the known part of the code, such as the part of the code (e.g., data-category="header") used to indicate element type information, only the autoregressive modeling based large language model can be used as a prompt word. The large language model also needs to predict and output tokens one by one. In this way, the number of tokens that the model needs to predict is large, and only one token can be predicted at a time, so that the efficiency of layout information generation is low.
[0049] The diffusion language model adopts a non-autoregressive paradigm, can divide the text into multiple blocks, can process multiple blocks at the same time, and can predict all the tokens that need to be predicted in the block at each time to achieve the denoising goal. That is, the diffusion language model can predict multiple tokens at each step, and the inference direction is no longer limited to a one-way process from left to right. The diffusion language model is a generative language model, and its core mechanism includes a forward process and a reverse process. In the forward process, the input sequence is randomly masked step by step until it is completely masked. In the reverse process, the complete sequence is denoised from the completely masked sequence to recover the predicted complete sequence. In the related art, the diffusion language model needs to recover the complete sequence from the completely masked sequence, that is, all the tokens in the sequence are <m>The code containing the layout information of the elements is recovered in the sequence, so that the diffusion language model needs to predict and denoise all the sequences.
[0050] In the present disclosure, on the one hand, the diffusion language model is used for predicting the layout information, the diffusion language model can predict multiple tokens at each step, the number of required prediction steps is less, and the efficiency of generating the layout information is higher.
[0051] On the other hand, the first code of the present disclosure contains code corresponding to the description information, and the layout information is a mask symbol in the first code. The code corresponding to the description information is not masked and can be used as known information of the diffusion language model. The diffusion language model does not need to predict the code corresponding to the description information again, but only needs to predict the layout information. The mask symbol in the first code is denoised according to the predicted layout information. The original mask symbol is processed into the predicted layout information after denoising, thereby obtaining the second code. In this way, the diffusion language model only needs to denoise the mask symbol in the first code, and does not need to predict and denoise the code corresponding to the description information as known information. Therefore, the number of tokens that need to be predicted by the diffusion language model is greatly reduced, thereby improving the efficiency of generating the layout information.
[0052] In addition, the code corresponding to the description information in the present disclosure contains element content information, and the diffusion language model can use the element content information as a prediction basis. The area of the region required by the element content in the target content display interface is small when the element content is small, and the area of the region required by the element content in the target content display interface is large when the element content is large, thereby enabling more accurate prediction of the layout information of the elements.
[0053] Figure 2B is a schematic diagram of the predicted layout information of each element, as Figure 2B As shown, the position information and size information of the rectangular frame 202 are layout information of the predicted page number 2. The position information and size information of the rectangular frame 203 are layout information of the predicted page header A. The position information and size information of the rectangular frame 204 are layout information of the predicted body text B. The position information and size information of the rectangular frame 205 are layout information of the predicted body text C. The position information and size information of the rectangular frame 206 are layout information of the predicted body text D. The position information and size information of the rectangular frame 207 are layout information of the predicted title E. The position information and size information of the rectangular frame 208 are layout information of the predicted sub-title F. The position information and size information of the rectangular frame 209 are layout information of the predicted body text G. The position information and size information of the rectangular frame 210 are layout information of the predicted sub-title H. The position information and size information of the rectangular frame 211 are layout information of the predicted body text J. The position information and size information of the rectangular frame 212 are layout information of the predicted sub-title K. The position information and size information of the rectangular frame 213 are layout information of the predicted body text L. The position information and size information of the rectangular frame 214 are layout information of the predicted page footer M.
[0054] For example, the layout information of the predicted body text D includes the horizontal coordinate x3 and the vertical coordinate y4 of the point 2061 in the document image 201, and the length w3 of the rectangular frame 206 in the horizontal direction and the length h3 of the rectangular frame 206 in the vertical direction.
[0055] For example, the layout information of the predicted page header A includes the horizontal coordinate x2 and the vertical coordinate y2 of the top-left corner of the rectangular frame 203, and the length w2 of the rectangular frame 203 in the horizontal direction and the length h2 of the rectangular frame 203 in the vertical direction. The layout information of the predicted title E includes the horizontal coordinate x4 and the vertical coordinate y4 of the top-left corner of the rectangular frame 207, and the length w4 of the rectangular frame 207 in the horizontal direction and the length h4 of the rectangular frame 207 in the vertical direction.
[0056] For example, according to the predicted layout information, the mask symbols in the first code are denoised to generate a second code, and an example of part of the second code is as follows: <svg width="w1" height="h1"> <rect data-category="header" ,x="x2,y=y2,width=w2,height=h2,text=AXXXX"> <rect data-category="para" ,x="x3,y=y3,width=w3,height=h3,text=DX……"> <rect data-category="sec" ,x="x4,y=y4,width=w4,height=h4,text=EXXXX"> From the value of x, the value of y, the value of width, the value of height in the second code, the layout information of each element can be parsed.
[0057] According to the description information of the target content display interface, the first code for determining the layout information of the element is generated, the first code contains the code corresponding to the description information, and the layout information is a mask symbol in the first code. Then, the second code is obtained through the diffusion language model, the diffusion language model is used to predict the layout information according to the first code, and the mask symbol is denoised according to the predicted layout information to obtain the second code. On the one hand, the diffusion language model is used to predict the layout information, and the diffusion language model can predict multiple word pieces at each step, the number of required prediction steps is less, and the efficiency of generating the layout information is higher. On the other hand, the code corresponding to the description information is not masked and can be used as known information of the diffusion language model, the diffusion language model does not need to predict and denoise the code corresponding to the description information, the number of word pieces that need to be predicted is greatly reduced, and the efficiency of generating the layout information is higher.
[0058] In the present disclosure, the implementation of obtaining the second code through the diffusion language model can be: The first code and the predicted position information are input into the diffusion language model to obtain the second code output by the diffusion language model, wherein the predicted position information is used to indicate that the layout information needs to be predicted and the code corresponding to the description information does not need to be predicted.
[0059] For example, the predicted position information can be a sequence containing 0 and 1 with the same number of characters as the first code, wherein the position corresponding to the mask symbol in the first code in the sequence is 1, indicating that the layout information needs to be predicted, and the position corresponding to the code of the description information in the sequence is 0, indicating that the code corresponding to the description information does not need to be predicted.
[0060] In this way, the code corresponding to the description information is used as known information of the diffusion language model, the predicted position information is used to indicate that the code corresponding to the description information does not need to be predicted, the number of word pieces that need to be predicted by the diffusion language model is greatly reduced, and the efficiency of generating the layout information is higher.
[0061] In the present disclosure, the diffusion language model can be trained in the following way: The layout information of the element in the sample content display interface is predicted by the diffusion language model to be trained; According to the difference information between the predicted layout information of the element in the sample content display interface and the labeled layout information of the element in the sample content display interface, the parameters of the diffusion language model to be trained are adjusted; When the training stop condition is met, a trained diffusion language model is obtained.
[0062] The diffusion language model to be trained can be an untrained diffusion language model or a diffusion language model that has been preliminarily trained. There are multiple sample content display interfaces.
[0063] The sample content display interface can be a sample document image, the elements of which are document elements, and the sample content interface can also be a sample page, the elements of which are page elements. Since the layout of a page and the layout of a document have certain differences, in order to enable the diffusion language model to better learn the layout information therein, during training, sample document images can be used for training in one training process, or sample pages can be used for training. In an optional embodiment, two diffusion language models can also be trained, one for predicting the layout information of a document and the other for predicting the layout information of a page. In step 13, if the target content display interface is a document image, the diffusion language model trained for predicting the layout information of a document is used to predict the layout information. If the target content display interface is a target page, the diffusion language model trained for predicting the layout information of a page is used to predict the layout information.
[0064] The description information of the sample content display interface may include element information of the elements in the sample content display interface, including element type information and element content information of the elements in the sample content display interface. In addition, the description information of the sample content display interface may also include size information of the sample content display interface, etc.
[0065] Based on the description information of the sample content display interface, a first sample code can be generated. The method for generating the first sample code can be referred to the method for generating the first code in step 12. The first sample code and the sample predicted position information are input into the diffusion language model to be trained, and a second sample code output by the diffusion language model is obtained. The sample predicted position information is used to indicate that the layout information of the elements in the sample content display interface needs to be predicted, and the code corresponding to the description information of the sample content display interface does not need to be predicted. The sample predicted position information can be a sequence containing 0s and 1s with the same number of characters as the first sample code. The layout information of the elements in the sample content display interface predicted by the diffusion language model to be trained can be parsed from the second sample code.
[0066] Assume that Figure 2A The document shown is a sample document. Figure 2C is a schematic diagram showing the layout information of each element of the annotation, such as Figure 2C As shown, the position information and size information of the rectangular box 217 are layout information of the annotated page number 2. The position information and size information of the rectangular box 216 are layout information of the annotated header A. The position information and size information of the rectangular box 218 are layout information of the annotated body text B. The position information and size information of the rectangular box 219 are layout information of the annotated body text C. The position information and size information of the rectangular box 220 are layout information of the annotated body text D. The position information and size information of the rectangular box 221 are layout information of the annotated title E. The position information and size information of the rectangular box 222 are layout information of the annotated sub-title F. The position information and size information of the rectangular box 223 are layout information of the annotated body text G. The position information and size information of the rectangular box 224 are layout information of the annotated sub-title H. The position information and size information of the rectangular box 225 are layout information of the annotated body text J. The position information and size information of the rectangular box 226 are layout information of the annotated sub-title K. The position information and size information of the rectangular box 227 are layout information of the annotated body text L. The position information and size information of the rectangular box 228 are layout information of the annotated footer M.
[0067] Taking the layout information of the annotated body text D as an example, the layout information of the annotated body text D can include the horizontal coordinate x5 and the vertical coordinate y5 of the point 2201 in the document image 215, and the length w5 of the rectangular box 220 in the horizontal direction and the length h5 in the vertical direction. The document image 215 is displayed as a sample content interface.
[0068] Suppose that the layout information of each element shown in FIG. 2 is predicted by the diffusion language model to be trained as the layout information, Figure 2B The document image 201 in FIG. 2 has the same size as the document image 215 in FIG. 3, and taking the body text D as an example, the difference information between the predicted layout information and the annotated layout information can include difference information between the horizontal coordinate x3 and the horizontal coordinate x5, difference information between the vertical coordinate y3 and the vertical coordinate y5, difference information between the length w3 and the length w5, and difference information between the length h3 and the length h5. The difference information corresponding to other elements is the same. Figure 2B Figure 2C The manner of adjusting the parameters of the diffusion language model to be trained according to the difference information can refer to related technologies, for example, determining a loss value according to the difference information and a preset loss function, and adjusting the parameters of the model according to the loss value. The parameter adjustment can be in the form of fine-tuning part of the parameters or adjusting all the parameters.
[0069] The manner of adjusting the parameters of the diffusion language model to be trained according to the difference information can refer to related technologies, for example, determining a loss value according to the difference information and a preset loss function, and adjusting the parameters of the model according to the loss value. The parameter adjustment can be in the form of fine-tuning part of the parameters or adjusting all the parameters.
[0070] The training stop condition can be preset, for example, the training iteration round reaches a preset number of times, and it is considered that the training stop condition is met, or the difference information between the predicted layout information and the labeled layout information is less than a preset threshold, and it is considered that the training stop condition is met. In the case where the training stop condition is met, the trained diffusion language model can be obtained.
[0071] In addition, in the training process, the sample predicted position information can make the trained diffusion language model more efficient in predicting the layout information, thereby improving the training efficiency of the model. In other embodiments, the trained diffusion language model can also be based on the first sample code of the full mask to predict the layout information.
[0072] In an embodiment, the element information of the element to be displayed in the target content display interface can be obtained by inputting the prompt word into the content generation model, and the content generation model generates and outputs the element information, wherein the prompt word is used to guide the content generation model to generate element content information and label element type information of the element.
[0073] The content generation model can be an arbitrary generative model, such as a large language model. The prompt word can be a natural language form of prompt information, for example, "Please generate an article about ZZ, and each paragraph of text needs to indicate the corresponding text type, such as article title, chapter title, subheading, author, body, header, footer, etc." ZZ is the theme of the article required by the user.
[0074] In this way, the element information of the element to be displayed in the target content display interface can be obtained by the content generation model, and the content generated by the content generation model can be laid out. The content generation model can generate element information of document elements, or can generate element information of page elements, for example, the prompt word includes a login page that needs to be designed, and the element information generated and output by the content generation model can include element information of elements such as control account input box, control password input box, and control confirmation button.
[0075] In an embodiment, the layout information generation method can further include: According to the element information and the layout information, the content is rendered in the target content display interface, wherein the element is a document element, and the target content display interface is a document image, or the element is a page element, and the target content display interface is a target page.
[0076] If the element is a document element and the target content display interface is a document image, the element content information can be rendered into a rectangular frame corresponding to the predicted layout information. If the element is a page element and the target content display interface is a target page, code for rendering the page can be generated based on the element information and layout information, and the target page can be rendered based on the code.
[0077] Figure 3 (a) is a schematic diagram showing an actual document image as an example. Figure 3 (b) is a schematic diagram showing a document image rendered according to the predicted layout information, where X is used to replace the specific text content. Figure 3 In the case of the text content and text type in the document shown in (a), according to the layout information generation method disclosed in the present invention, it is possible to generate Figure 3 (b) The document image shown.
[0078] It should be noted that Figure 3 (a) and Figure 3 The rectangular boxes in (b) are only used to reflect the layout of the document elements. The actual document image and the document image rendered based on the predicted layout information do not contain the rectangular boxes.
[0079] The above embodiment takes the element as a document element and the target content display interface as a document image as an example to introduce the implementation method. This is only for explanation. The layout information generation method disclosed in the present invention is also applicable to the layout design of the target page.
[0080] It is worth noting that, referring to Figure 2B , based on the predicted layout information, the document is divided into two columns, referring to Figure 3 (b) According to the predicted layout information, the document is in a single column format. During the training process of the diffusion language model, the sample document images used may include sample document images in a column format. For example, when text content is relatively large, it can be displayed in columns. During the training process, the diffusion language model can learn the column layout. Thus, when predicting the layout information of elements in the target content display interface, if the target content display interface contains a large amount of text, the predicted layout information may be in a column format. Similarly, the sample document images used may include sample document images in a single column format. During the training process, the diffusion language model can also learn such document layouts.
[0081] In addition, for document images, if there is a need for column division, it can also be reflected in the description information of the target content display interface, that is, the description information may include the structural information of the document image, such as information indicating that it is divided into two columns or three columns, which can be used as the basis for the diffusion language model to predict layout information.
[0082] Based on the same inventive concept, the disclosure also provides a layout information generation apparatus, Figure 4 FIG. 1 is a block diagram of a layout information generation apparatus according to an example embodiment, as shown, the layout information generation apparatus 40 comprises: Figure 4 An obtaining module 43 is configured to obtain a second code by using a diffusion language model, the diffusion language model being configured to predict the layout information according to the first code, and to denoise the mask symbol according to the predicted layout information, so as to obtain the second code. A parsing module 44 is configured to parse the layout information from the second code. Optionally, the obtaining module 43 comprises: A sub-module is configured to input the first code and prediction position information into the diffusion language model, so as to obtain the second code output by the diffusion language model, wherein the prediction position information is used to indicate that the layout information needs to be predicted and the code corresponding to the description information does not need to be predicted.
[0083] Optionally, the element information is obtained by the following module: An information input module is configured to input a prompt word into a content generation model, and to generate and output the element information by using the content generation model, wherein the prompt word is used to guide the content generation model to generate the element content information and to label the element type information of the element.
[0084] Optionally, the diffusion language model is obtained by training the following module: A prediction module is configured to predict the layout information of elements in a sample content display interface by using a diffusion language model to be trained.
[0085] A training module is configured to adjust parameters of the diffusion language model to be trained according to difference information between the predicted layout information of elements in the sample content display interface and labeled layout information of elements in the sample content display interface. A model obtaining module is configured to obtain the trained diffusion language model when a training stop condition is met.
[0086] Optionally, the apparatus 40 further comprises: a rendering module, configured to perform content rendering on the target content display interface according to the element information and the layout information, wherein the element is a document element and the target content display interface is a document image, or the element is a page element and the target content display interface is a target page.
[0087] With regard to the apparatus in the above-described embodiments, specific manners in which various modules perform operations have been described in detail in the embodiments of the method, and thus will not be described here in detail.
[0088] Reference is made below to Figure 5 which shows a structural schematic diagram of an electronic device 600 suitable for use in implementing embodiments of the present disclosure. The terminal device in embodiments of the present disclosure can include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet PCs), PMPs (Portable Multimedia Players), vehicle-mounted terminals (e.g., vehicle-mounted navigation terminals), and the like, as well as fixed terminals such as digital TVs, desktop computers, and the like. Figure 5 The electronic device shown is merely an example and should not impose any limitation on the functions and use range of embodiments of the present disclosure.
[0089] As shown in Figure 5 , the electronic device 600 can include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601 that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 602 or loaded into a random access memory (RAM) 603 from a storage device 608. Various programs and data required for operation of the electronic device 600 are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0090] Generally, the following devices can be connected to the I / O interface 605: input devices 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; output devices 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; storage devices 608 including, for example, a magnetic tape, a hard disk, and the like; and communication devices 609. The communication devices 609 can allow the electronic device 600 to communicate with other devices wirelessly or via wires to exchange data. Although Figure 5 The electronic device 600 is shown with various devices, but it should be understood that all of the shown devices are not required to be implemented or present. More or fewer devices can alternatively be implemented or present.
[0091] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program comprising program code for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the methods of embodiments of the present disclosure are performed.
[0092] It should be noted that the computer-readable medium described above in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination thereof. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in baseband or propagated as a carrier wave in a propagated data signal, in which the computer-readable program code is carried. Such a propagated data signal can take a variety of forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium that can be used to carry or store program code for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to wire, cable, RF (radio frequency), or any suitable combination thereof.
[0093] In some embodiments, the terminals, servers can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communications (e.g., communications networks) of any form or medium, such as the Internet. Examples of communications networks include local area networks ("LANs"), wide area networks ("WANs"), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future developed networks.
[0094] The computer readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device and be not assembled into the electronic device.
[0095] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: acquire description information of a target content display interface, the description information including element information of an element to be displayed in the target content display interface; According to the description information, generate first code for determining layout information of the element, the first code containing code corresponding to the description information, and the layout information being a mask symbol in the first code; Obtain second code through a diffusion language model, the diffusion language model being used for predicting the layout information according to the first code, and performing denoising processing on the mask symbol according to the predicted layout information, to obtain the second code; Parse the layout information from the second code.
[0096] Computer program code for carrying out operations of the present disclosure can be written in one or more programming languages or combinations of languages including object oriented programming languages such as Java, Smalltalk, C++ or conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network ("LAN") or a wide area network ("WAN"), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0097] The flow and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow and block diagrams can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0098] The modules involved in the embodiments of the present disclosure can be implemented in the form of software or in the form of hardware. In some cases, the name of the module does not constitute a limitation on the module itself, for example, the acquisition module can also be described as "a module for acquiring description information".
[0099] The functions described above in the present document can be performed, at least in part, by one or more hardware logic components. For example, non-limiting examples of exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SOCs), complex programmable logic devices (CPLDs), etc.
[0100] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0101] According to one or more embodiments of the present disclosure, example 1 provides a layout information generation method, the method comprising: obtaining description information of a target content display interface, the description information including element information of an element to be displayed in the target content display interface; generating, according to the description information, first code for determining layout information of the element, the first code containing code corresponding to the description information, and the layout information being a mask symbol in the first code; obtaining second code through a diffusion language model, the diffusion language model being used to predict the layout information according to the first code and to denoise the mask symbol according to the predicted layout information, so as to obtain the second code; parsing the layout information from the second code.
[0102] According to one or more embodiments of the present disclosure, example 2 provides the method of example 1, and the obtaining second code through a diffusion language model comprises: inputting the first code and prediction position information into the diffusion language model to obtain the second code output by the diffusion language model, wherein the prediction position information is used to indicate that the layout information needs to be predicted and the code corresponding to the description information does not need to be predicted.
[0103] According to one or more embodiments of the present disclosure, example 3 provides the method of example 1, and the element information includes element type information and element content information.
[0104] According to one or more embodiments of the present disclosure, example 4 provides the method of example 3, and the element information is obtained by: inputting a prompt word into a content generation model to generate and output the element information by the content generation model, wherein the prompt word is used to guide the content generation model to generate the element content information and label the element type information of the element.
[0105] According to one or more embodiments of the present disclosure, example 5 provides the method of example 1, and the diffusion language model is obtained by training in the following manner: predicting, by a to-be-trained diffusion language model, layout information of an element in a sample content display interface; adjusting parameters of the to-be-trained diffusion language model according to difference information between the predicted layout information of the element in the sample content display interface and labeled layout information of the element in the sample content display interface; obtaining the trained diffusion language model in the case of meeting a training stop condition.
[0106] According to one or more embodiments of the present disclosure, example 6 provides the method of example 1, and the method further comprises: According to the element information and the layout information, content rendering is performed on the target content display interface, wherein the element is a document element, and the target content display interface is a document image, or the element is a page element, and the target content display interface is a target page.
[0107] According to one or more embodiments of the present disclosure, example 7 provides the method of example 1, wherein the layout information comprises position information and size information of the element in the target content display interface.
[0108] According to one or more embodiments of the present disclosure, example 8 provides a layout information generation apparatus, comprising: An obtaining module configured to obtain description information of a target content display interface, the description information comprising element information of an element to be displayed in the target content display interface; A generating module configured to generate, according to the description information, a first code for determining layout information of the element, the first code comprising a code corresponding to the description information, and the layout information being a mask symbol in the first code; An obtaining module configured to obtain a second code by using a diffusion language model, the diffusion language model being configured to predict the layout information according to the first code, and to perform denoising processing on the mask symbol according to the predicted layout information, so as to obtain the second code; An analyzing module configured to analyze the layout information from the second code.
[0109] According to one or more embodiments of the present disclosure, example 9 provides a computer readable medium having stored thereon a computer program, which, when executed by a processing apparatus, implements the steps of the method of any one of examples 1-7.
[0110] According to one or more embodiments of the present disclosure, example 10 provides an electronic device, comprising: A storage device having stored thereon a computer program; A processing apparatus configured to execute the computer program in the storage device to implement the steps of the method of any one of examples 1-7.
[0111] According to one or more embodiments of the present disclosure, example 11 provides a computer program product comprising a computer program, which, when executed by a processor, implements the steps of the method of any one of examples 1-7.
[0112] The above description merely illustrates the preferred embodiment of the disclosure and a principle of applied technologies. It should be understood by those skilled in the art that the disclosed range of the disclosure is not limited to the technical solutions formed by the specific combinations of the technical features described above, and should also cover other technical solutions formed by the combinations of the technical features described above or their equivalent features without departing from the disclosed concept. For example, the technical solutions formed by the mutual replacement of the above-described features and the technical features with similar functions disclosed in the disclosure (but not limited to) can be formed.
[0113] Furthermore, although operations are depicted in a particular, sequential order, this should not be understood as requiring or implying that the operations are performed in the order illustrated or sequentially. In certain circumstances, multitasking and parallel processing can be advantageous. Likewise, although specific implementation details are contained in the above discussion, these should not be construed as limiting the scope of the disclosure. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.
[0114] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely illustrative of the example forms of implementing the claims. As to the means for performing the operations of the apparatus in the above-described embodiments, the specific manner in which the various modules perform the operations has been described in detail in the embodiments related to the method, and will not be described here in detail.< / rect> < / rect> < / rect> < / svg> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / svg>
Claims
1. A layout information generation method, characterized in that: The method comprises: Acquire description information of a target content display interface, the description information including element information of elements to be displayed in the target content display interface; generating, according to the description information, a first code for determining layout information of the element, wherein the first code includes a code corresponding to the description information, and the layout information is a mask symbol in the first code; Obtaining a second code through a diffusion language model, wherein the diffusion language model is used to predict the layout information according to the first code, and denoising the mask symbol according to the predicted layout information to obtain the second code; The layout information is parsed from the second code.
2. The method according to claim 1, characterized in that The obtaining of the second code by using the diffusion language model includes: The first code and the predicted position information are input into the diffusion language model to obtain the second code output by the diffusion language model, wherein the predicted position information is used to indicate that the layout information needs to be predicted and the code corresponding to the description information does not need to be predicted.
3. The method according to claim 1, characterized in that The element information includes element type information and element content information.
4. The method according to claim 3, characterized in that The element information is obtained in the following way: The prompt word is input into the content generation model, and the content generation model generates and outputs the element information, wherein the prompt word is used to guide the content generation model to generate the element content information and mark the element type information of the element.
5. The method according to claim 1, wherein The diffusion language model is trained in the following way: Predicting the layout information of elements in the sample content display interface through the diffusion language model to be trained; adjusting parameters of a diffusion language model to be trained based on difference information between the predicted layout information of elements in the sample content display interface and the marked layout information of elements in the sample content display interface; When the training stop condition is met, the diffusion language model after training is obtained.
6. The method according to claim 1, characterized in that The method further comprises: Content rendering is performed on the target content display interface according to the element information and the layout information, wherein the element is a document element and the target content display interface is a document image, or the element is a page element and the target content display interface is a target page.
7. The method according to claim 1, characterized in that The layout information includes position information and size information of the element in the target content display interface.
8. A layout information generating device, characterized in that: The device comprises: An acquisition module, configured to acquire description information of a target content display interface, wherein the description information includes element information of elements to be displayed in the target content display interface; a generating module, configured to generate, based on the description information, a first code for determining layout information of the element, wherein the first code includes a code corresponding to the description information, and the layout information is a mask symbol in the first code; an obtaining module, configured to obtain a second code by using a diffusion language model, wherein the diffusion language model is configured to predict the layout information according to the first code, and to perform denoising processing on the mask symbol according to the predicted layout information to obtain the second code; A parsing module is used to parse the second code to obtain the layout information.
9. A computer-readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processing device, the steps of the method according to any one of claims 1 to 7 are implemented.
10. An electronic device, characterized in that: include: a storage device having a computer program stored thereon; A processing device, configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1 to 7.
11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.