Layout method and device for multi-modal content generated by model, equipment and medium

By acquiring the structural information of multimodal content, breaking it down into image and text combinations, and optimizing the layout based on matching degree, the structural and readability issues in multimodal content layout are resolved, improving the rationality and aesthetics of the layout.

CN121414907APending Publication Date: 2026-01-27BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511525439.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

In existing technologies, the layout of multimodal content lacks structure and readability, resulting in the separation of text and image content, which disrupts the logical hierarchy, leads to a poor reading experience, low space utilization, and poor adaptability.

Method used

By acquiring the content structure information of multimodal content, it is broken down into image and text combinations. A reasonable layout method is determined based on the matching degree between the candidate layout method and the image and text combination, and then merged and optimized to achieve cross-level and cross-structure image and text layout.

Benefits of technology

It achieves a clear and visually appealing multimodal content layout, improving the user's reading experience and space utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121414907A_ABST
    Figure CN121414907A_ABST
Patent Text Reader

Abstract

One or more embodiments of the present disclosure provide a layout method and apparatus for model-generated multi-modal content, a device and a medium, the layout method for model-generated multi-modal content comprising: obtaining content structure information of a first content, the content structure information being used for representing at least one image-text combination, each image-text combination in the image-text combinations comprises a text content and an image content which have an association relationship; for each image-text combination, determining a first layout mode corresponding to the image-text combination according to the matching degree between the plurality of candidate layout modes and the image-text combination; and according to a first layout mode corresponding to the image-text combinations, merging at least part of the image-text combinations in the image-text combinations, and determining a layout mode of the first content. A logic structure of multi-modal heterogeneous data content is converted into a plurality of image-text combinations, image-text layout is carried out by taking the image-text combinations as units, and then cross-level and cross-structure combination is carried out to carry out overall optimization, so that image-text layout with a clear structure is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this disclosure relate to a method for laying out model-generated multimodal content, a means for laying out model-generated multimodal content, an electronic device, and a computer-readable storage medium. Background Technology

[0002] With the continuous development of artificial intelligence generated content (AIGC) technology, AI models are able to automatically generate multimodal content. For example, AI models can generate rich text content and image content related to the text content.

[0003] In order to better provide feedback to users, the layout and arrangement of the aforementioned multimodal content are necessary. How to make the layout more reasonable has become an urgent problem to be solved. Summary of the Invention

[0004] This summary section is provided to briefly introduce the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0005] At least one embodiment of this disclosure provides a layout method for multimodal content generated by a model, comprising: obtaining content structure information of a first content, wherein the first content includes multiple text contents and multiple image contents, the content structure information being used to characterize at least one text-image combination, each text-image combination including at least one text content and at least one image content having an association relationship; for each text-image combination, determining a first layout method corresponding to the text-image combination from the multiple candidate layout methods based on the degree of matching between the text-image combination and the multiple candidate layout methods; and merging at least some text-image combinations in the at least one text-image combination according to the first layout method corresponding to the at least one text-image combination to determine the layout method of the first content.

[0006] At least another embodiment of this disclosure provides a layout apparatus for multimodal content generated by a model, comprising: an acquisition module configured to: acquire content structure information of a first content, wherein the first content includes multiple text contents and multiple image contents, the content structure information being used to characterize at least one text-image combination, each text-image combination including at least one text content and at least one image content having an association relationship; a determination module configured to: for each text-image combination, determine a first layout method corresponding to the text-image combination from the multiple candidate layout methods based on the degree of matching between the text-image combination and the multiple candidate layout methods; and an optimization module configured to: merge at least some text-image combinations in the at least one text-image combination according to the first layout method corresponding to the at least one text-image combination to determine the layout method of the first content.

[0007] At least one further embodiment of this disclosure provides an electronic device, including: a processing device; and a storage device including one or more computer program instructions; wherein the one or more computer program instructions are executed by the processing device to perform a layout method for multimodal content generated by a model provided in at least one embodiment of this disclosure.

[0008] At least one further embodiment of this disclosure provides a computer-readable storage medium that non-transitory stores computer-readable instructions, wherein, when executed by a processor, the computer-readable instructions implement the layout method for multimodal content generated by the model provided in at least one embodiment of this disclosure.

[0009] At least one embodiment of this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements a method for laying out multimodal content generated by a model according to at least one embodiment of this disclosure. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0011] Figure 1 This illustration schematically depicts an application scenario of a model-generated multimodal content layout system provided in at least one embodiment of this disclosure;

[0012] Figure 2 The illustration shows a flowchart of a method for laying out multimodal content generated by a model, provided in at least one embodiment of the present disclosure.

[0013] Figure 3This illustration schematically shows a process diagram for constructing an image forest according to at least one embodiment of the present disclosure;

[0014] Figure 4A and Figure 4B The illustration shows a schematic diagram of a candidate layout provided by at least one embodiment of the present disclosure;

[0015] Figure 5 The illustration shows a schematic diagram of a combined text and image combination provided in at least one embodiment of the present disclosure;

[0016] Figure 6 The schematic diagram illustrates the structure of a layout device for multimodal content generated by a model, provided in at least one embodiment of this disclosure; and

[0017] Figure 7 The schematic diagram illustrates a structure suitable for implementing at least one embodiment of the present disclosure of an electronic device. Detailed Implementation

[0018] One or more embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0019] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0020] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0021] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0022] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0023] The names of the messages or information exchanged between the various devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0024] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition, use, storage or deletion of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0025] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, relevant users should be informed of the type, scope of use, and usage scenarios of the information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and authorization should be obtained from the relevant users. Among them, relevant users may include any type of rights holder, such as individuals, enterprises, and groups.

[0026] For example, in response to receiving an active request from a user, a prompt message is sent to the relevant user to clearly indicate that the operation requested by the user will require obtaining and using the user's information. This allows the relevant user to choose whether to provide information to the software or hardware such as the electronic device, application, server, or storage medium that performs the operation of any embodiment of the present disclosure based on the prompt message.

[0027] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, such as a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide information to the electronic device.

[0028] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0029] Multimodal heterogeneous data content, also known as multimodal content, can be understood as content that includes multiple different types of data modalities. For example, multimodal heterogeneous data content can include text content and image content.

[0030] In some scenarios, it is necessary to layout and format multimodal heterogeneous data content, including text and images. For example, layout and formatting user-inputted text and uploaded images can make the multimodal heterogeneous data content more aesthetically pleasing; another example is layout and formatting text and images automatically generated by artificial intelligence (AI) models to improve the aesthetics of the responses the AI ​​models provide to users.

[0031] In some examples, the layout of multimodal heterogeneous data content, including text and images, is often simple and linear. For example, text content is placed first, followed by images. This simple, sequential layout has the following drawbacks: First, it lacks structure and readability. Text and images are usually related, but they are separated in the layout, disrupting the original logical hierarchy of the text content. Users find it difficult to connect the text and images, resulting in a poor reading experience. Second, this layout is simplistic and lacks adaptability. For text and images of different sizes, it easily leads to a large amount of unsightly blank space, reducing the information density of the layout, making it rigid, and resulting in low space utilization.

[0032] In some examples, the layout can be determined based on fixed rules, such as "all image content is centered and fills the line width" or "image content is placed after the corresponding text content." These fixed rules are usually pre-defined and non-context-aware, treating the text content merely as a flat sequence without understanding its structure. Therefore, layouts determined by fixed rules are unlikely to reflect the nested hierarchical relationships within the text content.

[0033] To at least partially solve the above-mentioned technical problems, at least one embodiment of this disclosure provides a layout method for multimodal content generated by a model. The method includes: obtaining content structure information of a first content, the first content including multiple text contents and multiple image contents, the content structure information being used to characterize at least one text-image combination, each text-image combination including at least one text content and at least one image content with an association relationship; for each text-image combination, determining a first layout method corresponding to the text-image combination from multiple candidate layout methods based on the degree of matching between multiple candidate layout methods and the text-image combination; and merging at least some text-image combinations in the at least one text-image combination according to the first layout method corresponding to the at least one text-image combination to determine the layout method of the first content.

[0034] In a method for laying out multimodal content generated by a model, provided in at least one embodiment of this disclosure, the structure of multimodal heterogeneous data content (i.e., the first content), including text and image content, is analyzed. The first content is then divided into multiple text-image combinations. After identifying the first layout method matching each text-image combination from candidate layout methods, the different text-image combinations are merged and optimized as a whole to achieve text-image layout and typesetting. Thus, when laying out text-image content, the inherent logical hierarchy is fully considered. The logical structure of the multimodal heterogeneous data content is converted into multiple text-image combinations. Text-image layout is performed on a unit basis, and then cross-level and cross-structure merging is carried out to optimize the overall text-image layout, achieving a clear text-image layout.

[0035] Based on the model-generated multimodal content layout method provided in at least one embodiment of this disclosure, at least one embodiment of this disclosure also provides a model-generated multimodal content layout apparatus, electronic device, computer-readable storage medium, and computer program product.

[0036] The present disclosure and some examples thereof will now be described in detail with reference to the accompanying drawings.

[0037] Figure 1 The illustration shows an application scenario of a layout system for model-generated multimodal content provided in at least one embodiment of the present disclosure.

[0038] like Figure 1 As shown, the application scenario provided in this embodiment may include a model-generated multimodal content layout system 100. One or more embodiments of this disclosure do not limit the form of the model-generated multimodal content layout system 100. In some embodiments, the model-generated multimodal content layout system 100 may be an independent software system. For example, the model-generated multimodal content layout system 100 may provide graphic and text layout services as an independent application (APP). In other embodiments, the model-generated multimodal content layout system 100 may also be integrated into other systems as a functional module of those systems, providing graphic and text layout services. For example, the model-generated multimodal content layout system 100 may be integrated into a content generation system as a plugin, cloud service, etc., to perform graphic and text layout for multimodal heterogeneous data generated by the content generation system.

[0039] In the multimodal content layout system 100 generated by the model, a first content 101 can be obtained. The first content 101 may include multiple text contents and multiple image contents. For the first content 101, content structure information 102 of the first content 101 is obtained. For example, the content structure information 102 can be used to represent at least one text-image combination 104. Each text-image combination 104 may include at least one text content and at least one image content with an association relationship. In this way, the multimodal content layout system 100 generated by the model can realize the parsing of the logical hierarchy of the first content 101, avoiding text-image mismatch and structural confusion.

[0040] Furthermore, the multimodal content layout system 100 generated by the model can also provide candidate layout methods 103. There can be multiple candidate layout methods 103, such as candidate layout method 1031, candidate layout method 1032, ..., candidate layout method 103N, where N is an integer greater than 1. For each image-text combination 104, based on the degree of matching between the image-text combination 104 and the candidate layout method 103, the first layout method 105 corresponding to the image-text combination 104 is determined. In this way, the multimodal content layout system 100 generated by the model can perform preliminary image-text layout based on image-text combinations, finding a matching first layout method 105 for each image-text combination 104.

[0041] Furthermore, the multimodal content layout system 100 generated by the model can also merge at least some of the image-text combinations 104 according to the first layout method 105 corresponding to at least one image-text combination 104, thereby determining the layout method 106 of the first content. In this way, the multimodal content layout system 100 generated by the model can perform collaborative layout and global optimization across different content levels, realize the overall adjustment of the image-text layout of the first content, and obtain a layout method of the first content that is clear in structure and visually appealing.

[0042] In some embodiments, the layout method for multimodal content generated by the above model can be applied in video conferencing scenarios. Some video conferencing systems provide a function that automatically generates meeting minutes using an artificial intelligence model. The AI ​​model analyzes the video conference and its corresponding meeting text to automatically generate meeting minutes that include both text and image content. For example, the text content can be extracted and integrated from the meeting text corresponding to the video conference, and the image content can be screen-shared content from the video conference.

[0043] In the aforementioned video conferencing scenario, it is necessary to format the initial content (including text and images) generated by the AI ​​model. For example, the content structure information of the initial content is obtained. This structure information can be used to represent at least one text-image combination, such as related meeting text and screen-shared content. For each text-image combination, the most suitable initial layout is determined from multiple candidate layouts, and at least some text-image combinations are merged to determine the final layout of the initial content. In this way, the layout of the multimodal initial content is formatted using the final layout, resulting in a more reasonable overall layout of the generated meeting minutes and enhancing the user's visual experience.

[0044] The following will combine Figures 2 to 5 The present disclosure provides a detailed description of a method for laying out multimodal content generated by a model, based on at least one embodiment.

[0045] Figure 2 The illustration shows a flowchart of a method for laying out multimodal content generated by a model, provided in at least one embodiment of the present disclosure.

[0046] like Figure 2 As shown, the model-generated multimodal content layout method of this embodiment includes steps S201 to S203. In some embodiments, the executing entity of the model-generated multimodal content layout method can be an electronic device with a client deployed, an electronic device with a server deployed, or any electronic device that communicates between the client and the server; one or more embodiments of this disclosure do not limit this. The model-generated multimodal content layout method includes:

[0047] Step S201: Obtain the content structure information of the first content.

[0048] In one or more embodiments of this disclosure, the first content may include multiple text contents and multiple image contents. For example, the first content may include multiple text segments and multiple images. In other words, the first content can be understood as multimodal heterogeneous data.

[0049] One or more embodiments of this disclosure do not limit the method of obtaining the first content. In some embodiments, the first content may be input by the user, for example, the user may upload text content and image content, and the first content may be obtained through a content upload operation triggered by the user.

[0050] In other embodiments, in artificial intelligence generated content (AIGC), the artificial intelligence model can provide rich text content and related image content according to user instructions. Therefore, the first content may include: content generated by the artificial intelligence model based on user instructions.

[0051] Artificial intelligence models can provide first-hand content in different scenarios. For example, an artificial intelligence model can connect with a content generation service. The user sends a user command to the content generation service, and the content generation service calls the artificial intelligence model to obtain the first-hand content.

[0052] In other words, the multimodal content layout method for model-generated content provided in at least one embodiment of this disclosure can be applied to the first content generated by the artificial intelligence model, providing graphic layout capabilities in scenarios where graphic content is automatically generated by AI, intelligently laying out and arranging the graphic content automatically generated by the artificial intelligence model, and improving the aesthetics of the response information fed back to the user by the artificial intelligence model.

[0053] The content structure information of the first content can be understood as information related to the logical hierarchy of the text content and the relationship between the text content and the image content in the first content, obtained after structural understanding of the first content. For example, the content structure information can be used to represent at least one text-image combination, where each text-image combination may include at least one text content and at least one image content that are related.

[0054] The logical hierarchy of the text content in the first content can be represented in different ways. For example, the logical hierarchy of the text content in the first content can be represented by headings of different levels. Or, the text content in the first content can be lightweight markup language (markdown) text, which supports the representation of logical hierarchy in a nested form (such as the text content in the first content can be a nested unordered list).

[0055] Related text and image content can be understood as image content being related to text content. For example, text content may describe image content, or image content may be an illustration corresponding to text content. In some embodiments, multiple image contents in the first content can be associated with text content through image metadata, such as being associated with a certain line of text content. The related text and image content can be determined through the mapping relationship of "image content - line number".

[0056] In other words, by understanding the structure of the text content in the first content, multiple text contents with hierarchical relationships in the first content are separated, and the relationship between multiple text contents and multiple image contents is determined to form multiple image-text combinations. On the one hand, the text content, which was originally a whole, is decomposed from the logical hierarchy level to obtain multiple text contents belonging to different levels and having nested relationships. On the other hand, the originally independent text contents and image contents are associated to establish logical relationships between images and text, and the first content is mapped according to the logical structure to the basic data structure of the subsequent image-text layout (i.e. image-text combinations).

[0057] In this way, while preserving the logical hierarchy and text-image relationship of the first content, a data foundation is provided for subsequent context-aware text-image layout.

[0058] In some possible implementations, content structure information can be presented in the form of a forest. For example, an image forest can be constructed based on the hierarchical relationship of multiple text contents in the first content, and the association relationship between multiple image contents and nodes in the image forest can be established based on the position information of the text contents associated with multiple image contents in the first content.

[0059] In an image forest, nodes can represent text content, and the connections between nodes represent hierarchical relationships.

[0060] In other words, multiple text contents are converted into nodes. By parsing the hierarchical relationship of the text contents in the first content (such as the structure of a nested unordered list of Markdown text), the connection relationship between the nodes is determined, so that the multiple text contents of the first content can be converted into a tree-structured image forest. Furthermore, the topology of the image forest corresponds one-to-one with the hierarchical relationship of the multiple text contents in the first content. For example, each node in the image forest represents a text content in the first content.

[0061] Since multiple images in the first content have related text content, for example, the line numbers of the text content can be used to indicate the related text content of the images, after constructing the image forest, the image content can be mounted one by one onto the nodes of the image forest. For example, for image content A, the text content associated with image content A is the text content of line 15. The text content of line 15 is located at node B in the image forest. Therefore, the association relationship between image content A and node B is established, and image content A is mounted onto the image forest. By using the association relationship between nodes and image content in the image forest, it can be shown that "image content" and "text content" are logically related.

[0062] Figure 3 The illustration shows a schematic diagram of a process for constructing an image forest according to at least one embodiment of the present disclosure.

[0063] like Figure 3 As shown, the first content includes multiple text contents and multiple image contents. The multiple text contents have a hierarchical relationship. For example, the text contents corresponding to "Title A" and "Title B" are at the first level, and the text contents corresponding to "Point A.1" and "Point A.2" are at the second level below the text contents corresponding to "Title A". The multiple image contents include Image 1 and Image 2.

[0064] Based on the hierarchical relationship of multiple text contents, an image forest is constructed. The image forest includes node A corresponding to title A, node B corresponding to title B, node A.1 corresponding to point A.1, and node A.2 corresponding to point A.2. Nodes A.1 and A.2 are child nodes of node A. The parent-child nodes in the image forest are used to represent the hierarchical relationship of multiple text contents in the first content.

[0065] Next, multiple image contents are mounted to the image forest. For example, the text content associated with image 1 belongs to the text content corresponding to title A. Therefore, image 1 is mounted to node A, establishing the association between image 1 and node A. Similarly, the text content associated with image 2 belongs to the text content corresponding to point A.2. Therefore, image 2 is mounted to node A.2, establishing the association between image 2 and node A.2.

[0066] In this way, by constructing an image forest and mounting images, the linear, unstructured primary content is transformed into a data structure with clear hierarchy, explicit text-image relationships, and the ability to be subsequently laid out and analyzed by computing devices.

[0067] Step S202: For each of the at least one image and text combination, determine the first layout method corresponding to the image and text combination from the multiple candidate layout methods based on the degree of matching between the multiple candidate layout methods and the image and text combination.

[0068] In one or more embodiments of this disclosure, the graphic and text layout is first performed on a unit basis, that is, a corresponding first layout method is determined for each graphic and text combination.

[0069] Candidate layout options can be understood as pre-defined, selectable layout options. For example, candidate layout options may include row layout options with different parameters and column layout options with different parameters.

[0070] Figure 4A and Figure 4B The illustration shows a schematic diagram of a candidate layout provided by at least one embodiment of the present disclosure.

[0071] like Figure 4AAs shown, in the line layout mode, the image content is placed below the text content. In the line layout mode, parameters can include the spacing between the image content and the text content, the spacing between multiple images, the display size of the image content, etc. Different line layout modes under different parameters can correspond to different candidate layout modes.

[0072] like Figure 4B As shown, in the column layout, the image content and text content are placed side by side. In the column layout, the parameters can include the column width ratio, the spacing between the image content and the text content, the spacing between multiple images, the display size of the image content, and the arrangement of multiple images (such as horizontal or vertical). Different column layouts under different parameters can correspond to different candidate layouts.

[0073] The matching method between candidate layout methods and text-image combinations can be understood as the degree of reasonableness of applying candidate layout methods to text-image combinations for layout. The first layout method can be understood as the layout method whose degree of matching with the text-image combination meets the set conditions (such as being the most matching). In other words, compared with other candidate layout methods, the first layout method is more matching with the text-image combination.

[0074] In some possible implementations, a matching degree evaluation model is used to quantitatively evaluate the matching degree between candidate layout methods and image-text combinations. For example, for each candidate layout method among multiple candidate layout methods, the candidate layout method and image-text combination are sent to the matching degree evaluation model, and the matching degree returned by the matching degree evaluation model is received. This matching degree represents the matching degree between the candidate layout method and the image-text combination. Based on the matching degrees corresponding to multiple candidate layout methods, the first layout method corresponding to the image-text combination is determined from the multiple candidate layout methods.

[0075] In other words, the matching degree evaluation model is used to objectively measure the layout quality of each candidate layout method when applied to the combination of text and images, so as to achieve a quantitative evaluation of the matching degree between the candidate layout method and the combination of text and images, and thus provide a basis for determining the first layout method.

[0076] One or more embodiments of this disclosure do not limit the training method of the matching degree evaluation model. For example, the training data of the matching degree evaluation model may include candidate layout methods, training image and text combinations and matching degree as labels. The initial matching degree evaluation model is trained using the training data, and the model parameters of the matching degree evaluation model are updated to obtain the matching degree evaluation model.

[0077] In some embodiments, the degree of matching can be measured by multiple dimensions. For example, the degree of matching can be measured based on at least one of the following: the clarity of the image content in the text and image combination under the candidate layout method, the space utilization of the text and image combination under the candidate layout method, and the compositional balance of the text and image combination under the candidate layout method.

[0078] In this way, by measuring the degree of matching between candidate layout methods and image-text combinations from different dimensions, we can break away from fixed layout rules and flexibly adjust the first layout method corresponding to different image-text combinations based on evidence.

[0079] The following sections will explain how to determine the three measurement dimensions mentioned above.

[0080] In at least one embodiment of this disclosure, the matching degree can be measured based on the clarity of the image content in the text and image combination under the candidate layout method. The clarity can be used to evaluate whether the size of the image content is appropriate and whether the image content is clear and readable. That is, the clarity can characterize whether the individual image content is within a reasonable visual range, avoiding the problem of poor readability due to excessively large image content or wasted space due to excessively small image content.

[0081] Multiple candidate layout options may include a first candidate layout option. The clarity of the image content in the text and image combination under the first candidate layout option can be determined as follows: determine the average area of ​​a single image content in the text and image combination under the first candidate layout option, and determine the clarity of the image content in the text and image combination under the first candidate layout option based on the average area of ​​a single image content in the text and image combination.

[0082] For example, the clarity of the text and image combination under the candidate layout can be calculated using the following formula:

[0083]

[0084] In the above formula, For clarity, The average area of ​​a single image content. The square root of the average area of ​​a single image content can be understood as the "equivalent side length" of the image content. The function... Other parameters (such as 15, 50) are used to limit the sharpness to a range of 0 to 10.

[0085] In other words, for any first candidate layout, the display size of a single image content under the first candidate layout is calculated to determine whether the size of the image content under the first candidate layout is reasonable, and then the matching between the first candidate layout and the image and text combination is evaluated.

[0086] In at least one embodiment of this disclosure, the matching degree can be measured based on the space utilization degree of the graphic and text combination under the candidate layout method. The space utilization degree can be used to evaluate the space utilization efficiency of the candidate layout method on the layout interface. That is, the space utilization degree can reward compact and full layout methods and punish excessively empty layout methods.

[0087] Multiple candidate layout options may include a first candidate layout option. The space utilization of the text and image combination under the first candidate layout option can be determined as follows: determine the overall area ratio of text content and image content in the text and image combination under the first candidate layout option, and determine the blank area under the first candidate layout option. Based on the overall area ratio of text content and image content and the blank area of ​​the text and image combination, determine the space utilization of the text and image combination under the first candidate layout option.

[0088] For example, the space utilization of image content in a text-image combination under candidate layout methods can be calculated using the following formula:

[0089]

[0090]

[0091]

[0092]

[0093]

[0094] In the above formula, For the degree of space utilization, To determine the proportion of text and images, This represents the ratio of the sum of the areas of text content and image content to the total area. The function represents the difference between the total area and the sum of the areas of the text content and image content, i.e., the blank area. Other parameters (such as 100, 50, 7, 3, 250) are used to limit the space utilization rate within the range of 0 to 10.

[0095] In other words, for any first candidate layout, the proportion of text and images and the absolute blank area under the first candidate layout are calculated to determine whether the first candidate layout is full and whether there is absolute waste of space, and then to evaluate whether the first candidate layout matches the text and image combination.

[0096] In at least one embodiment of this disclosure, the degree of matching can be measured based on the compositional balance of the text and image combination under the candidate layout mode. The compositional balance can be used to evaluate whether the overall visual center of the text and image combination is aligned with the geometric center of the layout page. That is, the compositional balance can characterize the layout stability and harmony under the candidate layout mode, and avoid the overall visual center of the text and image combination from deviating too much from the geometric center of the layout page.

[0097] Multiple candidate layout methods may include a first candidate layout method. The compositional balance of the text and image combination under the first candidate layout method can be determined as follows: determine the centroid of the text content and the image content in the text and image combination under the first candidate layout method, determine the geometric center under the first candidate layout method, determine the vertical offset of the centroid and geometric center of the text content and the image content in the text and image combination, and determine the compositional balance of the text and image combination under the first candidate layout method based on the vertical offset of the centroid and geometric center of the text content and the image content in the text and image combination.

[0098] The centroid, also known as the center of mass, can be understood as an imaginary point where mass is concentrated. When applying the first candidate layout method to a text and image combination, if the area formed by the text content and the image content in the text and image combination is centrally symmetrical, then the centroid can be located at the center of symmetry. If the area formed by the text content and the image content in the text and image combination is not centrally symmetrical, then the centroid can be located at a point that can maintain the balance of the area formed by the text content and the image content in the text and image combination.

[0099] The geometric center can be understood as the central position of a shape with a certain degree of symmetry. The layout interface of the first candidate layout is usually a symmetrical shape, such as a rectangle. Therefore, the geometric center under the first candidate layout can be the intersection of the two diagonals of the layout interface of the first candidate layout.

[0100] For example, the compositional balance of text and image combinations under candidate layout methods can be calculated using the following formula:

[0101]

[0102] In the above formula, For the sake of compositional balance, Let be the coordinates of the geometric center in the vertical direction under the candidate layout method. The function represents the vertical coordinates of the centroids of the text and image content in the combined text and image content. This is used to ensure that the compositional balance is not less than 0.

[0103] In other words, for any first candidate layout, by calculating the offset between the centroid of the text and image combination under the first candidate layout and the geometric center under the first candidate layout, it is determined whether the first candidate layout can achieve compositional stability and balance, and then the matching between the first candidate layout and the text and image combination is evaluated.

[0104] Thus, by determining the degree of matching between candidate layout methods and image-text combinations, for each image-text combination, an objective and comparable quantitative matching degree is assigned to each candidate layout method, and the candidate layout method with the highest matching degree is taken as the first candidate layout method corresponding to the image-text combination, thereby determining the initial optimal layout method of the image-text combination.

[0105] It should be noted that when the matching degree is measured based on more than two dimensions, the matching degree can be obtained by weighted summation of the two or more dimensions. For example, the matching degree is measured based on the following three dimensions: the clarity of the image content in the text-image combination under the candidate layout, the space utilization of the text-image combination under the candidate layout, and the compositional balance of the text-image combination under the candidate layout. In this case, different weights are assigned to the clarity of the image content in the text-image combination under the candidate layout, the space utilization of the text-image combination under the candidate layout, and the compositional balance of the text-image combination under the candidate layout, and the matching degree is determined by weighted summation.

[0106] In some embodiments, considering that the matching degree of the first layout method corresponding to the image and text combination may be low, that is, it is impossible to find a first layout method that meets the set conditions for the matching degree of the image and text combination, in this case, the text content in the image and text combination can be adjusted to optimize the first layout method corresponding to the image and text combination.

[0107] For example, a text-image combination includes a first text-image combination. The degree of matching between the first layout method corresponding to the first text-image combination and the first text-image combination is a first matching degree. In response to the first matching degree being less than a set threshold, the text content in the first text-image combination is adjusted according to multiple text contents of the first content to obtain a second text-image combination. The degree of matching between multiple candidate layout methods and the second text-image combination is determined. In response to the fact that there is a second matching degree greater than the first matching degree among the multiple candidate layout methods and the second text-image combination, the first text-image combination is updated to the second text-image combination. The first layout method corresponding to the second text-image combination is the candidate layout method corresponding to the second matching degree.

[0108] In other words, if the first layout method corresponding to the first image and text combination has a low degree of matching with the image and text combination, the text content in the first image and text combination is adjusted to obtain a second image and text combination based on the first image and text combination. For the updated second image and text combination, the degree of matching between multiple candidate layout methods and the second image and text combination is recalculated. If the degree of matching between a candidate layout method and the second image and text combination is higher than the degree of matching between the first image and text combination and the first layout method, it indicates that the updated second image and text combination has a more matching candidate layout method. That is, by updating the image and text combination, the rationality of the image and text layout can be improved. Therefore, the candidate layout method with a higher degree of matching corresponding to the second image and text combination is adopted, and the first layout method corresponding to the second image and text combination is the candidate layout method with a higher degree of matching.

[0109] In this way, by adjusting the text content in the combination of text and images, the matching degree between the layout and the combination of text and images is improved, making the layout and typesetting of text and images more reasonable and beautiful.

[0110] In some possible implementations, adjusting the text content in a text-image combination can be done by expanding the scope of the text content. For example, from multiple text contents of the first content, a first text content adjacent to the text content in the first text-image combination can be determined, and this first text content can be added to the text content of the first text-image combination to obtain a second text-image combination.

[0111] In other words, by associating adjacent first text content, such as associating the first text content corresponding to the parent node, the range of text content in the image-text combination is expanded. Based on a larger range of text content, the degree of matching between the layout method and the image-text combination is calculated, and the first layout method corresponding to the image-text combination is optimized.

[0112] Step S203: Based on the first layout method corresponding to at least one graphic and text combination, merge at least some graphic and text combinations in the at least one graphic and text combination to determine the layout method of the first content.

[0113] After determining the layout method based on the combination of text and images, you can also merge adjacent and cross-level elements. That is, try to merge the first layout methods corresponding to multiple independent text and image combinations to seek the overall effect of "1+1>2".

[0114] In some possible implementations, based on the content structure information of the first content, at least one third and fourth image-text combination with a positional relationship are determined. The matching degree between the first layout method corresponding to the third image-text combination and the third image-text combination is the third matching degree, and the matching degree between the first layout method corresponding to the fourth image-text combination and the fourth image-text combination is the fourth matching degree. Then, the text content in the third image-text combination and the text content in the fourth image-text combination are merged, as are the image content in the third image-text combination and the image content in the fourth image-text combination, to obtain a fifth image-text combination.

[0115] Determine the degree of matching between multiple candidate layout methods and the fifth image and text combination. In response to the fact that the fifth matching degree is greater than the third and fourth matching degrees among the multiple candidate layout methods and the fifth image and text combination, determine the first layout method corresponding to the fifth image and text combination as the candidate layout method corresponding to the fifth matching degree. Based on the first layout method corresponding to at least one image and text combination including the fifth image and text combination, determine the layout method of the first content.

[0116] In other words, for the third and fourth image-text combinations with positional relationships (such as image-text combinations corresponding to sibling nodes or image-text combinations corresponding to parent and child nodes), the image content and text content in the third and fourth image-text combinations are merged respectively to form a new fifth image-text combination, realizing the merging of adjacent image-text combinations or cross-level image-text combinations.

[0117] For the merged fifth graphic combination, the matching degree between multiple candidate layout methods and the fifth graphic combination is recalculated. If the matching degree between a candidate layout method and the fifth graphic combination is higher than the matching degree between the original third graphic combination and the corresponding first layout method and the original fourth graphic combination and the corresponding first layout method, it indicates that a more suitable graphic layout can be achieved after merging the graphic combinations. That is, by merging graphic combinations, the rationality of the graphic layout can be improved. Therefore, the candidate layout method with a higher matching degree corresponding to the fifth graphic combination is adopted, and the first layout method corresponding to the fifth graphic combination is the candidate layout method with a higher matching degree.

[0118] After determining the first layout method corresponding to each image and text combination (including the first layout method corresponding to the updated image and text combination and the first layout method corresponding to the merged image and text combination), the layout method of the first content can include the first layout method corresponding to each image and text combination.

[0119] Figure 5 The illustration shows a schematic diagram of a combined graphic and text combination provided in at least one embodiment of the present disclosure.

[0120] like Figure 5 As shown, before the merger, the first layout method corresponding to the third and fourth image-text combinations was a column layout method. That is, the text content A and image content A in the third image-text combination were arranged side by side, and the text content B and image content B in the fourth image-text combination were arranged side by side.

[0121] By merging the third and fourth image-text combinations, a unified fifth image-text combination is formed. The first layout method corresponding to the fifth image-text combination is a line layout method, that is, the image content (including image content A and image content B) in the fifth image-text combination is placed below the text content (including text content A and text content B).

[0122] In this way, by merging text and image combinations, and crossing the logical boundaries of text and image combinations during the text and image layout process, some layout methods with poor layout effects can be optimized as a whole. The layout of text and image elements scattered in different logical levels in the first content can be coordinated and adjusted to find the global optimal solution and improve the layout effect of text and image elements.

[0123] Furthermore, layout applications can be applied to the first content. For example, a block element corresponding to the layout of the first content can be created in a blank document, and the text and image content of the first content can be filled into the block element.

[0124] For example, according to the parameters (such as proportion, spacing, etc.) of the first layout method corresponding to each combination of text and images in the first content layout method, corresponding grid and grid column block-level elements are created in the blank document. The text content and image content in the combination of text and images are used as elements to fill the block elements. In the block-based collaborative document, the first content is laid out according to the first content layout method, thus achieving automated and intelligent typesetting.

[0125] Thus, through four stages—"structural understanding—candidate layout evaluation—global layout optimization—layout implementation"—the system automatically finds and applies the best text and image layout for the primary content, ensuring a close correspondence between image content and related text content, optimizing visual flow, reducing page white space, and improving readability. Furthermore, it is applicable to various complex document structures, image content of different sizes, and text content of different lengths, exhibiting high flexibility and adaptability.

[0126] Based on the model-generated multimodal content layout method provided in at least one embodiment of this disclosure, at least one embodiment of this disclosure also provides a model-generated multimodal content layout apparatus. The following will be combined with... Figure 6 The layout device for the multimodal content generated by the model is described in detail.

[0127] Figure 6 The schematic diagram illustrates the structure of a layout device for multimodal content generated by a model, provided in at least one embodiment of the present disclosure.

[0128] like Figure 6 As shown, the multimodal content layout apparatus 600 for model generation in this embodiment includes an acquisition module 601, a determination module 602, and an optimization module 603. For example, the acquisition module 601, determination module 602, and optimization module 603 can be implemented using hardware (e.g., circuit) modules or software modules. The following embodiments are similar and will not be described again. For example, the acquisition module 601, determination module 602, and optimization module 603 can be implemented using a central processing unit (CPU), a general-purpose graphics processor (GPGPU), a graphics processing unit (GPU), a tensor processor (TPU), a field-programmable gate array (FPGA), or other processing units with data processing capabilities and / or instruction execution capabilities, along with corresponding computer instructions.

[0129] The acquisition module 601 is configured to: acquire content structure information of a first content, wherein the first content includes multiple text contents and multiple image contents, and the content structure information is used to characterize at least one text-image combination, wherein each text-image combination includes at least one text content and at least one image content that are related. For example, the acquisition module 601 can be configured to execute step S201 described above; its specific implementation principle can be found in the relevant description of step S201, and will not be repeated here.

[0130] The determining module 602 is configured to: for each of the at least one image-text combination, determine a first layout method corresponding to the image-text combination from the multiple candidate layout methods based on the degree of matching between the image-text combination and the multiple candidate layout methods. For example, the determining module 602 can be configured to execute step S202 described above; its specific implementation principle can be found in the relevant description of step S202, and will not be repeated here.

[0131] The optimization module 603 is configured to: merge at least a portion of the graphic elements in the at least one graphic element combination according to the first layout method corresponding to the at least one graphic element combination, and determine the layout method of the first content. For example, the optimization module 603 can be configured to execute step S203 described above; its specific implementation principle can be found in the relevant description of step S203, and will not be repeated here.

[0132] In at least one embodiment of this disclosure, the acquisition module 601 is further configured to: construct an image forest based on the hierarchical relationship of multiple text contents in the first content, wherein the nodes in the image forest represent the text contents, and the connection relationship between the nodes in the image forest represents the hierarchical relationship; and establish the association relationship between the multiple image contents and the nodes in the image forest based on the position information of the text contents associated with the multiple image contents in the first content.

[0133] In at least one embodiment of this disclosure, the determining module 602 is further configured to: for each of the plurality of candidate layout methods, send the candidate layout method and the image-text combination to a matching degree evaluation model, receive the matching degree returned by the matching degree evaluation model, wherein the matching degree characterizes the matching degree between the candidate layout method and the image-text combination; and determine the first layout method corresponding to the image-text combination from the plurality of candidate layout methods based on the matching degree corresponding to the plurality of candidate layout methods.

[0134] In at least one embodiment of this disclosure, the degree of matching is based on at least one of the following measures: the clarity of the image content in the image-text combination under the candidate layout, the degree of space utilization of the image-text combination under the candidate layout, and the degree of compositional balance of the image-text combination under the candidate layout.

[0135] In at least one embodiment of this disclosure, the matching degree is measured based on the clarity of the image content in the image-text combination under the candidate layout mode. The plurality of candidate layout modes include a first candidate layout mode. The determining module 602 is further configured to: determine the average area of ​​a single image content in the image-text combination under the first candidate layout mode; and determine the clarity of the image content in the image-text combination under the first candidate layout mode based on the average area of ​​a single image content in the image-text combination.

[0136] In at least one embodiment of this disclosure, the matching degree is measured based on the space utilization degree of the text and image combination under the candidate layout mode. The plurality of candidate layout modes include a first candidate layout mode. The determining module 602 is further configured to: determine the overall area ratio of text content and image content in the text and image combination under the first candidate layout mode, and determine the blank area under the first candidate layout mode; and determine the space utilization degree of the text and image combination under the first candidate layout mode based on the overall area ratio of text content and image content in the text and image combination and the blank area.

[0137] In at least one embodiment of this disclosure, the matching degree is measured based on the compositional balance of the text and image combination under the candidate layout mode. The plurality of candidate layout modes includes a first candidate layout mode. The determining module 602 is further configured to: determine the centroid of the text content and the image content in the text and image combination under the first candidate layout mode, and determine the geometric center under the first candidate layout mode; determine the offset of the centroid of the text content and the image content and the geometric center in the vertical direction in the text and image combination; and determine the compositional balance of the text and image combination under the first candidate layout mode based on the offset of the centroid of the text content and the image content and the geometric center in the vertical direction in the text and image combination.

[0138] In at least one embodiment of this disclosure, the image-text combination includes a first image-text combination, and the matching degree between the first layout method corresponding to the first image-text combination and the first image-text combination is a first matching degree. The optimization module 603 is further configured to: in response to the first matching degree being less than a set threshold, adjust the text content in the first image-text combination according to multiple text contents of the first content to obtain a second image-text combination; determine the matching degree between multiple candidate layout methods and the second image-text combination; in response to the fact that there is a second matching degree greater than the first matching degree among the matching degrees between the multiple candidate layout methods and the second image-text combination, update the first image-text combination to the second image-text combination, and the first layout method corresponding to the second image-text combination is the candidate layout method corresponding to the second matching degree.

[0139] In at least one embodiment of this disclosure, the optimization module 603 is further configured to: determine a first text content adjacent to the text content in the first image-text combination from a plurality of text contents of the first content; add the first text content to the text content in the first image-text combination to obtain the second image-text combination.

[0140] In at least one embodiment of this disclosure, the optimization module 603 is further configured to: determine, based on the content structure information of the first content, a third and a fourth image-text combination with a positional relationship among the at least one image-text combination, wherein the matching degree between the first layout method corresponding to the third image-text combination and the third image-text combination is a third matching degree, and the matching degree between the first layout method corresponding to the fourth image-text combination and the fourth image-text combination is a fourth matching degree; merge the text content in the third image-text combination and the text content in the fourth image-text combination, and merge the image content in the third image-text combination and the image content in the fourth image-text combination to obtain a fifth image-text combination; determine the matching degree between multiple candidate layout methods and the fifth image-text combination; in response to the fact that among the multiple candidate layout methods and the fifth image-text combination, there is a fifth matching degree greater than the third matching degree and the fourth matching degree, determine the first layout method corresponding to the fifth image-text combination as the candidate layout method corresponding to the fifth matching degree; and determine the layout method of the first content based on the first layout method corresponding to the at least one image-text combination including the fifth image-text combination.

[0141] In at least one embodiment of this disclosure, the first content includes: content generated by an artificial intelligence model based on user instructions.

[0142] In at least one embodiment of this disclosure, the layout device 600 for multimodal content generated by the model further includes a layout module, which is configured to: create block elements in a blank document corresponding to the layout mode of the first content; and fill the block elements with text content and image content from the first content.

[0143] It should be noted that, for clarity and brevity, at least one embodiment of this disclosure does not show all the constituent units of the layout device 600 for multimodal content generated by the model. To achieve the necessary functions of the layout device 600 for multimodal content generated by the model, those skilled in the art can provide and set other constituent units (not shown) according to specific needs, and one or more embodiments of this disclosure do not limit this.

[0144] At least one embodiment of this disclosure also provides an electronic device, including a processing device and a storage device, the storage device including one or more computer program modules; wherein the one or more computer program modules are stored in the storage device and configured to be executed by the processing device, the one or more computer program modules being used to implement the layout method for multimodal content generated by the model provided in any embodiment of this disclosure.

[0145] For example, the processing device may be a processor, such as a central processing unit (CPU), digital signal processor (DSP), image processor (GPU), general-purpose graphics processor (GPGPU), or other form of processing unit with data processing capabilities and / or instruction execution capabilities. It may be a general-purpose processor or a dedicated processor and may control other components in the electronic device to perform the desired functions.

[0146] For example, the storage device may be a memory, which may include one or more computer program products. These computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory. The non-volatile memory may, for example, include read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and a processing device may execute these program instructions to implement the functions (implemented by the processing device) in at least one embodiment of this disclosure and / or other desired functions. Various application programs and various data may also be stored in the computer-readable storage medium, which is not limited by one or more embodiments of this disclosure.

[0147] The following is for reference. Figure 7 The diagram illustrates a structural schematic of an electronic device (e.g., a terminal device or a server) 700 suitable for implementing at least one embodiment of the present disclosure. The terminal device in at least one embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of at least one embodiment of this disclosure.

[0148] like Figure 7 As shown, the electronic device 700 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. The RAM 703 also stores various programs and data required for the operation of the electronic device 700. The processing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0149] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 709. Communication device 709 allows electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 An electronic device 700 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0150] In particular, according to one or more embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, one or more embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 709, or installed from storage device 708, or installed from ROM 702. When the computer program is executed by processing device 701, it performs the functions defined in the methods of at least one embodiment of this disclosure.

[0151] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0152] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0153] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0154] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform a layout method for the multimodal content generated by the aforementioned model.

[0155] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0156] One or more embodiments of this disclosure also provide a computer program product comprising one or more computer instructions. When the computer instructions are loaded and executed on a computing device, all or part of the processes or functions described in any embodiment of this disclosure are generated.

[0157] The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, or data center to another website, computer, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.

[0158] When the computer program product is executed by a computer, the computer executes any of the methods for layout of multimodal content generated by the aforementioned model. The computer program product can be a software installation package; when any of the methods for layout of multimodal content generated by the aforementioned model is required, the computer program product can be downloaded and executed on the computer.

[0159] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0160] The units or modules described in at least one embodiment of this disclosure can be implemented in software or hardware. The names of the units or modules do not necessarily limit the specific unit or module itself.

[0161] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0162] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0163] According to one or more embodiments of this disclosure, Example 1 provides a method for laying out multimodal content generated by a model, including:

[0164] Obtain content structure information of the first content, wherein the first content includes multiple text contents and multiple image contents, and the content structure information is used to characterize at least one text-image combination, wherein each text-image combination includes at least one text content and at least one image content that are related to each other;

[0165] For each of the at least one image and text combination, based on the degree of matching between the image and text combination and the multiple candidate layout methods, a first layout method corresponding to the image and text combination is determined from the multiple candidate layout methods;

[0166] Based on the first layout method corresponding to the at least one graphic combination, at least some graphic combinations in the at least one graphic combination are merged to determine the layout method of the first content.

[0167] According to one or more embodiments of this disclosure, Example 2 provides the content structure information for obtaining the first content in Example 1, including:

[0168] Based on the hierarchical relationship of multiple text contents in the first content, an image forest is constructed, wherein the nodes in the image forest represent the text contents, and the connection relationship between the nodes in the image forest represents the hierarchical relationship;

[0169] Based on the location information of the text content associated with multiple images in the first content, an association relationship is established between the multiple images and nodes in the image forest.

[0170] According to one or more embodiments of this disclosure, Example 3 provides the method of determining a first layout corresponding to the image and text combination from the plurality of candidate layout methods based on the degree of matching between the candidate layout methods and the image and text combination, as described in Example 1, including:

[0171] For each of the multiple candidate layout methods, the candidate layout method and the image-text combination are sent to the matching degree evaluation model, and the matching degree returned by the matching degree evaluation model is received, wherein the matching degree represents the degree of matching between the candidate layout method and the image-text combination;

[0172] Based on the matching degree of the multiple candidate layout methods, the first layout method corresponding to the image and text combination is determined from the multiple candidate layout methods.

[0173] According to one or more embodiments of this disclosure, Example 4 provides that the matching degree in Example 1 is based on at least one of the following measures: the clarity of the image content in the image-text combination under the candidate layout method, the space utilization degree of the image-text combination under the candidate layout method, and the compositional balance degree of the image-text combination under the candidate layout method.

[0174] According to one or more embodiments of this disclosure, Example 5 provides a measurement of the matching degree in Example 4 based on the clarity of the image content in the image-text combination under the candidate layout method. The plurality of candidate layout methods includes a first candidate layout method, and the clarity of the image content in the image-text combination under the first candidate layout method is determined by the following method:

[0175] Determine the average area of ​​a single image content in the image-text combination under the first candidate layout method;

[0176] The clarity of the image content in the image-text combination under the first candidate layout is determined based on the average area of ​​the individual image content in the image-text combination.

[0177] According to one or more embodiments of this disclosure, Example Six provides a measurement of the matching degree in Example Four based on the space utilization degree of the graphic-text combination under the candidate layout methods. The plurality of candidate layout methods includes a first candidate layout method, and the space utilization degree of the graphic-text combination under the first candidate layout method is determined by the following method:

[0178] Determine the overall area ratio of text content and image content in the text-image combination under the first candidate layout method, and determine the blank area under the first candidate layout method;

[0179] Based on the overall area ratio of text content and image content in the graphic combination and the blank area, the space utilization degree of the graphic combination under the first candidate layout method is determined.

[0180] According to one or more embodiments of this disclosure, Example 7 provides a measurement of the compositional balance of the text and image combination under candidate layout methods, as in Example 4. The plurality of candidate layout methods includes a first candidate layout method, and the compositional balance of the text and image combination under the first candidate layout method is determined in the following manner:

[0181] Determine the centroids of the text content and image content in the image-text combination under the first candidate layout method, and determine the geometric center under the first candidate layout method;

[0182] Determine the offset of the centroid of the text content and the image content in the combined text and image, and the offset of the geometric center in the vertical direction.

[0183] The compositional balance of the text and image combination under the first candidate layout is determined based on the centroid of the text content and the image content in the combination and the offset of the geometric center in the vertical direction.

[0184] According to one or more embodiments of this disclosure, Example 8 provides that the graphic and text combination in Example 1 includes a first graphic and text combination, wherein the degree of matching between the first layout method corresponding to the first graphic and text combination and the first graphic and text combination is a first matching degree, and after determining the first layout method corresponding to the graphic and text combination, the method further includes:

[0185] In response to the first matching degree being less than a set threshold, the text content in the first image-text combination is adjusted according to the multiple text contents of the first content to obtain a second image-text combination;

[0186] Determine the degree of matching between multiple candidate layout methods and the second image and text combination;

[0187] In response to the fact that the second matching degree among the multiple candidate layout methods and the second graphic combination is greater than the first matching degree, the first graphic combination is updated to the second graphic combination, and the first layout method corresponding to the second graphic combination is the candidate layout method corresponding to the second matching degree.

[0188] According to one or more embodiments of this disclosure, Example 9 provides the method of adjusting the text content in the first image-text combination based on multiple text contents of the first content, as in Example 8, to obtain a second image-text combination, including:

[0189] From the multiple text contents of the first content, determine the first text content that is adjacent to the text content in the first image-text combination;

[0190] The first text content is added to the text content of the first image-text combination to obtain the second image-text combination.

[0191] According to one or more embodiments of this disclosure, Example 10 provides the method of merging at least a portion of the graphic elements in the at least one graphic element combination according to a first layout method corresponding to the at least one graphic element combination, as in Example 1, to determine the layout method of the first content, including:

[0192] Based on the content structure information of the first content, a third and fourth image-text combination with a positional relationship are determined among the at least one image-text combination. The degree of matching between the first layout method corresponding to the third image-text combination and the third image-text combination is the third matching degree, and the degree of matching between the first layout method corresponding to the fourth image-text combination and the fourth image-text combination is the fourth matching degree.

[0193] The text content in the third image-text combination and the text content in the fourth image-text combination are merged, and the image content in the third image-text combination and the image content in the fourth image-text combination are merged to obtain the fifth image-text combination;

[0194] Determine the degree of matching between multiple candidate layout methods and the fifth graphic combination;

[0195] In response to the fact that among the multiple candidate layout methods and the fifth graphic combination, there is a fifth matching degree that is greater than the third matching degree and the fourth matching degree, the first layout method corresponding to the fifth graphic combination is determined to be the candidate layout method corresponding to the fifth matching degree.

[0196] The layout of the first content is determined based on the first layout method corresponding to the at least one graphic combination including the fifth graphic combination.

[0197] According to one or more embodiments of this disclosure, Example 11 provides that the first content in any of Examples 1 to 10 includes: content generated by an artificial intelligence model based on user instructions.

[0198] According to one or more embodiments of this disclosure, Example Twelve provides a method from any of Examples One to Ten, further comprising:

[0199] Create a block element in the blank document that corresponds to the layout of the first content;

[0200] Fill the block element with the text and image content from the first content.

[0201] According to one or more embodiments of this disclosure, Example Thirteen provides a layout apparatus for model-generated multimodal content, comprising:

[0202] The acquisition module is configured to: acquire content structure information of a first content, wherein the first content includes multiple text contents and multiple image contents, the content structure information is used to characterize at least one text-image combination, and each text-image combination includes at least one text content and at least one image content that are related to each other;

[0203] The determining module is configured to: for each of the at least one image-text combination, determine a first layout method corresponding to the image-text combination from the multiple candidate layout methods based on the degree of matching between the multiple candidate layout methods and the image-text combination;

[0204] The optimization module is configured to: merge at least some of the graphic elements in the at least one graphic element combination according to the first layout method corresponding to the at least one graphic element combination, and determine the layout method of the first content.

[0205] According to one or more embodiments of this disclosure, Example Fourteen provides an electronic device, including:

[0206] Processing device; and

[0207] Storage device, including one or more computer program instructions;

[0208] The one or more computer program instructions are executed by a processing device to perform a layout method for multimodal content generated by a model according to at least one embodiment of the present disclosure.

[0209] According to one or more embodiments of the present disclosure, Example Fifteen provides a computer-readable storage medium that non-transitory stores computer-readable instructions, wherein the computer-readable instructions, when executed by a processor, implement a layout method for multimodal content generated by a model provided in at least one embodiment of the present disclosure.

[0210] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0211] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0212] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A method for laying out multimodal content generated by a model, comprising: Obtain content structure information of the first content, wherein the first content includes multiple text contents and multiple image contents, and the content structure information is used to characterize at least one text-image combination, wherein each text-image combination includes at least one text content and at least one image content that are related to each other; For each of the at least one image and text combination, based on the degree of matching between the image and text combination and the multiple candidate layout methods, a first layout method corresponding to the image and text combination is determined from the multiple candidate layout methods; Based on the first layout method corresponding to the at least one graphic combination, at least some graphic combinations in the at least one graphic combination are merged to determine the layout method of the first content.

2. The method according to claim 1, wherein, The acquisition of the content structure information of the first content includes: Based on the hierarchical relationship of multiple text contents in the first content, an image forest is constructed, wherein the nodes in the image forest represent the text contents, and the connection relationship between the nodes in the image forest represents the hierarchical relationship; Based on the location information of the text content associated with multiple images in the first content, an association relationship is established between the multiple images and nodes in the image forest.

3. The method according to claim 1, wherein, The step of determining the first layout method corresponding to the image and text combination from the multiple candidate layout methods based on the degree of matching between the multiple candidate layout methods and the image and text combination includes: For each of the multiple candidate layout methods, the candidate layout method and the image-text combination are sent to the matching degree evaluation model, and the matching degree returned by the matching degree evaluation model is received, wherein the matching degree represents the degree of matching between the candidate layout method and the image-text combination; Based on the matching degree of the multiple candidate layout methods, the first layout method corresponding to the image and text combination is determined from the multiple candidate layout methods.

4. The method according to claim 1, wherein, The matching degree is based on at least one of the following measures: the clarity of the image content in the image-text combination under the candidate layout method, the space utilization of the image-text combination under the candidate layout method, and the compositional balance of the image-text combination under the candidate layout method.

5. The method according to claim 4, wherein, The matching degree is measured based on the clarity of the image content in the image-text combination under the candidate layout method. The multiple candidate layout methods include the first candidate layout method. The clarity of the image content in the image-text combination under the first candidate layout method is determined by the following method: Determine the average area of ​​a single image content in the image-text combination under the first candidate layout method; The clarity of the image content in the image-text combination under the first candidate layout is determined based on the average area of ​​the individual image content in the image-text combination.

6. The method according to claim 4, wherein, The matching degree is measured based on the space utilization of the image and text combination under the candidate layout methods. The multiple candidate layout methods include the first candidate layout method. The space utilization of the graphic and text combination under the first candidate layout method is determined in the following way: Determine the overall area ratio of text content and image content in the text-image combination under the first candidate layout method, and determine the blank area under the first candidate layout method; Based on the overall area ratio of text content and image content in the graphic combination and the blank area, the space utilization degree of the graphic combination under the first candidate layout method is determined.

7. The method according to claim 4, wherein, The matching degree is measured based on the compositional balance of the text and image combination under the candidate layout methods. The multiple candidate layout methods include the first candidate layout method. The compositional balance of the text and image combination under the first candidate layout method is determined in the following way: Determine the centroids of the text content and image content in the image-text combination under the first candidate layout method, and determine the geometric center under the first candidate layout method; Determine the offset of the centroid of the text content and the image content in the combined text and image, and the offset of the geometric center in the vertical direction. The compositional balance of the text and image combination under the first candidate layout is determined based on the centroid of the text content and the image content in the combination and the offset of the geometric center in the vertical direction.

8. The method according to claim 1, wherein the image and text combination includes a first image and text combination, and the degree of matching between the first layout method corresponding to the first image and text combination and the first image and text combination is a first matching degree. After determining the first layout method corresponding to the graphic combination, the method further includes: In response to the first matching degree being less than a set threshold, the text content in the first image-text combination is adjusted according to the multiple text contents of the first content to obtain a second image-text combination; Determine the degree of matching between multiple candidate layout methods and the second image and text combination; In response to the fact that the second matching degree among the multiple candidate layout methods and the second graphic combination is greater than the first matching degree, the first graphic combination is updated to the second graphic combination, and the first layout method corresponding to the second graphic combination is the candidate layout method corresponding to the second matching degree.

9. The method according to claim 8, wherein, The step of adjusting the text content in the first image-text combination based on multiple text contents of the first content to obtain the second image-text combination includes: From the multiple text contents of the first content, determine the first text content that is adjacent to the text content in the first image-text combination; The first text content is added to the text content of the first image-text combination to obtain the second image-text combination.

10. The method according to claim 1, wherein, The step of merging at least a portion of the text and image combinations in the at least one text and image combination according to the first layout method corresponding to the at least one text and image combination, and determining the layout method of the first content, includes: Based on the content structure information of the first content, a third and fourth image-text combination with a positional relationship are determined among the at least one image-text combination. The degree of matching between the first layout method corresponding to the third image-text combination and the third image-text combination is the third matching degree, and the degree of matching between the first layout method corresponding to the fourth image-text combination and the fourth image-text combination is the fourth matching degree. The text content in the third image-text combination and the text content in the fourth image-text combination are merged, and the image content in the third image-text combination and the image content in the fourth image-text combination are merged to obtain the fifth image-text combination; Determine the degree of matching between multiple candidate layout methods and the fifth graphic combination; In response to the fact that among the multiple candidate layout methods and the fifth graphic combination, there is a fifth matching degree that is greater than the third matching degree and the fourth matching degree, the first layout method corresponding to the fifth graphic combination is determined to be the candidate layout method corresponding to the fifth matching degree. The layout of the first content is determined based on the first layout method corresponding to the at least one graphic combination including the fifth graphic combination.

11. The method according to any one of claims 1 to 10, wherein, The first content includes: content generated by an artificial intelligence model based on user instructions.

12. The method according to any one of claims 1 to 10, further comprising: Create a block element in the blank document that corresponds to the layout of the first content; Fill the block element with the text and image content from the first content.

13. A layout device for model-generated multimodal content, comprising: The acquisition module is configured to: acquire content structure information of a first content, wherein the first content includes multiple text contents and multiple image contents, the content structure information is used to characterize at least one text-image combination, and each text-image combination includes at least one text content and at least one image content that are related to each other; The determining module is configured to: for each of the at least one image-text combination, determine a first layout method corresponding to the image-text combination from the multiple candidate layout methods based on the degree of matching between the multiple candidate layout methods and the image-text combination; The optimization module is configured to: merge at least some of the graphic elements in the at least one graphic element combination according to the first layout method corresponding to the at least one graphic element combination, and determine the layout method of the first content.

14. An electronic device comprising: Processing device; as well as Storage device, including one or more computer program instructions; The one or more computer program instructions are executed by the processing device to perform the method according to any one of claims 1 to 12.

15. A computer-readable storage medium for non-transitory storage of computer-readable instructions, wherein, The method of any one of claims 1 to 12 is implemented when the computer-readable instructions are executed by a processor.