Document processing method and device
By acquiring and analyzing the characteristics of editable elements and determining adjustment parameters, the problem of unreasonable layout in document generation was solved, and the readability of the document was improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LENOVO (BEIJING) LTD
- Filing Date
- 2026-01-29
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies struggle to ensure the logical layout of editable elements within a document during document generation, resulting in poor readability.
By obtaining the element features and fusion features of each editable element, the target element features are determined, and based on this, adjustment parameters are determined. Each editable element is then adjusted to generate the target document.
It improves the layout rationality of editable elements within the document, thereby enhancing the document's readability.
Smart Images

Figure CN121900668A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a document processing method and apparatus. Background Technology
[0002] With the development of automatic document generation technology, AI-based document generation methods can automatically generate documents containing various editable elements such as text and images according to user needs. However, these technologies often struggle to ensure the logical layout of the editable elements within the document, resulting in poor readability. Summary of the Invention
[0003] In view of this, this application provides a document processing method and apparatus.
[0004] According to a first aspect of this application, a document processing method is provided, comprising: obtaining a first document to be processed; the first document comprising a plurality of first editable elements; obtaining a first element feature of each first editable element and a fusion feature of the plurality of first editable elements; determining a first target element feature of each first editable element based on the fusion feature and the first element feature; determining a first adjustment parameter of each first editable element based on the first target element feature; and adjusting each first editable element according to the first adjustment parameter to obtain a target document.
[0005] A second aspect of this application provides a document processing apparatus, comprising: an obtaining module for obtaining a first document to be processed; the first document including a plurality of first editable elements; an extraction module for obtaining a first element feature of each first editable element and a fusion feature of the plurality of first editable elements; a determining module for determining a first target element feature of each first editable element based on the fusion feature and the first element feature; a prediction module for determining a first adjustment parameter of each first editable element based on the first target element feature; and an adjusting module for adjusting each first editable element according to the first adjustment parameter to obtain a target document.
[0006] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description
[0007] The above and other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0008] Figure 1 A schematic diagram illustrating an example environment in which the methods according to embodiments of this application can be applied;
[0009] Figure 2 The flowchart illustrating a document processing method provided in an embodiment of this application is shown in the illustration;
[0010] Figure 3 A schematic diagram illustrating a feature extraction method provided in an embodiment of this application;
[0011] Figure 4 A flowchart illustrating the prediction of a second adjustment parameter using a prediction model, provided in an embodiment of this application;
[0012] Figure 5 This application provides an example of an interactive flow diagram of a document processing device.
[0013] Figure 6 This is a block diagram of a document processing apparatus provided in an embodiment of this application. Detailed Implementation
[0014] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.
[0015] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0016] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0017] When using expressions such as "at least one of A, B, and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having A alone, having B alone, having C alone, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0018] In the embodiments of this application, the collection, updating, analysis, processing, use, transmission, provision, disclosure, and storage of data (including but not limited to user personal information) comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures have been taken to prevent unauthorized access to user personal information data and to safeguard user personal information security, network security, and national security.
[0019] Figure 1 A schematic diagram of an example environment in which the method according to an embodiment of this application can be applied is shown. In this example environment, application 125 is installed on terminal device 110. User 140 can interact with application 125 via terminal device 110 and / or an attached device of terminal device 110.
[0020] In some embodiments, application 125 can be downloaded and installed on terminal device 110. In some embodiments, application 125 can also be accessed in other ways, such as through a web page. Figure 1 In this environment, in response to the launch of application 125, terminal device 110 can display the interface 150 of application 125.
[0021] In some embodiments, terminal device 110 can communicate with server 130 to provide services to application 125. Terminal device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, terminal device 110 can also support any type of user-facing interface.
[0022] In some embodiments, application 125 may provide document processing functionality. Application 125 may include an application dedicated to providing document processing services, or an application integrated with document processing functionality. Although Figure 1 The image shows a single application, but in reality, multiple applications can be installed on the terminal device 110.
[0023] In embodiments of this application, agent 160 can be deployed locally on terminal device 110 or remotely. When remotely deployed, terminal device 110 can directly invoke agent 160, or it can invoke agent 160 via server 130. Exemplarily, agent 160 may have document generation and document processing capabilities. Terminal device 110 provides an interface 150 that can present interactions with agent 160. In interface 150, user 140 can initiate document processing requests to agent 160 by inputting natural language (e.g., text input or voice input). Optionally, user 140 can upload online or offline documents to instruct agent 160 to generate or process the documents.
[0024] In embodiments of this application, during interaction with user 140, agent 160 can respond to user 140's document processing request and process the document indicated by the user. For example, user 140 can request agent 160 to generate a document based on a topic, or upload a first document to be processed to request agent 160 to process the first document. The first document may include multiple first editable elements, which may include various types of elements such as text boxes, images, shapes, and tables.
[0025] In some embodiments, during document processing, the agent 160 may invoke one or more tools 165 as needed to assist in the execution of document processing tasks. These tools 165 can be of any type, such as document parsing tools, feature extraction tools, document editing tools, document rendering tools, etc. For example, the agent 160 may invoke a document parsing tool to obtain attribute information for each first editable element in the first document, invoke a feature extraction tool to obtain the first element features of each first editable element and the fused features of multiple first editable elements, and invoke a document editing tool to adjust the first editable elements according to determined adjustment parameters, thereby generating the target document.
[0026] In some embodiments, the intelligent agent 160 can work collaboratively with the server 130 to process the first document. For example, the terminal device 110 can be responsible for parsing and extracting attributes from the first document, the server 130 can be responsible for feature encoding and determining adjustment parameters, and the terminal device 110 can further be responsible for adjusting the first editable element according to the adjustment parameters. This collaborative approach between the cloud and local systems reduces the computational burden on the terminal device 110 while ensuring processing effectiveness.
[0027] In some embodiments, agent 160 may be constructed based on one or more machine learning models. In some embodiments, the machine learning model upon which agent 160 is based may include at least a language model, such as a large language model. In some embodiments, the machine learning model upon which agent 160 is based may include a multimodal model capable of handling multiple modalities of input, such as text input, visual input (e.g., images, videos), audio input, etc. Exemplarily, agent 160 may extract text features of text content in a first editable element using a text encoder, and extract image features of images in the first editable element using an image encoder. These machine learning models may include content-generative models capable of generating corresponding outputs based on model inputs. In some embodiments, the machine learning model may receive text-modal model inputs (e.g., natural language and / or machine language) and / or non-text-modal model inputs (e.g., images, speech, videos, etc.), and may obtain corresponding model outputs based on the model inputs, thereby completing the document processing task.
[0028] It should be understood that the structure and function of the various elements in the environment are described for illustrative purposes only and are not intended to limit the scope of this application in any way.
[0029] The following will be based on Figure 1 The following describes the document processing method of this application embodiment in detail, based on the described scenario.
[0030] Figure 2 A flowchart illustrating a document processing method provided in an embodiment of this application is shown.
[0031] like Figure 2 As shown, this document processing method may specifically include the following operations.
[0032] Operation S210 obtains the first document to be processed; the first document includes multiple first editable elements;
[0033] Operation S220 obtains the first element feature of each first editable element, as well as the fusion feature of multiple first editable elements;
[0034] Operation S230 determines the first target element feature of each first editable element based on the fusion feature and the feature of each first element;
[0035] Operation S240 determines the first adjustment parameter for each first editable element based on the features of the first target element;
[0036] Operation S250 adjusts each of the first editable elements according to the first adjustment parameters to obtain the target document.
[0037] In operation S210, the first document to be processed refers to the document that needs to be adjusted in layout. For example, the first document may include, but is not limited to: PowerPoint documents, web page documents, typesetting documents, and documents with mixed text and images.
[0038] In some examples, the first document may be an automatically generated document based on a user request, where the layout of the editable elements may be flawed. In other examples, the first document may be an existing document uploaded by the user, whose layout the user wishes to optimize.
[0039] Optionally, the first document may include multiple first editable elements. A first editable element refers to an element object in the document that can be individually identified and adjusted. It can be understood as a content unit in the document with independent attributes and positional information, used to carry various types of content in the document and as the basic unit for layout adjustment.
[0040] For example, the first editable element may include, but is not limited to, text boxes, images, shapes, tables, charts, title boxes, body text boxes, etc. Different types of first editable elements may have different attribute characteristics; for example, text boxes have attributes such as text content, font size, and line spacing, while images have attributes such as image content and aspect ratio.
[0041] In one feasible implementation, in response to a user-initiated document processing request, a first document to be processed can be obtained from the file uploaded by the user. In this implementation, the user can select a document file from local storage or cloud storage to upload via a document processing application. After receiving the document file, the system will use it as the first document to be processed for subsequent processing.
[0042] In another feasible implementation, a first document to be processed can be obtained from a document automatically generated based on a user request. In this implementation, the user can provide a document generation request, such as providing information like the topic and content requirements. The system generates a document containing multiple first editable elements based on the user request and uses this generated document as the first document to be processed.
[0043] Optionally, when obtaining the first document, the identification information of each first editable element in the first document can be obtained simultaneously to facilitate subsequent identification and processing of each first editable element. For example, the identification information may include the element's number, the element's type label, etc.
[0044] It should be noted that the first document may include one or more pages, and each page may include one or more first editable elements. When the first document includes multiple pages, subsequent feature extraction and adjustment operations can be performed on each page separately, or the multiple pages can be processed as a whole.
[0045] In operation S220, the first element feature of each first editable element refers to the feature information used to characterize the attributes of the first editable element itself. It can be understood as a feature vector or feature representation describing the content, type, position, size and other attributes of a single first editable element, which is used to reflect the individual characteristics of the first editable element in the document.
[0046] For example, the first element features may include one or more of the following: content features, type features, position features, and size features of the first editable element. The content features may characterize the specific content information contained in the first editable element; the type features may characterize the category of the first editable element; the position features may characterize the coordinate position of the first editable element on the document page; and the size features may characterize the width and height information of the first editable element.
[0047] Similarly, the fusion feature of multiple first editable elements refers to the overall feature information that comprehensively considers the relationship between multiple first editable elements. It can be understood as a feature representation that reflects the relationship and overall layout state between multiple first editable elements, and is used to capture the overall information such as spatial relationship and semantic relationship between each first editable element in the document.
[0048] For example, fusion features may include information such as relative positional relationships, distance relationships, hierarchical relationships, and semantic relationships among multiple first editable elements. Fusion features can reflect the overall layout and structural characteristics of a document.
[0049] In one feasible implementation, feature extraction can be performed on each first editable element separately to obtain corresponding first element features, and overall feature extraction can be performed on multiple first editable elements to obtain fused features. In this implementation, the first document can be parsed first to identify each first editable element in the first document, and then the attribute information of each first editable element can be extracted and encoded as first element features. At the same time, the attribute information of multiple first editable elements can be comprehensively processed to extract fused features that reflect the relationship between elements.
[0050] Optionally, when obtaining the fusion feature, the fusion process can be performed based on the first element features of multiple first editable elements. For example, the first element features of each first editable element can be fused through attention calculation, feature aggregation, etc., to obtain the fusion feature that reflects the relationship between elements.
[0051] In operation S230, the first target element feature of each first editable element refers to the comprehensive feature information that integrates the element's own characteristics and the relationship between elements. It can be understood as a feature representation that combines the first element feature with the integrated feature and contains both individual element information and overall layout information. It is used to comprehensively reflect the adjustment needs of the first editable element in the overall layout of the document.
[0052] It should be noted that the first target element features include not only the content, type, position, size and other attribute information of the first editable element itself, but also the overall layout information such as the relative relationship and spatial distribution between the first editable element and other first editable elements.
[0053] In one feasible implementation, the first element feature and the fusion feature of each first editable element can be concatenated to obtain the first target element feature of the first editable element. In this implementation, for each first editable element, its corresponding first element feature and fusion feature are concatenated and combined in a specific order or manner to form a first target element feature containing information from both.
[0054] For example, if the first element feature is in vector form and the fused feature is also in vector form, the two can be concatenated to obtain a first target element feature vector with a higher dimension.
[0055] Optionally, when determining the features of the first target element, different fusion methods can be used for different first editable elements. For example, for a first editable element located in the center of the document, the fusion features can be given a higher weight so that its first target element features reflect more of the relationship with other elements; for a first editable element located at the edge of the document, the first element features can be given a higher weight so that its first target element features retain more of the element's own attribute information.
[0056] In operation S240, the first adjustment parameter refers to the parameter information used to adjust the layout of the first editable element. It can be understood as a parameter value that describes how to adjust the position, size and other attributes of the first editable element, and is used to guide the specific adjustment operation of the first editable element.
[0057] In one feasible implementation, the first target element features of each first editable element can be input into a parameter prediction model, which then outputs a first adjustment parameter for that first editable element. In this implementation, the parameter prediction model can predict suitable adjustment parameters for the first editable element based on the input first target element features. This parameter prediction model can be trained to learn the mapping relationship from the first target element features to appropriate adjustment parameters.
[0058] As an optional embodiment, the first adjustment parameter includes at least one of the following: position adjustment amount, size adjustment ratio, rotation angle adjustment amount, and transparency adjustment amount for the first editable element.
[0059] The position adjustment amount indicates the distance the first editable element needs to move in both the horizontal and vertical directions. For example, the position adjustment amount can be expressed as a combination of a horizontal offset and a vertical offset, which moves the first editable element from its current position to a target position.
[0060] The scaling factor indicates the percentage by which the width and height of the first editable element need to be adjusted. For example, the scaling factor can be expressed as a width scaling factor and a height scaling factor, or as a uniform scaling factor. This scaling factor allows the size of the first editable element to be enlarged or reduced.
[0061] The rotation angle adjustment amount indicates the angle by which the first editable element needs to be rotated. For example, the rotation angle adjustment amount can be expressed as a rotation in degrees relative to the current angle. This rotation angle adjustment amount can be used to adjust the orientation of the first editable element.
[0062] The transparency adjustment amount indicates the value by which the transparency of the first editable element needs to be adjusted. For example, the transparency adjustment amount can be expressed as an increase or decrease in transparency. This transparency adjustment amount allows for adjustment of the visibility of the first editable element, thereby achieving visual optimization of the element hierarchy.
[0063] Optionally, the parameter prediction model can output multiple adjustment parameters simultaneously, or it can selectively output some adjustment parameters based on the type and current state of the first editable element. For example, for some first editable elements that do not require rotation adjustment, the parameter prediction model can output only the position adjustment amount and the size adjustment ratio.
[0064] It should be noted that the various parameters in the first adjustment parameters can be used in combination to achieve comprehensive adjustments to the first editable element. For example, position adjustment, size adjustment, and transparency adjustment can be applied to a first editable element simultaneously, thereby optimizing its visual presentation while adjusting the element's position and size.
[0065] By adopting the technical solution of this application, after obtaining the first document to be processed, by obtaining the first element feature of each first editable element and the fusion feature of multiple first editable elements, it is possible to simultaneously capture the individual characteristics of a single element and the overall correlation between multiple elements. Furthermore, based on the fusion feature and the first element feature of each first editable element, the first target element feature of each first editable element is determined, so that the first target element feature not only contains the element's own information, but also integrates the correlation between the element and other elements, thereby more accurately reflecting the adjustment needs of each first editable element in the overall layout of the document. The first adjustment parameter determined based on the first target element feature can comprehensively consider the element's own characteristics and the interrelationship between elements, so that the layout of each editable element in the target document obtained after adjusting each first editable element according to the first adjustment parameter is more reasonable, improving the readability of the document and solving the problem that related technologies cannot guarantee the reasonableness of the layout of each editable element in the document.
[0066] Based on the above embodiments, as an optional embodiment, in order to more accurately characterize the attribute information of each first editable element, the above operation S220 may further include the following operations.
[0067] Operation S310: For each first editable element, obtain the text features and image features of the first editable element;
[0068] Operation S320 obtains the first element feature of the first editable element based on the text features and image features of the first editable element.
[0069] In operation S310, text features refer to the feature information that characterizes the text content in the first editable element. It can be understood as the feature representation obtained after encoding the text information contained in the first editable element, which is used to reflect the text semantics of the first editable element.
[0070] Similarly, image features refer to the feature information that characterizes the image content in the first editable element. They can be understood as the feature representation obtained by encoding the image information contained in the first editable element, which is used to reflect the visual content of the first editable element.
[0071] It should be noted that different types of first editable elements can have different types of features extracted. For example, for a first editable element of type text box, its text features can be extracted; for a first editable element of type image, its image features can be extracted; and for a first editable element that contains both text and image, its text features and image features can be extracted simultaneously.
[0072] In one feasible implementation, the text features of the first editable element can be obtained through a text encoder. In this implementation, the text content can first be extracted from the first editable element, and then the text content can be input into the text encoder. The text encoder encodes the text content to obtain text features that represent the semantic information of the text content.
[0073] In another feasible implementation, the image features of the first editable element can be obtained through an image encoder. In this implementation, image content can first be extracted from the first editable element, and then the image content can be input into the image encoder. The image encoder encodes the image content to obtain image features that characterize the visual information of the image content.
[0074] Optionally, the text encoder and image encoder can employ pre-trained encoding models. For example, a text encoder and image encoder trained using contrastive learning can be used, so that text features and image features reside in the same semantic space, thereby facilitating the unified processing of features of different types of first editable elements in subsequent processing.
[0075] In operation S320, the text features and image features of the first editable element can be fused to obtain the first element features that comprehensively represent the content information of the first editable element.
[0076] In one feasible implementation, the text features and image features of the first editable element can be concatenated to obtain the first element feature of the first editable element. In this implementation, for a first editable element that has both text features and image features, its text features and image features are vector-concatenated in a specific order to form a first element feature containing information from both.
[0077] In another feasible implementation, for a first editable element that has only text features, its text features can be used as the first element features; for a first editable element that has only image features, its image features can be used as the first element features.
[0078] Optionally, when fusing text features and image features, feature mapping can be performed on the text features and image features separately to ensure that their feature dimensions are consistent before concatenation. For example, a linear mapping layer or a multilayer perceptron can be used to perform dimensionality transformation on the text features and image features respectively, and then the transformed features can be concatenated to obtain the first element feature with unified dimensions.
[0079] By adopting the technical solution of this embodiment, text features and image features are obtained for each first editable element, and the content information of the first editable element can be captured from different modalities. Furthermore, the first element features are obtained based on the text features and image features, so that the first element features can comprehensively reflect the text semantics and visual content of the first editable element, thereby providing a more comprehensive feature representation for subsequently determining accurate adjustment parameters.
[0080] Based on the above embodiments, as an optional embodiment, in order to capture the relationships between the first editable elements, the above operation S320 may further include the following operations.
[0081] Operation S410 obtains semantic features based on the text and image features of the first editable element;
[0082] Operation S420 obtains the attribute features of the first editable element, the attribute features representing the geometric parameters of the first editable element within the corresponding first page of the first document;
[0083] Operation S430 concatenates the attribute features and semantic features of the first editable element to obtain the first element feature of the first editable element.
[0084] In operation S410, semantic features refer to the feature information that characterizes the meaning of the content of the first editable element. It can be understood as the feature representation obtained after encoding the semantic level of the content carried by the first editable element.
[0085] In one feasible implementation, for a first editable element having textual features, its textual features can be used as semantic features. In this implementation, the textual features themselves already contain semantic information of the text content, so the textual features can be directly used as the semantic features of the first editable element.
[0086] In another feasible implementation, for a first editable element having image features, its image features can be used as semantic features. In this implementation, the image features already contain semantic information of the image content, so the image features can be directly used as the semantic features of the first editable element.
[0087] Optionally, for a first editable element that simultaneously possesses textual and image features, the textual and image features can be fused to obtain semantic features. For example, the textual and image features can be weighted and fused or concatenated to obtain semantic features that comprehensively reflect the semantic content of the first editable element.
[0088] In operation S420, attribute features refer to the feature information that characterizes the layout attributes of the first editable element. It can be understood as the feature representation obtained after encoding the geometric parameters such as the position and size of the first editable element on the page, which is used to reflect the layout state of the first editable element.
[0089] It should be noted that the geometric parameters include the position and size information of the first editable element within the first page. For example, the geometric parameters may include the coordinates, width, and height of the first editable element, as well as the canvas width and height of the first page. The coordinates may represent the top-left corner position of the first editable element, the width and height may represent the size of the first editable element, and the canvas width and height may represent the overall size of the first page.
[0090] In one feasible implementation, the geometric parameters of each first editable element can be parsed from the first document, and the geometric parameters are input into an attribute encoder for encoding to obtain attribute features. In this implementation, the attribute encoder can convert the geometric parameters into attribute features in vector form. Exemplarily, the attribute encoder can employ a multilayer perceptron structure, and obtain attribute features representing geometric attributes by performing nonlinear transformations on the geometric parameters.
[0091] Optionally, when encoding geometric parameters, the geometric parameters can be normalized. For example, the coordinates and size of the first editable element can be divided by the canvas width and canvas height of the first page, respectively, to obtain normalized geometric parameters. Then, the normalized geometric parameters are encoded, thereby making the geometric parameters in pages of different sizes comparable.
[0092] In operation S430, attribute features and semantic features can be concatenated into vectors in a specific order to obtain the first element feature that simultaneously contains geometric attribute information and content semantic information.
[0093] Optionally, before concatenating attribute features and semantic features, dimensional alignment can be performed on both. For example, a linear mapping layer can be used to perform dimensional transformation on the attribute features and semantic features respectively, so that their dimensions are consistent before concatenation, thereby obtaining a first element feature with unified dimensions.
[0094] By adopting the technical solution of this embodiment, semantic features are obtained based on text features and image features, which can accurately capture the content meaning of the first editable element; attribute features representing geometric parameters are obtained, which can accurately reflect the layout state of the first editable element; the attribute features and semantic features are concatenated to obtain the first element features, so that the first element features simultaneously contain the semantic information and geometric information of the first editable element, thereby providing a more comprehensive and accurate feature representation for subsequent layout adjustments.
[0095] Based on the above embodiments, as an optional embodiment, in order to capture the interrelationships between the first editable elements, the above operation S230 may further include the following operations.
[0096] Operation S510 encodes the fused features to obtain an encoded sequence; the encoded sequence includes the associated features corresponding to each first editable element; the associated features represent the relationship between the corresponding first editable element and other first editable elements;
[0097] Operation S520 concatenates each first element feature with its corresponding associated feature to obtain the first target element feature of each first editable element.
[0098] In operation S510, the encoded sequence refers to the feature sequence obtained after encoding the fused features. It can be understood as a serialized representation containing the associated features of multiple first editable elements, used to reflect the association status of each first editable element in the overall layout of the document.
[0099] Association features refer to the feature information that characterizes the association relationship between a first editable element and other first editable elements. They can be understood as feature representations that describe the relative positional relationship, distance relationship, semantic relationship, etc. of the first editable element with other elements on the page, and are used to reflect the contextual information of the first editable element in the overall layout.
[0100] In one feasible implementation, the fused features can be input into a sequence encoder, which encodes the fused features to obtain an encoded sequence. In this implementation, the sequence encoder can encode the feature information of multiple first editable elements contained in the fused features, thereby generating corresponding associated features for each first editable element.
[0101] Optionally, the sequence encoder may employ an attention-based encoding structure. For example, the sequence encoder may employ a transformer encoder structure, capturing the relationships between different first editable elements through self-attention computation, thereby enabling the association features corresponding to each first editable element to reflect the relationship between that first editable element and other first editable elements.
[0102] It should be noted that the number of associated features contained in the encoding sequence corresponds to the number of first editable elements in the first document. For example, if a page of the first document contains N first editable elements, then the corresponding encoding sequence contains N associated features, and each associated feature corresponds to one first editable element.
[0103] In operation S520, for each first editable element, its corresponding first element feature can be concatenated with the corresponding associated feature in the encoding sequence to obtain the first target element feature of the first editable element.
[0104] In one feasible implementation, each first element feature can be vector-concatenated with the associated features at the same index position in the encoding sequence, according to the index order of the first editable elements. In this implementation, the correspondence between the first element features and the associated features is maintained, ensuring that the first target element feature of each first editable element contains the element's own attribute information and the association information between the element and other elements.
[0105] Optionally, before concatenating the first element feature and the associated feature, feature transformation processing can be performed on both. For example, a mapping layer can be used to perform feature transformation on the first element feature and the associated feature respectively, so that their feature dimensions or feature distribution meet the concatenation requirements, and then they can be concatenated to obtain the first target element feature.
[0106] By adopting the technical solution of this embodiment, the fused features are encoded to obtain an encoding sequence containing the associated features of each first editable element, which can accurately capture the mutual relationship between the first editable elements; the first element features are concatenated with the corresponding associated features to obtain the first target element features, so that the first target element features not only contain the attribute information of the element itself, but also integrate the association relationship between the element and other elements, thereby providing a more comprehensive feature representation for the prediction of subsequent adjustment parameters.
[0107] Please refer to Figure 3 , Figure 3 This is a schematic diagram of a feature extraction method provided in an embodiment of this application.
[0108] like Figure 3 As shown, for editable elements, their attribute features, semantic features, and first element features can be obtained separately. Specifically, text element attributes and image element attributes can be extracted from the editable elements. The text element attributes are input into a text encoder to obtain text features, and the image element attributes are input into an image encoder to obtain image features. The text features and image features can be used as semantic features. At the same time, attribute parameters can be extracted from the editable elements and input into an attribute feature MLP to obtain attribute features. Furthermore, the attribute features and semantic features can be concatenated to obtain the first element features.
[0109] For the first element features of multiple editable elements, a fused feature can be obtained through concatenation. The fused feature is then input into a Transformer encoder for encoding to obtain an encoded sequence. The encoded sequence is then concatenated with each first element feature, and the first target element feature of each editable element is obtained through an MLP. Finally, the first target element feature is input into an MLP decoder to obtain the first adjustment parameter.
[0110] based on Figure 3 The processing flow shown illustrates the feature extraction and parameter adjustment prediction process. In the feature extraction stage, text element attributes and image element attributes are extracted for each editable element. Text element attributes are processed by a text encoder to obtain text features, and image element attributes are processed by an image encoder to obtain image features. The text features and image features together constitute semantic features reflecting the semantic meaning of the element's content. Simultaneously, attribute parameters of the editable element are processed using an attribute feature MLP to obtain attribute features reflecting the element's geometric layout. Subsequently, the semantic features and attribute features are concatenated to obtain the first element feature, which contains both content and layout information.
[0111] In the association encoding stage, the first element features (including first element features 1 to n) corresponding to multiple editable elements are concatenated to obtain a fused feature containing the feature information of all editable elements. This fused feature is then input into a Transformer encoder, where self-attention calculation captures the association relationships between different editable elements, resulting in an encoded sequence containing the association features of each element. Further, the encoded sequence is concatenated with each first element feature, and feature fusion processing is performed using an MLP to obtain a first target element feature that integrates the element's own information and the association relationships between elements.
[0112] In the parameter prediction stage, the first target element features are input into the MLP decoder. The MLP decoder decodes the first target element features and predicts the first adjustment parameters for layout adjustment. These first adjustment parameters can guide the adjustment of the position, size and other attributes of the editable elements.
[0113] It should be noted that, Figure 3 The dashed boxes represent intermediate processing results, while solid boxes represent processing units or operations. In the diagram, MLP refers to a multilayer perceptron, used for non-linear transformations of features; the Transformer encoder is used to capture the relationships between different editable elements. This processing flow extracts feature information from editable elements at multiple levels, including the attributes and semantic features of individual elements, as well as the relationships between elements, ultimately predicting accurate adjustment parameters.
[0114] Based on the above embodiments, as an optional embodiment, in order to further improve the accuracy of document layout adjustment, the above document processing method may also include the following operations.
[0115] Operation S610 adjusts each first editable element according to the first adjustment parameter to obtain a second document; the second document includes multiple second editable elements;
[0116] Operation S620 predicts the second adjustment parameter for each second editable element using a prediction model; wherein the prediction model is trained based on sample documents and a reward function; the reward function is used to evaluate the layout quality of sample editable elements in the sample document.
[0117] Operation S630 adjusts each of the second editable elements according to the second adjustment parameters to obtain the target document.
[0118] In operation S610, the second editable element refers to the editable element object in the second document, which can be understood as the first editable element after being adjusted by the first adjustment parameters, and is used for the second stage of layout optimization adjustment.
[0119] It should be noted that the second editable element is consistent with the first editable element in terms of type and content. The difference is that the geometric properties of the second editable element, such as position and size, have been adjusted according to the first adjustment parameters.
[0120] In operation S620, the second adjustment parameter refers to the parameter information used to further adjust the second editable element. It can be understood as an adjustment parameter that performs fine-grained optimization based on the first adjustment parameter, and is used to achieve fine-grained adjustment of the document layout.
[0121] A predictive model is a model that can predict and adjust parameters based on the document layout state. It can be understood as a parameter prediction model trained by learning the layout quality evaluation results of sample documents, and used to predict reasonable adjustment parameters for editable elements in the document.
[0122] It should be noted that the prediction model is trained based on sample documents and a reward function. The sample documents provide training data for the prediction model, while the reward function evaluates the layout quality of editable elements within the sample documents, thus providing an optimization objective for the model's training. Through training based on the reward function, the prediction model learns how to adjust editable elements to achieve higher layout quality scores.
[0123] In one feasible implementation, the feature information of each second editable element in the second document can be input into a prediction model, and the prediction model can output a second adjustment parameter for that second editable element. In this implementation, the prediction model can predict a second adjustment parameter suitable for fine-grained adjustment based on the current layout state of the second editable element.
[0124] Optionally, the adjustment range of the second adjustment parameter can be smaller than that of the first adjustment parameter. For example, if the position adjustment amount in the first adjustment parameter has a large value range to achieve a larger position movement, then the position adjustment amount in the second adjustment parameter can have a smaller value range to achieve fine position adjustment.
[0125] In operation S630, each second editable element in the second document can be adjusted according to the second adjustment parameter output by the prediction model, thereby obtaining the final target document.
[0126] In one feasible implementation, each second editable element can be adjusted sequentially by applying the corresponding second adjustment parameters according to the order of the second editable elements in the second document. In this implementation, the adjustment operation is similar to the adjustment method for the first editable element in operation S610, and the layout adjustment of the second editable elements is achieved by applying parameters such as position adjustment amount and size adjustment ratio in the second adjustment parameters.
[0127] It should be noted that by performing a two-stage adjustment process—first making preliminary adjustments based on the first set of adjustment parameters to obtain the second document, and then making fine adjustments based on the second set of adjustment parameters to obtain the target document—the layout quality of the document can be gradually optimized, resulting in a more reasonable element layout in the final target document.
[0128] Optionally, after adjusting according to the second adjustment parameter, the layout quality of the target document can be evaluated. For example, the same reward function used when training the prediction model can be used to score the layout quality of the target document. If the score does not reach the expected threshold, the prediction operation of the prediction model can be re-executed to obtain a new second adjustment parameter and adjust it until the layout quality of the target document meets the requirements.
[0129] By adopting the technical solution of this embodiment, the first editable element is adjusted according to the first adjustment parameter to obtain the second document, which can complete the initial optimization of the document layout. Furthermore, the second adjustment parameter is predicted by the prediction model trained based on the sample document and the reward function, and the second editable element is adjusted according to the second adjustment parameter to obtain the target document. This allows the target document to undergo two stages of layout adjustment, which not only completes the overall layout optimization at the coarse-grained level, but also achieves the fine-grained adjustment at the fine-grained level, thereby effectively improving the accuracy of the document layout adjustment and the rationality of the final layout.
[0130] Based on the above embodiments, as an optional embodiment, in order to enable the prediction model to comprehensively consider the overall page layout and the characteristics of individual elements, the above operation S620 may further include the following operations.
[0131] Operation S710 obtains the page state features of the second page and the second target element features corresponding to each second editable element for the second page of the second document; the page state features characterize the overall layout features of the second page.
[0132] Operation S720 inputs the page state features and the second target element features into the prediction model to obtain the second adjustment parameters of the second editable element within the second page.
[0133] In operation S710, the second page refers to a page in the second document. The second document may include one or more second pages. When the second document includes multiple second pages, feature acquisition and parameter adjustment prediction operations can be performed separately for each second page.
[0134] Page state features refer to the feature information that characterizes the overall layout state of the second page. They can be understood as feature representations that reflect the overall layout structure, density distribution, spatial relationships, and other information of multiple second editable elements on the second page, and are used to provide page-level layout background information for the prediction model.
[0135] For example, page state features may include information such as the number of second editable elements in the second page, the overall distribution density among the second editable elements, and the distribution of white space on the second page.
[0136] In one feasible implementation, the second page can be processed by a feature extraction module to extract its page state features. In this implementation, the feature extraction module can perform a holistic analysis of multiple editable elements on the second page and extract feature information reflecting the overall layout characteristics of the page as page state features.
[0137] In another feasible implementation, such as Figure 3As shown, the page state features of the second page can also be obtained by aggregating the second target element features of multiple second editable elements within the second page. In this embodiment, the second target element features of all second editable elements within the second page can be concatenated, then the concatenated features can be encoded by a transformer encoder, and finally processed by a multilayer perceptron to obtain the page state features of the second page.
[0138] It should be noted that, Figure 3 The processing flow shown can be used for processing from the first editable element to the first adjustment parameter, or for processing from the second editable element to the second adjustment parameter.
[0139] The second target element feature refers to the comprehensive feature information that integrates the characteristics of the second editable element itself and the relationship between the element and other elements. It can be understood as a feature representation that simultaneously includes the attribute information, semantic information, and contextual information of the element in the overall layout of the page.
[0140] In one feasible implementation, it can be adopted with Figure 3 The processing method shown extracts the second target element features of the second editable element.
[0141] In another feasible implementation, the second target element features can be extracted for each second editable element in the second page. In this implementation, for each second editable element, its geometric parameters and content information can be obtained, and combined with the association between the element and other second editable elements in the second page, the second target element features of the second editable element can be obtained.
[0142] Optionally, the extraction method for the second target element features can be similar to the extraction method for the first target element features. For example, the attribute features and semantic features of the second editable element can be obtained first, and then combined with the association features of this element with other elements to obtain the second target element features through feature fusion.
[0143] It should be noted that the page state features and the second target element features describe the layout information of the second page from different levels. The page state features provide the overall layout background, while the second target element features provide the adjustment requirements for the perspective of individual elements. The combination of the two can provide more comprehensive input information for the prediction model.
[0144] In operation S720, the obtained page state features and the second target element features corresponding to each second editable element can be input into the prediction model, processed by the prediction model, and the second adjustment parameters of each second editable element in the second page can be output.
[0145] In one feasible implementation, for each second editable element on the second page, the second target element features of that element and the page state features of the second page can be input into the prediction model, and the prediction model can output the second adjustment parameters of that second editable element. In this implementation, the prediction model can simultaneously consider the characteristics of individual elements and the overall layout state of the page, thereby predicting more reasonable adjustment parameters.
[0146] Optionally, when inputting features into the prediction model, the page state features and the second target element features can be concatenated or fused. For example, the page state features and the second target element features can be vector-concatenated to form an input vector containing information from both, and then this input vector can be input into the prediction model.
[0147] It should be noted that the prediction model has learned during training how to predict reasonable second adjustment parameters based on page state features and second target element features, so that the predicted second adjustment parameters can meet the adjustment needs of individual elements while maintaining the overall layout coordination of the page.
[0148] Please refer to Figure 4 , Figure 4 This is a schematic diagram illustrating the process of a prediction model predicting a second adjustment parameter, as provided in an embodiment of this application.
[0149] like Figure 4 As shown, the input to the prediction model includes the page state features of the second page and the second target element features of the second editable element, and the output is the second adjustment parameter.
[0150] Specifically, when predicting the second adjustment parameter, the page state features and the second target element features are first fused to obtain a fused feature representation, which is then input into the prediction model. Based on the input feature representation and the optimization objective provided by the reward function, the prediction model outputs the second adjustment parameter.
[0151] After obtaining the second adjustment parameter, such as Figure 4 As shown, iterative adjustments can be made through a feedback loop. During this iterative adjustment process, the second editable element, adjusted according to the second adjustment parameters, is re-input into the prediction process. The adjusted page state features and second target element features are then obtained again, and the prediction model outputs new second adjustment parameters. Through this iterative process, the layout of the second editable element can be continuously fine-tuned, gradually optimizing the layout effect of the second page.
[0152] In one feasible implementation, a termination condition for iterative adjustment can be set. For example, a maximum number of iterations can be set, and iteration can stop when the number of iterations reaches a preset threshold; or the change in the second adjustment parameter output between two adjacent iterations can be monitored, and the layout is considered to have converged when the change is less than a preset value, and iterative adjustment can stop.
[0153] By adopting the technical solution of this embodiment, page state features and second target element features are obtained for the second page, and the layout information of the second page can be captured from both the overall and local levels. Furthermore, the page state features and second target element features are input into the prediction model to obtain the second adjustment parameters, so that the prediction model can comprehensively consider the overall layout structure of the page and the adjustment requirements of individual elements, thereby predicting more accurate second adjustment parameters.
[0154] Based on the above embodiments, as an optional embodiment, the above prediction model is trained through the following operations.
[0155] Operate S810 to obtain a sample document; the sample document includes multiple editable sample elements.
[0156] Operate S820 to input the sample document into the initial prediction model and obtain the candidate adjustment parameters for each editable element of the sample;
[0157] The S830 is operated to update the parameters of the initial model based on the reward function to obtain the prediction model.
[0158] Here, the prediction model refers to a neural network model capable of predicting adjustment parameters based on the document's layout. The prediction model can employ a policy network structure, which outputs the probability of selecting different adjustment parameters. In this implementation, the policy network maps the input document features to a probability distribution of adjustment parameters, thereby achieving the prediction of the adjustment parameters.
[0159] Optionally, the input to the prediction model may include page state features and target element features of each editable element. For example, the prediction model may concatenate or fuse the page state features with the target element features of each editable element before inputting them into a policy network for processing to obtain the adjustment parameters for that editable element.
[0160] It should be noted that the output of the prediction model consists of adjustment parameters, which may include position adjustment amount and size adjustment ratio. For example, the position adjustment amount may include horizontal position adjustment amount and vertical position adjustment amount, and the size adjustment ratio is used to characterize the size change ratio of the editable element.
[0161] In operation S810, a sample document refers to a document sample used to train a prediction model, which can be understood as document data containing multiple editable sample elements.
[0162] The editable elements in a sample document are editable element objects, which can be understood as element objects of the same type as the first or second editable elements mentioned above, and are used as adjustment objects during the training process.
[0163] In one feasible implementation, sample documents can be obtained from a document dataset, containing sample editable elements with different layout states. In this implementation, the sample documents can include documents with good layout quality and documents with poor layout quality, thereby providing diverse sample data for training the predictive model.
[0164] In operation S820, the initial prediction model refers to the prediction model whose parameters have not been trained or have been randomly initialized. It can be understood as the prediction model state at the beginning of training, which is used to gradually optimize it through parameter updates during training.
[0165] Candidate adjustment parameters refer to the adjustment parameters predicted by the initial prediction model based on the sample documents. They can be understood as the adjustment parameters generated by the initial prediction model for the editable elements of the sample under the current parameter state.
[0166] In one feasible implementation, the feature information of each editable element in the sample document can be input into an initial prediction model, which then outputs candidate adjustment parameters for that editable element. In this implementation, the initial prediction model processes the data similarly to how the aforementioned prediction model processes the second editable element, by receiving feature input and outputting adjustment parameters.
[0167] Optionally, before inputting the sample document into the initial prediction model, feature extraction processing can be performed on the sample document. For example, the same feature extraction method as in the foregoing embodiments can be used to obtain the page state features of the sample document and the target element features of each editable element of the sample, and then these features can be input into the initial prediction model.
[0168] In operating S830, the parameters of the initial prediction model can be updated based on the reward function to obtain the trained prediction model.
[0169] The reward function is a function used to evaluate the layout quality of a sample document. It can be understood as a function that calculates a score based on the layout state of the editable elements in the sample, and is used to guide the optimization direction of the parameters of the prediction model.
[0170] In one feasible implementation, editable elements in a sample document can be adjusted based on candidate adjustment parameters to obtain an adjusted sample document. Then, a reward function is used to evaluate the layout quality of the adjusted sample document, yielding a reward score. In this implementation, the reward score reflects the quality of the candidate adjustment parameters; a higher score indicates better layout quality after adjustment.
[0171] Optionally, reinforcement learning can be used to update the parameters of the initial prediction model. For example, the obtained reward score can be used as the reward signal for reinforcement learning, and the parameters of the initial prediction model can be updated through reinforcement learning algorithms such as policy gradient, so that the adjusted parameters output by the prediction model can obtain a higher reward score.
[0172] It should be noted that the parameter update process can be iterative. In each iteration, sample documents are input into the current prediction model to obtain candidate adjustment parameters, a reward score is calculated based on the reward function, and then the parameters of the prediction model are updated according to the reward score. Through multiple iterations, the parameters of the prediction model are gradually optimized until the training termination condition is met, thus obtaining the final prediction model.
[0173] By adopting the technical solution of this embodiment, a sample document containing multiple editable elements can be obtained, which can provide a data foundation for training the prediction model. Furthermore, the sample document is input into the initial prediction model to obtain candidate adjustment parameters, and the parameters of the initial prediction model are updated based on the reward function, so that the prediction model can learn how to predict reasonable adjustment parameters according to the layout state of the document, thereby improving the prediction accuracy of the prediction model for the second adjustment parameter.
[0174] Based on the above embodiments, as an optional embodiment, in order to evaluate the layout quality of sample documents from multiple dimensions, the above reward function includes a combination of at least one of the following.
[0175] The alignment reward function is used to characterize the deviation value of the boundary position of each editable element of the same type within the sample document. The deviation value of the boundary position is negatively correlated with the function value of the alignment reward function.
[0176] The consistency reward function is used to characterize the differences in style parameters among editable elements of the same type within a sample document. The differences in style parameters are negatively correlated with the value of the consistency reward function. Style parameters include at least one of font size, width, and line spacing.
[0177] The visual hierarchical reward function is used to characterize the difference in transparency between the categories of editable elements in each sample and the corresponding transparency values; the difference in transparency values is negatively correlated with the function value of the visual hierarchical reward function.
[0178] The overlap penalty function is used to characterize the area of the overlapping region between any two editable elements in a sample document. The area of the overlapping region is negatively correlated with the value of the overlap penalty function.
[0179] The out-of-bounds penalty function is used to characterize the area of the editable element in each sample that exceeds the page boundary. The area of the area exceeding the page boundary is negatively correlated with the function value of the out-of-bounds penalty function.
[0180] In this embodiment, the alignment reward function refers to a function used to evaluate the degree of alignment between editable elements of the same type. It can be understood as a function that calculates a reward value based on the boundary position deviation of elements of the same type.
[0181] Boundary position refers to the location of the boundary of the editable element of the sample, which may include at least one of the following: left boundary position, right boundary position, top boundary position, and bottom boundary position.
[0182] In one feasible implementation, the deviation values of the boundary positions of editable elements of the same type within a sample document can be calculated, with smaller deviation values indicating higher alignment. In this implementation, the alignment reward function increases as the deviation value decreases, thereby awarding a higher reward to layouts with high alignment.
[0183] Optionally, editable sample elements of the same type can include elements with the same hierarchical relationship. For example, for a sample document containing heading elements and body elements, the boundary position deviation values between all heading elements and the boundary position deviation values between all body elements can be calculated separately.
[0184] The consistency reward function is a function used to evaluate the style consistency between editable elements of the same type. It can be understood as a function that calculates a reward value based on the differences in style parameters of elements of the same type.
[0185] Style parameters refer to parameters that characterize the visual style of the editable elements in the sample, and may include at least one of font size, width, and line spacing.
[0186] In one feasible implementation, the difference in style parameters between editable elements of the same type within a sample document can be calculated; a smaller difference value indicates higher style consistency. In this implementation, the consistency reward function increases as the difference value decreases, thereby awarding a higher reward to layouts with high style consistency.
[0187] The visual hierarchy reward function refers to a function used to evaluate the reasonableness of the visual hierarchy of sample editable elements.
[0188] Transparency refers to a parameter that characterizes the visual prominence of editable elements in a sample. It can be understood as a numerical value that reflects the visual salience of an element on a page and is used to distinguish elements at different levels.
[0189] In one feasible implementation, the target transparency that an editable element should have can be determined based on the category of the sample, and then the difference between the actual transparency and the target transparency can be calculated. In this implementation, the function value of the visual hierarchy reward function increases as the difference decreases, thereby giving a higher reward to a reasonable layout in terms of visual hierarchy.
[0190] Optionally, different categories of editable elements can correspond to different target transparency ranges. For example, title elements can correspond to higher target transparency for a more prominent visual effect, body text elements can correspond to medium target transparency, and annotation elements can correspond to lower target transparency.
[0191] The overlap penalty function is a function used to evaluate the degree of overlap between sample editable elements.
[0192] The overlapping area refers to the area of the intersection of the bounding boxes of any two editable sample elements.
[0193] In one feasible implementation, the area of the overlapping region between any two editable elements in the sample document can be calculated; a larger overlapping area indicates a more severe degree of overlap. In this implementation, the value of the overlap penalty function decreases as the overlapping area increases, thereby assigning a lower score to layouts with severe overlap.
[0194] The out-of-bounds penalty function is a function used to evaluate how far a sample editable element exceeds the page boundaries.
[0195] In one feasible implementation, the area of each editable element that extends beyond the page boundary can be calculated; a larger area indicates a more severe boundary violation. In this implementation, the boundary violation penalty function decreases as the area extends beyond the boundary, thus assigning a lower score to layouts exhibiting boundary violations.
[0196] It should be noted that the reward function can be constructed by a weighted combination of the above-mentioned sub-functions. For example, weight coefficients can be set for the alignment reward function, consistency reward function, visual hierarchy reward function, overlap penalty function, and out-of-bounds penalty function, and the final reward function value can be obtained by weighted summation, thereby comprehensively evaluating the layout quality of the sample document.
[0197] Optionally, the weight coefficients of each sub-function can be adjusted according to actual application requirements. For example, if more attention is paid to the element alignment effect, the weight coefficient of the alignment reward function can be increased; if more attention is paid to avoiding element overlap, the weight coefficient of the overlap penalty function can be increased.
[0198] By adopting the technical solution of this embodiment, the alignment and style consistency of the editable elements of the sample can be evaluated through the alignment reward function and the consistency reward function, the visual hierarchy of the elements can be evaluated through the visual hierarchy reward function, and the overlap and out-of-bounds penalty function can be evaluated through the overlap and out-of-bounds penalty function. Thus, the reward function can comprehensively evaluate the layout quality of the sample document from multiple dimensions, providing an accurate optimization target for the training of the prediction model.
[0199] Please refer to Figure 5 , Figure 5 This is a schematic diagram of the interactive flow of a document processing device provided in an embodiment of this application.
[0200] like Figure 5 As shown, the document processing device is deployed on electronic devices and servers, and the electronic devices and servers interact with each other via a network.
[0201] In this embodiment, the server refers to a cloud server device with strong computing power, which can be understood as a backend server capable of performing complex feature encoding and parameter prediction tasks, used to complete computationally intensive processing operations such as document layout adjustment.
[0202] Electronic devices refer to user-side terminal devices, which can be understood as local devices capable of document parsing and editing, used to complete local document processing and interaction with the user. For example, electronic devices may include at least one of a personal computer, mobile terminal, tablet computer, or workstation.
[0203] In one feasible implementation, such as Figure 5 As shown, the electronic device receives a document query instruction and uploads it to the server. In this embodiment, the user inputs a document query instruction on the electronic device, which describes the content requirements or topic information of the document to be generated.
[0204] After receiving a document query command, the server generates HTML content through a large language model proxy, generates an initial slideshow based on the HTML content, and downloads the initial slideshow to the electronic device. In this implementation, as... Figure 5 As shown, the server generates an HTML document containing text content and structural tags based on the document query command through a large language model agent. The HTML document is then converted into a slide format to obtain an initial slide containing multiple editable elements, which serves as the first PPT.
[0205] After downloading the first PowerPoint presentation, the electronic device parses the presentation to obtain the attribute information of each editable element and uploads this information to the server. In this embodiment, as shown... Figure 5 As shown, the electronic device opens the first PPT in the local environment, identifies each of the first editable elements in the first PPT, and extracts the geometric parameters and content information of each first editable element.
[0206] After receiving the attribute information, the server processes the attribute information using an encoder and a predictor, predicts the first adjustment parameter, and downloads it to the electronic device. In this embodiment, as... Figure 5 As shown, the server performs feature encoding on the attribute information and predicts the first adjustment parameter of each first editable element based on the first target element features obtained from the encoding.
[0207] After receiving the first adjustment parameter, the electronic device performs a first editing operation on the first PPT according to the first adjustment parameter to obtain the second PPT. In this embodiment, as follows... Figure 5 As shown, the electronic device performs adjustment operations on each of the first editable elements in the first PPT, adjusting the position and size of the elements according to the first adjustment parameters.
[0208] Optionally, such as Figure 5 As shown, the electronic device parses the second PPT to obtain its attribute information and uploads this information to the server. The server then predicts second adjustment parameters using a reinforcement learning predictor and downloads these parameters to the electronic device. Based on these second adjustment parameters, the electronic device performs a second editing operation on the second PPT to obtain the target PPT.
[0209] Optionally, such as Figure 5 As shown, the electronic device uploads the target PPT to the server. The server generates a preview template through the rendering module and downloads the preview template to the electronic device for display. In this embodiment, the server renders the target PPT to generate a template file for preview display. After receiving the preview template, the electronic device displays it locally, allowing the user to view the final document layout effect.
[0210] It should be noted that by performing document parsing and editing operations on the electronic device and feature encoding and parameter prediction operations on the server, a reasonable allocation of computational tasks can be achieved. Compared to the solution of using the electronic device alone for processing, when performing feature encoding and parameter prediction operations locally on the electronic device, there are performance issues when processing complex slides due to the limited computing power of the electronic device's processor, especially for slides containing a large number of editable elements. The local processor struggles to complete complex neural network calculations within a reasonable time, leading to processing delays or even failures. In contrast, this embodiment utilizes the server's high-performance computing resources to perform feature encoding and parameter prediction operations, ensuring processing efficiency.
[0211] Compared to solutions that use a dedicated server for processing, when document parsing and editing operations are performed on the server, the differences between the server environment and the local environment of the electronic device, including differences in operating system version, document processing software version, font library configuration, etc., may lead to inconsistencies between the parsing results and the actual document display, resulting in formatting errors or garbled characters. In contrast, this embodiment ensures that the processing results are fully compatible with the document processing environment of the electronic device by performing document parsing and editing operations locally on the electronic device, adapting to the display styles of different client applications.
[0212] Furthermore, compared to solutions that directly adjust document layout using large language models, which primarily process text semantics and have a poor understanding of geometric relationships such as spatial position and size ratios between editable elements, making it difficult to accurately predict element position and size adjustment parameters, this embodiment uses a specially trained encoder and predictor on the server to perform precise calculations based on the geometric parameters and semantic features of elements, accurately capturing the spatial relationships between elements, thereby significantly improving the understanding of element spatial relationships and the accuracy of predicting adjustment parameters.
[0213] By adopting the technical solution of this embodiment, the document layout adjustment process can be reasonably distributed to the cloud and local through the collaborative work of the server and electronic devices. Furthermore, the electronic devices perform document parsing and editing operations locally, which can ensure the compatibility of the processing results with the local environment and the adaptability of the application style. The server performs feature encoding and parameter prediction operations, which can utilize high-performance computing resources to improve processing efficiency. Thus, while ensuring environmental adaptability, the accuracy and processing speed of document layout adjustment are effectively improved.
[0214] Please refer to Figure 6 , Figure 6 This is a block diagram of a document processing apparatus provided in an embodiment of this application.
[0215] This application also discloses a document processing apparatus, including:
[0216] The module is used to obtain the first document to be processed; the first document includes multiple first editable elements.
[0217] The extraction module is used to obtain the first element features of each first editable element, as well as the fused features of multiple first editable elements;
[0218] A determination module is used to determine the first target element feature of each first editable element based on the fusion feature and the feature of each first element;
[0219] The prediction module is used to determine the first adjustment parameter for each first editable element based on the features of the first target element;
[0220] The adjustment module is used to adjust each first editable element according to the first adjustment parameters to obtain the target document.
[0221] Based on the above embodiments, as an optional embodiment, the extraction module is further configured to obtain the text features and image features of each first editable element; and to obtain the first element features of the first editable element based on the text features and image features of the first editable element.
[0222] Based on the above embodiments, as an optional embodiment, the extraction module is further configured to obtain semantic features based on the text features and image features of the first editable element; obtain attribute features of the first editable element, wherein the attribute features characterize the geometric parameters of the first editable element within the corresponding first page of the first document; and concatenate the attribute features and semantic features of the first editable element to obtain the first element feature of the first editable element.
[0223] Based on the above embodiments, as an optional embodiment, the determining module is further configured to encode the fusion features to obtain an encoding sequence; the encoding sequence includes the associated features corresponding to each first editable element; the associated features characterize the association relationship between the corresponding first editable element and other first editable elements; and each first element feature is concatenated with the corresponding associated features to obtain the first target element feature of each first editable element.
[0224] Based on the above embodiments, as an optional embodiment, the adjustment module is further configured to adjust each first editable element according to the first adjustment parameters to obtain a second document; the second document includes multiple second editable elements; predict the second adjustment parameters of each second editable element through a prediction model; and adjust each second editable element according to the second adjustment parameters to obtain a target document; wherein, the prediction model is trained based on sample documents and a reward function; the reward function is used to evaluate the layout quality of sample editable elements in the sample document.
[0225] Based on the above embodiments, as an optional embodiment, the adjustment module is further configured to obtain, for the second page of the second document, the page state features of the second page and the second target element features corresponding to each second editable element; the page state features characterize the overall layout features of the second page; and input the page state features and the second target element features into the prediction model to obtain the second adjustment parameters of the second editable elements within the second page.
[0226] The embodiments of this application have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of this application. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. The scope of this application is defined by the appended claims and their equivalents. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this application, and all such substitutions and modifications should fall within the scope of this application.
Claims
1. A document processing method, comprising: Obtain the first document to be processed; The first document includes a plurality of first editable elements; Obtain the first element feature of each of the first editable elements, and the fusion feature of the plurality of first editable elements; Based on the fusion features and each of the first element features, determine the first target element features for each of the first editable elements; Based on the features of the first target element, determine the first adjustment parameter for each of the first editable elements; The target document is obtained by adjusting each of the first editable elements according to the first adjustment parameters.
2. The method according to claim 1, wherein obtaining the first element feature of each first editable element comprises: For each of the first editable elements, obtain the text features and image features of the first editable element; Based on the text features and image features of the first editable element, the first element feature of the first editable element is obtained.
3. The method according to claim 2, wherein obtaining the first element feature of the first editable element based on the text features and image features of the first editable element includes: Based on the text and image features of the first editable element, semantic features are obtained; Obtain the attribute features of the first editable element, wherein the attribute features characterize the geometric parameters of the first editable element within the corresponding first page of the first document; The attribute features and semantic features of the first editable element are concatenated to obtain the first element feature of the first editable element.
4. The method according to claim 1, wherein determining the first target element feature of each first editable element based on the fusion feature and each first element feature comprises: The fused features are encoded to obtain an encoded sequence; the encoded sequence includes the associated features corresponding to each of the first editable elements. The association feature represents the association relationship between the corresponding first editable element and other first editable elements; Each of the first element features is concatenated with the corresponding associated feature to obtain the first target element feature of each of the first editable elements.
5. The method according to claim 1, wherein adjusting each of the first editable elements according to the first adjustment parameter to obtain the target document comprises: Based on the first adjustment parameter, each of the first editable elements is adjusted to obtain a second document; the second document includes multiple second editable elements; The second adjustment parameter for each of the second editable elements is predicted using a predictive model; The target document is obtained by adjusting each of the second editable elements according to the second adjustment parameters. in, The prediction model is trained based on sample documents and a reward function; the reward function is used to evaluate the layout quality of editable elements in the sample documents.
6. The method of claim 5, wherein predicting the second adjustment parameter for each of the second editable elements using a prediction model comprises: For the second page of the second document, obtain the page state features of the second page and the second target element features corresponding to each second editable element; The page state features characterize the overall layout features of the second page; The page state features and the second target element features are input into the prediction model to obtain the second adjustment parameters of the second editable element within the second page.
7. The method according to claim 5, wherein the prediction model is trained in the following manner: Obtain a sample document; the sample document includes multiple editable sample elements; The sample documents are input into the initial prediction model to obtain candidate adjustment parameters for each editable element of the sample; The prediction model is obtained by updating the parameters of the initial model based on the reward function.
8. The method of claim 5, wherein the reward function comprises a combination of at least one of the following: The alignment reward function is used to characterize the deviation value of the boundary position of each editable element of the same type in the sample document, and the deviation value of the boundary position is negatively correlated with the function value of the alignment reward function. A consistency reward function is used to characterize the difference in style parameters of editable elements of the same type within the sample document. The difference in style parameters is negatively correlated with the function value of the consistency reward function. The style parameters include at least one of font size, width, and line spacing. A visual hierarchical reward function is used to characterize the difference in transparency between the categories of editable elements in each sample and the corresponding transparency values; the difference in transparency values is negatively correlated with the function value of the visual hierarchical reward function. An overlap penalty function is used to characterize the area of the overlapping region between any two editable elements in the sample document, and the area of the overlapping region is negatively correlated with the function value of the overlap penalty function. The out-of-bounds penalty function is used to characterize the area of the editable element of each sample that exceeds the page boundary. The area of the area exceeding the page boundary is negatively correlated with the function value of the out-of-bounds penalty function.
9. The method according to claim 1, wherein the first adjustment parameter comprises at least one of the following: Adjust the position, size, rotation angle, and transparency of the first editable element.
10. A document processing apparatus, comprising: The `get` module is used to obtain the first document to be processed. The first document includes a plurality of first editable elements; An extraction module is used to obtain the first element features of each of the first editable elements, and the fusion features of the plurality of first editable elements; The determining module is configured to determine a first target element feature for each first editable element based on the fusion feature and each first element feature; The prediction module is used to determine a first adjustment parameter for each of the first editable elements based on the features of the first target element; The adjustment module is used to adjust each of the first editable elements according to the first adjustment parameters to obtain the target document.