Office document rapid arrangement method, medium and system
Document features are extracted through multi-branch neural network and contribution evaluation model, and combined with rendering optimization strategies, the problem of low rendering performance in the existing technology is solved, and efficient document rendering and optimization orchestration effects are achieved.
Patent Information
- Application Number
- CN202510130869.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to improve rendering performance while ensuring orchestration quality, especially when processing complex documents, which has a long response time and affects the user experience.
A multi-branch neural network is used to construct a feature processing model, extract multimodal features of layout elements such as text, tables, and pictures in the document, and analyze the contribution of each element to document layout through the contribution evaluation model. Combined with the rendering optimization equation system, block size calculation, hardware task allocation, calculation overhead estimate and execution sequence optimization are carried out, and rendering calculation is adopted using a multi-threaded parallel mechanism.
It achieves the realization of significantly improving rendering performance while ensuring orchestration quality, shortening the response time of document rendering, and improving user experience.
Smart Images

Figure CN120146005A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of document layout, and more particularly, relates to a method, medium, and system for quickly laying out office documents. Background Art
[0002] As an important tool for improving work efficiency, office automation software is widely used in enterprise daily management and personal office work. In scenarios such as document writing, editing, and typesetting, users need to manually adjust many details such as page layout, font style, and paragraph format. This manual operation method is not only inefficient but also prone to problems of inconsistent layout caused by subjective preferences and experience differences. Therefore, there is an urgent need to develop an intelligent office document layout technology that can automatically analyze the document structure and content characteristics and give an optimized typesetting plan.
[0003] Existing document layout methods mainly focus on two directions: one is rule-based layout optimization, and the other is intelligent recommendation using machine learning. The former automatically typesets the document through preset layout rules, paragraph formats, etc., but the rule design is cumbersome and it is difficult to cover various complex situations. The latter attempts to use models such as neural networks to learn from a large number of manually typeset samples and output typesetting suggestions, but these methods are limited to single layout elements and it is difficult to comprehensively consider the characteristics of various mixed elements such as text, tables, and pictures. In addition, existing document layout methods generally ignore the optimization of rendering efficiency. As the document content becomes more and more rich and the page layout becomes more and more complex, using the traditional sequential rendering method will result in a long response time, seriously affecting the user experience. Therefore, how to improve the rendering performance while ensuring the layout quality is also a key problem to be solved urgently. Summary of the Invention
[0004] In view of this, the present invention provides a method, medium, and system for quickly laying out office documents, which can solve the technical problem in the prior art that it is difficult to improve the rendering performance while ensuring the layout quality.
[0005] The present invention is implemented as follows:
[0006] The present invention provides a method for quickly arranging office documents, which includes: parsing the OpenXML tags of the document, extracting the node data in document.xml for feature extraction, and constructing the original feature data; constructing a feature processing model including a table branch neural network, a text branch neural network, a picture branch neural network, a feature fusion layer, and a scoring output layer; constructing a contribution degree evaluation model to analyze the contribution degree value of the layout elements in the document to the document layout score; obtaining the contribution degree evaluation model through training; determining the layout priority level and layout interval value according to the contribution degree value; generating the OpenXML tag tree of the document; performing dimensionality reduction mapping on the document feature vector and mapping it to a three-dimensional display space to generate a three-dimensional layout mapping model; performing partition rendering calculation on the typeset result document through a rendering optimization equation set including a block size equation, a hardware allocation equation, a calculation overhead equation, and an execution order equation, determining the graphics card rendering block, the processor rendering block, the memory storage block and their processing order, and performing rendering calculation using a multi-threaded parallel mechanism.
[0007] Among them, the table branch neural network in the feature processing model includes a horizontal table model, a vertical table model, and a table feature fusion layer. The horizontal table model analyzes the row data attributes of the w:tr node and the cell data attributes of the w:tc node. The vertical table model analyzes the column data attributes of the w:gridCol node and the merged data attributes of the w:vMerge node. The table feature fusion layer fuses the row and column feature data output by the horizontal table model and the merged feature data output by the vertical table model, and outputs the table feature vector.
[0008] Furthermore, the text branch neural network in the feature processing model includes a text extraction layer, a text semantics layer, and a text structure layer. The text extraction layer analyzes the text content data and format data. The text semantics layer calculates the semantic connection value of the text paragraph. The text structure layer analyzes the paragraph data attributes of the w:pPr node, the interval data attributes of the w:spacing node, and the indent data attributes of the w:ind node, and outputs the text feature vector.
[0009] Furthermore, the picture branch neural network in the feature processing model includes a picture extraction layer, a picture positioning layer, and a picture-text connection layer. The picture extraction layer analyzes the picture size data and clarity data. The picture positioning layer analyzes the picture coordinate position data. The picture-text connection layer analyzes the positioning data attributes of the w:anchor node, the horizontal data attributes of the wp:positionH node, and the vertical data attributes of the wp:positionV node, and outputs the picture feature vector.
[0010] Further, the feature fusion layer uses a feature dimensionality reduction algorithm to map the table feature vector, the text feature vector, and the picture feature vector into a three-dimensional feature space. The feature dimensionality reduction algorithm includes a principal feature analysis method and a random embedding method. The principal feature analysis method is used to extract the main direction of high-dimensional features, and the random embedding method is used to maintain the local structure of features. The output of the feature fusion layer is the document feature vector.
[0011] Further, the rendering optimization equation set includes a block size equation, a hardware allocation equation, a calculation overhead equation, and an execution order equation. The block size equation is used to calculate the size value of the rendering calculation block. The hardware allocation equation is used to calculate the task allocation ratio of hardware devices. The calculation overhead equation is used to calculate the calculation consumption of the rendering calculation block. The execution order equation is used to optimize the calculation order of the rendering calculation block.
[0012] Further, the feature data acquisition step includes: collecting manually arranged documents as basic data; extracting the OpenXML tags of the basic data and recording the scoring data; dividing the basic data into a standard sample data set and an optimization sample data set; extracting the position data, spacing data, and style data of layout elements; and constructing the original feature data.
[0013] Further, the model training step includes: training a feature processing model using the original feature data; constructing training data for a contribution degree evaluation model based on the calculation results of the feature processing model; training the contribution degree evaluation model using the original feature data; evaluating performance indicators through a cross-validation method; and saving the training parameter data.
[0014] The block size equation is used to calculate the size value of the rendering calculation block. The input data of the block size equation includes the display device resolution value, the video card memory capacity value, and the system memory capacity value. The output data of the block size equation is the rendering block pixel size value;
[0015] The hardware allocation equation is used to calculate the task allocation ratio of hardware devices. The input data of the hardware allocation equation includes the video card calculation performance value, the number of processor cores value, the system memory capacity value, and the document content complexity value. The output data of the hardware allocation equation is the video card task ratio, the processor task ratio, and the memory task ratio;
[0016] The calculation overhead equation is used to calculate the calculation consumption of the rendering calculation block. The input data of the calculation overhead equation includes the number of rendering characters value, the rendering table complexity value, and the rendering picture resolution value. The output data of the calculation overhead equation is the estimated rendering calculation duration value;
[0017] The execution order equation is used to optimize the calculation order of the rendering calculation block. The input data of the execution order equation includes the estimated value of the rendering calculation duration, the graphics card task ratio, the processor task ratio, the memory task ratio, and the rendering data dependency relationship value. The output data of the execution order equation is the rendering calculation block execution order value;
[0018] The table branch neural network includes a horizontal table model, a vertical table model, and a table feature fusion layer. The horizontal table model analyzes the row data attributes of the w:tr node and the cell data attributes of the w:tc node. The vertical table model analyzes the column data attributes of the w:gridCol node and the merged data attributes of the w:vMerge node. The table feature fusion layer fuses the row-column feature data output by the horizontal table model and the merged feature data output by the vertical table model to output the table feature vector;
[0019] The text branch neural network includes a text extraction layer, a text semantics layer, and a text structure layer. The text extraction layer analyzes the text content data and format data. The text semantics layer calculates the semantic connection value of the text paragraph. The text structure layer analyzes the paragraph data attributes of the w:pPr node, the spacing data attributes of the w:spacing node, and the indent data attributes of the w:ind node to output the text feature vector;
[0020] The picture branch neural network includes a picture extraction layer, a picture positioning layer, and a picture-text connection layer. The picture extraction layer analyzes the picture size data and clarity data. The picture positioning layer analyzes the picture coordinate position data. The picture-text connection layer analyzes the positioning data attributes of the w:anchor node, the horizontal data attributes of the wp:positionH node, and the vertical data attributes of the wp:positionV node to output the picture feature vector;
[0021] The feature fusion layer uses a feature dimensionality reduction algorithm to map the table feature vector, the text feature vector, and the picture feature vector to a three-dimensional feature space. The feature dimensionality reduction algorithm includes a principal feature analysis method and a random embedding method. The principal feature analysis method is used to extract the main direction of high-dimensional features, and the random embedding method is used to maintain the local structure of features. The output of the feature fusion layer is the document feature vector;
[0022] The scoring output layer calculates the document layout score based on the document feature vector. The document layout score represents the quality level of the document layout;
[0023] The document content complexity value is obtained by calculating the weighted sum of the number of text layout elements, the number of table layout elements, and the number of picture layout elements;
[0024] The numerical value of the rendering data dependency is calculated according to the spatial position relationship and the logical sequence relationship between the layout elements;
[0025] The numerical value of the number of rendering characters, the numerical value of the complexity of the rendering table, and the numerical value of the resolution of the rendering picture are directly extracted from the typesetting result document.
[0026] The present invention provides a computer-readable storage medium, wherein program instructions are stored in the computer-readable storage medium, and when the program instructions run on a computer, they are used to execute the above-mentioned method for quickly typesetting an office document.
[0027] The present invention provides an office document quick typesetting system, which includes the above-mentioned computer-readable storage medium.
[0028] Compared with the prior art, the beneficial effects of the office document quick typesetting method, medium and system provided by the present invention are as follows: The intelligent office document typesetting method proposed by the present invention has made remarkable progress in solving the problems existing in the prior art;
[0029] First of all, this method can comprehensively extract the multi-modal features of the document content, including the visual, structural and semantic features of various layout elements such as text, tables, and pictures, and perform deep fusion through a multi-branch neural network to generate an overall feature representation of the document. Compared with the existing typesetting methods based on a single layout element, this multi-modal feature fusion method can more accurately reflect the overall typesetting characteristics of the document;
[0030] Secondly, this method constructs a contribution degree evaluation model, which can accurately analyze the actual contribution degree of each layout element to the document typesetting quality, providing a reliable basis for subsequent layout optimization. Existing typesetting methods based on empirical rules or machine learning are difficult to accurately capture the importance of different layout elements and are prone to improper typesetting results. The contribution degree evaluation method of the present invention can objectively quantify the priority of each element, thereby generating a more reasonable layout scheme;
[0031] In addition, aiming at the rendering efficiency problem of complex documents, this method proposes a series of optimization strategies. Through measures such as block size calculation, hardware task allocation, calculation overhead estimation, and execution order optimization, efficient parallel processing of rendering tasks is achieved. Compared with the traditional sequential rendering method, this method significantly shortens the response time of document rendering and effectively improves the user experience;
[0032] Generally speaking, the intelligent office document layout method proposed by the present invention makes full use of technical means such as multi-modal feature analysis, contribution degree evaluation, and rendering optimization, and has made remarkable progress in two key dimensions of layout quality and rendering performance, solving the technical problem in the prior art that it is difficult to improve the rendering performance while ensuring the layout quality; it provides an innovative solution for the further development of office automation software. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is a flowchart of the present invention;
[0034] Figure 2 is a flowchart of the feature data acquisition step;
[0035] Figure 3 is a flowchart of the model training step;
[0036] Figure 4 is a schematic diagram of the official document assistant module in Embodiment 4;
[0037] Figure 5 is a schematic diagram of the page settings in Embodiment 4;
[0038] Figure 6 is a schematic diagram of approaching the center of the page in Embodiment 4;
[0039] Figure 7 is a schematic diagram of a single character on a single line in Embodiment 4;
[0040] Figure 8 is a schematic diagram of other settings in Embodiment 4;
[0041] Figure 9 is a schematic diagram of one-key proofreading in Embodiment 4;
[0042] Figure 10 is a schematic diagram of generating a QR code in Embodiment 4. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0044] As Figure 1 shown, it is a flowchart of the present invention, including the following steps:
[0045] S01. Analyze the OpenXML tags of the document, extract the text nodes, table nodes, picture nodes, and style nodes in document.xml, perform feature extraction on the node data, obtain the text block features, table block features, and picture block features, and construct the original feature data;
[0046] S02. Construct a feature processing model, where the feature processing model includes a table branch neural network, a text branch neural network, a picture branch neural network, a feature fusion layer, and a scoring output layer;
[0047] S03. Input the original feature data into the feature processing model, and perform multi-modal feature fusion on the table feature vector, the text feature vector, and the picture feature vector through the feature fusion layer to generate a document feature vector;
[0048] S04. Construct a contribution degree evaluation model, where the contribution degree evaluation model analyzes the contribution degree values of each layout element in the document to the document layout score, and the layout elements include table layout elements, text layout elements, and picture layout elements;
[0049] S05. Construct a loss function based on the similarity between the document layout score and the contribution degree value, and use the loss function to train the contribution degree evaluation model until the convergence condition is met;
[0050] S06. Use the trained contribution degree evaluation model to analyze the contribution degree values of each layout element in the target document, where the contribution degree values include the element's own contribution degree value and the influence degree value between elements;
[0051] S07. Determine the layout priority level of each layout element according to the element's own contribution degree value, and calculate the layout interval value between layout elements according to the influence degree value between elements, where the layout interval value is positively correlated with the influence degree value between elements;
[0052] S08. Sort the layout elements based on the layout priority level to generate the page layout attributes of the w:sectPr node and the style layout attributes of the w:style node;
[0053] S09. Convert the layout interval value into the spacing layout attributes of the w:spacing node to control the spacing between layout elements;
[0054] S10. Construct an OpenXML tag tree of the document based on the page layout attributes, the style layout attributes, and the spacing layout attributes to generate a document layout scheme;
[0055] S11. Serialize the OpenXML tag tree into a document.xml file to obtain a typeset result document;
[0056] S12. Perform dimensionality reduction mapping on the document feature vector, map the document feature vector to a three-dimensional display space, and generate a three-dimensional layout mapping model;
[0057] S13. Perform partition rendering calculations on the typesetting result document through a rendering optimization equation set, divide the typesetting result document into multiple rendering calculation blocks, and the rendering optimization equation set includes a block size equation, a hardware allocation equation, a calculation overhead equation, and an execution order equation;
[0058] S14. Determine a graphics card rendering block, a processor rendering block, and a memory storage block according to the calculation result of the hardware allocation equation;
[0059] S15. Determine the processing order of the graphics card rendering block, the processor rendering block, and the memory storage block according to the calculation result of the execution order equation, and perform rendering calculations using a multi-threaded parallel mechanism.
[0060] The following describes the specific implementation manners of the above steps in detail:
[0061] As Figure 2 shown, it is a flowchart of the feature data acquisition step, including the following steps:
[0062] T01. Collect a manually typeset document as basic data, and the basic data includes the text layout element, the table layout element, and the picture layout element;
[0063] T02. Extract the OpenXML tags of the basic data and record the scoring data of professional typesetting personnel on the document typesetting quality;
[0064] T03. Divide the basic data into a standard sample data set and an optimization sample data set according to the scoring data;
[0065] T04. For the standard sample data set, extract the position data, spacing data, and style data of the layout elements;
[0066] T05. Construct the original feature data, including the document feature vector, the layout element feature data, and the document typesetting score;
[0067] As Figure 3 shown, it is a flowchart of the model training step, including the following steps:
[0068] M01. Train the feature processing model using the original feature data so that it accurately outputs the document typesetting score;
[0069] M02. Based on the calculation result of the feature processing model, construct the training data of the contribution degree evaluation model;
[0070] M03. Train the contribution degree evaluation model using the original feature data and establish the contribution degree calculation rule of the layout element;
[0071] M04. Evaluate the performance metrics of the feature processing model and the contribution degree evaluation model through a cross-validation method, and adjust the neural network parameters until a stable state is reached;
[0072] M05. Save the training parameter data of the feature processing model and the contribution degree evaluation model.
[0073] The specific implementation of step S01 is as follows: First, parse the OpenXML tags of the document, and extract the text nodes, table nodes, picture nodes, and style nodes in document.xml. For these node data, perform feature extraction to obtain text block features, table block features, and picture block features, and finally construct the original feature data. The main purpose of this step is to extract the basic feature data that needs to be analyzed from the document structure, laying a foundation for subsequent feature processing and layout optimization.
[0074] Next is step S02, and the specific implementation is as follows: Construct a feature processing model, which includes a table branch neural network, a text branch neural network, a picture branch neural network, a feature fusion layer, and a scoring output layer. The table branch neural network is responsible for analyzing the table structure features, the text branch neural network is responsible for analyzing the text content and format features, and the picture branch neural network is responsible for analyzing the picture visual features. The feature fusion layer fuses these three types of features to generate an overall feature vector of the document. The scoring output layer calculates the layout score of the document based on this feature vector, reflecting the layout quality level of the document. The purpose of this step is to construct a model that can comprehensively analyze the features of various layout elements in the document and give an overall layout quality assessment.
[0075] The specific implementation of step S03 is as follows: Input the original feature data constructed in step S01 into the feature processing model in step S02. Through the feature fusion layer, perform multi-modal feature fusion on the table feature vector, text feature vector, and picture feature vector, and finally generate an overall feature vector of the document. The purpose of this step is to comprehensively consider the feature information of different types of layout elements, obtain an overall feature representation of the document, and provide a basis for subsequent layout optimization.
[0076] The specific implementation of step S04 is as follows: Construct a contribution degree evaluation model. The purpose of this model is to analyze the contribution degree of each layout element in the document to the layout score of the document. Here, the layout elements include table layout elements, text layout elements, and picture layout elements.
[0077] The specific implementation of step S05 is as follows: Based on the similarity between the document layout score and the contribution degrees of each layout element, a loss function is constructed, and this loss function is used to train the contribution degree evaluation model in step S04 until the convergence condition is met. The purpose of this step is to enable the contribution degree evaluation model to accurately predict the contribution degree values of each layout element through optimized training, providing a basis for subsequent layout optimization.
[0078] The specific implementation of step S06 is as follows: Using the trained contribution degree evaluation model, analyze the contribution degree values of each layout element in the target document. Here, the contribution degree values include the contribution degree of the element itself and the influence degree between elements. The purpose of this step is to obtain the importance degree of each layout element in document layout, providing a basis for subsequent determination of layout priorities.
[0079] The specific implementation of step S07 is as follows: According to the contribution degree value of the element itself, determine the layout priority levels of each layout element. At the same time, according to the influence degree value between elements, calculate the layout interval value between layout elements. The layout interval value is positively correlated with the influence degree between elements. The purpose of this step is to determine the positions and spacings of each layout element in the document, preparing for subsequent generation of OpenXML tags.
[0080] The specific implementation of step S08 is as follows: Based on the layout priority levels determined in step S07, sort the layout elements to generate the page layout attributes of the w:sectPr node and the style layout attributes of the w:style node. The purpose of this step is to determine the page layout and style layout according to the element importance, laying a foundation for the construction of the final OpenXML tag tree.
[0081] The specific implementation of step S09 is as follows: Convert the layout interval value calculated in step S07 into the spacing layout attributes of the w:spacing node, which are used to control the spacing between layout elements. The purpose of this step is to set appropriate spacings according to the influence relationship between elements, further optimizing the document layout.
[0082] The specific implementation of step S10 is as follows: Based on the page layout attributes, style layout attributes, and spacing layout attributes generated in the previous steps, construct the OpenXML tag tree of the document to generate the final document layout scheme. The purpose of this step is to integrate the layout information obtained in the previous steps into the OpenXML tags to form a complete document layout.
[0083] The specific implementation of step S11 is as follows: Serialize the constructed OpenXML tag tree into a document.xml file to obtain the final typeset result document. The purpose of this step is to output the typeset document, preparing for subsequent rendering optimization.
[0084] The specific implementation of step S12 is as follows: perform dimensionality reduction mapping on the document feature vector, map it to a three-dimensional display space, and generate a three-dimensional layout mapping model. The purpose of this step is to compress the high-dimensional features into a three-dimensional space to visually display the layout information of the document and provide a visual reference for users.
[0085] The specific implementation of step S13 is as follows: perform partition rendering calculation on the typeset result document through a set of rendering optimization equations. It includes:
[0086] Block size equation:
[0087]
[0088] In the formula, S b is the pixel size of the rendering block, R d is the display device resolution, M g is the video memory capacity, M s is the system memory capacity, N b is the number of blocks, η 1 , η 2 , η 3 is the resource utilization coefficient. This equation is based on the bucket principle, considering three limiting factors: display resolution, video memory capacity, and system memory, and takes the minimum value as the block size.
[0089] Hardware allocation equation:
[0090]
[0091] In the formula, P g , P p , P m are the task ratios of the graphics card, processor, and memory respectively, G is the graphics card performance index, P is the number of processor cores, M is the memory capacity, C is the document complexity, k 1 , k 2 , k 3 is the hardware weight coefficient, and λ is the complexity impact factor. This set of equations uses an exponential saturation function to describe the relationship between the task allocation ratio and the document complexity, considering the relative weights of the hardware performance indicators.
[0092] Computing overhead equation:
[0093]
[0094] In the formula, T is the estimated value of the rendering calculation duration, N t is the number of characters, N p is the table complexity, N i is the number of pictures, R i is the average resolution of the pictures, ω1 , ω 2 , ω 3 where ω is the time weight coefficient and μ is the basic calculation cost. This equation takes into account the logarithmic complexity of text layout, the quadratic complexity of table layout, and the linear complexity of image rendering.
[0095] Execution order equation:
[0096]
[0097] In the formula, O i is the execution priority of block i, T i is the calculation cost of the block, P i is the hardware allocation ratio, Dep(i) is the set of dependent blocks of block i, and D ij is the dependency value. This equation determines the execution order based on the principle of critical path analysis, taking into account three factors: calculation cost, hardware allocation, and dependency relationship.
[0098] The purpose of this step is to reasonably divide the rendering tasks according to the hardware resources and document complexity, optimize the execution order of the rendering calculations, and improve the rendering efficiency.
[0099] The specific implementation of step S14 is as follows: Determine the graphics card rendering block, processor rendering block, and memory storage block according to the calculation results of the hardware allocation equation in step S13. The purpose of this step is to reasonably allocate the rendering tasks to different hardware devices and make full use of the system resources.
[0100] The specific implementation of step S15 is as follows: Determine the processing order of each rendering block according to the calculation results of the execution order equation in step S13, and use the multi-thread parallel mechanism for rendering calculations. The purpose of this step is to optimize the execution order of the rendering calculations and improve the rendering efficiency.
[0101] Calculation model for document content complexity:
[0102] C = α 1 N t + α 2 N p + α 3 N i + β;
[0103] In the formula, C is the document content complexity value, N t is the number of text layout elements, N p is the number of table layout elements, N i is the number of image layout elements, α 1 , α 2 , α 3is the weight coefficient, and β is the basic complexity constant. This model considers the contribution weights of different types of layout elements to the document complexity and calculates the overall complexity using a linear weighting method. The parameter α 1 , α 2 , α 3 ranges from 0 to 1, reflecting the processing difficulty of different types of layout elements. The parameter β ranges from 1 to 10 and is used to adjust the overall complexity level.
[0104] Computation model for rendering data dependencies:
[0105] D ij = γ 1 S ij + γ 2 L ij + δ;
[0106] In the formula, D ij is the dependency relationship value between layout elements i and j, S ij is the spatial position correlation degree, L ij is the logical order correlation degree, γ 1 , γ 2 is the weight coefficient, and δ is the basic dependency constant.
[0107] The spatial position correlation degree is calculated as follows:
[0108]
[0109] In the formula, d ij is the Euclidean distance between elements i and j, and σ is the distance influence factor.
[0110] The logical order correlation degree is calculated as follows:
[0111]
[0112] In the formula, p i , p j are the serial numbers of elements i and j in the document respectively. This model comprehensively considers the dependency relationships in two dimensions of spatial distance and logical order, uses a Gaussian kernel function to describe the spatial correlation, and uses a reciprocal function to describe the logical correlation. The parameter γ 1 , γ 2 ranges from 0 to 1, the parameter σ ranges from 10 to 100 pixels, and the parameter δ ranges from 0.1 to 0.5.
[0113] A specific Embodiment 1 of the present invention is provided below. The specific implementation methods of each step in this Embodiment 1 are described in detail as follows: In this embodiment, Figure 1 is the flowchart of this embodiment, where:
[0114] The specific implementation of step S01 is to first parse the OpenXML tags of the document, extract the text nodes, table nodes, picture nodes, and style nodes in document.xml. Feature extraction is performed on these node data to obtain text block features, table block features, and picture block features, and the original feature data is constructed. The main purpose of this step is to extract the basic feature data that needs to be analyzed from the document structure.
[0115] The specific implementation of step S02 is to construct a feature processing model, including a table branch neural network, a text branch neural network, a picture branch neural network, a feature fusion layer, and a score output layer. The purpose of this model is to perform feature extraction and fusion on different types of layout elements through a multi-branch neural network, and finally output the layout score of the document.
[0116] The specific implementation of step S03 is to input the original feature data into the feature processing model. Through the feature fusion layer, multi-modal feature fusion is performed on the table feature vector, text feature vector, and picture feature vector to generate the overall feature vector of the document. The purpose of this step is to comprehensively consider the feature information of different layout elements to obtain the overall representation of the document.
[0117] The specific implementation of step S04 is to construct a contribution degree evaluation model. The purpose of this model is to analyze the contribution degree of each layout element in the document to the document layout score. The layout elements here include table layout elements, text layout elements, and picture layout elements.
[0118] The specific implementation of step S05 is to construct a loss function based on the similarity between the document layout score and the contribution degree of each layout element, and use this loss function to train the contribution degree evaluation model until the convergence condition is met. The purpose of this step is to make the contribution degree evaluation model accurately predict the contribution degree values of each layout element through optimized training.
[0119] The specific implementation of step S06 is to use the trained contribution degree evaluation model to analyze the contribution degree values of each layout element in the target document. The contribution degree values here include the contribution degree of the element itself and the influence degree between elements. The purpose of this step is to obtain the importance degree of each layout element in the document layout.
[0120] The specific implementation of step S07 is to determine the layout priority level of each layout element according to the contribution degree value of the element itself, and calculate the layout interval value between the layout elements according to the influence degree value between the elements. The layout interval value is positively correlated with the influence degree between the elements. The purpose of this step is to determine the position and spacing of each layout element in the document.
[0121] The specific implementation of step S08 is to sort the layout elements based on the layout priority level, generating the page layout attributes of the w:sectPr node and the style layout attributes of the w:style node. The purpose of this step is to determine the page layout and style layout according to the element importance.
[0122] The specific implementation of step S09 is to convert the calculated layout interval value into the spacing layout attributes of the w:spacing node, which is used to control the spacing between layout elements. The purpose of this step is to set appropriate spacing according to the influence relationship between elements.
[0123] The specific implementation of step S10 is to construct the OpenXML markup tree of the document based on the page layout attributes, style layout attributes, and spacing layout attributes generated in the previous steps, and generate the final document layout scheme. The purpose of this step is to integrate the layout information obtained in the previous steps into the OpenXML markup to form a complete document layout.
[0124] The specific implementation of step S11 is to serialize the constructed OpenXML markup tree into a document.xml file to obtain the final typeset result document. The purpose of this step is to output the typeset document.
[0125] The specific implementation of step S12 is to perform dimensionality reduction mapping on the document feature vector, map it to a three-dimensional display space, and generate a three-dimensional layout mapping model. The purpose of this step is to compress the high-dimensional features into a three-dimensional space for intuitive display of the document layout information.
[0126] The specific implementation of step S13 is to perform partition rendering calculation on the typeset result document through a set of rendering optimization equations. Specifically, it includes the block size equation, hardware allocation equation, calculation overhead equation, and execution order equation. The purpose of this step is to reasonably divide the rendering tasks according to the hardware resources and document complexity to improve the rendering efficiency.
[0127] The specific implementation of step S14 is to determine the graphics card rendering block, processor rendering block, and memory storage block according to the calculation result of the hardware allocation equation. The purpose of this step is to reasonably allocate the rendering tasks to different hardware devices.
[0128] The specific implementation of step S15 is to determine the processing order of each rendering block according to the calculation result of the execution order equation, and use a multi-threaded parallel mechanism for rendering calculation. The purpose of this step is to optimize the execution order of the rendering calculation and make full use of the hardware resources to improve the rendering efficiency.
[0129] The calculation of the numerical value of the document content complexity adopts a linear weighted model, considering the quantities of three types of layout elements: text layout elements, table layout elements, and image layout elements, as well as their contribution weights to complexity. The parameter α 1 、α 2 and α 3 range from 0 to 1, reflecting the processing difficulty of different types of layout elements. The basic complexity constant β ranges from 1 to 10 and is used to adjust the overall complexity level.
[0130] The calculation of the numerical value of the rendering data dependency adopts a model that combines spatial position and logical order. The spatial position correlation is characterized by a Gaussian kernel function, and the value range of the distance influence factor σ is from 10 to 100 pixels. The logical order correlation is characterized by a reciprocal function, and the weight coefficients γ 1 and γ 2 range from 0 to 1. The basic dependency constant δ ranges from 0.1 to 0.5.
[0131] The block size equation is based on the limiting factors of system resources, including display resolution, video memory capacity, and system memory capacity. It adopts the bucket principle and takes the minimum value of the three factors as the size of the rendering block. The resource utilization coefficients η 1 、η 2 and η 3 range from 0.6 to 0.9.
[0132] The hardware allocation equation set adopts an exponential saturation function to describe the relationship between the task allocation ratio and the document complexity. The hardware weight coefficients k 1 、k 2 and k 3 are determined through experimental optimization, and the complexity influence factor λ ranges from 0.001 to 0.01.
[0133] The calculation overhead equation considers the logarithmic complexity of text typesetting, the quadratic complexity of table layout, and the linear complexity of image rendering. The time weight coefficients ω 1 、ω 2 and ω 3 are determined through regression analysis, and the basic calculation overhead μ ranges from 0.1 to 1 second.
[0134] The execution order equation is based on the principle of critical path analysis and determines the execution order of the rendering block by comprehensively considering three factors: calculation overhead, hardware allocation, and dependency.
[0135] 1. Document content complexity calculation model:
[0136] The calculation of the numerical value of the document content complexity is specifically expressed as follows:
[0137] C = α1 N t +α 2 N p +α 3 N i +β;
[0138] In the formula, C is the numerical value of the document content complexity; N t is the number of text layout elements; N p is the number of table layout elements; N i is the number of picture layout elements; α 1 , α 2 , α 3 , α are weight coefficients; β is the basic complexity constant.
[0139] Among them, the method for obtaining parameters is:
[0140] N t , N p , N i is directly obtained by parsing OpenXML tags;
[0141] α 1 , α 2 , α 3 is determined by regression analysis, and the range is 0 to 1;
[0142] The value range of β is 1 to 10.
[0143] This model considers the contribution weights of different types of layout elements to the document complexity, calculates the overall complexity using the linear weighting method, and the weight coefficients reflect the processing difficulty of various elements.
[0144] 2. Rendering data dependency calculation model:
[0145] The calculation of the rendering data dependency value is specifically expressed as follows:
[0146] D ij = γ 1 S ij + γ 2 L ij + δ;
[0147] In the formula, D ij is the dependency value between layout elements i and j; S ij is the spatial position correlation degree; L ij is the logical order correlation degree; γ 1 , γ 2 are weight coefficients; δ is the basic dependency constant.
[0148] The calculation of the spatial position correlation degree is as follows:
[0149]
[0150] In the formula, d ij is the Euclidean distance between elements i and j; σ is the distance influence factor.
[0151] The calculation of the logical order correlation degree is as follows:
[0152]
[0153] In the formula, p i , p j are the serial numbers of elements i and j in the document respectively.
[0154] Among them, the method for obtaining parameters is:
[0155] d ij is calculated through the element coordinates;
[0156] p i , p j is obtained through OpenXML tag parsing;
[0157] γ 1 , γ 2 is determined through experiments, and the range is 0 to 1;
[0158] The value range of σ is 10 to 100 pixels;
[0159] The value range of δ is 0.1 to 0.5.
[0160] This model comprehensively considers the dependency relationships in two dimensions of spatial distance and logical order, uses a Gaussian kernel function to describe spatial correlation, and uses a reciprocal function to describe logical correlation.
[0161] 3. Block size equation:
[0162] The specific expression of the block size equation is as follows:
[0163]
[0164] In the formula, S b is the pixel size of the rendering block; R d is the display device resolution; M g is the video memory capacity; M s is the system memory capacity; N b is the number of blocks; η 1 , η 2 , η 3 is the resource utilization rate coefficient.
[0165] Among them, the method for obtaining parameters is:
[0166] R d , Mg , M s Obtained through the system API;
[0167] N b Determined by a heuristic algorithm;
[0168] η 1 , η 2 , η 3 The value range is 0.6 to 0.9.
[0169] This equation is based on the bucket principle, considering three limiting factors: display resolution, video memory capacity, and system memory, and taking the minimum value as the block size.
[0170] 4. Hardware allocation equation:
[0171] The specific expression of the hardware allocation equation is as follows:
[0172]
[0173] In the formula, P g , P p , P m are the task ratios of the graphics card, processor, and memory respectively; G is the graphics card performance index; P is the number of processor cores; M is the memory capacity; C is the document complexity; k 1 , k 2 , k 3 is the hardware weight coefficient; λ is the complexity impact factor.
[0174] Among them, the parameter acquisition method is:
[0175] G, P, M are obtained through the system API;
[0176] C is calculated through the aforementioned complexity model;
[0177] k 1 , k 2 , k 3 Determined by experimental optimization;
[0178] The value range of λ is 0.001 to 0.01.
[0179] This system of equations uses an exponential saturation function to describe the relationship between the task allocation ratio and the document complexity, considering the relative weights of the hardware performance indicators.
[0180] 5. Computational overhead equation:
[0181] The specific expression of the computational overhead equation is as follows:
[0182]
[0183] Where, T is the estimated rendering calculation duration; N t is the number of characters; N p is the table complexity; N i is the number of images; R i is the average resolution of the images; ω 1 , ω 2 , ω 3 is the time weight coefficient; μ is the basic calculation overhead.
[0184] Among them, the parameter acquisition method is:
[0185] N t , N p , N i , R i is obtained by parsing OpenXML tags;
[0186] ω 1 , ω 2 , ω 3 is determined by regression analysis;
[0187] The value range of μ is 0.1 to 1 second.
[0188] This equation takes into account the logarithmic complexity of text layout, the quadratic complexity of table layout, and the linear complexity of image rendering.
[0189] 6. Execution order equation:
[0190] The specific expression of the execution order equation is as follows:
[0191]
[0192] Where, O i is the execution priority of block i; T i is the calculation overhead of the block; P i is the hardware allocation ratio; Dep(i) is the set of dependent blocks of block i; D ij is the dependency value.
[0193] Among them, the parameter acquisition method is:
[0194] T i is obtained through the calculation overhead equation;
[0195] P i is obtained through the hardware allocation equation;
[0196] D ij is calculated through the dependency relationship model.
[0197] This equation determines the execution order based on the critical path analysis principle, taking into account three factors: calculation overhead, hardware allocation, and dependency relationship.
[0198] To better understand and implement the present invention, an embodiment 2 of a specific application scenario of the present invention is provided below: The office automation department of an enterprise plans to develop an intelligent document layout software to improve the work efficiency of employees. This software needs to be able to automatically analyze the document structure and content features and generate an optimized layout plan. For this purpose, they decided to adopt the intelligent layout method based on a multi-branch neural network proposed by the present invention. The following is the implementation process of this method in a specific application scenario.
[0199] One day, Xiao Li from the finance department of a company was writing a monthly financial report. This report contains a large amount of text description, complex financial data tables, and related chart analyses. Xiao Li hopes to quickly generate a neatly typeset and well-structured document for the superior leaders to read and review. So, he used the above intelligent layout software.
[0200] First, the software automatically parsed the OpenXML format document written by Xiao Li and extracted the text content, table data, and picture elements contained therein. For different types of layout elements, the software respectively constructed a table branch neural network, a text branch neural network, and a picture branch neural network to extract their unique visual, structural, and semantic features. Taking the table branch neural network as an example, it analyzed the row data attributes of the w:tr node and the cell data attributes of the w:tc node in the document, and extracted the horizontal features of the table; at the same time, it also analyzed the column data attributes of the w:gridCol node and the merged data attributes of the w:vMerge node, and extracted the vertical features of the table. These feature data were input into the table feature fusion layer to generate the final table feature vector. The working principles of the text branch neural network and the picture branch neural network are similar, both extracting their unique features for different layout elements.
[0201] Next, the software integrated these heterogeneous features through the feature fusion layer to generate an overall feature vector of the document. This feature vector reflects the mutual relationship and overall layout characteristics of various layout elements in the document.
[0202] With the overall feature representation of the document, the software then constructed a contribution degree evaluation model to analyze the specific contributions of each layout element to the document layout quality. Taking the monthly financial report as an example, this model first calculated the relative weights of the text content, table data, and chart analysis in the entire document, and obtained the results shown in Table 1 below:
[0203] Table 1 Relative Contribution Degree Table
[0204] Layout element Relative contribution degree Text content 40% Table data 45% Chart analysis 15%
[0205] It can be seen that the tabular data occupies a relatively large proportion in this financial report and has the highest contribution. Next, the model also calculated the mutual influence between these layout elements. For example, a key table may affect the layout of the surrounding text paragraphs, and the position of a certain chart may affect the arrangement of adjacent tables. After comprehensively considering these factors, the contribution evaluation model finally quantified the contribution values of each layout element.
[0206] With these contribution data, the software then determined the typesetting priorities of these layout elements in the document according to their importance. For this case of the financial report, since the tabular data has the highest contribution, they are arranged at the front of the document. The text content and chart analysis are respectively arranged at the back of the tables and the end of the document according to their relative importance. At the same time, the software also calculated the appropriate spacing that they should maintain according to the mutual influence between the layout elements. Specifically as shown in Table 2 below:
[0207] Table 2 Layout Spacing Table
[0208] Layout element Typesetting priority Layout spacing Table data 1 - Text content 2 2 paragraphs Chart analysis 3 1 paragraph
[0209] So far, the software has obtained a preliminary document layout plan. However, for a complex financial report, if the traditional sequential rendering method is adopted, it may lead to a long response time and affect the user experience. Therefore, the software also carried out further rendering optimization according to the specific situation of this report.
[0210] First of all, the software calculated the content complexity of this financial report: C = 92;
[0211] Among them, N t = 200 is the number of text layout elements, N p = 50 is the number of table layout elements, N i = 10 is the number of picture layout elements, α 1 = 0.4, α 2 = 0.45, α 3 = 0.15, β = 2. It can be seen that the content of this financial report is relatively complex.
[0212] Next, the software calculated the rendering data dependency relationship between each layout element according to the specific layout characteristics of the report:
[0213] D ij = 0.6×S ij + 0.4×L ij + 0.2;
[0214] Among them, represents the spatial position correlation degree, L ij = 1 / (|p i - pj |+1) represents the logical order relevance. Through this model, the software quantifies the degree of mutual influence between different layout elements.
[0215] With the content complexity and rendering data dependencies, the software then calculated the optimal allocation plan for the rendering tasks for this financial report: P g = 0.75; P p = 0.15; P m = 0.10;
[0216] Among them, G = 6400 is the graphics card performance index, P = 8 is the number of processor cores, and M = 32GB is the memory capacity. It can be seen that the rendering tasks of this report are mainly borne by the graphics card, followed by the processor and memory.
[0217] With the task allocation ratio, the software further optimized the execution order of the rendering calculations:
[0218]
[0219] Among them, T i is the rendering calculation overhead of block i, P i is the proportion of hardware resources allocated to this block, and Dep(i) is the set of dependent blocks of block i. Through this equation, the software determined the execution priorities of each rendering block to maximize the utilization of hardware resources and improve the overall rendering efficiency.
[0220] Finally, based on the above optimization measures, the software generated the typesetting plan for the financial report. The whole process is as follows:
[0221] 1. Parse the OpenXML tags, extract the text content, table data, and chart elements, and respectively construct the corresponding neural network models to extract their visual, structural, and semantic features.
[0222] 2. Through the feature fusion layer, integrate these heterogeneous features to generate the overall feature vector of the document.
[0223] 3. Use the contribution degree evaluation model to analyze the contribution degree of each layout element to the overall quality of the report, and determine their typesetting priorities and layout spacing in the document.
[0224] 4. According to the content complexity of the report and the dependencies between layout elements, adopt optimization strategies such as rendering task partitioning and parallel computing to improve the rendering efficiency of the report.
[0225] The finally generated typesetting plan for the financial report is shown in Table 3 below:
[0226] Table 3 Typesetting Plan for Financial Report
[0227] Layout element Typesetting priority Layout spacing Rendering task allocation Execution order Table data 1 - Graphics card 75% 1 Text content 2 2 paragraphs Processor 15% 2 Chart analysis 3 1 paragraph Memory 10% 3
[0228] It can be seen that this intelligent layout method not only reasonably arranges the layout of the report, but also fully considers the optimization of rendering performance, significantly improving the overall processing efficiency.
[0229] Compared with the traditional document layout method, the main improvements of the solution of the present invention are reflected in the following aspects:
[0230] Table 4 Document Layout Method Table
[0231] Comparison item Traditional method The solution of the present invention Feature extraction Only for a single layout element Integrate multi-modal features to comprehensively reflect the characteristics of the document Contribution degree evaluation Based on subjective experience rules Use a data-driven evaluation model for objective quantification Layout optimization Empirical adjustment with unstable effects Generate a reasonable layout plan based on contribution degree analysis Rendering efficiency Sequential processing with long response time Task partitioning and parallelism to greatly improve efficiency
[0232] The following provides Example 3 of a specific application scenario of the present invention: This example selects the actual application scenario of the document service department of a large enterprise for detailed description. This department needs to process about 5,000 office documents with complex layouts every month, involving various types such as financial statements, technical documents, and meeting records, which is highly representative. During the implementation process, a typical quarterly financial report was selected as an example, and a complete document intelligent layout process was realized by using a well-configured hardware environment and a professional software platform. The hardware environment uses an Intel i9-12900K processor (16 cores and 32 threads), an NVIDIA RTX 4080 graphics card (16GB video memory), 64GB DDR5 memory, 2TB NVMe SSD storage, and a 4K resolution (3840×2160 pixels) monitor; the software environment includes the Windows11 Pro operating system, the.NET 6.0 development framework, the PyTorch 2.1.0 deep learning framework, and the OpenXML SDK 2.20.0 version.
[0233] In the document data collection and preprocessing stage, a comprehensive analysis and feature extraction were carried out on the financial report for the fourth quarter of 2023. The specific content composition of this report is shown in Table 5 below:
[0234] Table 5 Specific Content Composition Table
[0235] Content type Quantity statistics Detailed description Text paragraph 156 paragraphs Including the main text, title, notes, etc Table 12 Including financial data tables, comparative analysis tables, etc Picture 8 Including statistical charts such as trend charts and pie charts
[0236] Through the parsing of the OpenXML SDK, rich original feature data was obtained: in terms of text, the average paragraph length is 320 characters, including 4 levels of heading hierarchies, using 3 font types, and the paragraph spacing ranges from 1.0 to 1.5 times the line spacing; in terms of tables, the maximum number of rows reaches 28 rows, the maximum number of columns is 12 columns, including 46 cell merges, and 3 table border styles are used; in terms of pictures, the average resolution is 1024×768 pixels, and both PNG and JPEG formats are used at the same time, adopting two embedding methods of floating and embedding, involving 4 wrapping methods.
[0237] During the training process of the feature processing model, a carefully designed multi-branch neural network structure is adopted. The table branch neural network includes a horizontal and vertical table model with three convolutional layers (64 feature maps in each layer), and a table feature fusion layer that outputs a 256-dimensional feature vector; the text branch neural network uses a text extraction layer with 300-dimensional word vectors, a text semantic layer with 512 LSTM units, and a fully connected text structure layer that outputs 128 dimensions; the image branch neural network combines an image extraction layer of the ResNet50 pre-trained model, an image localization layer of spatial pyramid pooling, and a text-image connection layer based on the attention mechanism. The training dataset contains 30,000 labeled enterprise documents, of which 24,000 are used for training and 6,000 are used for validation. 100 rounds of training are carried out with a batch size of 32, a learning rate of 0.001, and the Adam optimizer, and the mean squared error is used as the loss function.
[0238] In terms of complexity calculation, according to the document content complexity calculation model C = α 1 N t + α 2 N p + α 3 N i + β, set the text weight α 1 = 0.4, the table weight α 2 = 0.35, the image weight α 3 = 0.25, the basic complexity constant β = 5, substitute the number of text paragraphs N t = 156, the number of tables N p = 12, the number of images N i = 8, and the calculated document complexity value is 71.9. For the calculation of the rendering dependency, an adjacent table and an image element in the document are selected for analysis, and according to the rendering data dependency calculation model D ij = γ 1 S ij + γ 2 L ij + δ, set the spatial position weight γ 1 = 0.6, the logical order weight γ 2 = 0.4, the basic dependency constant δ = 0.2, the distance influence factor σ = 50, combined with the measured distance of 120 pixels and the document serial numbers of 25 and 26, the calculated final dependency value is 0.534.
[0239] In the block division and hardware allocation link, based on the block size equation Calculations were carried out. Considering a 4K display resolution (3840×2160), a 16GB video memory capacity, a 64GB system memory, 16 block quantities were set, as well as resource utilization rates of 0.8, 0.7, and 0.75. Finally, the size of each block was determined to be 960×540 pixels. In terms of hardware task allocation, by setting hardware weight coefficients (graphics card 1.2, processor 1.0, memory 0.8), a complexity impact factor of 0.005, combining hardware parameters (16384 graphics card computing units, 32 processor threads, 65536MB memory capacity) and a document complexity of 71.9, the calculated graphics card task ratio was 48.2%, the processor task ratio was 32.5%, and the memory task ratio was 19.3%.
[0240] In the rendering calculation and execution optimization stage, through the calculation overhead equation Time weight coefficients (0.001, 0.005, 0.0001) and a basic calculation overhead of 0.5 seconds were set. Combining with an average picture resolution of 1024×768, the total rendering duration was estimated to be 2.86 seconds. Based on the execution order equation, 16 rendering blocks were optimized and sorted. The calculation results of the execution priorities of some blocks are shown in Table 6 below:
[0241] Table 6 Calculation Results Table of Execution Priorities
[0242]
[0243] The method of the present invention has achieved a significant improvement in actual application. The comparison results with traditional technical means are shown in Table 7 below:
[0244] Table 7 Comparison Results Table of Technical Means
[0245]
[0246]
[0247] In terms of traditional technical means, the layout of office documents mainly relies on rule-based template matching methods. By using preset fixed layout rules, a single layout optimization algorithm cannot effectively handle the collaborative layout of multiple types of elements. At the same time, the serial rendering processing method fails to fully utilize the parallel capabilities of modern computing hardware and mainly relies on manual adjustment, resulting in a low degree of automation, which severely restricts the improvement of processing efficiency. In contrast, through the innovative adoption of multi-branch neural networks and intelligent rendering optimization technologies, the present invention has successfully achieved the collaborative layout optimization of multiple types of elements such as text, tables, and pictures, significantly improving the layout quality; by introducing a contribution evaluation model, the layout decision-making has better interpretability; by adopting a multi-thread parallel rendering mechanism, the processing efficiency has been greatly improved; and at the same time, a complete mathematical model system has been established, providing a solid theoretical foundation for intelligent layout.
[0248] This embodiment fully verifies the feasibility and practical effects of the intelligent document arrangement method based on multi-branch neural networks. Through the processing practice of enterprise financial reports, it is proved that this method can effectively improve the efficiency and quality of document arrangement, with significant application value and broad promotion prospects. In the actual application process, this method can flexibly adjust the parameter configuration according to specific needs to adapt to different types of document processing scenarios, fully demonstrating its strong adaptability and practical value.
[0249] The following provides an embodiment 4 of a specific application scenario of the present invention: This embodiment 4 is applied to the embedded official document assistant module of WPS, as Figure 4 shown:
[0250] Function module:
[0251] 1. In [Page Setup], it is divided into default settings and custom settings. The default settings include [Ordinary Official Document] and [Letter Official Document]. When the required official document format is not available in the default settings, you can click the [Customize] button to enter the custom settings and set the official document format according to actual needs, as Figure 5 shown. In the [Customize] module, you can set each element of the page margins, paper, layout, document grid, and column layout.
[0252] 2. In [Header and Footer], after clicking, select the corresponding function in the drop-down list to delete the header and footer content at once, or delete the header and footer content separately.
[0253] 3. Page orientation: [Insert Landscape Page]: Insert a landscape page after the current page.
[0254] [This Section Landscape]: Set the page orientation of this section to landscape.
[0255] [This Section Portrait]: Set the page orientation of this section to portrait.
[0256]
Current content horizontal
[0257] Note: A section is a document structure unit.
[0258] 4. One-click typesetting: Process a document that has not been typeset at all. The system has preset the corresponding document formats in
One-click typesetting
One-click typesetting
[0259] After confirming that the content of
One-click typesetting
OK
[0260] 5. Align with the typesetting area: Typesetting area size: 156mm * 225mm; For the page where the mouse is located or the entire selected document, as Figure 6 shown, click Align with the typesetting area: The typesetting of the text aligns with the top of the typesetting area. To cancel the typesetting, the arrangement of aligning with the typesetting area can be cancelled.
[0261] 6. Single character on a single line:
Single character does not form a line
[0262]
Single line does not form a page
[0263] 7. Name line break: Name line break includes:
Increase character spacing
Decrease character spacing
Auto line break
Clear character spacing
[0264]
Increase character spacing
[0265]
Decrease character spacing
[0266]
Automatic line wrapping
[0267]
Clear character spacing
[0268] 8. Other settings:
Font preset
Name preset
[0269] 9. One-click proofreading | Clear proofreading: As Figure 9 shown, click the
One-click proofreading
Clear proofreading
[0270] 10. Thesaurus management: Click
Thesaurus management
[0271] 11. Main title | Subtitle | Main text | First-level heading | Second-level heading | Third-level heading | Fourth-level heading: Main title, subtitle, main text, first-level heading, second-level heading, third-level heading, fourth-level heading, document serial number / classification / urgency level, heading serial number: Select the heading type according to the content requirements;
[0272] 12. Left alignment / Center alignment / Right alignment / Justify / Distribute / First line indent.
[0273] 13. Increase line spacing | Decrease line spacing: Generally used to fine-tune the line spacing of paragraphs in special cases. For example, when the last line of text on the first page is too far from the page margin, or when the end of the text cannot accommodate the signature and imprint, you can fine-tune the line spacing in front.
[0274] 14. Paragraph spacing:
Paragraph spacing
[0275] 15. Paragraph tools: Paragraph tools include:
Cancel space before and after paragraphs
Delete blank paragraphs
Delete manual line breaks
Convert manual line breaks to carriage returns
Merge paragraphs
Delete leading and trailing blanks in paragraphs
Delete full-text page breaks
Delete full-text section breaks
Show / hide paragraph marks
[0276] 16. Paragraph Indentation:
Paragraph Indentation
First line indented by two characters
Left indented by one character
Left indented by two characters
Right indented by one character
Right indented by two characters
Indented by one character on both left and right
Remove all indents
[0277] 17. Character Conversion
[0278] (1)
Convert Auto-numbering to Text
[0279] (2)
Convert the Punctuation Mark after the Serial Number to a Period | Convert the Period after the Serial Number to a Punctuation Mark
[0280] (3)
Convert the Full-width Period after the Serial Number to a Half-width Period | Convert the Half-width Period after the Serial Number to a Full-width Period
[0281] (4) [] to 〔〕: Convert the square brackets with years to hexagonal brackets. For example, convert
[2022] to 〔2022〕 and [2] to 〔2〕.
[0282] (5) 【】 to 〔〕: Convert the black brackets with years to hexagonal brackets. For example, convert
[2022] to 〔2022〕 and 【2】 to 〔2〕.
[0283] (6) ﹝﹞ to 〔〕: Convert the non-standard hexagonal brackets “﹝﹞” typed by some input methods to the standard hexagonal brackets “〔〕”.
[0284] (7) <> to 〈〉: Convert the angled brackets <> to single book titles 〈〉.
[0285] (8) 〈〉 to <> : Convert the single book titles 〈〉 to angled brackets <>.
[0286] (9) [Convert short horizontal lines to long horizontal lines | Convert long horizontal lines to short horizontal lines]: Convert the long and short horizontal lines in the selected range (when there is selected content) or the whole document (when there is no selected content), such as converting "-" to "—" and "—" to "-".
[0287] (10) Convert <> to 《》: Convert single angle brackets <> to double angle brackets.
[0288] (11) [Convert long horizontal lines, short horizontal lines, and wavy lines between numbers]: Convert the long horizontal lines, short horizontal lines, and wavy lines between numbers in the selected range (when there is selected content) or the whole document (when there is no selected content), such as converting "1-5" to "1—5" or "1~5".
[0289] (12) [Convert the full stop before a number to a decimal point]: Convert the full stop before a number in the selected range (when there is selected content) or the whole document (when there is no selected content) to a decimal point, such as converting "12。3" to "12.3".
[0290] (13) [Convert common punctuation marks to Chinese punctuation marks]: Convert the common punctuation marks in the selected range (when there is selected content) or the whole document (when there is no selected content) to Chinese punctuation marks, such as converting "." to "。", except for the decimal point after a number.
[0291] (14) [Convert English punctuation marks to Chinese punctuation marks]: Convert the English punctuation marks in the selected range (when there is selected content) or the whole document (when there is no selected content) to Chinese punctuation marks, such as converting ";" to ";", "." to "。", except for the decimal point after a number.
[0292] (15) [Convert Chinese punctuation marks to English punctuation marks]: Convert the Chinese punctuation marks in the selected range (when there is selected content) or the whole document (when there is no selected content) to English punctuation marks, such as converting "。" to ".".
[0293] (16) [Convert full-width numbers and letters to half-width | Convert half-width numbers and letters to full-width]: Convert the numbers and letters in the selected range (when there is selected content) or the whole document (when there is no selected content) between full-width and half-width. For example, convert "2" to "2" and "A" to "A".
[0294] (17) [Delete the punctuation mark between book titles or quotation marks]: Delete the punctuation mark between two book titles or two quotation marks. For example, change 《XX》、《XX》 to 《XX》《XX》, and "XX"、"XX" to "XX" "XX".
[0295] (18) [Add a punctuation mark between book titles or quotation marks]: Add a punctuation mark between book titles or quotation marks. Add a punctuation mark between two book titles or two quotation marks. For example, change 《XX》《XX》 to 《XX》、《XX》, and "XX" "XX" to "XX"、"XX".
[0296] (19) [Add Thousand Separator to Numbers]: Add a thousand separator to the numbers in the selected range (when there is selected content) or the entire document (when there is no selected content). For example, 1212345.123 becomes 1,212,345.123.
[0297] (20) [Remove Thousand Separator from Numbers]: Remove the thousand separator from the numbers in the selected range (when there is selected content) or the entire document (when there is no selected content). For example, 1,212,345.123 becomes 1212345.123.
[0298] (21) [Convert Lowercase Amount to Uppercase]: Replace the selected lowercase amount number with the uppercase amount. For example: When the mouse selects "12.5" and clicks the conversion button, "12.5" will be replaced with "Twelve Yuan and Five Jiao Only".
[0299] (22) [Delete Hyperlinks]: Delete all hyperlinks in the document.
[0300] 18. Chart Tools: Chart tools include: [Delete All Images in the Document], [Export All Images in the Document], [Image Formatting], [Convert Table to Text], [Delete All Tables in the Document], [Insert Watermark], [Delete All Watermarks in the Document]. When processing official document content, corresponding settings can be made according to actual needs.
[0301] 19. Attachment | Note: Insert information about [Single Attachment Description], [Multiple Attachment Descriptions], [Meeting Participants], [Notes], [Attachment Header].
[0302] 20. Signature | Date: [Signature | Date] includes: [Short Signature], [Long Signature], [Single Organization with Official Seal], [Multiple Organizations with Official Seal], [Personal Seal], [Date].
[0303] 21. Letter Page: When you click [Letter Page], the page number is not displayed on the first page of the document, and the page numbering starts from the second page.
[0304] 22. Separator Lines: Separator lines include: [Ordinary Separator Line], [Five-Pointed Star Separator Line], [Upper Letter Line], [Lower Letter Line]. The separator line of the footer is a graphic line and is in the embedded type. Its position will change with the change of the number of lines of the main text. It is inserted at the cursor position by default and needs to be manually adjusted to the bottom of the page. (Note: It is recommended to use the table form to ensure that the footer is at the bottom of the page and does not move with the main text).
[0305] 23. Footer: Existing footer modes: [CC | Issued], [Distributed | Issued], [Issued], [Sent].
[0306] 24. Document Merging: You can select two documents for merging operations.
[0307] 25. Page Number: The page number templates are divided into:
General
Letter
[0308] 26. QR Code: As Figure 10 shown, the official document QR code is automatically generated by capturing data. * is a required item. After confirming that the captured data is correct, click OK to generate the QR code.
[0309] 27. About: Click the
About
[0310] The variables and explanations of the present invention are shown in Table 8 below:
[0311] Table 8 Variable and Explanation Table
[0312]
[0313]
[0314] Specifically, the principle of the present invention is: The core technical idea of the present invention is to extract multi-modal features of document layout elements by constructing a feature processing model based on a multi-branch neural network, and use a contribution evaluation model to analyze the influence degree of each element on the overall typesetting quality, so as to generate an optimized document layout scheme. At the same time, combined with optimization strategies such as rendering task partitioning and parallel computing, a significant improvement in document rendering efficiency has been achieved;
[0315] Specifically, in the feature extraction stage, the method first parses the OpenXML tags of the document to extract the basic data of layout elements such as text, tables, and pictures, including information such as content, format, and position. Then, for different types of layout elements, a table branch neural network, a text branch neural network, and a picture branch neural network are respectively constructed to extract their unique visual, structural, and semantic features. Through the feature fusion layer, these heterogeneous features are integrated into the overall feature vector of the document, laying a foundation for subsequent layout optimization;
[0316] In the contribution evaluation stage, the method constructs a special evaluation model, and calculates their contribution values by analyzing the influence degree of each layout element on the document typesetting score. This includes the importance of the element itself and its mutual influence with other elements. With the help of this quantified contribution information, the layout priority and layout spacing of each layout element in the document can be determined, and an optimized layout scheme can be generated;
[0317] In the rendering optimization stage, this method proposes a set of rendering optimization equations, which involve block size calculation, hardware task allocation, calculation overhead estimation, and execution order optimization, etc. Through these mathematical models, rendering tasks can be reasonably divided according to actual factors such as document complexity and hardware resources, and a multi-threaded parallel mechanism can be adopted for efficient calculation, ultimately greatly improving the performance of document rendering.
[0318] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention.
Claims
1. A method for rapid office document arrangement, characterized in that: include: Parse the OpenXML tags of the document, extract the node data in document.xml for feature extraction, and construct the original feature data; construct a feature processing model including a table branch neural network, a text branch neural network, an image branch neural network, a feature fusion layer, and a scoring output layer; construct a contribution evaluation model to analyze the contribution value of the layout elements in the document to the document typesetting score; obtain the contribution evaluation model through training; determine the layout priority and layout interval value according to the contribution value; generate the OpenXML tag tree of the document; perform dimensionality reduction mapping on the document feature vector and map it to the three-dimensional display space to generate a three-dimensional layout mapping model; perform partition rendering calculation on the typesetting result document through a rendering optimization equation group including block size equation, hardware allocation equation, computational overhead equation, and execution order equation, determine the graphics card rendering block, processor rendering block, memory storage block and their processing order, and use a multi-threaded parallel mechanism for rendering calculation.
2. A method for rapid office document arrangement according to claim 1, characterized in that: The table branch neural network in the feature processing model includes a horizontal table model, a vertical table model, and a table feature fusion layer. The horizontal table model analyzes the row data attributes of the w:tr node and the cell data attributes of the w:tc node. The vertical table model analyzes the column data attributes of the w:gridCol node and the merged data attributes of the w:vMerge node. The table feature fusion layer fuses the row and column feature data output by the horizontal table model with the merged feature data output by the vertical table model to output the table feature vector.
3. A method for rapid office document arrangement according to claim 2, characterized in that: The text branch neural network in the feature processing model includes a text extraction layer, a text semantic layer, and a text structure layer. The text extraction layer analyzes text content data and format data, the text semantic layer calculates the semantic connection value of the text paragraph, and the text structure layer analyzes the paragraph data attribute of the w:pPr node, the spacing data attribute of the w:spacing node, and the indentation data attribute of the w:ind node, and outputs the text feature vector.
4. A method for rapid office document arrangement according to claim 3, characterized in that: The image branch neural network in the feature processing model includes an image extraction layer, an image positioning layer, and an image-text connection layer. The image extraction layer analyzes image size data and clarity data, the image positioning layer analyzes image coordinate position data, and the image-text connection layer analyzes the positioning data attribute of the w:anchor node, the horizontal data attribute of the wp:positionH node, and the vertical data attribute of the wp:positionV node, and outputs the image feature vector.
5. A method for rapid office document arrangement according to claim 4, characterized in that: The feature fusion layer uses a feature dimension reduction algorithm to map the table feature vector, the text feature vector, and the image feature vector to a three-dimensional feature space. The feature dimension reduction algorithm includes a main feature analysis method and a random embedding method. The main feature analysis method is used to extract the main direction of high-dimensional features, and the random embedding method is used to maintain the local structure of features. The output of the feature fusion layer is the document feature vector.
6. A method for rapid office document arrangement according to claim 5, characterized in that: The rendering optimization equation group includes a block size equation, a hardware allocation equation, a computational overhead equation, and an execution order equation. The block size equation is used to calculate the size value of the rendering calculation block, the hardware allocation equation is used to calculate the task allocation ratio of the hardware device, the computational overhead equation is used to calculate the computational consumption of the rendering calculation block, and the execution order equation is used to optimize the calculation order of the rendering calculation block.
7. A method for rapid office document arrangement according to claim 6, characterized in that: The feature data collection steps include: collecting manually arranged documents as basic data; extracting OpenXML tags of the basic data and recording scoring data; dividing the basic data into a standard sample data set and an optimized sample data set; extracting the position data, spacing data, and style data of layout elements; and constructing original feature data.
8. A method for rapid office document arrangement according to claim 7, characterized in that: The model training steps include: using the original feature data to train the feature processing model; constructing the training data of the contribution evaluation model based on the calculation results of the feature processing model; using the original feature data to train the contribution evaluation model; evaluating the performance indicators through the cross-validation method; and saving the training parameter data.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program instructions, and when the program instructions are executed in a computer, they are used to execute the office document rapid layout method according to any one of claims 1 to 8.
10. A system for rapid office document arrangement, characterized in that: A computer-readable storage medium comprising the computer-readable storage medium of claim 9.
Citation Information
Cited By
DTP service complexity auxiliary evaluation method based on AI and typesetting analysis
CN120911402A
DTP service complexity aided evaluation method based on ai and layout analysis
CN120911402B