A method and system for automatic document typesetting
By performing semantic understanding analysis and replanning adjustments on documents, the problem of traditional typesetting affecting reading continuity is solved, optimizing the reading experience and achieving a typesetting effect that conforms to strict standards.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING KONOS TECH CO LTD
- Filing Date
- 2026-02-28
- Publication Date
- 2026-06-05
AI Technical Summary
Traditional layout and typesetting mechanically apply rules, which can easily affect the reader's reading flow, especially when the article format is adjusted, resulting in a poor reading experience.
By analyzing the current layout, the physical segmentation of document content at the reading area jump is identified, and the degree of semantic coherence disruption is assessed. Based on predefined assessment rules, replanning and adjustments are made to reduce semantic coherence disruption and meet predetermined layout constraints.
Without altering the semantic meaning of the document content, optimize the layout, improve the reading experience, reduce disruptions to semantic coherence, minimize manual work for users, and ensure that the document conforms to strict specifications.
Smart Images

Figure CN122154630A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of document editing, and in particular to a method and system for automatic document formatting. Background Technology
[0002] In the writing process of articles, especially academic articles, the formatting usually needs to be manually adjusted according to specifications. However, some editors can achieve semi-automatic or fully automated layout and typesetting by introducing relevant templates or typesetting rules. Traditional layout and typesetting often mechanically apply rules, and in layout optimization, it can only adjust obvious layout defects such as "orphan lines." Mechanical typesetting and pagination can easily affect the reader's reading flow. Summary of the Invention
[0003] The purpose of this application is to propose a method and system for automatic document typesetting, in order to solve the problem that traditional layout and typesetting often mechanically applies rules, which can easily affect the reader's reading flow.
[0004] The document automatic formatting method in this application includes:
[0005] The current typesetting result is analyzed to identify the physical segmentation of the document content at the jump point in the reading area, and the degree of semantic coherence disruption of the segmented document content is assessed.
[0006] If the degree of semantic coherence disruption exceeds the preset limit, the document layout will be replanned and adjusted while satisfying the predetermined layout constraints, so as to determine the layout adjustment method that can reduce the degree of semantic coherence disruption.
[0007] The document content is formatted and adjusted based on the determined formatting adjustment method.
[0008] Optionally, identifying the physical segmentation of document content at the jump point in the reading area and assessing the degree of semantic coherence disruption of the segmented document content includes:
[0009] Identify the semantic type and logical structure of the segmented document content;
[0010] Based on the semantic type, the logical structure, and the location where the document content is segmented, the degree of semantic coherence disruption is determined according to predefined evaluation rules.
[0011] Optionally, determining the degree of semantic coherence violation according to predefined evaluation rules includes:
[0012] Based on the logical structure of the content before and after the segmented part of the document, identify the type of logical relationship that was cut off due to the segmentation, and determine the basic score assigned to the type of logical relationship in the evaluation rules;
[0013] A first adjustment factor for the score is determined from the evaluation rules based on the proportion of a single content block that is divided into the next reading area;
[0014] A second adjustment coefficient is determined from the evaluation rule based on the semantic type;
[0015] The score for the degree of semantic coherence disruption is determined based on the basic score, the first adjustment coefficient, and the second adjustment coefficient.
[0016] Optionally, when determining the first adjustment coefficient for the score from the evaluation rules based on the proportion of the content block being segmented into the next reading area, the proportion of the content block being segmented into the next reading area is positively correlated with the first adjustment coefficient.
[0017] Optionally, determining the degree of semantic coherence violation according to predefined evaluation rules further includes:
[0018] Additional penalty scores are determined from the evaluation rules based on the type of content units into which the document content is segmented.
[0019] Optionally, determining additional penalty scores from the evaluation rules based on the type of content units into which the document content is segmented includes at least one of the following:
[0020] If the document content is split at the middle of a sentence, a first penalty score is added to the basic score;
[0021] If the document content is split between two related logical connectives, a second penalty score is added to the basic score.
[0022] If the document content is split in the middle of consecutive numbered items or list items, a third penalty score is added to the basic score.
[0023] Optionally, before parsing the current typesetting result, the process further includes:
[0024] Determine the reading environment for which the document was formatted;
[0025] When the target reading environment is the current display terminal, obtain the screen parameters of the display terminal, and divide the reading area jump point according to the display range determined by the screen parameters;
[0026] When the target reading environment is a publication, the corresponding basic layout style is determined, and the reading area jump point is determined according to the pagination and / or column positions of the basic layout style.
[0027] Optionally, parsing the current typesetting result further includes:
[0028] Based on the position and logical relationship between the document content, identify content that is adjacent in position and not suitable for splitting and form a group of related elements, wherein the group of related elements includes at least two different types of content.
[0029] When the target reading environment is the current display terminal, the process of re-planning and adjusting the document layout will also constrain the content within the related element group to the same screen.
[0030] When the target reading environment is a publication, the process of re-planning and adjusting the document layout will also constrain the content within the related element group to the same page or column.
[0031] Optionally, the step of re-planning and adjusting the document's layout while satisfying predetermined layout constraints to determine a layout adjustment method that can reduce the degree of semantic coherence disruption includes:
[0032] Once the location where the degree of semantic coherence violation exceeds the preset limit is identified, replanning and adjustment are performed backwards.
[0033] If forward replanning cannot eliminate all the aforementioned semantic coherence disruptions exceeding a preset limit, then full-text replanning is performed.
[0034] Choose the option that minimizes changes to the overall layout or the total disruption to semantic coherence during the replanning and adjustment of the entire text.
[0035] On the other hand, this application also provides an automatic document formatting system, including:
[0036] The semantic evaluation unit is used to parse the current typesetting result, identify the physical segmentation of the document content at the jump point in the reading area, and evaluate the degree of semantic coherence disruption of the segmented document content.
[0037] The layout planning unit is used to re-plan and adjust the layout of the document if the degree of semantic coherence disruption exceeds a preset limit, while satisfying predetermined layout constraints, so as to determine the layout adjustment method that can reduce the degree of semantic coherence disruption.
[0038] The typesetting application unit is used to adjust the content of the document based on the determined typesetting adjustment method.
[0039] The automatic document formatting method provided in this application analyzes the current document formatting based on semantic understanding. It can identify instances where semantic coherence is disrupted due to formatting issues. While meeting predetermined formatting constraints, it re-plans and adjusts the document's formatting to find a way to reduce the degree of semantic coherence disruption. This eliminates the need for manual user intervention and reduces workload. Furthermore, while ensuring the document conforms to strict specifications, it further optimizes the formatting based on semantic understanding, making the author's work easier to understand and improving the reader's experience. Attached Figure Description
[0040] Figure 1 A basic flowchart illustrating the document automatic formatting method provided in this application embodiment;
[0041] Figure 2 This is a schematic diagram of the structure of the automatic document formatting system provided in the embodiments of this application. Detailed Implementation
[0042] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0043] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0044] Example:
[0045] This embodiment provides a method for automatic document formatting. Please refer to [link / reference]. Figure 1 This method includes, but is not limited to, the following steps:
[0046] S101. Analyze the current typesetting result, identify the physical segmentation of the document content at the jump point in the reading area, and evaluate the degree of semantic coherence disruption of the segmented document content.
[0047] In this embodiment, the reading area jump point refers to the location where the document content is divided during reading, requiring the user to jump to other non-contiguous positions to continue viewing the next content, such as pagination or column breaks. Currently, viewing documents on electronic terminals has become a common reading method. After turning pages on an electronic terminal, the user's viewing position jumps from the bottom to the top of the screen. Therefore, the bottom of the reading sequence on an electronic terminal can also be considered a reading area jump point.
[0048] It is understood that this embodiment will analyze the semantics of the document content. During this process, a pre-trained natural language processing model can be used to understand the document content. Semantic coherence in this embodiment refers to the degree of connection between the preceding and following text when the user is reading. For example, if the same sentence is directly split, semantic coherence will be disrupted. Of course, splitting the same paragraph or even other non-textual content within the document can also disrupt semantic coherence. In a reading process with poor semantic coherence, the user's reading experience is unsatisfactory.
[0049] S102. If the degree of semantic coherence disruption exceeds the preset limit, the document layout will be replanned and adjusted while meeting the predetermined layout constraints, so as to determine the layout adjustment method that can reduce the degree of semantic coherence disruption.
[0050] In this embodiment, the degree of semantic coherence disruption is quantitatively assessed. The preset limit is set based on actual needs. For example, the preset limit can be lowered in the first draft stage to avoid frequent changes in the layout during the revision process. In the final draft stage, the preset limit can be raised to form a final layout result with better semantic coherence.
[0051] In this embodiment, the pre-defined formatting constraints are rules pre-entered by the user or editor. These constraints typically conform to the strict formatting requirements of certain journals or academic articles and can constrain parameters such as font size, line spacing, and headers. This ensures that even after automatically adjusting the document's formatting, this embodiment still meets the user's actual needs.
[0052] In this embodiment, possible layout adjustments include, but are not limited to, adjusting parameters such as font size and line spacing that can change the display position of document content while meeting predetermined layout constraints. For example, if a graduation thesis requires a font size of "size 4" or "small 4," then the font size originally set to "size 4" can be adjusted to "small 4." Of course, in practical applications, these adjustments to font size and line spacing, which apply throughout the entire text, often cannot always effectively ensure semantic coherence. In this embodiment, the document content can also be directly modified without changing the semantics to adjust its final layout. For example, based on semantic understanding, text can be segmented or paragraphs merged without changing the semantics, which can change the local layout of the document content within a relatively small range. Although such adjustments make minor adjustments to the original document content, they improve the user's reading experience and understanding of the article because they enhance the continuity of reading. Furthermore, this embodiment can determine the textual association range of non-textual document content during semantic analysis and adjust the position of the non-textual document content within this range to adjust the layout.
[0053] This embodiment uses a natural language model to provide a preliminary interpretation of the content semantics and segmentation, and combines preset rules to determine the degree of semantic coherence disruption. This avoids directly using the model to obtain the evaluation results of the degree of semantic coherence disruption, and greatly simplifies the model building and training process and reduces the amount of computation.
[0054] S103. Adjust the layout of the document content based on the determined layout adjustment method;
[0055] Understandably, if there is no layout adjustment method that can reduce the degree of disruption to semantic coherence based on the currently available adjustment means, then no layout adjustment will be made to the content of the document.
[0056] The automatic document layout method in this embodiment can automatically adjust the layout of the document content based on the current content distribution, eliminating the need for manual user adjustments. The calculation of the layout adjustment method in this embodiment can be processed in the background, and after determining the final usable layout adjustment method, it is applied automatically or based on user instructions.
[0057] This embodiment's automatic document formatting method analyzes the current document formatting based on semantic understanding. It can identify instances where semantic coherence is disrupted due to formatting issues. While meeting predetermined formatting constraints, it re-plans and adjusts the document's formatting to find a way to reduce the degree of semantic coherence disruption. This eliminates the need for manual user intervention and reduces workload. Traditional formatting aims merely to make documents conform to rigid specifications, but mechanical formatting can easily lead to unreasonable content segmentation. This embodiment introduces semantic understanding, further optimizing the formatting based on the semantic understanding experience while ensuring document compliance with rigid specifications. This makes the author's work easier to understand and improves the reader's reading experience.
[0058] In some implementations, identifying the physical segmentation of document content at reading area transitions and assessing the degree of semantic coherence disruption of the segmented document content includes:
[0059] S201. Identify the semantic type and logical structure of the segmented document content;
[0060] In this embodiment, semantic type identification is performed on a content block basis. A content block can typically be a text paragraph, an image, a table, and corresponding figure / table titles. In this embodiment, if the content block is plain text, it can be categorized into formulas, lists, theorems / proofs, and other types based on its presentation or content logic. If the content block includes images, the images or tables, along with their corresponding figure / table titles, can be categorized into chart / graph types. This process can be performed by checking for inserted images in the document, or in some examples, visual detection can be introduced for accurate identification.
[0061] S202. Based on semantic type, logical structure, and the location where the document content is segmented, determine the degree of semantic coherence disruption according to predefined evaluation rules.
[0062] The impact of segmentation on semantic coherence may vary depending on the semantic type. For example, segmenting the image and title of a chart significantly affects semantic coherence. In practice, text paragraph segmentation is more common, but the impact varies depending on the semantic content and the segmentation location. For instance, segmenting a text paragraph that primarily describes a theorem and explains its definition and proof will significantly impact the user's reading coherence; however, segmentation of background information or overviews has less impact on comprehension. Therefore, in predefined evaluation rules, different evaluation methods (e.g., assigning different scores or weighting coefficients) can be configured for different semantic types to differentiate the potential damage to semantic coherence after segmentation. In this embodiment, semantic types include, but are not limited to, at least one of charts, definitions, theorems, conclusions, formulas, lists, and plain text.
[0063] In this embodiment, the logical structure refers to the logical structure of the content before and after the segmented part. For plain text, its logical structure can be the logical structure between two parts of the same paragraph, such as causal relationships / progressive proofs. For charts, its logical structure can be the reference structure of the chart. In practical applications, the reference logical structure can be determined based on the textual semantics of the preceding and following paragraphs and the chart's title or table title. That is, the logical structure may reflect the logic between different parts within a content block, or it may be the logic between a content block and other content blocks. In practical applications, the method for determining its logical structure can be specified based on different content formats.
[0064] Similarly, the impact of segmentation on semantic coherence can vary depending on the logical structure. For example, even when splitting a list, the degree of impact will differ depending on the logical hierarchy between the list items. When the hierarchy between list items is simple, progressive, or parallel, the impact of segmentation on semantic coherence is relatively small. However, if there are multiple complex nested hierarchies among the list items, segmentation may severely disrupt semantic coherence, forcing users to repeatedly compare and contrast the items to clarify their relationships. Furthermore, for plain text descriptions, based on different logical structures, the logical structure can be identified by recognizing logical keywords (such as "because," "assuming," "then," "therefore," "hence," "Q.E.D.," etc.) during some implementation processes.
[0065] The placement of document content segments affects the display size of individual content blocks within the current and next reading areas. When less content is segmented into the next reading area, readers can usually grasp the main idea of the content block within the current reading area, with minimal impact on overall comprehension. Conversely, when a large amount of content is segmented into the next reading area, readers may need to combine information from two pages to fully understand it.
[0066] In practical applications, evaluation rules can be formulated based on the patterns observed in the above situations. This embodiment determines the degree of semantic coherence disruption based on predefined evaluation rules, and can comprehensively evaluate the loss of semantic coherence caused by segmentation from multiple different dimensions.
[0067] In some implementations of this embodiment, determining the degree of semantic coherence violation according to predefined evaluation rules includes:
[0068] S301. Based on the logical structure of the content before and after the segmented part of the document content, identify the type of logical relationship that was cut off due to the segmentation, and determine the basic score assigned to the type of logical relationship in the evaluation rules.
[0069] S302. Determine the first adjustment factor for the score from the evaluation rules based on the proportion of a single content block that is divided into the next reading area;
[0070] S303. Determine the second adjustment coefficient from the evaluation rules based on the semantic type;
[0071] S304. Determine the score value of the degree of semantic coherence disruption based on the basic score, the first adjustment coefficient, and the second adjustment coefficient.
[0072] In this embodiment, the score for the degree of semantic coherence disruption can be calculated by multiplying the base score by a first adjustment coefficient and then by a second adjustment coefficient. This embodiment is based on the scenario where the logical structure of the document content is broken, effectively ensuring logical clarity and continuity during reading comprehension. This significantly contributes to the assessment of semantic coherence. By adjusting the score in conjunction with the segmentation location and the type of semantics expressed by the content, it can more accurately reflect the overall state of semantic coherence, achieving an effective quantitative assessment of the disruption to semantic coherence.
[0073] In some implementations of this embodiment, when determining the first adjustment coefficient for the score based on the proportion of content blocks segmented into the next reading area from the evaluation rules, the proportion of content blocks segmented into the next reading area is positively correlated with the first adjustment coefficient. It should be understood that although both are 3 / 7 content segments, the first adjustment coefficient when a content block displays 70% of its content in the current reading area will be lower than the first adjustment coefficient when a content block displays 30% of its content in the current reading area. That is, in this embodiment, the more content blocks displayed in the current reading area, the less disruption to semantic coherence is considered. Although low and high proportions of segmentation are mathematically equivalent (e.g., both are 3 / 7 distributions), they differ in reading experience. Readers are more likely to use the content in the current reading area as the primary anchor point for cognition; therefore, the two segmentation scenarios do not have the same degree of cognitive interruption for the user. Based on this, this embodiment sets the proportion of content blocks segmented into the next reading area to be positively correlated with the first adjustment coefficient to better ensure a coherent reading experience for the user.
[0074] Based on the specific location where the document content is segmented, content units smaller than a single content block may be segmented. These content units include, but are not limited to, a sentence, several closely related sentences, or a list. Segmenting such content units typically causes greater disruption to semantic coherence. For example, segmenting within a sentence usually causes greater disruption to semantic coherence than segmenting before or after a sentence; similarly, segmenting within a list item causes greater coherence loss than segmenting before or after a list item. To better assess the disruption to semantic coherence caused by page segmentation, some implementations, in addition to determining the degree of semantic coherence disruption according to predefined evaluation rules, also include determining additional penalty scores from the evaluation rules based on the type of content unit segmented from the document content.
[0075] In some implementations, additional penalty scores are determined from the evaluation rules based on the type of content units into which the document content is segmented, including at least one of the following:
[0076] If the document content is split at the middle of a sentence, a first penalty score will be added to the basic score;
[0077] If the document content is split between two related logical connectors, a second penalty score will be added to the basic score.
[0078] If the document content is split in the middle of consecutive numbered items or list items, a third penalty score will be added to the basic score.
[0079] The first, second, and third penalty scores can be set according to actual needs and can be the same or different. In this embodiment, the score for the degree of semantic coherence disruption can be the sum of the basic score and the penalty score multiplied by the first adjustment coefficient and the second adjustment coefficient.
[0080] In actual reading, authors have different needs for semantic coherence depending on the reading environment. In the reading environment of publications, the constraints on typesetting are stricter, so a lower preset limit can be set for the degree of semantic coherence disruption. In reading environments such as electronic terminals, readers primarily prioritize ease of reading and viewing, so a higher preset limit can be set for the degree of semantic coherence disruption. It is understandable that users' actual reading methods differ depending on the reading environment. To ensure a better reading experience for users across different platforms and media, some implementations include, before parsing the current typesetting result, determining the reading environment for which the document was typeset. When the target reading environment is the current display terminal, the screen parameters of the display terminal are obtained, and the reading area jump points are divided according to the display range determined by the screen parameters. When the target reading environment is a publication, the corresponding basic typesetting style is determined, and the reading area jump points are determined according to the pagination and / or column positions of the basic typesetting style. It is understandable that the typesetting style for different publications is written into predetermined typesetting constraints.
[0081] In some implementations, parsing the current layout result further includes: determining adjacent content that is not suitable for splitting based on the position and logical correlation between document content, and forming related element groups, where each related element group includes at least two different types of content. When the target reading environment is the current display terminal, the content within the related element group is also constrained to the same screen during the document layout replanning and adjustment process. When the target reading environment is a publication, the content within the related element group is also constrained to the same page or column during the document layout replanning and adjustment process. In this embodiment, the display terminal includes, but is not limited to, devices with display functions such as computers, mobile phones, and smartwatches.
[0082] It is understandable that this embodiment requires the application of a natural language model for parsing, and therefore the associated element groups can be determined together with the results of the natural language model. Of course, this requires the corresponding training and configuration of the natural language model.
[0083] For well-defined and easily identifiable associations, matching can be performed using predefined rules, such as matching images and titles to form associated element groups. For more complex and implicit associations, natural language models can be used to determine them. In this embodiment, the natural language model is primarily determined based on the logical correlation between document content.
[0084] As an example, the suitability of content for splitting can be determined by analyzing the content similarity, structural relevance, and functional relevance between adjacent document content. Determining content similarity involves calculating semantic similarity; for example, content mentioning the same entity or having a clear reference relationship typically has high semantic similarity. Determining structural relevance involves analyzing the relationship between content within the document structure, such as whether they are in the same section or at the same level. Determining functional relevance involves analyzing the function of content within the document; based on a defined degree of functional association, a functional relevance score can be obtained. For example, a strong functional association between "theorem statement" and "its proof process" will be assigned a higher weight. Finally, the results of the above three dimensions are weighted and fused to obtain a comprehensive "logical relevance" score. When this score exceeds a preset threshold, these contents are identified as a group of related elements. It should be understood that a group of related elements includes at least two different types of content, such as an image and explanatory text referencing that image.
[0085] Traditional terminal reading experiences suffer greatly from zooming or simple sequential formatting, resulting in a poor reading experience and easily disrupting the original text's logical structure. This embodiment analyzes the reading environment and actual reading area before formatting, ensuring a good reading experience in various environments and adapting to different reading terminals. Furthermore, in some implementations, content within related element groups is constrained to the same screen, allowing users to view closely related content on a single screen, significantly improving the reading experience.
[0086] To reduce the scope and computational load of global layout adjustments, some implementations re-plan and adjust the document's layout while meeting predetermined layout constraints. The layout adjustment methods that can reduce the disruption of semantic coherence include:
[0087] S501. Identify the location where the last instance of semantic coherence disruption exceeds the preset limit, and then replan and adjust the layout forward.
[0088] S502. If forward replanning cannot eliminate all semantic coherence disruptions exceeding the preset limit, perform full-text replanning.
[0089] S503. Select the option that minimizes changes to the overall layout or the total disruption to semantic coherence during the replanning and adjustment of the entire text.
[0090] In practical applications, you can try a variety of different layout adjustment methods and select the applicable solution from them.
[0091] This embodiment also provides an automatic document formatting system 100, see [link to documentation]. Figure 2 As shown, it includes, but is not limited to, a semantic evaluation unit 101, a typesetting planning unit 102, and a typesetting application unit 103.
[0092] The semantic evaluation unit 101 is used to parse the current typesetting result, identify the physical segmentation of the document content at the jump point in the reading area, and evaluate the degree of semantic coherence disruption of the segmented document content.
[0093] The layout planning unit 102 is used to re-plan and adjust the layout of the document while meeting the predetermined layout constraints if there is a situation where the degree of semantic coherence disruption exceeds the preset limit, so as to determine the layout adjustment method that can reduce the degree of semantic coherence disruption.
[0094] The typesetting application unit 103 is used to adjust the content of a document based on the determined typesetting adjustment method.
[0095] The specific execution process of each unit in the document automatic formatting system 100 can also refer to the steps of the document automatic formatting method provided above in this embodiment, which will not be repeated in this embodiment.
[0096] Furthermore, although exemplary embodiments have been described herein, their scope includes any and all embodiments based on this disclosure that have equivalents, modifications, omissions, combinations (e.g., schemes involving intersections of various embodiments), adaptations, or changes. They are not limited to the examples described in this specification or during the implementation of this application, and such examples are to be interpreted as non-exclusive.
[0097] The above description is intended to be illustrative and not restrictive. For example, the above examples (or one or more of them) can be used in combination with each other. Other embodiments can be used by those skilled in the art when reading the above description.
[0098] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various modifications and variations can be made to the embodiments of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for automatic document formatting, characterized in that, include: The current typesetting result is analyzed to identify the physical segmentation of the document content at the jump point in the reading area, and the degree of semantic coherence disruption of the segmented document content is assessed. If the degree of semantic coherence disruption exceeds the preset limit, the document layout will be replanned and adjusted while satisfying the predetermined layout constraints, so as to determine the layout adjustment method that can reduce the degree of semantic coherence disruption. The document content is formatted and adjusted based on the determined formatting adjustment method.
2. The document automatic formatting method as described in claim 1, characterized in that, The process of identifying the physical segmentation of document content at the jump points in the reading area and assessing the degree of semantic coherence disruption of the segmented document content includes: Identify the semantic type and logical structure of the segmented document content; Based on the semantic type, the logical structure, and the location where the document content is segmented, the degree of semantic coherence disruption is determined according to predefined evaluation rules.
3. The document automatic formatting method as described in claim 2, characterized in that, The determination of the degree of semantic coherence violation according to predefined evaluation rules includes: Based on the logical structure of the content before and after the segmented part of the document, identify the type of logical relationship that was cut off due to the segmentation, and determine the basic score assigned to the type of logical relationship in the evaluation rules; A first adjustment factor for the score is determined from the evaluation rules based on the proportion of a single content block that is divided into the next reading area; A second adjustment coefficient is determined from the evaluation rule based on the semantic type; The score for the degree of semantic coherence disruption is determined based on the basic score, the first adjustment coefficient, and the second adjustment coefficient.
4. The document automatic formatting method as described in claim 3, characterized in that, When determining the first adjustment coefficient for the score from the evaluation rules based on the proportion of the content block being divided into the next reading area, the proportion of the content block being divided into the next reading area is positively correlated with the first adjustment coefficient.
5. The document automatic formatting method as described in claim 3, characterized in that, The determination of the degree of semantic coherence violation based on predefined evaluation rules also includes: Additional penalty scores are determined from the evaluation rules based on the type of content units into which the document content is segmented.
6. The document automatic formatting method as described in claim 5, characterized in that, The determination of additional penalty scores from the evaluation rules based on the type of content units segmented from the document content includes at least one of the following: If the document content is split at the middle of a sentence, a first penalty score is added to the basic score; If the document content is split between two related logical connectives, a second penalty score is added to the basic score. If the document content is split in the middle of consecutive numbered items or list items, a third penalty score is added to the basic score.
7. The document automatic formatting method as described in claim 1, characterized in that, Before parsing the current layout result, the process also includes: Determine the reading environment for which the document was formatted; When the target reading environment is the current display terminal, obtain the screen parameters of the display terminal, and divide the reading area jump point according to the display range determined by the screen parameters; When the target reading environment is a publication, the corresponding basic layout style is determined, and the reading area jump point is determined according to the pagination and / or column positions of the basic layout style.
8. The document automatic formatting method as described in claim 7, wherein parsing the current formatting result further includes: Based on the position and logical relationship between the document content, identify content that is adjacent in position and not suitable for splitting and form a group of related elements, wherein the group of related elements includes at least two different types of content. When the target reading environment is the current display terminal, the process of re-planning and adjusting the document layout will also constrain the content within the related element group to the same screen. When the target reading environment is a publication, the process of re-planning and adjusting the document layout will also constrain the content within the related element group to the same page or column.
9. The document automatic formatting method according to any one of claims 1-8, characterized in that, The step of re-planning and adjusting the document's layout while satisfying predetermined layout constraints to determine a layout adjustment method that can reduce the degree of semantic coherence disruption includes: Once the location where the degree of semantic coherence violation exceeds the preset limit is identified, replanning and adjustment are performed backwards. If forward replanning cannot eliminate all the aforementioned semantic coherence disruptions exceeding a preset limit, then full-text replanning is performed. Choose the option that minimizes changes to the overall layout or the total disruption to semantic coherence during the replanning and adjustment of the entire text.
10. An automatic document formatting system, characterized in that, include: The semantic evaluation unit is used to parse the current typesetting result, identify the physical segmentation of the document content at the jump point in the reading area, and evaluate the degree of semantic coherence disruption of the segmented document content. The layout planning unit is used to re-plan and adjust the layout of the document if the degree of semantic coherence disruption exceeds a preset limit, while satisfying predetermined layout constraints, so as to determine the layout adjustment method that can reduce the degree of semantic coherence disruption. The typesetting application unit is used to adjust the content of the document based on the determined typesetting adjustment method.