A method for extracting the outline of a Word document
By extracting and preprocessing the format information of Word documents, combining outline extraction and verification modules, automatically identifying and marking the outline, it solves the problem that writers have difficulty in building a clear outline, and quickly generate outlines, improving writing efficiency.
Patent Information
- Application Number
- CN202210794259.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-05
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-07-05
AI Technical Summary
When writing high-quality application articles, many people find it difficult to build a clear and logical outline, and it takes a lot of time due to data collection and outline modification.
Through a Word document outline extraction method, the Word document is extracted and preprocessed using the analysis module and the extraction unit to generate structured data, and through the outline extraction module, verification module and other units, the outline is automatically identified and marked, and the hierarchical errors, missing and duplicate problems are corrected, and the outline can be generated for reference.
It realizes rapid and accurate extraction and generation of Word documents outlines, saving writers time on building outlines and improving business processing efficiency.
Smart Images

Figure CN115129817B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of document extraction, and specifically to a method for extracting the outline of a Word document. Background Art
[0002] Applied writing is a practical writing style commonly used by modern enterprises to handle daily business. With the popularization of information technology, the business volume faced by enterprises and institutions is increasing day by day. In order to improve the quality and efficiency of handling daily business work, employees are required to write more and better applied writing.
[0003] Due to the characteristics of applied writing itself, in order to write high-quality applied writing, the writer needs to build a clear-level and logically rigorous article outline after referring to a large number of highly relevant materials.
[0004] However, in actual writing, many people cannot write an outline or are not used to writing an outline. Even those who are willing to write an outline will consume a lot of time due to actions such as collecting materials and repeatedly modifying the outline.
[0005] Therefore, the present invention provides a method for extracting the outline of a Word document to solve the problems raised in the above background art. Summary of the Invention
[0006] The purpose of the present invention is to provide a method for extracting the outline of a Word document to solve the problems raised in the above background art.
[0007] To achieve the above purpose, the present invention provides the following technical solutions:
[0008] A method for extracting the outline of a Word document, comprising the following steps:
[0009] Step 1: Import the Word document into the system;
[0010] Import the Word document used as reference materials into the system through the system terminal;
[0011] Step 2: Read the Word document format information;
[0012] Receive the Word document imported into the system through the parsing module, and then extract the text, font, font size, format attributes, and paragraph format of the Word document by the extraction unit to generate structured data in units of paragraphs, and transmit it together with the Word document to the document preprocessing module for processing;
[0013] Step 3: Preprocess the Word document;
[0014] The preprocessing module receives the Word document and structured data sent by the parsing module. The merging unit merges the font, font size, and their format attributes of each paragraph to make the text format consistent. Then it is handed over to the list unit, which divides the merged paragraphs based on punctuation marks to generate the first sentence of the paragraph and the remaining text.
[0015] After updating the unit structured data based on the generated first sentence of the paragraph and the remaining text, the unit structure data is added to the list, and finally the full text divided by paragraphs is transmitted to the outline extraction module.
[0016] The statistical unit in the preprocessing module reads the font size catalog attributes in the structured data, statistically analyzes them separately, groups the combinations of font sizes by paragraph unit and counts the quantity of each combination, and at the same time counts the catalog format attribute values of the paragraphs. The statistical results are respectively transmitted to the outline verification module and the outline extraction module.
[0017] Step 4: Extract the outline of the Word document.
[0018] The outline extraction module receives the full text divided by paragraphs and the statistical results transmitted by the preprocessing module. The title marking unit first reads the font and font size of the first paragraph of the full text, and then checks whether the statistical value of the same font size combination in the statistical results is unique. If it is unique, it is marked as a title.
[0019] The hierarchical rule matching unit, based on the document number library, sets the first paragraph or the first sentence of the paragraph that matches the rules within the range as the current hierarchical rule sample and transmits it to the outline marking unit.
[0020] And the outline marking unit receives the hierarchical rule sample sent by the hierarchical rule matching unit, extracts the paragraphs with the same rules within the range according to the sample rules and marks them as the current hierarchical outline, and finally sends the marked outline to the outline verification module.
[0021] When there is a hierarchical error problem:
[0022] It is transmitted to the hierarchical error correction unit for correction. The hierarchical error correction unit reads the outline, divides the possible upper-level outline area, calculates the scores of the paragraphs within the area based on the font, font size, and format attributes, and marks the paragraph with the highest score as the outline. If the marking is successful, the newly generated outline is submitted to the outline generation unit.
[0023] When there is an outline missing problem:
[0024] It is then transmitted to the outline missing correction unit for correction. The outline missing correction unit extracts the font, font size, and format attributes of the current-level outline, matches paragraphs that are exactly the same within the missing range, marks the missing outlines, and verifies the outline numbers again. If they are not incremented continuously, the newly generated outline is submitted to the outline generation unit;
[0025] When there is an outline duplication problem:
[0026] It is then transmitted to the outline duplication correction unit for correction. The outline duplication correction unit cancels the duplicate outline marks, extracts the font, font size, and format attributes of the remaining outlines, verifies the cancelled outlines, and if they are exactly the same, submits the new outline to the outline generation unit;
[0027] Outline generation unit:
[0028] It is used to receive the outlines transmitted by the judgment unit or other units and generate them. After receiving the outlines transmitted by other units, the outline generation unit marks the text content following the outline as belonging to the outline, and saves the generated content for the writer to refer to when establishing the outline.
[0029] Step Five: Verify whether the extracted outline is complete;
[0030] The outline verification module receives the outlines extracted by the outline extraction module, and the judgment unit determines what problems exist in the extracted outlines according to the outline numbers in the number library, and hands them over to the corresponding units for processing according to the types of problems;
[0031] When there is a hierarchical error problem:
[0032] It is then transmitted to the hierarchical error correction unit for correction. The hierarchical error correction unit reads the outline, divides the possible upper-level outline areas, calculates the scores of the paragraphs within the areas based on the font, font size, and format attributes, marks the paragraph with the highest score as the outline, and if the marking is successful, submits the newly generated outline to the outline generation unit;
[0033] When there is an outline missing problem:
[0034] It is then transmitted to the outline missing correction unit for correction. The outline missing correction unit extracts the font, font size, and format attributes of the current-level outline, matches paragraphs that are exactly the same within the missing range, marks the missing outlines, and verifies the outline numbers again. If they are not incremented continuously, the newly generated outline is submitted to the outline generation unit;
[0035] When there is an outline duplication problem:
[0036] Then it is transmitted to the outline duplicate correction unit for correction. The outline duplicate correction unit cancels the duplicate outline marks, extracts the remaining outline font, font size and format attributes, verifies the cancelled outline, and submits the new outline to the outline generation unit if they are exactly the same;
[0037] Outline generation unit:
[0038] It is used to receive the outline transmitted by the judgment unit or other units and generate it. After receiving the outline transmitted by other units, the outline generation unit marks the text content following the outline as the text content belonging to the outline, and saves the generated content for the writer to refer to when establishing the outline.
[0039] A method for extracting the outline of a Word document. The outline extraction module includes a title marking unit, an outline marking unit and a hierarchical rule matching unit for extracting the outline;
[0040] The title marking unit reads the font and font size of the first paragraph of the document and marks the ones that meet the standard as titles;
[0041] The hierarchical rule matching unit, according to the number library of the document, sets the paragraphs or the first sentences that meet the rules as rule samples and transmits them to the outline marking unit;
[0042] After receiving the hierarchical rule samples, the outline marking unit marks the hierarchical outline according to the rules and sends the generated outline to the outline verification module.
[0043] Preferably: The outline verification module includes a judgment unit, a hierarchical correction unit, an outline missing correction unit, an outline duplicate correction unit and an outline generation unit for generating the outline, marking the text content and saving it.
[0044] Preferably: The outline verification module includes a judgment unit, a hierarchical correction unit, an outline missing correction unit, an outline duplicate correction unit and an outline generation unit for generating the outline, marking the text content and saving it.
[0045] Preferably: The hierarchical error correction unit repairs the hierarchy of the outline with hierarchical errors by re-marking the outline. The outline missing correction unit complements the missing outline by re-extracting the missing outline.
[0046] Preferably: The outline duplicate correction unit deletes the duplicate outline by canceling the duplicate outline marks for the outline with duplicate problems.
[0047] Preferably: The outline generation unit receives the outline finally transmitted by other units, marks the text following the outline as the text content of the outline, and saves the generated result.
[0048] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0049] The present invention adopts technologies such as knowledge graph and NLP. The present invention can identify and extract the outline of a Word document, and the extracted outline can be used as a reference for users when establishing the article outline, saving time required for the writer to write applied articles and improving the business processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is a three-dimensional structural schematic diagram of a method for extracting the outline of a Word document provided in Embodiment 1 of the present application.
[0051] Figure 2 is a structural schematic diagram of the system in a method for extracting the outline of a Word document provided in Embodiment 1 of the present application.
[0052] In the figure: DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0054] Embodiment 1:
[0055] Please refer to Figures 1-2 , in the embodiment of the present invention, a method for extracting the outline of a Word document,
[0056] A method for extracting the outline of a Word document includes the following steps:
[0057] Step 1: Import the Word document into the system;
[0058] Import the Word document used as reference materials into the system through the system terminal;
[0059] Step 2: Read the Word document format information;
[0060] Receive the Word document imported into the system through the parsing module, and then the extraction unit extracts the text, font, font size, format attributes, and paragraph format of the Word document to generate structured data in units of paragraphs, which is transmitted to the document preprocessing module for processing together with the Word document;
[0061] Step 3: Preprocess the Word document;
[0062] Through the preprocessing module, receive the Word document and structured data sent by the parsing module. The merging unit merges the font, font size, and their format attributes of each paragraph to make the text format consistent; then it is handed over to the list unit, which divides according to the merged paragraphs based on punctuation marks to generate the first sentence of the paragraph and the remaining text;
[0063] And according to the generated first sentence of the paragraph and the remaining text, after updating the unit structured data, add the unit structure data to the list, and finally transmit the full text divided by paragraphs to the outline extraction module;
[0064] The statistical unit in the preprocessing module reads the font size catalog attributes in the structured data, statistically analyzes them respectively, groups the combinations of font sizes by paragraph unit and counts the quantity of each combination, and at the same time counts the catalog format attribute values of the paragraphs. Transmit the statistical results to the outline verification module and the outline extraction module respectively;
[0065] Step 4: Extract the outline of the Word document;
[0066] The outline extraction module receives the full text divided by paragraphs and the statistical results transmitted by the preprocessing module. The title marking unit first reads the font and font size of the first paragraph of the full text, and then checks whether the statistical value of the same font size combination in the statistical results is unique. If it is unique, it is marked as a title;
[0067] The hierarchical rule matching unit, based on the document number library, sets the first paragraph or the first sentence of the paragraph that matches the rule within the range as the current hierarchical rule sample and transmits it to the outline marking unit;
[0068] And the outline marking unit receives the hierarchical rule sample sent by the hierarchical rule matching unit, extracts the paragraphs with the same rule within the range according to the sample rule and marks them as the current hierarchical outline, and finally sends the marked outline to the outline verification module;
[0069] When there is a hierarchical error problem:
[0070] Then it is transmitted to the hierarchical error correction unit for correction. The hierarchical error correction unit reads the outline, divides the possible upper-level outline area, calculates the scores of the paragraphs within the area according to the font, font size, and format attributes, and marks the paragraph with the highest score as the outline. If the marking is successful, submit the newly generated outline to the outline generation unit;
[0071] When there is an outline missing problem:
[0072] It is then transmitted to the outline missing correction unit for correction. The outline missing correction unit extracts the font, font size, and format attributes of the outline at the current level, matches paragraphs that are exactly the same within the missing range, marks the missing outline, and verifies the outline numbers again. If they are not incremented continuously, the newly generated outline is submitted to the outline generation unit;
[0073] When there is an outline duplication problem:
[0074] It is then transmitted to the outline duplication correction unit for correction. The outline duplication correction unit cancels the duplication mark of the outline, extracts the font, font size, and format attributes of the remaining outlines, verifies the cancelled outline, and if it is exactly the same, submits the new outline to the outline generation unit;
[0075] Outline generation unit:
[0076] It is used to receive the outline transmitted by the judgment unit or other units and generate it. After receiving the outline transmitted by other units, the outline generation unit marks the text content following the outline as belonging to the outline, and saves the generated content for the writer to refer to when establishing the outline.
[0077] Step Five: Verify whether the extracted outline is complete;
[0078] The outline verification module receives the outline extracted by the outline extraction module, and the judgment unit determines what problems exist in the extracted outline according to the outline numbers in the number library, and hands it over to the corresponding unit for processing according to the type of problem;
[0079] When there is a hierarchical error problem:
[0080] It is then transmitted to the hierarchical error correction unit for correction. The hierarchical error correction unit reads the outline, divides the possible upper-level outline area, calculates the scores of the paragraphs in the area according to the font, font size, and format attributes, and marks the paragraph with the highest score as the outline. If the marking is successful, the newly generated outline is submitted to the outline generation unit;
[0081] When there is an outline missing problem:
[0082] It is then transmitted to the outline missing correction unit for correction. The outline missing correction unit extracts the font, font size, and format attributes of the outline at the current level, matches paragraphs that are exactly the same within the missing range, marks the missing outline, and verifies the outline numbers again. If they are not incremented continuously, the newly generated outline is submitted to the outline generation unit;
[0083] When there is an outline duplication problem:
[0084] It is then transmitted to the outline duplicate correction unit for correction. The outline duplicate correction unit cancels the duplicate outline marks, extracts the remaining outline font, font size and format attributes, verifies the cancelled outline, and submits the new outline to the outline generation unit if it is exactly the same.
[0085] Outline generation unit:
[0086] It is used to receive the outline transmitted by the judgment unit or other units and generate it. After receiving the outline transmitted by other units, the outline generation unit marks the text following the outline as the text content belonging to the outline and saves the generated content for the writer to refer to when establishing the outline.
[0087] A method for extracting the outline of a Word document. The outline extraction module includes a title marking unit, an outline marking unit, and a hierarchical rule matching unit for extracting the outline.
[0088] The title marking unit reads the font and font size of the first paragraph of the document and marks the ones that meet the standard as titles.
[0089] The hierarchical rule matching unit, according to the number library of the document, sets the paragraphs or the first sentences that meet the rules as rule samples and transmits them to the outline marking unit.
[0090] After receiving the hierarchical rule samples, the outline marking unit marks the hierarchical outline according to the rules and sends the generated outline to the outline verification module.
[0091] The outline verification module includes a judgment unit, a hierarchical correction unit, an outline missing correction unit, an outline duplicate correction unit, and an outline generation unit, and is used to generate the outline, mark the text content, and save it.
[0092] The outline verification module includes a judgment unit, a hierarchical correction unit, an outline missing correction unit, an outline duplicate correction unit, and an outline generation unit, and is used to generate the outline, mark the text content, and save it.
[0093] The hierarchical error correction unit repairs the hierarchy of the outline with hierarchical errors by re-marking the outline. The outline missing correction unit completes the missing outline by re-extracting the missing outline for the outline with missing problems.
[0094] The outline duplicate correction unit deletes the duplicate outline by canceling the duplicate outline marks for the outline with duplicate problems.
[0095] The outline generation unit receives the outline finally transmitted by other units, marks the text following the outline as the text content of the outline, and saves the generated result.
[0096] Working principle:
[0097] First, through the terminal, import all the Word documents of the reference papers related to the paper theme collected into the system;
[0098] Then, the system reads the information of the imported Word document and extracts relevant information such as the text, font, font size, and format attributes of the Word document to generate structured data;
[0099] And the system will preprocess the document according to the structured data generated in the previous step. Merge the fonts, font sizes, and their format attributes in each paragraph. Then divide the merged paragraph according to punctuation marks to generate the first sentence of the paragraph and the remaining text. Update the unit structure data according to the generated content, add the updated unit structure to the list, and finally generate the full text divided by paragraphs;
[0100] When extracting the outline of the full text divided by paragraphs, first mark the titles, and then set the paragraphs or sentences that meet the standards as the current hierarchical rule samples according to the number library. Finally, mark the hierarchical rule samples according to the sample rules to generate the outline;
[0101] When verifying the generated outline, first judge what problems exist in the outline of the previous step and process them. If there is a hierarchical error problem, re-divide the correct outline. If there is an outline missing problem, re-extract the correct outline. If there is an outline duplication problem, delete the duplicate outline. After correcting the problems in the outline, mark the text after the outline as the text content of the outline and save it.
[0102] As described above, it is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its inventive concept, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A method for extracting the outline of a Word document, characterized in that, it includes the following steps: Step 1: Import the Word document into the system; Import the Word document used as reference materials into the system through the system terminal; Step 2: Read the format information of the Word document; Receive the Word document imported into the system through the parsing module, and then extract the text, font, font size, format attributes and paragraph format of the Word document by the extraction unit; Step 3: Preprocess the Word document; Receive the Word document and structured data sent by the parsing module through the preprocessing module; The merging unit merges the font, font size and their format attributes of each paragraph to make the text format consistent; Then hand it over to the list unit, which divides according to the merged paragraphs with punctuation marks as the standard to generate the first sentence of the paragraph and the remaining text; And according to the generated first sentence of the paragraph and the remaining text, after updating the unit structured data, add the unit structure data to the list, and finally transmit the full text divided by paragraphs to the outline extraction module; Step 4: Extract the outline of the Word document; The outline extraction module receives the full text divided by paragraphs and the statistical results transmitted by the preprocessing module; The title marking unit first reads the font and font size of the first paragraph of the full text, and then checks whether the statistical value of the same font and font size combination in the statistical results is unique. If it is unique, it is marked as a title; the outline extraction module includes a title marking unit, an outline marking unit and a hierarchical rule matching unit for extracting the outline; the title marking unit reads the font and font size of the first paragraph of the document and marks the ones that meet the standards as titles; the hierarchical rule matching unit, according to the document number library, sets the paragraphs or the first sentences that meet the rules as rule samples and transmits them to the outline marking unit; after receiving the hierarchical rule samples, the outline marking unit marks the hierarchical outline according to the rules and sends the generated outline to the outline verification module; Step 5: Verify whether the extracted outline is perfect; The outline verification module receives the outline extracted by the outline extraction module, and the judgment unit judges what problems exist in the extracted outline according to the outline serial numbers in the number library, and hands it over to the corresponding unit for processing according to the type of problems; When there is a hierarchical error problem: It is transmitted to the hierarchical error correction unit for correction. The hierarchical error correction unit reads the outline, divides the possible upper-level outline area, calculates the scores of the paragraphs in the area according to the font, font size and format attributes, and marks the paragraph with the highest score as the outline. If the marking is successful, the newly generated outline is submitted to the outline generation unit; When there is an outline missing problem: It is transmitted to the outline missing correction unit for correction. The outline missing correction unit extracts the font, font size and format attributes of the current hierarchical outline, matches the paragraphs that are exactly the same within the missing range, marks them as the missing outline, and verifies the outline serial numbers again. If they are not continuously increasing, the newly generated outline is submitted to the outline generation unit; When there is an outline duplication problem: It is then transmitted to the outline duplication correction unit for correction; the outline duplication correction unit cancels the duplicate outline markers, extracts the remaining outline font, font size, and format attributes, verifies the cancelled outlines, and if they are exactly the same, submits the new outline to the outline generation unit; Outline generation unit: It is used to receive and generate the outline transmitted by the judgment unit or other units. After receiving the outline transmitted by other units, the outline generation unit marks the text following the outline as the text content belonging to the outline, and saves the generated content for the writer to refer to when establishing the outline.
2. A method for extracting the outline of a Word document according to claim 1, characterized in that, the second step further includes: the generated structured data in units of paragraphs is transmitted together with the Word document to the document preprocessing module for processing.
3. A method for extracting the outline of a Word document according to claim 1, characterized in that, the third step further includes: the statistical unit in the preprocessing module reads the font size catalog attributes in the structured data and statistically processes them respectively; groups the combinations of font sizes in units of paragraphs and counts the number of each combination, and at the same time counts the catalog format attribute values of the paragraphs, and transmits the statistical results to the outline verification module and the outline extraction module respectively.
Citation Information
Patent Citations
Method and equipment for labeling feature words in reference book
CN111274352A
Document title tree construction method and device, electronic equipment and storage medium
CN111460083A