A method for extracting complex tables based on xml format docx

By using global variables to mark the serial numbers of tables, rows and cells in the data extraction after docx is converted into xml, and by recursively traversing cell labels to extract text content, the problem of not being able to accurately associate component labels and text information in the existing technology is solved, and efficient and accurate data extraction is achieved.

CN119623440BActive Publication Date: 2025-06-20SSE INFORMATION NETWORK LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510151796.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-06-20
Estimated Expiration
2045-02-12

AI Technical Summary

Technical Problem

现有技术在处理docx转化为xml后的数据提取时,尤其是在处理包含<w:sdtPr>文本组件的表格结构时,无法准确关联组件标签与文本信息,导致数据提取的准确性和完整性受到影响。

Method used

通过获取word文档的底层xml信息,遍历xml信息中的标签,利用全局变量标记表格、行和单元格的序号,并通过递归遍历单元格标签<w:tc>,提取<w:t>标签的文本内容,建立标记标签<w:tag>与单元格标签<w:tc>的对应关系,以实现高效的数据提取。

Benefits of technology

It improves the accuracy and efficiency of data extraction, can complete the extraction of all tags and contents within O(1) time complexity, solves the problem of text component association in complex table structures, and enhances the integrity of data extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119623440B_ABST
    Figure CN119623440B_ABST
Patent Text Reader

Abstract

The present application provides a method for extracting complex tables based on xml format docx. The method includes: obtaining the underlying xml information of a word document, traversing the tags in the xml information, and for each table tag traversed, increasing the value of the first global variable tableFlag by 1; after traversing to a table tag, continue traversing, and for each row tag traversed, increasing the value of the second global variable trFlag by 1, and at the same time assigning the third global variable tcFlag to 0; for each cell tag traversed, execute tcFlag = tcFlag + 1, and record the row number and column number where the cell is located; perform independent recursive traversal on each cell tag, and whenever a <w:t>When extracting the label <w:t>The text content of the label, all within the cell <w:t>The text content of the labels is concatenated as the content of this cell. This application plays an important role in the process of templated Word writing and preparation, and solves the problem of content structuring in templated Word writing using components.< / w:t> < / w:t> < / w:t> is encountered during the traversal process <w:t>When extracting the label <w:t>The text content of the label, all within the cell <w:t>The text content of the labels is concatenated as the content of this cell. This application plays an important role in the process of templated Word writing and preparation, and solves the problem of content structuring in templated Word writing using components.< / w:t> < / w:t> < / w:t>
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and specifically relates to a method for extracting complex tables based on xml format docx. Background Art

[0002] The prior art simply relies on <w:tr>(table row) and <w:tc>The (table cell) tags are used to parse the table content, ignoring various text types that may exist in the table. In fact, the table not only contains ordinary text entered directly, but may also contain <w:sdtpr>Text of the text component type. This text component is different from ordinary <w:t>The label texts have structural differences, resulting in the inability to accurately associate component labels with their corresponding text information solely relying on row and cell labels.

[0003] When the prior art processes data extraction after converting docx to xml, especially when dealing with the inclusion of <w:sdtpr>When dealing with the table structure of text components, there is a problem that the component tags cannot be accurately associated with the text information, which limits the accuracy and integrity of data extraction and affects the subsequent further analysis and processing of the document content. Summary of the Invention

[0004] To help solve the above technical problems, the present application provides a method for extracting complex tables based on xml format docx, and adopts the following technical solutions:

[0005] A method for extracting complex tables based on xml format docx, wherein the method includes:

[0006] Step S1: Obtain the underlying xml information of the word document, traverse the tags in the xml information, and for each table tag traversed <w:tbl>, then increment the value of the first global variable tableFlag by 1. The first global variable tableFlag is used to record the table serial number;

[0007] Step S2: Traverse to the table label <w:tbl>After that, continue the traversal. Each time a line label is traversed <w:tr>, then increment the value of the second global variable trFlag by 1. The second global variable trFlag is used to mark the line number. At the same time, assign the value 0 to the third global variable tcFlag, indicating that no cell has been found in the current new line. The third global variable tcFlag is used to mark the column number;

[0008] Step S3: Traverse to the row label <w:tr>After that, continue the traversal. Every time a cell label is traversed <w:tc>, execute tcFlag = tcFlag + 1 to mark the cell, record the row number and column number where the cell is located, and at the same time assign an empty string to the fourth global variable txt, which is used to record the content of each cell in the table;

[0009] Step S4: Among the row labels that have been traversed <w:tr>In it, locate each cell label <w:tc>, and for each cell label <w:tc>Perform an independent recursive traversal. Whenever a <w:t>When extracting tags <w:t>The text content of the label, all within the cell <w:t>The text contents of the tags are concatenated as the content of this cell.

[0010] Preferably, the method further includes step S5: Locate the table tag <w:tbl>, for the table tag <w:tbl>Cell labels within <w:tc>And Marking Label <w:tag>Traverse, and record each marked label during the traversal <w:tag>and cell label <w:tc>The parallel following bit order during traversal, and establishing a marker label according to the parallel following bit order <w:tag>With cell label <w:tc>The corresponding relationship.

[0011] Preferably, the step S1 includes:

[0012] Input and decompress the Word document, read the document.xml file, and use an XML parsing library to load the xml information of the document.xml file. The xml information includes tags.

[0013] In summary, this application plays an important role in the process of templated word writing and advance, solves the problem of content structuring in templated word writing using components, and has great application significance in actual production:

[0014] 1. This application provides an effective method for parsing docx containing advanced components based on xml;

[0015] 2. Utilize the xml structure of docx <w:tag>Label and <w:t>The juxtaposed follow-up bit-order relationship of tags enables the extraction of all tag labels and contents within a time complexity of O(1). This method provides a reliable solution for processing complex docx document structures and helps improve the accuracy and efficiency of data extraction. Description of the Drawings

[0016] Figure 1 It is a schematic flowchart of an embodiment of a method for extracting complex tables from docx based on xml format in this application;

[0017] Figure 2 It is a partial schematic diagram of a docx format word file containing advanced components to be extracted;

[0018] Figure 3 is Figure 1 The partial code schematic diagram of the embodiment shown applied to Figure 2 is shown. Detailed Embodiment

[0019] The following further describes this application with reference to the drawings. The structure and principle of this application are very clear to those skilled in the art. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0020] Figure 1 It is a schematic flowchart of an embodiment of a method for extracting complex tables from docx based on xml format in this application.

[0021] Step S1: Obtain the underlying xml information of the word document, input and decompress the Word document, read the document.xml file, and use an XML parsing library to load the xml information of the document.xml file. The xml information includes tags. Traverse the tags in the xml information. For each table tag traversed <w:tbl>, then increment the value of the first global variable tableFlag by 1. The first global variable tableFlag is used to record the table sequence number.

[0022] Step S1 is the starting step for processing the underlying XML information of the Word document, aiming to obtain the XML structure of the Word document and traverse the <w:tbl>Labels are used to mark the serial numbers of tables. The following is a specific introduction to step S1:

[0023] Step S1 is used to parse the XML content of the Word document, identify and mark the serial number of each table. This is achieved by traversing the <w:tbl>It is implemented by a tag. Every time this tag is encountered, the value of the global variable tableFlag is incremented to assign a unique serial number to each table in the document. A Word document is actually a compressed package that contains multiple XML files, and document.xml is the file that contains the main content of the document. To obtain the XML content of a Word document, the Word document needs to be decompressed first, and then the document.xml file needs to be read. The above content can be implemented programmatically. For example, in Python, the zipfile library can be used to decompress the file, and libraries such as lxml or xml.etree.ElementTree can be used to parse the XML file. Furthermore, the XML parsing library is used to load the content of the document.xml file. The parsing library will build an internal representation based on the structure and tags of the XML, which is a tree structure. During the parsing process, the parsing library will identify all tags, attributes, and text content, and this information can be used for subsequent processing and analysis.

[0024] In addition, in step S1, the parsed XML document is also traversed to find all <w:tbl>Label. <w:tbl>Tags are used to mark the start of a table in the XML representation of a Word document. The traversal process can be implemented recursively or iteratively. Recursive traversal is suitable for XML documents with nested structures, while iterative traversal is suitable for documents with relatively simple structures. Whenever a <w:tbl>When encountering the label, increment the value of the global variable tableFlag by 1. tableFlag is used to record the serial number of the current table, so that each table will be assigned a unique serial number. This serial number can be used to identify and reference different tables in subsequent processing.

[0025] Step S2: Traverse to the table label <w:tbl>After that, continue the traversal. For each line label traversed <w:tr>, the value of the second global variable trFlag is incremented by 1. The second global variable trFlag is used to mark the line number. At the same time, the third global variable tcFlag is assigned the value 0, indicating that no cell has been found in the current new line. The third global variable tcFlag is used to mark the column number.

[0026] Step S2 is a process of further processing the XML content of the Word document based on the recognition and marking of the table serial number in Step S1. The main purpose is to mark the line number inside the table and initialize the column number marking for the new line. The following is a detailed introduction to Step S2:

[0027] Step S2 is used for the already recognized table (i.e., <w:tbl>Continue traversing within the label to find all line labels ( <w:tr>), and assign a unique line number to each line. At the same time, whenever a new line is encountered, initialize the column number marker to prepare for recording information about the cells in that line.

[0028] In step S1, identify the table ( <w:tbl>After the label), continue to traverse the content inside the table. In the XML representation of a Word document, the content of a table is usually contained in one or more <w:tr>Inside the label, each <w:tr>The label represents a row in the table. Find all in the table content <w:tr>tag, whenever encountering a new <w:tr>When the tag is encountered, it indicates the start of a new row in the table. Whenever a new <w:tr>When the label is encountered, increment the value of the second global variable trFlag by 1. trFlag is used to record the serial number of the current line, so that each line in the table will be assigned a unique line number, and this line number can be used to identify and reference different lines in subsequent processing. When starting to process a new line (i.e., when a new <w:tr>When the third global variable tcFlag is assigned a value of 0 (tag), tcFlag is used to record the serial number of the cell in the current row. At this time, the output cell and the row and column coordinates are output. Initializing tcFlag to 0 is to indicate that no cell has been found in the current new row yet, and it is ready to start recording the information of the cells in this row.

[0029] By assigning a unique row number to each row in the table and a unique column number to each cell in each row (although the column number is further processed in step S3), step S2 provides a structured identification for the table content. This identification helps to quickly locate and reference specific rows or cells in the table during subsequent processing. In addition, when processing a Word document, it is often necessary to extract data from the table for analysis. Step S2 provides convenience for subsequent data extraction and analysis by providing a structured identification for the table content.

[0030] Step S3: Traverse to the row tag <w:tr>After that, continue the traversal. Every time a cell label is traversed <w:tc>, execute tcFlag = tcFlag + 1 to mark the cell, record the row number and column number where the cell is located, and at the same time assign an empty string to the fourth global variable txt, which is used to record the content of each cell in the table.

[0031] Step S3 is based on Step S2 to process the table cells in the Word document ( <w:tc>The process of traversing and marking by using the label) aims to assign a unique column number to each cell in the table and record the content of the cell. The following is a detailed introduction to step S3:

[0032] Step S3 is used to identify the table rows that have been <w:tr>Based on the label), continue to traverse the table content and search for all cell labels ( <w:tc>), and assign a unique column number to each cell. At the same time, record the content of each cell for subsequent processing.

[0033] In step S2, the table rows have been identified ( <w:tr>Label) and marked the line numbers. In step S3, it is necessary to traverse the content of each line to find all cell labels ( <w:tc>), each <w:tc>The label represents a cell in the table. Whenever a new <w:tc>When the label is encountered, the value of the global variable tcFlag is incremented by 1. tcFlag is used to record the serial number of the cell in the current row, that is, the column number. In this way, each cell will be assigned a unique column number, which is unique in the current row. When traversing the cell label ( <w:tc>When ( ), record the content of each cell at the same time. The content of the cell may include various forms such as text, pictures, formulas, etc. When processing text cells, their content can be extracted and saved into corresponding data structures, such as lists, dictionaries, etc. In a Word document, a table may contain merged cells. When processing merged cells, special attention needs to be paid to the allocation of column numbers and the extraction of content. Merged cells may span multiple rows or columns, so corresponding logic is required to handle this situation during traversal and marking.

[0034] By assigning a unique column number to each cell in the table, step S3 provides a unique identifier for the table cells, which helps to quickly locate and reference specific cells in the table during subsequent processing. When processing table data in a Word document, it is often necessary to extract data from the table for analysis. Step S3 provides convenience for subsequent data extraction and analysis by providing a unique identifier for the table cells and recording their content.

[0035] Create an object to record the previous table, row number, and column number markings. Its structure is as follows:

[0036] {

[0037] String cellValue;

[0038] int tableId;

[0039] int rowNum;

[0040] int colNum;

[0041] int colCount;

[0042] }

[0043] In this embodiment, the structure can include:

[0044] String cellValue: The value of the cell.

[0045] int tableId: The serial number of the table.

[0046] int rowNum: The serial number of the row.

[0047] int colNum: The serial number of the column.

[0048] int colCount: May be used to record the total number of columns or other relevant information.

[0049] Step S4: Among the row labels that have been traversed <w:tr>Locate each cell label <w:tc>, and for each cell label <w:tc>Perform an independent recursive traversal. Whenever a <w:t>When tagging, extract <w:t>The text content of the label, all within the cell <w:t>The text contents of the labels are concatenated as the content of the cell.

[0050] Step S4 is used to process the extraction of the content of table cells in a Word document, especially for the cases of empty cells and the possible complex structures (such as multiple text components) within the cells. The following is a detailed explanation of Step S4:

[0051] In the XML representation of a Word document, a table cell is composed of <w:tc>Label representation. The text content within the cell is usually contained in <w:t>inside the label. However, when the cell is empty, <w:t>The label may not exist. Additionally, the cell may contain multiple <w:t>Label, indicating segmented or differently formatted text. The purpose of step S4 is to recursively traverse <w:tc>Label, accurately extract and splice all the text content within the cell.

[0052] In the already recognized table rows ( <w:tr>), locate each cell label ( <w:tc>). For each <w:tc>The tags are traversed independently and recursively. This means that it will go deep into <w:tc>All child nodes of the label, looking for possible <w:t>Label. During the recursive traversal, whenever a <w:t>When extracting the text content of a label. Since <w:t>Labels may contain text with paragraphs or different formats, so there may be multiple within a single cell <w:t>Label. All within the cell <w:t>The text contents of the tags are concatenated to form the final value of the cell. When concatenating, the order of the text contents is maintained to ensure the accuracy of the final value. If nothing is found during the recursive traversal <w:t>If the label is present, it is determined that the cell is empty. The value of the empty cell can be set to an empty string, a specific marker, or other processing can be performed as needed.

[0053] Step S4 ensures that all text content within the cell can be accurately extracted through recursive traversal and text concatenation, including segmented and differently formatted text, while clearly handling the case of empty cells to avoid errors or omissions in data analysis and processing. Through recursive traversal and concatenation processing, the code can handle various complex cell structures, enhancing the robustness and applicability of the code.

[0054] Step S4 is mainly to address the phenomenon of empty cells in word documents because empty cells do not appear <w:t>The label, so it is impossible to determine whether the value has been obtained. We can determine the cell label <w:tc>Must be <w:t>The superior of the label. Even if a text component or other complex structure is used in the cell, it does not affect <w:tc>As <w:t>The subordination relationship of the upper-level tags. Therefore, when performing independent recursion, if it is found that <w:t>Label, its value can be retrieved as the value of this cell. At the same time, if there are multiple <w:t>Labels, all of which belong to that cell <w:tc>, so just replace each occurrence of <w:t>The tag values are concatenated to obtain the final value of the cell.

[0055] Step S5: Locate the table tag <w:tbl>, for the table tag <w:tbl>Cell labels within <w:tc>And Marking Tags <w:tag>Traverse, and record each marked label during the traversal <w:tag>and cell label <w:tc>The parallel following order during traversal and establishing a marker label according to the parallel following order <w:tag>With cell label <w:tc>The corresponding relationship.

[0056] Step S5 is used to utilize the <w:tag>Label and <w:t>Steps for extracting tags and content based on the juxtaposed following bit-order relationship of tags: Step S5 utilizes the <w:tag>Label and <w:t>Label (or other text label, such as <w:r>under <w:t>), the parallel following sequence relationship, efficiently and accurately extract all <w:tag>The label and its content. The detailed steps are as follows:

[0057] In the word / document.xml file, locate the one containing <w:tag>Table or paragraph of labels. <w:tag>Labels are usually used to add metadata or tags to the content of a document. Traverse the XML elements in the located table or paragraph. During the traversal, pay attention to identifying <w:tag>Label and <w:t>The juxtaposed following ordinal relationship of labels (or other text labels). Since <w:tag>Label and <w:t>Labels (or other text tags) follow side by side, and the tags and their contents can be extracted by recording their ordinal relationships. Whenever a <w:tag>When the tag is encountered, check the first text tag that follows it (such as <w:t>), the content of this text label is combined with <w:tag>Associate the tags. A dictionary or list can be used to store this association relationship during the traversal process, where the key is <w:tag>The content or attribute of the label, and the value is the corresponding text content.

[0058] Since this method utilizes <w:tag>Label and <w:t>The juxtaposed following bit-order relationship of the tags can be completed for a single one within the time complexity of O(1). <w:tag>Extraction of labels and content (here O(1) means for each encountered <w:tag>For the label, the time complexity of extracting its corresponding text content is at the constant level). Overall, the time complexity of traversing the entire document depends on the size and complexity of the document, but the process of extracting the label and content itself is efficient.

[0059] By accurately identifying the bit order relationship in step S5, it can be ensured that the extracted <w:tag>The label and its content are correct. The time complexity of the extraction process is low and it is suitable for processing large-scale documents. This method can be applied to different types of <w:tag>Labels and text labels, with a certain degree of flexibility.

[0060] Figure 2 Partial schematic diagram of a docx-format word file containing high-level components to be extracted.

[0061] All components and their corresponding content can be obtained through extraction, including the components and content in the table:

[0062] [{"keyTag":"_SEC_e03362b794b84f94aa9e1cf0d75ab0a7","keyValue":"Stock Code: 601107 Stock Abbreviation: Sichuan Expressway Co., Ltd. Announcement No.: 2024-032"},{"keyTag":"_GBC_cc6fdf7dc2054e4f9e082ed74b6a5425","keyValue":"601107"},{"keyTag":"_GBC_77e8fad392474aa4be479414251ffb31","keyValue":"Sichuan Expressway Co., Ltd."},{"keyTag":"_GBC_fff01b59764149628ec7651b658cdfb6","keyValue":"2024-032"},{"keyTag":"_SEC_b41758ce170d4001bde4cc3e7b883f48","keyValue":"Sichuan Expressway Co., Ltd. Announcement on the Suspension of Trading in Connection with the Planning of Issuing Shares and Paying Cash to Purchase Assets and Related Party Transactions"},{"keyTag":"_GBC_469ed98c26544cde935109dfa7edca74","keyValue":"Sichuan Expressway Co., Ltd."},{"keyTag":"_SEC_36fd8164e5044892854b7c85141847a1","keyValue":"The board of directors of the company and all directors guarantee that the content of this announcement does not contain any false records, misleading statements or material omissions, and shall bear legal liability for the authenticity, accuracy and integrity of its content."},{"keyTag":"_SEC_bd115fdf57bb417bb207a7a3bcb3e7ae","keyValue":"Stock Suspension and Resumption Information: Applicable"},{"keyTag":"_GBC_7bc6e15c25be4af7a604115d7be3c1e0","keyValue":"Applicable"},{"keyTag":"_SEC_b999c2774ad04c7a8dbd224d5f4ce2bf","keyValue":"Due to the planning of issuing shares and paying cash to purchase assets, the suspension and resumption information of the company's relevant securities is as follows: Stock Code Stock Abbreviation Type of Suspension and Resumption Date of Suspension Start Date of Suspension Period Date of Suspension End Date of Resumption Date of Resumption 601107 Sichuan Expressway A-share Suspension 2024 / 8 / 1 Test Test"},{"keyTag":"_GBC_61b6f7d4484843e7833d3af4a0b1815a","keyValue": "Planning to issue shares and pay cash to purchase assets"}, {"keyTag": "_GBC_61b6f7d4484843e7833d3af4a0b1815a", "keyValue": "Planning to issue shares and pay cash to purchase assets"}, {"keyTag": "_PLD_91fd862c5ccb4038b8fab53000b60eab", "keyValue": "Stock Code"}, {"keyTag": "_PLD_b497811ce5314dcb881e3522117bd243", "keyValue": "Stock Abbreviation"}, {"keyTag": "_PLD_d81a043638844ba484a776d722b21bf9", "keyValue": "Type of Suspension / Resumption of Trading"}, {"keyTag": "_PLD_f4831a4233dd436abe6b7b3339f83673", "keyValue": "Start Date of Suspension"}, {"keyTag": "_PLD_b4e8cf5d3e45400ebe8f801e5d5dae30", "keyValue": "Suspension Period"}, {"keyTag": "_PLD_1cddf0503b3448558ba96b1bcf87f03e", "keyValue": "End Date of Suspension"}, {"keyTag": "_PLD_9bb8cdea07214f2996453d259535c186", "keyValue": "Resumption Date"}, {"keyTag": "_TUP_c8ea1c70c9b54a6787cb064245caf3b0", "keyValue": "601107 Sichuan Expressway A-share Suspended on August 1, 2024"}, {"keyTag": "_GBC_7d973a87f5b5431a917e529ec3d0e616", "keyValue": "601107"}, {"keyTag": "_GBC_701adc90556040dca1f1a8640d69bb50", "keyValue": "Sichuan Expressway"}, {"keyTag": "_GBC_217c8c8e470449ef8e5af2ebdb9d4d24", "keyValue": "A-share"}, {"keyTag": "_GBC_8dd9a1d861b846d7a792e9a2f570c689", "keyValue": " "},{"keyTag":"_GBC_f6393728d9ce4378a646226773f36ac6","keyValue":"Trading Halt"},{"keyTag":"_GBC_ed78751f42a54bb98cb77b2add4fc2e7","keyValue":"2024 / 8 / 1"},{"keyTag":"_GBC_324721512e9b463cab0882af78445285","keyValue":" "},{"keyTag":"_GBC_41f8881dc1884262859a4bd77e1db0ab","keyValue":" "},{"keyTag":"_GBC_8c1300bb25dc48e18ccf94d390189524","keyValue":" "},{"keyTag":"_GBC_a0dbe34339a344a896b553a3a318a794","keyValue":"Sichuan Chengyu Expressway Company Limited"},{"keyTag":"_GBC_b0649edb53524c19a256bbb6e780e07f","keyValue":"July 31, 2024"] This, Figure 2 Figure 2 shows the component property information displayed for DOCX when using development tools. By viewing the property information of the advanced text component with the content "Stock Code", it can be known that the marked value of this component is: "_PLD_91fd862c5ccb4038b8fab53000b60eab", and this information will be used as the unique primary key for xml extraction corresponding to the document.

[0063] This method plays an important role in the process of templated word writing and preparation, solves the problem of content structuring in templated word writing using components, and is widely applied in actual production.

[0064] Figure 3 For Figure 1 The shown embodiment is applied to Figure 2 A partial code schematic diagram of. Combining Figure 2 and Figure 3 It can be known that the content "Stock Code" is located in a cell of the table. According to the normal xml logic, this text component should be located at <w:tc>inside the label, but the actual situation is that this <w:sdtpr>And <w:sdtcontent>is a peer-to-peer relationship, and <w:tc>The label is located at <w:sdtcontent>In this case, there is no way to pass <w:tc>Positioning <w:tag>Content. Therefore, our method is to analyze the structure of the table <w:tag>There will definitely be one that appears after it appears <w:t>The tag. At this time, when we traverse to <w:tag w:val="_PLD_91fd862c5ccb4038b8fab53000b60eab" / > , record {"keyTag":"_PLD_91fd862c5ccb4038b8fab53000b60eab","keyValue":""}, and continue to traverse. When traversing to the nearest <w:t>Stock Code< / w:t> , supplement the previous structure to get {"keyTag":"_PLD_91fd862c5ccb4038b8fab53000b60eab","keyValue":"Security Code"}, thus completing the extraction of high-level components in a table.

[0065] This figure shows the positioning of cell content in a table using tags of high-level components in the underlying xml data of DOCX, where <w:tbl>For positioning a table, <w:tr>For positioning the line, <w:tc>For positioning the cell, it can be seen from the document structure that the component information in the cell is located at <w:tc>Outside the label belongs to a parallel relationship, and the content to be extracted is located in <w:tc>Within the label range <w:t>In the label.< / w:t> < / w:tc> < / w:tc> < / w:tc> < / w:tr> < / w:tbl> < / w:t> < / w:tag> < / w:tag> < / w:tc> < / w:sdtcontent> < / w:tc> < / w:sdtcontent> < / w:sdtpr> < / w:tc> < / w:tag> < / w:tag> < / w:tag> < / w:tag> < / w:t> < / w:tag> < / w:tag> < / w:tag> < / w:t> < / w:tag> < / w:t> < / w:tag> < / w:t> < / w:tag> < / w:tag> < / w:tag> < / w:tag> < / w:t> < / w:r> < / w:t> < / w:tag> < / w:t> < / w:tag> < / w:tc> < / w:tag> < / w:tc> < / w:tag> < / w:tag> < / w:tc> < / w:tbl> < / w:tbl> < / w:t> < / w:tc> < / w:t> < / w:t> < / w:t> < / w:tc> < / w:t> < / w:tc> < / w:t> < / w:t> < / w:t> < / w:t> < / w:t> < / w:t> < / w:t> < / w:tc> < / w:tc> < / w:tc> < / w:tr> < / w:tc> < / w:t> < / w:t> < / w:t> < / w:tc> < / w:t> < / w:t> < / w:t> is encountered during the traversal <w:t>When tagging, extract <w:t>The text content of the label, all within the cell <w:t>The text contents of the labels are concatenated as the content of the cell.

[0050] Step S4 is used to process the extraction of the content of table cells in a Word document, especially for the cases of empty cells and the possible complex structures (such as multiple text components) within the cells. The following is a detailed explanation of Step S4:

[0051] In the XML representation of a Word document, a table cell is composed of <w:tc>Label representation. The text content within the cell is usually contained in <w:t>inside the label. However, when the cell is empty, <w:t>The label may not exist. Additionally, the cell may contain multiple <w:t>Label, indicating segmented or differently formatted text. The purpose of step S4 is to recursively traverse <w:tc>Label, accurately extract and splice all the text content within the cell.

[0052] In the already recognized table rows ( <w:tr>), locate each cell label ( <w:tc>). For each <w:tc>The tags are traversed independently and recursively. This means that it will go deep into <w:tc>All child nodes of the label, looking for possible <w:t>Label. During the recursive traversal, whenever a <w:t>When extracting the text content of a label. Since <w:t>Labels may contain text with paragraphs or different formats, so there may be multiple within a single cell <w:t>Label. All within the cell <w:t>The text contents of the tags are concatenated to form the final value of the cell. When concatenating, the order of the text contents is maintained to ensure the accuracy of the final value. If nothing is found during the recursive traversal <w:t>If the label is present, it is determined that the cell is empty. The value of the empty cell can be set to an empty string, a specific marker, or other processing can be performed as needed.

[0053] Step S4 ensures that all text content within the cell can be accurately extracted through recursive traversal and text concatenation, including segmented and differently formatted text, while clearly handling the case of empty cells to avoid errors or omissions in data analysis and processing. Through recursive traversal and concatenation processing, the code can handle various complex cell structures, enhancing the robustness and applicability of the code.

[0054] Step S4 is mainly to address the phenomenon of empty cells in word documents because empty cells do not appear <w:t>The label, so it is impossible to determine whether the value has been obtained. We can determine the cell label <w:tc>Must be <w:t>The superior of the label. Even if a text component or other complex structure is used in the cell, it does not affect <w:tc>As <w:t>The subordination relationship of the upper-level tags. Therefore, when performing independent recursion, if it is found that <w:t>Label, its value can be retrieved as the value of this cell. At the same time, if there are multiple <w:t>Labels, all of which belong to that cell <w:tc>, so just replace each occurrence of <w:t>The tag values are concatenated to obtain the final value of the cell.

[0055] Step S5: Locate the table tag <w:tbl>, for the table tag <w:tbl>Cell labels within <w:tc>And Marking Tags <w:tag>Traverse, and record each marked label during the traversal <w:tag>and cell label <w:tc>The parallel following order during traversal and establishing a marker label according to the parallel following order <w:tag>With cell label <w:tc>The corresponding relationship.

[0056] Step S5 is used to utilize the <w:tag>Label and <w:t>Steps for extracting tags and content based on the juxtaposed following bit-order relationship of tags: Step S5 utilizes the <w:tag>Label and <w:t>Label (or other text label, such as <w:r>under <w:t>), the parallel following sequence relationship, efficiently and accurately extract all <w:tag>The label and its content. The detailed steps are as follows:

[0057] In the word / document.xml file, locate the one containing <w:tag>Table or paragraph of labels. <w:tag>Labels are usually used to add metadata or tags to the content of a document. Traverse the XML elements in the located table or paragraph. During the traversal, pay attention to identifying <w:tag>Label and <w:t>The juxtaposed following ordinal relationship of labels (or other text labels). Since <w:tag>Label and <w:t>Labels (or other text tags) follow side by side, and the tags and their contents can be extracted by recording their ordinal relationships. Whenever a <w:tag>When the tag is encountered, check the first text tag that follows it (such as <w:t>), the content of this text label is combined with <w:tag>Associate the tags. A dictionary or list can be used to store this association relationship during the traversal process, where the key is <w:tag>The content or attribute of the label, and the value is the corresponding text content.

[0058] Since this method utilizes <w:tag>Label and <w:t>The juxtaposed following bit-order relationship of the tags can be completed for a single one within the time complexity of O(1). <w:tag>Extraction of labels and content (here O(1) means for each encountered <w:tag>For the label, the time complexity of extracting its corresponding text content is at the constant level). Overall, the time complexity of traversing the entire document depends on the size and complexity of the document, but the process of extracting the label and content itself is efficient.

[0059] By accurately identifying the bit order relationship in step S5, it can be ensured that the extracted <w:tag>The label and its content are correct. The time complexity of the extraction process is low and it is suitable for processing large-scale documents. This method can be applied to different types of <w:tag>Labels and text labels, with a certain degree of flexibility.

[0060] Figure 2 Partial schematic diagram of a docx-format word file containing high-level components to be extracted.

[0061] All components and their corresponding content can be obtained through extraction, including the components and content in the table:

[0062] [{"keyTag":"_SEC_e03362b794b84f94aa9e1cf0d75ab0a7","keyValue":"Stock Code: 601107 Stock Abbreviation: Sichuan Expressway Co., Ltd. Announcement No.: 2024-032"},{"keyTag":"_GBC_cc6fdf7dc2054e4f9e082ed74b6a5425","keyValue":"601107"},{"keyTag":"_GBC_77e8fad392474aa4be479414251ffb31","keyValue":"Sichuan Expressway Co., Ltd."},{"keyTag":"_GBC_fff01b59764149628ec7651b658cdfb6","keyValue":"2024-032"},{"keyTag":"_SEC_b41758ce170d4001bde4cc3e7b883f48","keyValue":"Sichuan Expressway Co., Ltd. Announcement on the Suspension of Trading in Connection with the Planning of Issuing Shares and Paying Cash to Purchase Assets and Related Party Transactions"},{"keyTag":"_GBC_469ed98c26544cde935109dfa7edca74","keyValue":"Sichuan Expressway Co., Ltd."},{"keyTag":"_SEC_36fd8164e5044892854b7c85141847a1","keyValue":"The board of directors of the company and all directors guarantee that the content of this announcement does not contain any false records, misleading statements or material omissions, and shall bear legal liability for the authenticity, accuracy and integrity of its content."},{"keyTag":"_SEC_bd115fdf57bb417bb207a7a3bcb3e7ae","keyValue":"Stock Suspension and Resumption Information: Applicable"},{"keyTag":"_GBC_7bc6e15c25be4af7a604115d7be3c1e0","keyValue":"Applicable"},{"keyTag":"_SEC_b999c2774ad04c7a8dbd224d5f4ce2bf","keyValue":"Due to the planning of issuing shares and paying cash to purchase assets, the suspension and resumption information of the company's relevant securities is as follows: Stock Code Stock Abbreviation Type of Suspension and Resumption Date of Suspension Start Date of Suspension Period Date of Suspension End Date of Resumption Date of Resumption 601107 Sichuan Expressway A-share Suspension 2024 / 8 / 1 Test Test"},{"keyTag":"_GBC_61b6f7d4484843e7833d3af4a0b1815a","keyValue": "Planning to issue shares and pay cash to purchase assets"}, {"keyTag": "_GBC_61b6f7d4484843e7833d3af4a0b1815a", "keyValue": "Planning to issue shares and pay cash to purchase assets"}, {"keyTag": "_PLD_91fd862c5ccb4038b8fab53000b60eab", "keyValue": "Stock Code"}, {"keyTag": "_PLD_b497811ce5314dcb881e3522117bd243", "keyValue": "Stock Abbreviation"}, {"keyTag": "_PLD_d81a043638844ba484a776d722b21bf9", "keyValue": "Type of Suspension / Resumption of Trading"}, {"keyTag": "_PLD_f4831a4233dd436abe6b7b3339f83673", "keyValue": "Start Date of Suspension"}, {"keyTag": "_PLD_b4e8cf5d3e45400ebe8f801e5d5dae30", "keyValue": "Suspension Period"}, {"keyTag": "_PLD_1cddf0503b3448558ba96b1bcf87f03e", "keyValue": "End Date of Suspension"}, {"keyTag": "_PLD_9bb8cdea07214f2996453d259535c186", "keyValue": "Resumption Date"}, {"keyTag": "_TUP_c8ea1c70c9b54a6787cb064245caf3b0", "keyValue": "601107 Sichuan Expressway A-share Suspended on August 1, 2024"}, {"keyTag": "_GBC_7d973a87f5b5431a917e529ec3d0e616", "keyValue": "601107"}, {"keyTag": "_GBC_701adc90556040dca1f1a8640d69bb50", "keyValue": "Sichuan Expressway"}, {"keyTag": "_GBC_217c8c8e470449ef8e5af2ebdb9d4d24", "keyValue": "A-share"}, {"keyTag": "_GBC_8dd9a1d861b846d7a792e9a2f570c689", "keyValue": " "},{"keyTag":"_GBC_f6393728d9ce4378a646226773f36ac6","keyValue":"Trading Halt"},{"keyTag":"_GBC_ed78751f42a54bb98cb77b2add4fc2e7","keyValue":"2024 / 8 / 1"},{"keyTag":"_GBC_324721512e9b463cab0882af78445285","keyValue":" "},{"keyTag":"_GBC_41f8881dc1884262859a4bd77e1db0ab","keyValue":" "},{"keyTag":"_GBC_8c1300bb25dc48e18ccf94d390189524","keyValue":" "},{"keyTag":"_GBC_a0dbe34339a344a896b553a3a318a794","keyValue":"Sichuan Chengyu Expressway Company Limited"},{"keyTag":"_GBC_b0649edb53524c19a256bbb6e780e07f","keyValue":"July 31, 2024"] This, Figure 2 Figure 2 shows the component property information displayed for DOCX when using development tools. By viewing the property information of the advanced text component with the content "Stock Code", it can be known that the marked value of this component is: "_PLD_91fd862c5ccb4038b8fab53000b60eab", and this information will be used as the unique primary key for xml extraction corresponding to the document.

[0063] This method plays an important role in the process of templated word writing and preparation, solves the problem of content structuring in templated word writing using components, and is widely applied in actual production.

[0064] Figure 3 For Figure 1 The shown embodiment is applied to Figure 2 A partial code schematic diagram of. Combining Figure 2 and Figure 3 It can be known that the content "Stock Code" is located in a cell of the table. According to the normal xml logic, this text component should be located at <w:tc>inside the label, but the actual situation is that this <w:sdtpr>And <w:sdtcontent>is a peer-to-peer relationship, and <w:tc>The label is located at <w:sdtcontent>In this case, there is no way to pass <w:tc>Positioning <w:tag>Content. Therefore, our method is to analyze the structure of the table <w:tag>There will definitely be one that appears after it appears <w:t>The tag. At this time, when we traverse to <w:tag w:val="_PLD_91fd862c5ccb4038b8fab53000b60eab" / > , record {"keyTag":"_PLD_91fd862c5ccb4038b8fab53000b60eab","keyValue":""}, and continue to traverse. When traversing to the nearest <w:t>Stock Code< / w:t> , supplement the previous structure to get {"keyTag":"_PLD_91fd862c5ccb4038b8fab53000b60eab","keyValue":"Security Code"}, thus completing the extraction of high-level components in a table.

[0065] This figure shows the positioning of cell content in a table using tags of high-level components in the underlying xml data of DOCX, where <w:tbl>For positioning a table, <w:tr>For positioning the line, <w:tc>For positioning the cell, it can be seen from the document structure that the component information in the cell is located at <w:tc>Outside the label belongs to a parallel relationship, and the content to be extracted is located in <w:tc>Within the label range <w:t>In the label.< / w:t> < / w:tc> < / w:tc> < / w:tc> < / w:tr> < / w:tbl> < / w:t> < / w:tag> < / w:tag> < / w:tc> < / w:sdtcontent> < / w:tc> < / w:sdtcontent> < / w:sdtpr> < / w:tc> < / w:tag> < / w:tag> < / w:tag> < / w:tag> < / w:t> < / w:tag> < / w:tag> < / w:tag> < / w:t> < / w:tag> < / w:t> < / w:tag> < / w:t> < / w:tag> < / w:tag> < / w:tag> < / w:tag> < / w:t> < / w:r> < / w:t> < / w:tag> < / w:t> < / w:tag> < / w:tc> < / w:tag> < / w:tc> < / w:tag> < / w:tag> < / w:tc> < / w:tbl> < / w:tbl> < / w:t> < / w:tc> < / w:t> < / w:t> < / w:t> < / w:tc> < / w:t> < / w:tc> < / w:t> < / w:t> < / w:t> < / w:t> < / w:t> < / w:t> < / w:t> < / w:tc> < / w:tc> < / w:tc> < / w:tr> < / w:tc> < / w:t> < / w:t> < / w:t> < / w:tc> < / w:t> < / w:t> < / w:t> < / w:tc> < / w:tc> < / w:tr> < / w:tc> < / w:tc> < / w:tc> < / w:tc> < / w:tr> < / w:tc> < / w:tr> < / w:tc> < / w:tc> < / w:tr> < / w:tr> < / w:tr> < / w:tr> < / w:tr> < / w:tr> < / w:tr> < / w:tbl> < / w:tr> < / w:tbl> < / w:tr> < / w:tbl> < / w:tbl> < / w:tbl> < / w:tbl> < / w:tbl> < / w:tbl> < / w:tbl> < / w:t> < / w:tag> < / w:tc> < / w:tag> < / w:tc> < / w:tag> < / w:tag> < / w:tc> < / w:tbl> < / w:tbl> < / w:t> < / w:t> < / w:t> is encountered during the traversal <w:t>When extracting tags <w:t>The text content of the label, all within the cell <w:t>The text contents of the tags are concatenated as the content of this cell.

[0010] Preferably, the method further includes step S5: Locate the table tag <w:tbl>, for the table tag <w:tbl>Cell labels within <w:tc>And Marking Label <w:tag>Traverse, and record each marked label during the traversal <w:tag>and cell label <w:tc>The parallel following bit order during traversal, and establishing a marker label according to the parallel following bit order <w:tag>With cell label <w:tc>The corresponding relationship.

[0011] Preferably, the step S1 includes:

[0012] Input and decompress the Word document, read the document.xml file, and use an XML parsing library to load the xml information of the document.xml file. The xml information includes tags.

[0013] In summary, this application plays an important role in the process of templated word writing and advance, solves the problem of content structuring in templated word writing using components, and has great application significance in actual production:

[0014] 1. This application provides an effective method for parsing docx containing advanced components based on xml;

[0015] 2. Utilize the xml structure of docx <w:tag>Label and <w:t>The juxtaposed follow-up bit-order relationship of tags enables the extraction of all tag labels and contents within a time complexity of O(1). This method provides a reliable solution for processing complex docx document structures and helps improve the accuracy and efficiency of data extraction. Description of the Drawings

[0016] Figure 1 It is a schematic flowchart of an embodiment of a method for extracting complex tables from docx based on xml format in this application;

[0017] Figure 2 It is a partial schematic diagram of a docx format word file containing advanced components to be extracted;

[0018] Figure 3 is Figure 1 The partial code schematic diagram of the embodiment shown applied to Figure 2 is shown. Detailed Embodiment

[0019] The following further describes this application with reference to the drawings. The structure and principle of this application are very clear to those skilled in the art. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0020] Figure 1 It is a schematic flowchart of an embodiment of a method for extracting complex tables from docx based on xml format in this application.

[0021] Step S1: Obtain the underlying xml information of the word document, input and decompress the Word document, read the document.xml file, and use an XML parsing library to load the xml information of the document.xml file. The xml information includes tags. Traverse the tags in the xml information. For each table tag traversed <w:tbl>, then increment the value of the first global variable tableFlag by 1. The first global variable tableFlag is used to record the table sequence number.

[0022] Step S1 is the starting step for processing the underlying XML information of the Word document, aiming to obtain the XML structure of the Word document and traverse the <w:tbl>Labels are used to mark the serial numbers of tables. The following is a specific introduction to step S1:

[0023] Step S1 is used to parse the XML content of the Word document, identify and mark the serial number of each table. This is achieved by traversing the <w:tbl>It is implemented by a tag. Every time this tag is encountered, the value of the global variable tableFlag is incremented to assign a unique serial number to each table in the document. A Word document is actually a compressed package that contains multiple XML files, and document.xml is the file that contains the main content of the document. To obtain the XML content of a Word document, the Word document needs to be decompressed first, and then the document.xml file needs to be read. The above content can be implemented programmatically. For example, in Python, the zipfile library can be used to decompress the file, and libraries such as lxml or xml.etree.ElementTree can be used to parse the XML file. Furthermore, the XML parsing library is used to load the content of the document.xml file. The parsing library will build an internal representation based on the structure and tags of the XML, which is a tree structure. During the parsing process, the parsing library will identify all tags, attributes, and text content, and this information can be used for subsequent processing and analysis.

[0024] In addition, in step S1, the parsed XML document is also traversed to find all <w:tbl>Label. <w:tbl>Tags are used to mark the start of a table in the XML representation of a Word document. The traversal process can be implemented recursively or iteratively. Recursive traversal is suitable for XML documents with nested structures, while iterative traversal is suitable for documents with relatively simple structures. Whenever a <w:tbl>When encountering the label, increment the value of the global variable tableFlag by 1. tableFlag is used to record the serial number of the current table, so that each table will be assigned a unique serial number. This serial number can be used to identify and reference different tables in subsequent processing.

[0025] Step S2: Traverse to the table label <w:tbl>After that, continue the traversal. For each line label traversed <w:tr>, the value of the second global variable trFlag is incremented by 1. The second global variable trFlag is used to mark the line number. At the same time, the third global variable tcFlag is assigned the value 0, indicating that no cell has been found in the current new line. The third global variable tcFlag is used to mark the column number.

[0026] Step S2 is a process of further processing the XML content of the Word document based on the recognition and marking of the table serial number in Step S1. The main purpose is to mark the line number inside the table and initialize the column number marking for the new line. The following is a detailed introduction to Step S2:

[0027] Step S2 is used for the already recognized table (i.e., <w:tbl>Continue traversing within the label to find all line labels ( <w:tr>), and assign a unique line number to each line. At the same time, whenever a new line is encountered, initialize the column number marker to prepare for recording information about the cells in that line.

[0028] In step S1, identify the table ( <w:tbl>After the label), continue to traverse the content inside the table. In the XML representation of a Word document, the content of a table is usually contained in one or more <w:tr>Inside the label, each <w:tr>The label represents a row in the table. Find all in the table content <w:tr>tag, whenever encountering a new <w:tr>When the tag is encountered, it indicates the start of a new row in the table. Whenever a new <w:tr>When the label is encountered, increment the value of the second global variable trFlag by 1. trFlag is used to record the serial number of the current line, so that each line in the table will be assigned a unique line number, and this line number can be used to identify and reference different lines in subsequent processing. When starting to process a new line (i.e., when a new <w:tr>When the third global variable tcFlag is assigned a value of 0 (tag), tcFlag is used to record the serial number of the cell in the current row. At this time, the output cell and the row and column coordinates are output. Initializing tcFlag to 0 is to indicate that no cell has been found in the current new row yet, and it is ready to start recording the information of the cells in this row.

[0029] By assigning a unique row number to each row in the table and a unique column number to each cell in each row (although the column number is further processed in step S3), step S2 provides a structured identification for the table content. This identification helps to quickly locate and reference specific rows or cells in the table during subsequent processing. In addition, when processing a Word document, it is often necessary to extract data from the table for analysis. Step S2 provides convenience for subsequent data extraction and analysis by providing a structured identification for the table content.

[0030] Step S3: Traverse to the row tag <w:tr>After that, continue the traversal. Every time a cell label is traversed <w:tc>, execute tcFlag = tcFlag + 1 to mark the cell, record the row number and column number where the cell is located, and at the same time assign an empty string to the fourth global variable txt, which is used to record the content of each cell in the table.

[0031] Step S3 is based on Step S2 to process the table cells in the Word document ( <w:tc>The process of traversing and marking by using the label) aims to assign a unique column number to each cell in the table and record the content of the cell. The following is a detailed introduction to step S3:

[0032] Step S3 is used to identify the table rows that have been <w:tr>Based on the label), continue to traverse the table content and search for all cell labels ( <w:tc>), and assign a unique column number to each cell. At the same time, record the content of each cell for subsequent processing.

[0033] In step S2, the table rows have been identified ( <w:tr>Label) and marked the line numbers. In step S3, it is necessary to traverse the content of each line to find all cell labels ( <w:tc>), each <w:tc>The label represents a cell in the table. Whenever a new <w:tc>When the label is encountered, the value of the global variable tcFlag is incremented by 1. tcFlag is used to record the serial number of the cell in the current row, that is, the column number. In this way, each cell will be assigned a unique column number, which is unique in the current row. When traversing the cell label ( <w:tc>When ( ), record the content of each cell at the same time. The content of the cell may include various forms such as text, pictures, formulas, etc. When processing text cells, their content can be extracted and saved into corresponding data structures, such as lists, dictionaries, etc. In a Word document, a table may contain merged cells. When processing merged cells, special attention needs to be paid to the allocation of column numbers and the extraction of content. Merged cells may span multiple rows or columns, so corresponding logic is required to handle this situation during traversal and marking.

[0034] By assigning a unique column number to each cell in the table, step S3 provides a unique identifier for the table cells, which helps to quickly locate and reference specific cells in the table during subsequent processing. When processing table data in a Word document, it is often necessary to extract data from the table for analysis. Step S3 provides convenience for subsequent data extraction and analysis by providing a unique identifier for the table cells and recording their content.

[0035] Create an object to record the previous table, row number, and column number markings. Its structure is as follows:

[0036] {

[0037] String cellValue;

[0038] int tableId;

[0039] int rowNum;

[0040] int colNum;

[0041] int colCount;

[0042] }

[0043] In this embodiment, the structure can include:

[0044] String cellValue: The value of the cell.

[0045] int tableId: The serial number of the table.

[0046] int rowNum: The serial number of the row.

[0047] int colNum: The serial number of the column.

[0048] int colCount: May be used to record the total number of columns or other relevant information.

[0049] Step S4: Among the row labels that have been traversed <w:tr>Locate each cell label <w:tc>, and for each cell label <w:tc>Perform an independent recursive traversal. Whenever a <w:t>When tagging, extract <w:t>The text content of the label, all within the cell <w:t>The text contents of the labels are concatenated as the content of the cell.

[0050] Step S4 is used to process the extraction of the content of table cells in a Word document, especially for the cases of empty cells and the possible complex structures (such as multiple text components) within the cells. The following is a detailed explanation of Step S4:

[0051] In the XML representation of a Word document, a table cell is composed of <w:tc>Label representation. The text content within the cell is usually contained in <w:t>inside the label. However, when the cell is empty, <w:t>The label may not exist. Additionally, the cell may contain multiple <w:t>Label, indicating segmented or differently formatted text. The purpose of step S4 is to recursively traverse <w:tc>Label, accurately extract and splice all the text content within the cell.

[0052] In the already recognized table rows ( <w:tr>), locate each cell label ( <w:tc>). For each <w:tc>The tags are traversed independently and recursively. This means that it will go deep into <w:tc>All child nodes of the label, looking for possible <w:t>Label. During the recursive traversal, whenever a <w:t>When extracting the text content of a label. Since <w:t>Labels may contain text with paragraphs or different formats, so there may be multiple within a single cell <w:t>Label. All within the cell <w:t>The text contents of the tags are concatenated to form the final value of the cell. When concatenating, the order of the text contents is maintained to ensure the accuracy of the final value. If nothing is found during the recursive traversal <w:t>If the label is present, it is determined that the cell is empty. The value of the empty cell can be set to an empty string, a specific marker, or other processing can be performed as needed.

[0053] Step S4 ensures that all text content within the cell can be accurately extracted through recursive traversal and text concatenation, including segmented and differently formatted text, while clearly handling the case of empty cells to avoid errors or omissions in data analysis and processing. Through recursive traversal and concatenation processing, the code can handle various complex cell structures, enhancing the robustness and applicability of the code.

[0054] Step S4 is mainly to address the phenomenon of empty cells in word documents because empty cells do not appear <w:t>The label, so it is impossible to determine whether the value has been obtained. We can determine the cell label <w:tc>Must be <w:t>The superior of the label. Even if a text component or other complex structure is used in the cell, it does not affect <w:tc>As <w:t>The subordination relationship of the upper-level tags. Therefore, when performing independent recursion, if it is found that <w:t>Label, its value can be retrieved as the value of this cell. At the same time, if there are multiple <w:t>Labels, all of which belong to that cell <w:tc>, so just replace each occurrence of <w:t>The tag values are concatenated to obtain the final value of the cell.

[0055] Step S5: Locate the table tag <w:tbl>, for the table tag <w:tbl>Cell labels within <w:tc>And Marking Tags <w:tag>Traverse, and record each marked label during the traversal <w:tag>and cell label <w:tc>The parallel following order during traversal and establishing a marker label according to the parallel following order <w:tag>With cell label <w:tc>The corresponding relationship.

[0056] Step S5 is used to utilize the <w:tag>Label and <w:t>Steps for extracting tags and content based on the juxtaposed following bit-order relationship of tags: Step S5 utilizes the <w:tag>Label and <w:t>Label (or other text label, such as <w:r>under <w:t>), the parallel following sequence relationship, efficiently and accurately extract all <w:tag>The label and its content. The detailed steps are as follows:

[0057] In the word / document.xml file, locate the one containing <w:tag>Table or paragraph of labels. <w:tag>Labels are usually used to add metadata or tags to the content of a document. Traverse the XML elements in the located table or paragraph. During the traversal, pay attention to identifying <w:tag>Label and <w:t>The juxtaposed following ordinal relationship of labels (or other text labels). Since <w:tag>Label and <w:t>Labels (or other text tags) follow side by side, and the tags and their contents can be extracted by recording their ordinal relationships. Whenever a <w:tag>When the tag is encountered, check the first text tag that follows it (such as <w:t>), the content of this text label is combined with <w:tag>Associate the tags. A dictionary or list can be used to store this association relationship during the traversal process, where the key is <w:tag>The content or attribute of the label, and the value is the corresponding text content.

[0058] Since this method utilizes <w:tag>Label and <w:t>The juxtaposed following bit-order relationship of the tags can be completed for a single one within the time complexity of O(1). <w:tag>Extraction of labels and content (here O(1) means for each encountered <w:tag>For the label, the time complexity of extracting its corresponding text content is at the constant level). Overall, the time complexity of traversing the entire document depends on the size and complexity of the document, but the process of extracting the label and content itself is efficient.

[0059] By accurately identifying the bit order relationship in step S5, it can be ensured that the extracted <w:tag>The label and its content are correct. The time complexity of the extraction process is low and it is suitable for processing large-scale documents. This method can be applied to different types of <w:tag>Labels and text labels, with a certain degree of flexibility.

[0060] Figure 2 Partial schematic diagram of a docx-format word file containing high-level components to be extracted.

[0061] All components and their corresponding content can be obtained through extraction, including the components and content in the table:

[0062] [{"keyTag":"_SEC_e03362b794b84f94aa9e1cf0d75ab0a7","keyValue":"Stock Code: 601107 Stock Abbreviation: Sichuan Expressway Co., Ltd. Announcement No.: 2024-032"},{"keyTag":"_GBC_cc6fdf7dc2054e4f9e082ed74b6a5425","keyValue":"601107"},{"keyTag":"_GBC_77e8fad392474aa4be479414251ffb31","keyValue":"Sichuan Expressway Co., Ltd."},{"keyTag":"_GBC_fff01b59764149628ec7651b658cdfb6","keyValue":"2024-032"},{"keyTag":"_SEC_b41758ce170d4001bde4cc3e7b883f48","keyValue":"Sichuan Expressway Co., Ltd. Announcement on the Suspension of Trading in Connection with the Planning of Issuing Shares and Paying Cash to Purchase Assets and Related Party Transactions"},{"keyTag":"_GBC_469ed98c26544cde935109dfa7edca74","keyValue":"Sichuan Expressway Co., Ltd."},{"keyTag":"_SEC_36fd8164e5044892854b7c85141847a1","keyValue":"The board of directors of the company and all directors guarantee that the content of this announcement does not contain any false records, misleading statements or material omissions, and shall bear legal liability for the authenticity, accuracy and integrity of its content."},{"keyTag":"_SEC_bd115fdf57bb417bb207a7a3bcb3e7ae","keyValue":"Stock Suspension and Resumption Information: Applicable"},{"keyTag":"_GBC_7bc6e15c25be4af7a604115d7be3c1e0","keyValue":"Applicable"},{"keyTag":"_SEC_b999c2774ad04c7a8dbd224d5f4ce2bf","keyValue":"Due to the planning of issuing shares and paying cash to purchase assets, the suspension and resumption information of the company's relevant securities is as follows: Stock Code Stock Abbreviation Type of Suspension and Resumption Date of Suspension Start Date of Suspension Period Date of Suspension End Date of Resumption Date of Resumption 601107 Sichuan Expressway A-share Suspension 2024 / 8 / 1 Test Test"},{"keyTag":"_GBC_61b6f7d4484843e7833d3af4a0b1815a","keyValue": "Planning to issue shares and pay cash to purchase assets"}, {"keyTag": "_GBC_61b6f7d4484843e7833d3af4a0b1815a", "keyValue": "Planning to issue shares and pay cash to purchase assets"}, {"keyTag": "_PLD_91fd862c5ccb4038b8fab53000b60eab", "keyValue": "Stock Code"}, {"keyTag": "_PLD_b497811ce5314dcb881e3522117bd243", "keyValue": "Stock Abbreviation"}, {"keyTag": "_PLD_d81a043638844ba484a776d722b21bf9", "keyValue": "Type of Suspension / Resumption of Trading"}, {"keyTag": "_PLD_f4831a4233dd436abe6b7b3339f83673", "keyValue": "Start Date of Suspension"}, {"keyTag": "_PLD_b4e8cf5d3e45400ebe8f801e5d5dae30", "keyValue": "Suspension Period"}, {"keyTag": "_PLD_1cddf0503b3448558ba96b1bcf87f03e", "keyValue": "End Date of Suspension"}, {"keyTag": "_PLD_9bb8cdea07214f2996453d259535c186", "keyValue": "Resumption Date"}, {"keyTag": "_TUP_c8ea1c70c9b54a6787cb064245caf3b0", "keyValue": "601107 Sichuan Expressway A-share Suspended on August 1, 2024"}, {"keyTag": "_GBC_7d973a87f5b5431a917e529ec3d0e616", "keyValue": "601107"}, {"keyTag": "_GBC_701adc90556040dca1f1a8640d69bb50", "keyValue": "Sichuan Expressway"}, {"keyTag": "_GBC_217c8c8e470449ef8e5af2ebdb9d4d24", "keyValue": "A-share"}, {"keyTag": "_GBC_8dd9a1d861b846d7a792e9a2f570c689", "keyValue": " "},{"keyTag":"_GBC_f6393728d9ce4378a646226773f36ac6","keyValue":"Trading Halt"},{"keyTag":"_GBC_ed78751f42a54bb98cb77b2add4fc2e7","keyValue":"2024 / 8 / 1"},{"keyTag":"_GBC_324721512e9b463cab0882af78445285","keyValue":" "},{"keyTag":"_GBC_41f8881dc1884262859a4bd77e1db0ab","keyValue":" "},{"keyTag":"_GBC_8c1300bb25dc48e18ccf94d390189524","keyValue":" "},{"keyTag":"_GBC_a0dbe34339a344a896b553a3a318a794","keyValue":"Sichuan Chengyu Expressway Company Limited"},{"keyTag":"_GBC_b0649edb53524c19a256bbb6e780e07f","keyValue":"July 31, 2024"] This, Figure 2 Figure 2 shows the component property information displayed for DOCX when using development tools. By viewing the property information of the advanced text component with the content "Stock Code", it can be known that the marked value of this component is: "_PLD_91fd862c5ccb4038b8fab53000b60eab", and this information will be used as the unique primary key for xml extraction corresponding to the document.

[0063] This method plays an important role in the process of templated word writing and preparation, solves the problem of content structuring in templated word writing using components, and is widely applied in actual production.

[0064] Figure 3 For Figure 1 The shown embodiment is applied to Figure 2 A partial code schematic diagram of. Combining Figure 2 and Figure 3 It can be known that the content "Stock Code" is located in a cell of the table. According to the normal xml logic, this text component should be located at <w:tc>inside the label, but the actual situation is that this <w:sdtpr>And <w:sdtcontent>is a peer-to-peer relationship, and <w:tc>The label is located at <w:sdtcontent>In this case, there is no way to pass <w:tc>Positioning <w:tag>Content. Therefore, our method is to analyze the structure of the table <w:tag>There will definitely be one that appears after it appears <w:t>The tag. At this time, when we traverse to <w:tag w:val="_PLD_91fd862c5ccb4038b8fab53000b60eab" / > , record {"keyTag":"_PLD_91fd862c5ccb4038b8fab53000b60eab","keyValue":""}, and continue to traverse. When traversing to the nearest <w:t>Stock Code< / w:t> , supplement the previous structure to get {"keyTag":"_PLD_91fd862c5ccb4038b8fab53000b60eab","keyValue":"Security Code"}, thus completing the extraction of high-level components in a table.

[0065] This figure shows the positioning of cell content in a table using tags of high-level components in the underlying xml data of DOCX, where <w:tbl>For positioning a table, <w:tr>For positioning the line, <w:tc>For positioning the cell, it can be seen from the document structure that the component information in the cell is located at <w:tc>Outside the label belongs to a parallel relationship, and the content to be extracted is located in <w:tc>Within the label range <w:t>In the label.< / w:t> < / w:tc> < / w:tc> < / w:tc> < / w:tr> < / w:tbl> < / w:t> < / w:tag> < / w:tag> < / w:tc> < / w:sdtcontent> < / w:tc> < / w:sdtcontent> < / w:sdtpr> < / w:tc> < / w:tag> < / w:tag> < / w:tag> < / w:tag> < / w:t> < / w:tag> < / w:tag> < / w:tag> < / w:t> < / w:tag> < / w:t> < / w:tag> < / w:t> < / w:tag> < / w:tag> < / w:tag> < / w:tag> < / w:t> < / w:r> < / w:t> < / w:tag> < / w:t> < / w:tag> < / w:tc> < / w:tag> < / w:tc> < / w:tag> < / w:tag> < / w:tc> < / w:tbl> < / w:tbl> < / w:t> < / w:tc> < / w:t> < / w:t> < / w:t> < / w:tc> < / w:t> < / w:tc> < / w:t> < / w:t> < / w:t> < / w:t> < / w:t> < / w:t> < / w:t> < / w:tc> < / w:tc> < / w:tc> < / w:tr> < / w:tc> < / w:t> < / w:t> < / w:t> < / w:tc> < / w:t> < / w:t> < / w:t> is encountered during the traversal <w:t>When tagging, extract <w:t>The text content of the label, all within the cell <w:t>The text contents of the labels are concatenated as the content of the cell.

[0050] Step S4 is used to process the extraction of the content of table cells in a Word document, especially for the cases of empty cells and the possible complex structures (such as multiple text components) within the cells. The following is a detailed explanation of Step S4:

[0051] In the XML representation of a Word document, a table cell is composed of <w:tc>Label representation. The text content within the cell is usually contained in <w:t>inside the label. However, when the cell is empty, <w:t>The label may not exist. Additionally, the cell may contain multiple <w:t>Label, indicating segmented or differently formatted text. The purpose of step S4 is to recursively traverse <w:tc>Label, accurately extract and splice all the text content within the cell.

[0052] In the already recognized table rows ( <w:tr>), locate each cell label ( <w:tc>). For each <w:tc>The tags are traversed independently and recursively. This means that it will go deep into <w:tc>All child nodes of the label, looking for possible <w:t>Label. During the recursive traversal, whenever a <w:t>When extracting the text content of a label. Since <w:t>Labels may contain text with paragraphs or different formats, so there may be multiple within a single cell <w:t>Label. All within the cell <w:t>The text contents of the tags are concatenated to form the final value of the cell. When concatenating, the order of the text contents is maintained to ensure the accuracy of the final value. If nothing is found during the recursive traversal <w:t>If the label is present, it is determined that the cell is empty. The value of the empty cell can be set to an empty string, a specific marker, or other processing can be performed as needed.

[0053] Step S4 ensures that all text content within the cell can be accurately extracted through recursive traversal and text concatenation, including segmented and differently formatted text, while clearly handling the case of empty cells to avoid errors or omissions in data analysis and processing. Through recursive traversal and concatenation processing, the code can handle various complex cell structures, enhancing the robustness and applicability of the code.

[0054] Step S4 is mainly to address the phenomenon of empty cells in word documents because empty cells do not appear <w:t>The label, so it is impossible to determine whether the value has been obtained. We can determine the cell label <w:tc>Must be <w:t>The superior of the label. Even if a text component or other complex structure is used in the cell, it does not affect <w:tc>As <w:t>The subordination relationship of the upper-level tags. Therefore, when performing independent recursion, if it is found that <w:t>Label, its value can be retrieved as the value of this cell. At the same time, if there are multiple <w:t>Labels, all of which belong to that cell <w:tc>, so just replace each occurrence of <w:t>The tag values are concatenated to obtain the final value of the cell.

[0055] Step S5: Locate the table tag <w:tbl>, for the table tag <w:tbl>Cell labels within <w:tc>And Marking Tags <w:tag>Traverse, and record each marked label during the traversal <w:tag>and cell label <w:tc>The parallel following order during traversal and establishing a marker label according to the parallel following order <w:tag>With cell label <w:tc>The corresponding relationship.

[0056] Step S5 is used to utilize the <w:tag>Label and <w:t>Steps for extracting tags and content based on the juxtaposed following bit-order relationship of tags: Step S5 utilizes the <w:tag>Label and <w:t>Label (or other text label, such as <w:r>under <w:t>), the parallel following sequence relationship, efficiently and accurately extract all <w:tag>The label and its content. The detailed steps are as follows:

[0057] In the word / document.xml file, locate the one containing <w:tag>Table or paragraph of labels. <w:tag>Labels are usually used to add metadata or tags to the content of a document. Traverse the XML elements in the located table or paragraph. During the traversal, pay attention to identifying <w:tag>Label and <w:t>The juxtaposed following ordinal relationship of labels (or other text labels). Since <w:tag>Label and <w:t>Labels (or other text tags) follow side by side, and the tags and their contents can be extracted by recording their ordinal relationships. Whenever a <w:tag>When the tag is encountered, check the first text tag that follows it (such as <w:t>), the content of this text label is combined with <w:tag>Associate the tags. A dictionary or list can be used to store this association relationship during the traversal process, where the key is <w:tag>The content or attribute of the label, and the value is the corresponding text content.

[0058] Since this method utilizes <w:tag>Label and <w:t>The juxtaposed following bit-order relationship of the tags can be completed for a single one within the time complexity of O(1). <w:tag>Extraction of labels and content (here O(1) means for each encountered <w:tag>For the label, the time complexity of extracting its corresponding text content is at the constant level). Overall, the time complexity of traversing the entire document depends on the size and complexity of the document, but the process of extracting the label and content itself is efficient.

[0059] By accurately identifying the bit order relationship in step S5, it can be ensured that the extracted <w:tag>The label and its content are correct. The time complexity of the extraction process is low and it is suitable for processing large-scale documents. This method can be applied to different types of <w:tag>Labels and text labels, with a certain degree of flexibility.

[0060] Figure 2 Partial schematic diagram of a docx-format word file containing high-level components to be extracted.

[0061] All components and their corresponding content can be obtained through extraction, including the components and content in the table:

[0062] [{"keyTag":"_SEC_e03362b794b84f94aa9e1cf0d75ab0a7","keyValue":"Stock Code: 601107 Stock Abbreviation: Sichuan Expressway Co., Ltd. Announcement No.: 2024-032"},{"keyTag":"_GBC_cc6fdf7dc2054e4f9e082ed74b6a5425","keyValue":"601107"},{"keyTag":"_GBC_77e8fad392474aa4be479414251ffb31","keyValue":"Sichuan Expressway Co., Ltd."},{"keyTag":"_GBC_fff01b59764149628ec7651b658cdfb6","keyValue":"2024-032"},{"keyTag":"_SEC_b41758ce170d4001bde4cc3e7b883f48","keyValue":"Sichuan Expressway Co., Ltd. Announcement on the Suspension of Trading in Connection with the Planning of Issuing Shares and Paying Cash to Purchase Assets and Related Party Transactions"},{"keyTag":"_GBC_469ed98c26544cde935109dfa7edca74","keyValue":"Sichuan Expressway Co., Ltd."},{"keyTag":"_SEC_36fd8164e5044892854b7c85141847a1","keyValue":"The board of directors of the company and all directors guarantee that the content of this announcement does not contain any false records, misleading statements or material omissions, and shall bear legal liability for the authenticity, accuracy and integrity of its content."},{"keyTag":"_SEC_bd115fdf57bb417bb207a7a3bcb3e7ae","keyValue":"Stock Suspension and Resumption Information: Applicable"},{"keyTag":"_GBC_7bc6e15c25be4af7a604115d7be3c1e0","keyValue":"Applicable"},{"keyTag":"_SEC_b999c2774ad04c7a8dbd224d5f4ce2bf","keyValue":"Due to the planning of issuing shares and paying cash to purchase assets, the suspension and resumption information of the company's relevant securities is as follows: Stock Code Stock Abbreviation Type of Suspension and Resumption Date of Suspension Start Date of Suspension Period Date of Suspension End Date of Resumption Date of Resumption 601107 Sichuan Expressway A-share Suspension 2024 / 8 / 1 Test Test"},{"keyTag":"_GBC_61b6f7d4484843e7833d3af4a0b1815a","keyValue": "Planning to issue shares and pay cash to purchase assets"}, {"keyTag": "_GBC_61b6f7d4484843e7833d3af4a0b1815a", "keyValue": "Planning to issue shares and pay cash to purchase assets"}, {"keyTag": "_PLD_91fd862c5ccb4038b8fab53000b60eab", "keyValue": "Stock Code"}, {"keyTag": "_PLD_b497811ce5314dcb881e3522117bd243", "keyValue": "Stock Abbreviation"}, {"keyTag": "_PLD_d81a043638844ba484a776d722b21bf9", "keyValue": "Type of Suspension / Resumption of Trading"}, {"keyTag": "_PLD_f4831a4233dd436abe6b7b3339f83673", "keyValue": "Start Date of Suspension"}, {"keyTag": "_PLD_b4e8cf5d3e45400ebe8f801e5d5dae30", "keyValue": "Suspension Period"}, {"keyTag": "_PLD_1cddf0503b3448558ba96b1bcf87f03e", "keyValue": "End Date of Suspension"}, {"keyTag": "_PLD_9bb8cdea07214f2996453d259535c186", "keyValue": "Resumption Date"}, {"keyTag": "_TUP_c8ea1c70c9b54a6787cb064245caf3b0", "keyValue": "601107 Sichuan Expressway A-share Suspended on August 1, 2024"}, {"keyTag": "_GBC_7d973a87f5b5431a917e529ec3d0e616", "keyValue": "601107"}, {"keyTag": "_GBC_701adc90556040dca1f1a8640d69bb50", "keyValue": "Sichuan Expressway"}, {"keyTag": "_GBC_217c8c8e470449ef8e5af2ebdb9d4d24", "keyValue": "A-share"}, {"keyTag": "_GBC_8dd9a1d861b846d7a792e9a2f570c689", "keyValue": " "},{"keyTag":"_GBC_f6393728d9ce4378a646226773f36ac6","keyValue":"Trading Halt"},{"keyTag":"_GBC_ed78751f42a54bb98cb77b2add4fc2e7","keyValue":"2024 / 8 / 1"},{"keyTag":"_GBC_324721512e9b463cab0882af78445285","keyValue":" "},{"keyTag":"_GBC_41f8881dc1884262859a4bd77e1db0ab","keyValue":" "},{"keyTag":"_GBC_8c1300bb25dc48e18ccf94d390189524","keyValue":" "},{"keyTag":"_GBC_a0dbe34339a344a896b553a3a318a794","keyValue":"Sichuan Chengyu Expressway Company Limited"},{"keyTag":"_GBC_b0649edb53524c19a256bbb6e780e07f","keyValue":"July 31, 2024"] This, Figure 2 Figure 2 shows the component property information displayed for DOCX when using development tools. By viewing the property information of the advanced text component with the content "Stock Code", it can be known that the marked value of this component is: "_PLD_91fd862c5ccb4038b8fab53000b60eab", and this information will be used as the unique primary key for xml extraction corresponding to the document.

[0063] This method plays an important role in the process of templated word writing and preparation, solves the problem of content structuring in templated word writing using components, and is widely applied in actual production.

[0064] Figure 3 For Figure 1 The shown embodiment is applied to Figure 2 A partial code schematic diagram of. Combining Figure 2 and Figure 3 It can be known that the content "Stock Code" is located in a cell of the table. According to the normal xml logic, this text component should be located at <w:tc>inside the label, but the actual situation is that this <w:sdtpr>And <w:sdtcontent>is a peer-to-peer relationship, and <w:tc>The label is located at <w:sdtcontent>In this case, there is no way to pass <w:tc>Positioning <w:tag>Content. Therefore, our method is to analyze the structure of the table <w:tag>There will definitely be one that appears after it appears <w:t>The tag. At this time, when we traverse to <w:tag w:val="_PLD_91fd862c5ccb4038b8fab53000b60eab" / > , record {"keyTag":"_PLD_91fd862c5ccb4038b8fab53000b60eab","keyValue":""}, and continue to traverse. When traversing to the nearest <w:t>Stock Code< / w:t> , supplement the previous structure to get {"keyTag":"_PLD_91fd862c5ccb4038b8fab53000b60eab","keyValue":"Security Code"}, thus completing the extraction of high-level components in a table.

[0065] This figure shows the positioning of cell content in a table using tags of high-level components in the underlying xml data of DOCX, where <w:tbl>For positioning a table, <w:tr>For positioning the line, <w:tc>For positioning the cell, it can be seen from the document structure that the component information in the cell is located at <w:tc>Outside the label belongs to a parallel relationship, and the content to be extracted is located in <w:tc>Within the label range <w:t>In the label.< / w:t> < / w:tc> < / w:tc> < / w:tc> < / w:tr> < / w:tbl> < / w:t> < / w:tag> < / w:tag> < / w:tc> < / w:sdtcontent> < / w:tc> < / w:sdtcontent> < / w:sdtpr> < / w:tc> < / w:tag> < / w:tag> < / w:tag> < / w:tag> < / w:t> < / w:tag> < / w:tag> < / w:tag> < / w:t> < / w:tag> < / w:t> < / w:tag> < / w:t> < / w:tag> < / w:tag> < / w:tag> < / w:tag> < / w:t> < / w:r> < / w:t> < / w:tag> < / w:t> < / w:tag> < / w:tc> < / w:tag> < / w:tc> < / w:tag> < / w:tag> < / w:tc> < / w:tbl> < / w:tbl> < / w:t> < / w:tc> < / w:t> < / w:t> < / w:t> < / w:tc> < / w:t> < / w:tc> < / w:t> < / w:t> < / w:t> < / w:t> < / w:t> < / w:t> < / w:t> < / w:tc> < / w:tc> < / w:tc> < / w:tr> < / w:tc> < / w:t> < / w:t> < / w:t> < / w:tc> < / w:t> < / w:t> < / w:t> < / w:tc> < / w:tc> < / w:tr> < / w:tc> < / w:tc> < / w:tc> < / w:tc> < / w:tr> < / w:tc> < / w:tr> < / w:tc> < / w:tc> < / w:tr> < / w:tr> < / w:tr> < / w:tr> < / w:tr> < / w:tr> < / w:tr> < / w:tbl> < / w:tr> < / w:tbl> < / w:tr> < / w:tbl> < / w:tbl> < / w:tbl> < / w:tbl> < / w:tbl> < / w:tbl> < / w:tbl> < / w:t> < / w:tag> < / w:tc> < / w:tag> < / w:tc> < / w:tag> < / w:tag> < / w:tc> < / w:tbl> < / w:tbl> < / w:t> < / w:t> < / w:t> < / w:tc> < / w:tc> < / w:tr> < / w:tc> < / w:tr> < / w:tr> < / w:tbl> < / w:tbl> < / w:sdtpr> < / w:t> < / w:sdtpr> < / w:tc> < / w:tr>

Claims

1. A method for extracting complex tables based on XML format docx, characterized in that: The method comprises: Step S1: Get the underlying XML information of the word document, traverse the tags in the XML information, and traverse to a table tag each time <w:tbl> , then the value of the first global variable tableFlag is increased by 1, and the first global variable tableFlag is used to record the table number;< / w:tbl> Step S2: Traverse to the table tag <w:tbl>After that, continue to traverse, each time you traverse to a row label <w:tr> , then the value of the second global variable trFlag is increased by 1, and the second global variable trFlag is used to mark the row number. At the same time, the third global variable tcFlag is assigned a value of 0, indicating that a new row has not yet found a cell, and the third global variable tcFlag is used to mark the column number;< / w:tr> < / w:tbl> Step S3: Traverse to row labels <w:tr>After that, continue to traverse, each time you traverse to a cell label <w:tc> , execute tcFlag=tcFlag+1 to mark the column number of the cell and record the row number and column number of the cell. At the same time, assign the fourth global variable txt to an empty string. The fourth global variable txt is used to record the content of each cell in the table.< / w:tc> < / w:tr> Step S4: On the row labels that have been traversed <w:tr>, locate each cell label <w:tc>, and for each cell label <w:tc>Perform independent recursive traversal, and whenever you encounter a <w:t>When labeling, extract <w:t>The text content of the label will be <w:t> The text contents of the labels are concatenated as the content of the cell.< / w:t> < / w:t> < / w:t> < / w:tc> < / w:tc> < / w:tr> 2. The method according to claim 1, characterized in that The method further comprises step S5: locating the table tag <w:tbl>, for the table tag <w:tbl>Cell labels within <w:tc>and mark tags <w:tag>Traverse and record each mark label during the traversal <w:tag>and cell labels <w:tc>The parallel follow-up sequence during the traversal process, and the mark label is established according to the parallel follow-up sequence <w:tag>With cell label <w:tc> The corresponding relationship.< / w:tc> < / w:tag> < / w:tc> < / w:tag> < / w:tag> < / w:tc> < / w:tbl> < / w:tbl> 3. The method according to claim 1 or 2, characterized in that: The step S1 comprises: Input and decompress the Word document, read the document.xml file, and use the XML parsing library to load the XML information of the document.xml file, which includes tags.

Citation Information

Patent Citations

  • Method and device for processing word document

    CN112861489A

  • Document template processing method and device, electronic equipment and storage medium

    CN115759029A