Document analysis method, device and equipment and computer readable storage medium

By determining and filling the content of table cells and splitting the table, the formatting confusion and information loss problems when merging cells in spreadsheet files and converting them into Markdown documents are solved, and the accuracy and completeness of plain text documents are achieved.

CN120706410APending Publication Date: 2025-09-26CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510895013.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

In the prior art, when converting merged cells in a spreadsheet file into a Markdown document, problems such as formatting errors or information loss may occur.

Method used

By determining the content, row number and column number of each cell in the table, non-merged cells and merged cells are filled into the specified positions in the correction table respectively, and the correction table is embedded in a plain text document, or the table is split and converted into a plain text description.

Benefits of technology

It effectively avoids formatting errors and information loss, ensures that the content in the plain text document corresponds one-to-one with the original file table, and maintains the integrity and accuracy of the data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706410A_ABST
    Figure CN120706410A_ABST
Patent Text Reader

Abstract

The invention discloses a document analysis method, device and equipment and a computer readable storage medium. The method comprises the following steps: when the type of an original file is a spreadsheet file type and a table in the original file has merged cells, determining the content of each cell in the table, the row number i and the column number j of each non-merged cell in the table and the area coordinate of each merged cell in the table; for each non-merged cell, filling the content of the non-merged cell into the cells in the ith row and the jth column in the correction table; for each merged cell, filling the content of the merged cell into a target cell corresponding to the region coordinate in the correction table; and after all the cells are traversed, embedding the obtained correction table into the plain text document. By means of the method and device, the situation that the format of the content embedded into the plain text document is disordered or information is lost when compared with that of a table in an original file is avoided to a great extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a document parsing method, apparatus, device, and computer-readable storage medium. Background Art

[0002] Intelligent question-answering systems use large language models to understand natural language questions and rely on a knowledge base to provide answers, reducing the amount of knowledge research required by the questioner. The accuracy of the answers depends not only on the natural language understanding capabilities of the large language model but also on the data quality of the knowledge base.

[0003] Currently, building a knowledge base requires parsing different types of files using document parsing technology, and then storing the results in the knowledge base. Markdown documents, due to their plain text nature, explicit expression of structured semantics, and lack of redundant style interference, are a perfect fit for the underlying logic of large language models processing natural language. Parsing results are typically stored in the knowledge base as Markdown documents.

[0004] However, there are some problems with the existing knowledge base construction method, especially when parsing spreadsheet file types. If there are merged cells in the table of the file, it is easy to have formatting errors or information loss when converting to Markdown documents, thereby reducing the data quality of the knowledge base. Summary of the Invention

[0005] The present application provides a document parsing method, apparatus, device and computer-readable storage medium, which can solve the technical problem in the prior art that tables with merged cells are prone to formatting errors or information loss when converted to Markdown documents.

[0006] In a first aspect, an embodiment of the present application provides a document parsing method, the document parsing method comprising: When the original file is a spreadsheet file and a table in the original file contains merged cells, determining the content of each cell in the table, the row number i and column number j of each non-merged cell in the table, and the region coordinates of each merged cell in the table; For each non-merged cell, fill the content of the non-merged cell into the cell at row i and column j in the correction table; For each merged cell, fill the content of the merged cell into the target cell corresponding to the region coordinates in the correction table; After all cells are traversed, the obtained correction table is embedded in a plain text document.

[0007] In combination with the first aspect, in one embodiment, the region coordinates include the row number i of the cell q covered by the merged cell in the table. q and column number j q , where q ranges from 1 to N, and N is the number of cells covered by the merged cell. Filling the contents of the merged cell into the target cell corresponding to the region coordinate in the correction table includes: Fill the contents of the merged cell into the correction table at the i q Row j q The target cell of the column.

[0008] In conjunction with the first aspect, in one embodiment, the document parsing method further includes: When the type of the original file is a spreadsheet file type and there are no merged cells in the table of the original file, determining whether a splitting condition is met; If the splitting condition is met, the table is split into several sub-tables; Convert different subtables into plain text descriptions and record them in different plain text files.

[0009] In conjunction with the first aspect, in one embodiment, splitting the table into a plurality of sub-tables includes: The table is split according to a preset number of rows to obtain a plurality of sub-tables, wherein the number of rows included in each sub-table is not greater than the preset number of rows.

[0010] In conjunction with the first aspect, in one embodiment, splitting the table into a plurality of sub-tables includes: The table is split according to a preset number of bytes to obtain a plurality of sub-tables, wherein the size of each sub-table is not greater than the preset number of bytes.

[0011] In conjunction with the first aspect, in one embodiment, the document parsing method further includes: When the type of the original file is a rich element file type, identifying valid elements in the original file and determining the original position of each valid element in the original file; When the valid element is text, write the text to a plain text document; When the valid element is a picture, the storage address and description of the picture are written into a plain text document; When the valid element is a formula, the formula is converted into a preset format and written into a plain text document; When the valid element is a non-page-spanning table, the plain text description, storage address and table description of the non-page-spanning table are written into the plain text document; When the valid element is a cross-page table, the plain text description and storage address of the cross-page table are written into the plain text document; The position where the valid element is written into the plain text document is determined based on the original position of the valid element in the original file.

[0012] In conjunction with the first aspect, in one embodiment, when the type of the original file is a rich element file type, after identifying the valid elements in the original file and determining the original position of each valid element in the original file, the method further includes: For the target valid element of the table, a serial number is assigned to the target valid element according to the original position of the target valid element in the original file, wherein the continuity of the serial number is consistent with the continuity of the original position; Check whether two target elements with adjacent serial numbers belong to the same spreadsheet; Merge all target valid elements belonging to the same spread table into a spread table.

[0013] In a second aspect, an embodiment of the present application provides a document parsing device, the document parsing device comprising: a determination module, configured to, when the original file is a spreadsheet file type and a table in the original file contains merged cells, determine the content of each cell in the table, the row number i and column number j of each non-merged cell in the table, and the region coordinates of each merged cell in the table; A filling module is used to fill the content of each non-merged cell into the cell located in the i-th row and j-th column of the correction table; and for each merged cell, fill the content of the merged cell into the target cell corresponding to the region coordinates in the correction table; The embedding module is used to embed the obtained correction table into a plain text document after all cells are traversed.

[0014] In a third aspect, an embodiment of the present application provides a document parsing device, which includes a processor, a memory, and a document parsing program stored on the memory and executable by the processor, wherein when the document parsing program is executed by the processor, the steps of the document parsing method described in the first aspect are implemented.

[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a document parsing program is stored, wherein when the document parsing program is executed by a processor, the steps of the document parsing method described in the first aspect are implemented.

[0016] The beneficial effects of the technical solutions provided in the embodiments of the present application include: In an embodiment of the present application, when the type of the original file is a spreadsheet file type and there are merged cells in the table of the original file, the content of each cell in the table, the row number i and column number j of each non-merged cell in the table, and the area coordinates of each merged cell in the table are determined; for each non-merged cell, the content of the non-merged cell is filled into the cell located in the i-th row and j-th column in the correction table; for each merged cell, the content of the merged cell is filled into the target cell corresponding to the area coordinates in the correction table; after all cells are traversed, the obtained correction table is embedded in a plain text document. Through the embodiment of the present application, the table with merged cells is optimized into a correction table with a one-to-one correspondence between the header row and any record row, and then embedded in a plain text document, which greatly avoids the situation where the content embedded in the plain text document is formatted incorrectly or information is lost compared to the table in the original file. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is a flowchart of the first embodiment of the document parsing method of this application; Figure 2 This is a diagram of a table with merged cells; Figure 3 A schematic diagram of the calibration table; Figure 4 This is a flowchart of the second embodiment of the document parsing method of this application; Figure 5 This is a flowchart of the third embodiment of the document parsing method of this application; Figure 6 Schematic diagram of the process of merging tables across pages; Figure 7 This is a functional module diagram of an embodiment of a document parsing device of the present application; Figure 8 This is a schematic diagram of the hardware structure of the document parsing device involved in the embodiment of the present application. DETAILED DESCRIPTION

[0018] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0019] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0020] In a first aspect, an embodiment of the present application provides a document parsing method.

[0021] In one embodiment, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the document parsing method of this application. Figure 1 As shown, the document parsing method includes: Step S10, when the original file is a spreadsheet file type and a table in the original file contains merged cells, determining the content of each cell in the table, the row number i and column number j of each non-merged cell in the table, and the region coordinates of each merged cell in the table; In this embodiment, after receiving an original file, the document parsing method determines the type of the original file based on its file extension. For example, if the original file extension is any of XLS, XLSX, and CSV, the original file type is determined to be a spreadsheet file. XML parsing techniques are then used to extract the table's structured attribute information, thereby determining whether the table in the original file contains merged cells based on this structured attribute information.

[0022] Reference Figure 2 , Figure 2 This is a diagram of a table with merged cells. Figure 2 A~H in the table represent columns 1~8 respectively. Figure 2 As shown, when there are merged cells, the content of each cell in the table is determined, that is, the content of each non-merged cell and the content of each merged cell are determined. Figure 2 As described above, non-merged cells include cells located in the 1st row and 1st column, cells in the 2nd row and 3rd column, etc., which are not listed here exhaustively; merged cells include cells that record "power system", "electric power system", "acceleration weakness", "MAF waveform detection" and other contents, which are not listed here exhaustively.

[0023] The row number i and column number j of each non-merged cell in the table are also determined. For example, the cell recording "faulty system" is a non-merged cell, and its row number and column number in the table are both 1.

[0024] The regional coordinates of each merged cell in the table are also determined. For example, its boundary coordinates are used as the regional coordinates. Taking the merged cell recording "power system" as an example, its upper boundary is located in the area of ​​the 2nd row and 1st column, and its lower boundary is located in the area of ​​the 5th row and 1st column. Its regional coordinates include the coordinates of its upper boundary (2,1) and the coordinates of its lower boundary (5,1).

[0025] Step S201 , for each non-merged cell, fill the content of the non-merged cell into the cell located at the i-th row and j-th column in the correction table; In this embodiment, taking the cell recording "faulty system" as an example, the row number i and column number j in the table are 1 and 1 respectively, then "faulty system" is filled into the cell located in the 1st row and 1st column in the correction table; and so on, the same processing is performed on each non-merged cell.

[0026] Step S202 , for each merged cell, filling the content of the merged cell into the target cell corresponding to the region coordinates in the correction table; In this embodiment, referring to the above explanation, taking the merged cell recording "Power System" as an example, its regional coordinates include the coordinates of its upper boundary (2,1) and the coordinates of the lower boundary (5,1). Then, in the correction table, the target cells corresponding to its regional coordinates include the cell in the 2nd row and 1st column, the cell in the 3rd row and 1st column, the cell in the 4th row and 1st column, and the cell in the 5th row and 1st column, and "Power System" is filled into these cells in the correction table.

[0027] Furthermore, in one embodiment, the region coordinates include the row number i of the cell q covered by the merged cell in the table. q and column number j q , where q ranges from 1 to N, and N is the number of cells covered by the merged cell. Filling the contents of the merged cell into the target cell corresponding to the region coordinate in the correction table includes: Fill the contents of the merged cell into the correction table at the i q Row j q The target cell for the column.

[0028] In this embodiment, continue to refer to Figure 2 Taking the merged cell recording "Power System" as an example, the cells covered include the cell in the 2nd row and 1st column, the cell in the 3rd row and 1st column, the cell in the 4th row and 1st column, and the cell in the 5th row and 1st column. The area coordinates include (2,1), (3,1), (4,1), and (5,1). Then "Power System" is filled into the cell in the 2nd row and 1st column, the cell in the 3rd row and 1st column, the cell in the 4th row and 1st column, and the cell in the 5th row and 1st column in the correction table.

[0029] Step S30: After all cells are traversed, the obtained correction table is embedded into a plain text document.

[0030] In this embodiment, after traversing all cells in the table, the following can be obtained: Figure 3 The correction table shown, Figure 3 A schematic diagram of the calibration table.

[0031] After obtaining the correction table, it can be embedded in a plain text document, wherein the plain text document can be a Markdown document.

[0032] In an embodiment of the present application, when the type of the original file is a spreadsheet file type and there are merged cells in the table of the original file, the content of each cell in the table, the row number i and column number j of each non-merged cell in the table, and the area coordinates of each merged cell in the table are determined; for each non-merged cell, the content of the non-merged cell is filled into the cell located in the i-th row and j-th column in the correction table; for each merged cell, the content of the merged cell is filled into the target cell corresponding to the area coordinates in the correction table; after all cells are traversed, the obtained correction table is embedded in a plain text document. Through the embodiment of the present application, the table with merged cells is optimized into a correction table with a one-to-one correspondence between the header row and any record row, and then embedded in a plain text document, which greatly avoids the situation where the content embedded in the plain text document is formatted incorrectly or information is lost compared to the table in the original file.

[0033] Furthermore, in one embodiment, referring to Figure 4 , Figure 4 This is a flow chart of the second embodiment of the document parsing method of this application. Figure 4 As shown, the document parsing method further includes: Step S40: When the original file is a spreadsheet file and there are no merged cells in the table of the original file, determining whether a splitting condition is met; In this embodiment, according to the embodiment of step S10, when it is determined based on the suffix of the original file and XML parsing technology that the type of the original file is a spreadsheet file type and there are no merged cells in the table of the original file, it is determined whether the splitting condition is met.

[0034] Among them, considering that a plain text document in Markdown format cannot visually display a table with too large content, it is necessary to split the long table.

[0035] For example, based on Markdown format restrictions, a row threshold or a field threshold can be set. When the actual number of rows in the table is greater than the row threshold and / or the actual number of fields is greater than the field threshold, it is determined that the splitting condition is met.

[0036] Step S50: if the splitting condition is met, split the table into several sub-tables; In this embodiment, if the splitting condition is met, the table is split into a plurality of sub-tables according to a preset number of rows or a preset number of bytes.

[0037] Furthermore, in one embodiment, splitting the table into a plurality of sub-tables includes: The table is split according to a preset number of rows to obtain a plurality of sub-tables, wherein the number of rows included in each sub-table is not greater than the preset number of rows.

[0038] In this embodiment, assuming that the table includes 1 header row and 77 data rows, and the preset number of rows is 10, the table is split into 9 sub-tables, where sub-table 1 includes the header and rows 1 to 9 of the data rows, sub-table 2 includes the header and rows 10 to 18 of the data rows, and so on.

[0039] Furthermore, in one embodiment, splitting the table into a plurality of sub-tables includes: The table is split according to a preset number of bytes to obtain a plurality of sub-tables, wherein the size of each sub-table is not greater than the preset number of bytes.

[0040] In this example, assume that the table includes one header row and 77 data rows, and the preset number of bytes is max_size. To ensure that the size of the Markdown string generated by each sub-table does not exceed max_size, the steps are as follows: 1. Save the table header separately (each subtable requires a header).

[0041] 2. Initialize a data list for the current subtable (put it in the header first).

[0042] 3. Initialize the size of the current subtable (initially the size of the table header converted to a Markdown string).

[0043] 4. Traverse the data rows: a. Convert the current line to a Markdown table line (string) and calculate its length (number of bytes).

[0044] b. If the current subtable size plus the length of the current row is less than or equal to max_size, then the current row is added to the current subtable and the current subtable size is updated.

[0045] c. Otherwise, save the current subtable as a subtable, create a new subtable (containing the header and the current row), and reset the current subtable size (to the header size + the current row size).

[0046] 5. The last subtable is saved after the data row traversal is completed.

[0047] Step S60 : Convert different subtables into plain text descriptions and record them into different plain text documents.

[0048] In this embodiment, for example, the subtable i is converted into HTML format and recorded into a Markdown document i.

[0049] Furthermore, in one embodiment, referring to Figure 5 , Figure 5 This is a flow chart of the third embodiment of the document parsing method of this application. Figure 5 As shown, the document parsing method further includes: Step S70: When the original file is a rich element file, identifying valid elements in the original file and determining the original position of each valid element in the original file; In this embodiment, the original file extension is identified. If the extension is any of PDF, DOC, DOCX, PPT, or PPTX, the original file type is determined to be a rich element file type. The original file is then analyzed using the LayoutLMv3 model and optical character recognition (OCR) technology to perform layout detection and content recognition. Valid elements in the original file, such as text, formulas, tables, and images, are identified. Contaminants such as headers, footers, and footnotes are removed. The bounding box and corresponding two-dimensional position information for each valid element are obtained. Layout feature representations are generated through bounding box angle correction and normalization.

[0050] Step S801, when the valid element is text, write the text into a plain text document; In this embodiment, the bounding boxes of each page in the original document are sorted according to the reading order from left to right and from top to bottom, and the text results are saved in the Markdown document in order.

[0051] Step S802: When the valid element is a picture, the storage address and description of the picture are written into a plain text document; In this embodiment, images are stored in an image library, and their descriptions can be obtained using the lightweight Qwen-VL large model. First, the image is decoded using the CLIP visual encoder to extract visual features and generate a fixed-length feature sequence. The visual feature sequence is then concatenated with a text prompt (e.g., "Describe this image") as a joint input. This allows the cross-modal attention mechanism to dynamically fuse the image and text information to generate a text description of the image. Since the storage address corresponding to the image's unique label has been inserted as an index into the corresponding position of the Markdown document, the image is referenced in a mapping manner. On this basis, the image description corresponding to the image is inserted near the storage address.

[0052] Step S803: When the valid element is a formula, the formula is converted into a preset format and written into a plain text document; In this embodiment, inline formulas and interline formulas are converted into LaTeX format using the formula recognition model UniMERNet, and the LaTeX formulas are inserted into the corresponding positions of the parsed document according to the position sequence of their bounding boxes in the original document.

[0053] Step S804: When the valid element is a non-spread-page table, the plain text description, storage address, and table description of the non-spread-page table are written into the plain text document; In this embodiment, non-cross-page tables use the table recognition model StructEqTable to extract the table structure and convert it into a structured HTML format; HTML tables use 、 、 Tags fully describe the table structure, support advanced features such as merging cells (colspan / rowspan), style definitions, etc. After converting the table to HTML, it is embedded in the corresponding position of the Markdown document according to the original position of the non-spread table in the original file. In addition, the non-spread table is also saved in the image library in image format, and its storage address and table description are also written into the plain text document, specifically near the HTML text. Among them, the implementation of obtaining the table description refers to the above-mentioned embodiment of obtaining the image description, which will not be repeated here.

[0054] Step S805: When the valid element is a cross-page table, the plain text description and storage address of the cross-page table are written into the plain text document; In this embodiment, for a spreadsheet, only the plain text description and storage address of the spreadsheet are written into a plain text document, such as a Markdown document. The plain text description is HTML text, and the spreadsheet is also saved in a picture library in picture format, and its storage address is also written into the plain text document.

[0055] The position where the valid element is written into the plain text document is determined based on the original position of the valid element in the original file.

[0056] The layout style and arrangement order of each valid element in the original file are retained in the plain text document.

[0057] Furthermore, in one embodiment, after step S70, the method further includes: For the target valid element of the table, a serial number is assigned to the target valid element according to the original position of the target valid element in the original file, wherein the continuity of the serial number is consistent with the continuity of the original position; Check whether two target elements with adjacent serial numbers belong to the same spreadsheet; Merge all target valid elements belonging to the same spread table into a spread table.

[0058] In this embodiment, for example, the table includes 7 target valid elements, which are recorded as target valid element 1 to target valid element 7.

[0059] If the original positions of the two target valid elements are continuous, their serial numbers are also continuous, for example, according to the order of appearance, the serial numbers are 1 to 7 respectively.

[0060] Check whether two target elements with adjacent serial numbers belong to the same cross-page table, that is, check whether tables i and i+1 corresponding to the two target elements with serial numbers i and i+1 belong to the same cross-page table.

[0061] Among them, if table i is located at the end of page k, table i+1 is located at the beginning of page k+1, and there are no other valid elements (such as text, pictures and formulas) between table i and table i+1, then it is determined that the two target elements with adjacent detection numbers belong to the same cross-page table.

[0062] For example, if it is determined according to the above detection that target valid element 1 and target valid element 2 belong to the same cross-page table, and target valid element 2 and target valid element 3 belong to the same cross-page table, then target valid elements 1 to 3 are merged into a cross-page table, that is, Table 1 to Table 3 are merged.

[0063] In another optional embodiment, referring to Figure 6 , Figure 6 This is a flowchart for merging tables across multiple pages. Figure 6 As shown, first, an empty cross-page table memory is designed to store cross-page tables. The table memory contains the table information in the original document, including the page number, content, position, etc. of the table. The tables in the table memory are judged one by one to see whether they meet the requirements of cross-page tables. Then, it is judged whether the mth table is empty. If it is empty, it means there are two possibilities. One is that there is no table in the table memory itself, that is, no table element is recognized in the original document, and no table processing operation is required. The other is that all tables in the table memory have been traversed and no table processing operation is required. If it is not empty, it is necessary to judge whether the mth table is at the end of the page. Because the polluting elements such as page numbers and headers have been discarded, when a table is detected at When there is a table at the bottom of a page and at the top of the next page, and there are no other non-table elements (such as paragraph text, pictures, etc.) between the two tables, it can be preliminarily determined to be a cross-page table; for further confirmation, the structured attribute information of the table is extracted in combination with XML parsing technology. When the key attributes such as the number of columns and column width of the two tables match, and the position meets the cross-page characteristics, it is determined to be a cross-page table; the contents of the two tables confirmed to be cross-page tables are physically spliced ​​and stored in the cross-page table memory and the position index of the previous table is recorded to facilitate the insertion of the merged table in the corresponding position of the Markdown document; then continue to determine whether the next table meets the cross-page merging requirements. For tables that span multiple pages continuously, it is necessary to continue to determine the next page until the end.

[0064] In a second aspect, an embodiment of the present application also provides a document parsing device.

[0065] In one embodiment, referring to Figure 7 , Figure 7 This is a functional module diagram of an embodiment of the document parsing device of this application. Figure 7 As shown, the document parsing device includes: a determination module 10 for, when the original file is a spreadsheet file type and a table in the original file contains merged cells, determining the content of each cell in the table, the row number i and column number j of each non-merged cell in the table, and the region coordinates of each merged cell in the table; A filling module 20 is configured to fill, for each non-merged cell, the contents of the non-merged cell into the cell at row i and column j in the correction table; and, for each merged cell, fill the contents of the merged cell into the target cell corresponding to the region coordinates in the correction table; The embedding module 30 is used to embed the obtained correction table into a plain text document after all cells are traversed.

[0066] Furthermore, in one embodiment, the region coordinates include the row number i of the cell m covered by the merged cell in the table. q and column number j q , where q ranges from 1 to N, and N is the number of cells covered by the merged cells. The filling module 20 is used to: Fill the contents of the merged cell into the correction table at the i q Row j q The target cell of the column.

[0067] Furthermore, in one embodiment, the document parsing apparatus further includes: A judgment module, configured to judge whether a splitting condition is satisfied when the original file is a spreadsheet file and there are no merged cells in the table of the original file; A splitting module, configured to split the table into a plurality of sub-tables if a splitting condition is met; The recording module is used to convert different subtables into plain text descriptions and record them in different plain text documents.

[0068] Furthermore, in one embodiment, the splitting module is used to: The table is split according to a preset number of rows to obtain a plurality of sub-tables, wherein the number of rows included in each sub-table is not greater than the preset number of rows.

[0069] Furthermore, in one embodiment, the splitting module is used to: The table is split according to a preset number of bytes to obtain a plurality of sub-tables, wherein the size of each sub-table is not greater than the preset number of bytes.

[0070] Furthermore, in one embodiment, the document parsing apparatus further includes a rich element file processing module configured to: When the type of the original file is a rich element file type, identifying valid elements in the original file and determining the original position of each valid element in the original file; When the valid element is text, write the text to a plain text document; When the valid element is a picture, the storage address and description of the picture are written into a plain text document; When the valid element is a formula, the formula is converted into a preset format and written into a plain text document; When the valid element is a non-page-spanning table, the plain text description, storage address and table description of the non-page-spanning table are written into the plain text document; When the valid element is a cross-page table, the plain text description and storage address of the cross-page table are written into the plain text document; The position where the valid element is written into the plain text document is determined based on the original position of the valid element in the original file.

[0071] Furthermore, in one embodiment, the rich element file processing module is further configured to: For the target valid element of the table, a serial number is assigned to the target valid element according to the original position of the target valid element in the original file, wherein the continuity of the serial number is consistent with the continuity of the original position; Check whether two target elements with adjacent serial numbers belong to the same spreadsheet; Merge all target valid elements belonging to the same spread table into a spread table.

[0072] Among them, the functional implementation of each module in the above-mentioned document parsing device corresponds to each step in the above-mentioned document parsing method embodiment, and its functions and implementation processes are no longer repeated here.

[0073] In a third aspect, an embodiment of the present application provides a document parsing device, which may be a personal computer (PC), a laptop computer, a server, or other device with data processing capabilities.

[0074] Reference Figure 8 , Figure 8 Schematic diagram of the hardware structure of the document parsing device involved in the embodiment of the present application. In the embodiment of the present application, the document parsing device may include a processor, a memory, a communication interface and a communication bus.

[0075] The communication bus may be of any type and is used to interconnect the processor, memory, and communication interface.

[0076] Communication interfaces include input / output (I / O) interfaces, physical interfaces, and logical interfaces. These interfaces interconnect components within the document parsing device, as well as interfaces that connect the document parsing device to other devices (such as other computing devices or user devices). Physical interfaces can include Ethernet, fiber, or ATM interfaces; user devices can include displays and keyboards.

[0077] The memory can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), flash memory, optical storage, hard disk, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.

[0078] The processor may be a general-purpose processor that can invoke a document parsing program stored in a memory and execute the document parsing method provided in the embodiments of the present application. For example, the general-purpose processor may be a central processing unit (CPU). The method executed when the document parsing program is invoked can be referenced in the various embodiments of the document parsing method of the present application and will not be further described here.

[0079] Those skilled in the art will understand that Figure 8 The hardware structure shown in the figure does not constitute a limitation to the present application and may include more or fewer components than shown in the figure, or a combination of certain components, or a different arrangement of components.

[0080] In a fourth aspect, an embodiment of the present application also provides a computer-readable storage medium.

[0081] The computer-readable storage medium of the present application stores a document parsing program, wherein when the document parsing program is executed by a processor, the steps of the document parsing method described above are implemented.

[0082] Among them, the method implemented when the document parsing program is executed can refer to the various embodiments of the document parsing method of this application, and will not be repeated here.

[0083] It should be noted that the serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0084] The terms "including" and "having," and any variations thereof, in the specification and claims of this application and the accompanying drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to the process, method, product, or apparatus. The terms "first," "second," and "third" are used to distinguish between different objects, etc., and do not indicate a sequential order, nor do they limit the "first," "second," and "third" to different types.

[0085] In the description of the embodiments of this application, words such as "exemplary," "for example," or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary," "for example," or "for example" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary," "for example," or "for example" is intended to present the relevant concepts in a concrete manner.

[0086] In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in the text is merely a description of the association relationship of associated objects, indicating that three relationships may exist, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, "multiple" refers to two or more than two.

[0087] In some processes described in the embodiments of the present application, multiple operations or steps are included that appear in a specific order. However, it should be understood that these operations or steps may not be performed in the order in which they appear in the embodiments of the present application or may be performed in parallel. The sequence numbers of the operations are only used to distinguish between different operations, and the sequence numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations or steps may be performed in sequence or in parallel, and these operations or steps may be combined.

[0088] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, or the part that contributes to the existing technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above and includes a number of instructions for enabling a terminal device to execute the methods described in each embodiment of this application.

[0089] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A document parsing method, characterized in that: The document parsing method comprises: When the original file is a spreadsheet file and a table in the original file contains merged cells, determining the content of each cell in the table, the row number i and column number j of each non-merged cell in the table, and the region coordinates of each merged cell in the table; For each non-merged cell, fill the content of the non-merged cell into the cell at row i and column j in the correction table; For each merged cell, fill the content of the merged cell into the target cell corresponding to the region coordinates in the correction table; After all cells are traversed, the obtained correction table is embedded in a plain text document.

2. The document parsing method according to claim 1, wherein: The area coordinates include the row number i of the cell q covered by the merged cell in the table q and column number j q , where q ranges from 1 to N, and N is the number of cells covered by the merged cell. Filling the contents of the merged cell into the target cell corresponding to the region coordinate in the correction table includes: Fill the contents of the merged cell into the correction table at the i q Row j q The target cell for the column.

3. The document parsing method according to claim 1, wherein: The document parsing method further includes: When the type of the original file is a spreadsheet file type and there are no merged cells in the table of the original file, determining whether a splitting condition is met; If the splitting condition is met, the table is split into several sub-tables; Convert different subtables into plain text descriptions and record them in different plain text files.

4. The document parsing method according to claim 3, wherein: Splitting the table into several sub-tables includes: The table is split according to a preset number of rows to obtain a plurality of sub-tables, wherein the number of rows included in each sub-table is not greater than the preset number of rows.

5. The document parsing method according to claim 3, wherein: Splitting the table into several sub-tables includes: The table is split according to a preset number of bytes to obtain a plurality of sub-tables, wherein the size of each sub-table is not greater than the preset number of bytes.

6. The document parsing method according to claim 1, wherein: The document parsing method further includes: When the type of the original file is a rich element file type, identifying valid elements in the original file and determining the original position of each valid element in the original file; When the valid element is text, write the text to a plain text document; When the valid element is a picture, the storage address and description of the picture are written into a plain text document; When the valid element is a formula, the formula is converted into a preset format and written into a plain text document; When the valid element is a non-page-spanning table, the plain text description, storage address and table description of the non-page-spanning table are written into the plain text document; When the valid element is a cross-page table, the plain text description and storage address of the cross-page table are written into the plain text document; The position where the valid element is written into the plain text document is determined based on the original position of the valid element in the original file.

7. The document parsing method according to claim 6, wherein: When the type of the original file is a rich element file type, after identifying the valid elements in the original file and determining the original position of each valid element in the original file, the method further includes: For the target valid element of the table, a serial number is assigned to the target valid element according to the original position of the target valid element in the original file, wherein the continuity of the serial number is consistent with the continuity of the original position; Check whether two target elements with adjacent serial numbers belong to the same spreadsheet; Merge all target valid elements belonging to the same spread table into a spread table.

8. A document parsing device, characterized in that: The document parsing device comprises: a determination module, configured to, when the original file is a spreadsheet file type and a table in the original file contains merged cells, determine the content of each cell in the table, the row number i and column number j of each non-merged cell in the table, and the region coordinates of each merged cell in the table; A filling module is used to fill the content of each non-merged cell into the cell located in the i-th row and j-th column of the correction table; and for each merged cell, fill the content of the merged cell into the target cell corresponding to the region coordinates in the correction table; The embedding module is used to embed the obtained correction table into a plain text document after all cells are traversed.

9. A document parsing device, characterized in that: The document parsing device includes a processor, a memory, and a document parsing program stored in the memory and executable by the processor, wherein when the document parsing program is executed by the processor, the steps of the document parsing method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a document parsing program, wherein when the document parsing program is executed by a processor, the steps of the document parsing method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Method for positioning, reading and writing special-shaped table of document

    CN121708617A