Document analysis method and device
By associating the visual order of the presentation page elements and adjusting the large language model, the semantic fragmentation problem caused by the partitioning of presentations by page is solved, and more complete analytical data is achieved, improving the accuracy of vector encoding and reducing data redundancy.
Patent Information
- Application Number
- CN202510708196.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-02
AI Technical Summary
When analyzing a presentation, the partition by page leads to semantic fragmentation, and the semantic information between pages cannot be effectively retained, and the title and text are processed equally, losing structural information.
By determining the visual presentation order of elements of the demonstration page, establish the data structure relationship between the reference title and the child elements, associate the child element data of the same reference title, adjust the title level using a large language model, and retain the semantic association of important elements in the analytical data, degrade or delete unimportant elements.
Effectively preserve semantic information between pages in the presentation, improve the structure and semantic integrity of the analytical data, improve the accuracy of vector encoding, and reduce data redundancy.
Smart Images

Figure CN120579537A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a document parsing method and device. Background Art
[0002] Vectorized retrieval technology is a technology that encodes document chunks into high-dimensional vectors and achieves fast retrieval by calculating the similarity between vectors.
[0003] Generally, document chunking is performed based on the document's underlying data parsing structure. For presentations represented in extensible markup language (XML) format (PowerPointXML, PPTX), the underlying data parsing structure is a compressed package containing multiple XML files and resource files, organized directly by page. This data parsing approach can lead to semantic fragmentation in presentation chunking, as a complete logic flow (e.g., "Market Analysis -> Conclusion") often spans multiple pages. This approach can lead to fragmented presentation chunking, as chunking by page can disrupt semantics.
[0004] Therefore, how to provide a parsing method for presentation documents, effectively retain the semantic information between pages in the presentation, and provide parsed data with more complete semantic information for block coding is a key research topic for those skilled in the art. Summary of the Invention
[0005] In a first aspect, the present application provides a document parsing method, the method comprising: obtaining a presentation to be parsed, the presentation to be parsed comprising at least two presentation pages; determining first parsing data comprising target structure data based on elements contained in the at least two presentation pages, the target structure data comprising a data structure relationship between a reference title and a reference sub-element; wherein the reference title is a title element contained in the at least two presentation pages, the sub-elements of the reference title are located after the reference title in a visual presentation order, the sub-elements comprising at least one of the following elements: sub-title, text, picture, table, link, the title level of the sub-title is lower than that of the reference title, and the reference sub-element comprises an element belonging to a different presentation page from the reference title.
[0006] By adopting the document parsing method provided in the present application, an association is established between the sub-element data belonging to the same reference title in the same page or different pages of the page to which the reference title belongs and the reference title through the target structure data in the first parsed data, thereby effectively retaining the semantic information between pages in the presentation, rather than simply parsing by page, thereby providing the first parsed data with complete semantic information for vector encoding of document segmentation.
[0007] In some possible implementations, before determining the first parsed data containing the target structure data, the method further includes: obtaining elements contained in the at least two demonstration pages in a visual presentation order; obtaining the elements contained in the at least two demonstration pages includes: obtaining the first text content contained in the first text box in the first page, and determining whether the first text content meets a preset title condition, the first page being any one of the at least two demonstration pages; if it is determined that the first text content meets the preset title condition, adding a first preset mark to the first text content, the first preset mark being used to indicate that the first text content belongs to a title element; if it is determined that the first text content does not meet the preset title condition, determining that the first text content is the text data content corresponding to the first title, the first title being the title closest to the first text content in the visual presentation order, the page to which the first title belongs is the first page, or the page to which the first title belongs is before the first page in the visual presentation order.
[0008] This approach effectively preserves the structural information between the title and the text, rather than treating the title and the text equally, thereby improving the structural integrity of the parsed data.
[0009] In some possible implementations, the preset title condition includes that the font attributes of the text contain title features, and the text length is less than or equal to a preset length threshold.
[0010] In this way, title attributes are determined based on two conditions: font features and text features, thereby improving the accuracy of identifying title elements.
[0011] In some possible implementations, obtaining the elements contained in the at least two demonstration pages also includes: obtaining the first image contained in the first page, converting the first image into a lightweight markup language markdown syntax, and determining that the first image is the image data content corresponding to the second title, the second title is the title closest to the first image in the visual presentation order, the page to which the second title belongs is the first page, or, the page to which the second title belongs is before the first page in the visual presentation order.
[0012] In this way, the first image is associated with the second title, effectively retaining the intrinsic association between the first image and the second title, and ensuring the semantic integrity of the parsed data.
[0013] In some possible implementations, determining first parsed data containing target structure data based on elements contained in the at least two demonstration pages includes: generating second parsed data based on elements contained in the at least two demonstration pages, wherein the elements in the second parsed data correspond sequentially to the elements contained in the at least two demonstration pages in a visual presentation order, and the title level of any title element contained in the second parsed data is a preset initial level; determining a first title set based on the second parsed data, wherein the first title set is a title set included in the second parsed data; obtaining a second title set based on the first title set, a preset prompt word, and a large language model, wherein the preset prompt word is used to Instructing the large language model to divide the titles in the first title set into title levels on the basis of keeping the semantics of the titles unchanged, the title levels include at least two levels; when it is determined that the first edit distance between the text of the first output title and the text of the first original title is less than or equal to a preset distance threshold, replacing the preset initial level of the first original title in the second parsed data with the title level corresponding to the first output title to obtain updated second parsed data, the first original title is any one title in the first title set, and the first output title is any one title in the second title set; based on the updated second parsed data, determining the first parsed data.
[0014] In this way, when the first edit distance is less than or equal to the preset distance threshold, the level of the first output title output by the large language model is determined, thereby improving the accuracy of title level classification.
[0015] In some possible implementations, a second preset marker exists between elements that are adjacent to the page in the first parsed data, and the second preset marker is used to distinguish different elements that are adjacent to the page.
[0016] This approach preserves as many data features as possible within the document, further effectively retains the contextual semantic relationships between pages, and ensures the semantic integrity of the parsed data.
[0017] In some possible implementations, a second preset marker exists between elements adjacent to the page in the first parsed data, and the second preset marker is used to distinguish different elements adjacent to the page. After determining the first parsed data containing the target structure data, the method further includes: determining N title elements contained in the elements of the same demonstration page based on the second preset marker; when N is greater than or equal to 2, retaining the target title with the highest title level among the N title elements, and setting the other titles among the N titles except the target title as the body elements of the target title, and updating the first parsed data.
[0018] Using this method, for the case of multiple titles in a single slide in the first parsed data, the highest-level target title is retained, and other titles are downgraded to the body elements of the target title. This can enhance the relevance of the data content in the same presentation page and improve the problem of title structure fragmentation in the parsed data.
[0019] In some possible implementations, the obtaining of the elements contained in the at least two demonstration pages includes: when the attribute of the second text box containing the second text content in the first page is a header attribute, adding a third preset mark to the second text content, the third preset mark being used to indicate that the second text content belongs to the header content; when the attribute of the third text box containing the third text content in the first page is a footer attribute, adding a fourth preset mark to the third text content, the fourth preset mark being used to indicate that the third text content belongs to the footer content; after determining the first parsed data containing the target structure data, the method further includes: after determining that the elements of each demonstration page contain the target header content, and the target When the target header content is not recognized as a title, all the target header content originally contained in the first parsed data is deleted, and the target header content is added at the end of the first parsed data, and the target header content is set to a title element belonging to a target level, and the target level is lower than the title level of other formal title elements contained in the first parsed data; when it is determined that the elements of each presentation page contain target footer content, and the target footer content is not recognized as a title, all the target footer content originally contained in the first parsed data is deleted, and the target footer content is added at the end of the first parsed data, and the target footer content is set to a title element belonging to the target level.
[0020] In some possible implementations, after the electronic device determines the first parsed data containing the target structure data, the above method also includes: when the electronic device determines that the elements of each page of the demonstration page contain the target image content, deleting all the target image content originally contained in the first parsed data, adding the target image content at the end of the first parsed data, and setting the target image content as a title element belonging to the target level.
[0021] By adopting this method, for unimportant repeated elements (headers, footers, logos), only one record is retained at the end of the document. On the one hand, it can reduce the data content redundancy problem of unimportant repeated elements in the parsed data; on the other hand, by setting the header, footer, and logo elements to a target level lower than the title level of other formal title elements, the semantic relationship between them and other title elements is retained, which can better protect the integrity and originality of the parsed data.
[0022] In a second aspect, the present application further provides a document parsing device, comprising a unit for executing any one of the document parsing methods in the first aspect.
[0023] Exemplarily, the document parsing device includes: a first acquisition unit, used to acquire a presentation to be parsed, the presentation to be parsed includes at least two pages of presentation pages; a determination unit, used to determine first parsing data including target structure data based on elements included in the at least two pages of presentation pages, so as to perform block encoding based on the target structure data in the first parsing data, the target structure data including a data structure relationship between a reference title and a reference sub-element; wherein the reference title is a title element included in the at least two pages of presentation pages, the sub-elements of the reference title are located after the reference title in the visual presentation order, the sub-elements include at least one of the following elements: sub-title, text, picture, table, link, the title level of the sub-title is lower than that of the reference title, and the reference sub-element includes elements belonging to a different presentation page from the reference title.
[0024] In a third aspect, the present application further provides a computer storage medium, which can store multiple instructions, and the instructions are suitable for being loaded by a processor and executing any one of the document parsing methods in the first aspect.
[0025] In a fourth aspect, an embodiment of the present application further provides a computer program product comprising instructions, which, when run on an electronic device, enables the electronic device to execute any one of the document parsing methods in the first aspect.
[0026] In a fifth aspect, an embodiment of the present application further provides a chip module, comprising a transceiver component and a chip, wherein the chip is used to execute any one of the document parsing methods in the first aspect.
[0027] It is understood that the document parsing device, computer storage medium, computer program, computer program product, and chip system provided above are all used to perform the method described in any implementation of the first aspect of the embodiments of this application. Therefore, the beneficial effects that can be achieved can be referenced to the beneficial effects of the corresponding method and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 This is a flowchart of a document parsing method provided by an embodiment of the present application;
[0029] Figure 2 is a schematic diagram of a second parsed data provided in an embodiment of the present application;
[0030] Figure 3This is a schematic diagram of a comparison between an original title set and a title set after title level conversion provided by an embodiment of the present application;
[0031] Figure 4 is a schematic diagram of first parsed data provided by an embodiment of the present application;
[0032] Figure 5 This is a schematic diagram of first parsed data after single-slide multi-title processing provided by an embodiment of the present application;
[0033] Figure 6 This is a flowchart of another document parsing method provided by an embodiment of the present application;
[0034] Figure 7 is a schematic diagram of a document parsing device provided in an embodiment of the present application;
[0035] Figure 8 is a schematic diagram of another document parsing device provided in an embodiment of the present application;
[0036] Figure 9 This is a schematic diagram of another document parsing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0037] In order to make the purpose, technical solutions and advantages of this application clearer, this application will be further described below in conjunction with the accompanying drawings.
[0038] Example 1:
[0039] See also Figure 1 , Figure 1 This is a flow chart of a document parsing method provided in an embodiment of the present application. Figure 1 As shown, the document parsing method includes the following steps:
[0040] S101: The electronic device obtains a presentation to be parsed, where the presentation to be parsed includes at least two presentation pages.
[0041] In the embodiment of the present application, each presentation page may include one or more elements of a title, text, picture, or table.
[0042] S102: The electronic device determines first parsed data including target structure data based on elements included in at least two presentation pages.
[0043] In an embodiment of the present application, the target structure data includes a data structure relationship between a reference title and a reference sub-element.
[0044] The reference title is a title element contained in the at least two demonstration pages.
[0045] The reference sub-element is located after the reference title in the visual presentation order, and includes at least one of the following elements: sub-title, text, picture, table, link. The title level of the sub-title is lower than that of the reference title, and the reference sub-element contains elements that belong to different presentation pages from the reference title. The reference sub-element contains elements that belong to different pages from the reference title. Specifically, the reference sub-element may contain elements that belong to the same page and / or different pages as the reference title. If the reference sub-element includes a title element, the title element belongs to the sub-title of the reference title. If the reference sub-element includes a text element, the text element belongs to the text element of the reference title. The same applies to pictures and table data.
[0046] In an embodiment of the present application, in terms of the presentation order between pages, the visual presentation order is from the first page to the last page of the presentation; in terms of the presentation order of the same page, the visual presentation order is from top to bottom and from left to right.
[0047] In some possible implementations, the electronic device uses a depth-first traversal algorithm to process the elements in each page (slide), and extracts all text boxes, picture objects, table data, and link data in the order of visual presentation. In some possible implementations, after the electronic device obtains the text box object, it performs a classification judgment on the text box object. Specifically, when the first text content in the first text box meets the preset title condition, it is determined that the first text content belongs to the title element. When the first text content in the first text box does not meet the preset title condition, it is determined that the first text content belongs to the text.
[0048] In some possible implementations, the preset title condition includes that the font attribute of the text contains the title feature, or the preset title condition includes that the font attribute of the text contains the title feature and the text length is less than a preset length threshold.
[0049] In some possible implementations, after obtaining the elements of each slide, the electronic device determines the second parsed data based on a preset conversion rule. The preset conversion rule includes:
[0050] 1) Add a first preset mark to the title text content corresponding to the title element, where the first preset mark is used to indicate the title element.
[0051] In some possible implementations, the first preset mark can be a mark for indicating the initial level of the title, and the initial level can be a preset lowest level of the title level (for example, level three). For example, the first preset mark is '###', and adding the first preset mark to the title text content corresponding to the title element means adding the '###' mark before the title text content.
[0052] This approach sets the initial level of the original title to the lowest level, improving the compatibility of the original title level. Compared to setting the initial level to the highest level, setting the initial level to the lowest level can reduce the influence of the original title's level on the hierarchical structure of the title in the first parsed data when the large language model cannot accurately determine the original title's level and the original title remains at the initial level, thereby improving the accuracy of the hierarchical relationships in the first parsed data.
[0053] It should be noted that the initial level is level three and the first preset mark is the '###' mark for example only. The first preset mark can also be other marks, which is not limited in this document. For example, the initial level is level one and the first preset mark is the '#' mark.
[0054] 2) Convert the image object, table object, and link object into the lightweight markup language markdown syntax respectively.
[0055] For example, the markdown image syntax is: [image alt](image link "image title"). Image alt is the alternative text, which means the text displayed when the image cannot be displayed. Image link is the source URL or local path of the image. Image title is the text displayed when the mouse hovers over the image.
[0056] For example, the markdown table syntax uses a vertical bar '|' to separate different cells and a hyphen '-' to separate the header and other rows. For example, |header1|header2||-----|-----||cell1|cell2||cell3|cell4|.
[0057] For example, the markdown link syntax is [hyperlink display name](hyperlink address "hyperlink title").
[0058] Generally speaking, vector encoding is difficult to implement for image data, so it generally doesn't encode image data. Specifically, when parsing presentation data by page, the document is chunked by page. If a presentation page contains only image data, since vector encoding doesn't encode image data and the chunk corresponding to that page only contains images, the chunk corresponding to that page won't be encoded, resulting in the loss of the image data.
[0059] However, by adopting the method provided in the present application, the first image is associated with the second title. When the document is segmented, the first image and the second title will be divided into associated blocks (for example, divided into the same chunk, or divided into two different chunks, but there is an association relationship between the two chunks and sub-elements belonging to the same title), that is, the vector of the block corresponding to the second title has an index relationship with the first image. Then, even if the first image or the chunk corresponding to the first image is not encoded, the data content of the first image can be retrieved through the second title during data retrieval, thereby ensuring the integrity of the parsed data and vector encoding, and avoiding the data loss problem of page parsing.
[0060] 3) The main content remains the same.
[0061] In an embodiment of the present application, the elements in the second parsed data correspond in sequence to the elements contained in the visual presentation order of the presentation page of the above-mentioned presentation document to be parsed. So as to determine that the first text content is the text data content belonging to the first title, the first title is the title that is closest to the first text content in the visual presentation order, and the page to which the first title belongs is the page to which the first text content belongs, or, in the visual presentation order, the page to which the first title belongs is before the page to which the first text content belongs. And, so as to determine that the first picture belongs to the picture data content corresponding to the second title, the second title is the title that is closest to the first picture in the visual presentation order, and the page to which the second title belongs is the page to which the first picture belongs, or, in the visual presentation order, the page to which the second title belongs is before the page to which the first picture belongs. The same applies to table data and link data, which will not be described in detail here.
[0062] As an example, the second parsed data includes the following content: Figure 2 shown.
[0063] In some possible implementations, after obtaining the second parsed data, the electronic device determines a first title set, which is a set of title elements in the second parsed data that include a first preset tag. The electronic device determines the title hierarchy of the first title set based on a preset prompt word and a large language model. The preset prompt word instructs the large language model to classify the titles in the first title set into title levels while maintaining the title text unchanged.
[0064] As an example, the title hierarchy of the first title set contains three levels. The first-level title is represented by '#', the second-level title is represented by '##', and the third-level title is represented by '###'. The priority of the first-level title is higher than that of the second-level title, and the priority of the second-level title is higher than that of the third-level title. The first title set contains the following content: Figure 3 As shown in 3A in FIG, the title hierarchy relationship of the first title set is as follows Figure 3 As shown in 3B.
[0065] The electronic device determines the first parsed data containing the target structure data based on the title hierarchical relationship of the first title set and the second parsed data. Specifically, the electronic device merges the title hierarchical relationship of the first title set into the second parsed data to obtain the target structure data.
[0066] Exemplarily, the reference sub-element in the target structure data includes an element that is located after the reference title and before the reference element in the visual presentation order. The reference element is a title element with a title level higher than or equal to the reference title, or, if there is no title element with a title level higher than or equal to the reference title after the reference title in the visual presentation order, then the reference element can be an element in the last page of the presentation to be parsed (for example, the last element of the last page), and the reference element and the reference title belong to the same or different presentation pages. In other words, the target structure data refers to associating all data after the reference title and before the reference element to the reference title.
[0067] As an example, the first parsed data includes the following content: Figure 4 shown.
[0068] Generally, because a complete logic flow (e.g., "Market Analysis -> Conclusion") spans multiple pages, parsing PPTX documents by page can disrupt semantics and lead to fragmented presentation encoding. Furthermore, this parsing method treats title and body text equally within a page, losing structural information. Images and tables may also be separated from their corresponding titles, disrupting their inherent connection to other data.
[0069] By adopting the document parsing method provided in the embodiment of the present application, the sub-element data belonging to the same reference title in the same or different pages are associated with the reference title through the target structure data in the first parsed data. On the one hand, the semantic information between pages in the presentation can be effectively retained, rather than simply parsing by page. On the other hand, the structural information between the title and subtitle, and the title and the text is effectively extracted, rather than treating the title and the text equally, thereby improving the structural integrity of the parsed data. On another hand, the pictures and table data are associated with the titles to which they belong, effectively retaining their intrinsic association with other data, thereby improving the semantic integrity of the parsed data. In this way, parsed data with complete semantic information can be provided for vector encoding of document blocks.
[0070] Generally, a large language model receives an input text containing a preset prompt word and generates an output text that meets the preset prompt word requirements based on the input text. Although the preset prompt word is assigned, in actual applications, the output text may deviate from the requirements.
[0071] In some possible implementations, the electronic device determines the first parsed data including the target structure data, specifically including:
[0072] The electronic device generates second parsed data based on elements included in at least two demonstration pages, wherein the elements in the second parsed data correspond sequentially to the elements included in the at least two demonstration pages in a visual presentation order, and the title level of any title element included in the second parsed data is a preset initial level; determines a first title set based on the second parsed data, wherein the first title set is the title set included in the second parsed data; obtains a second title set based on the first title set, a preset prompt word, and a large language model, wherein the preset prompt word is used to instruct the large language model to classify the titles in the first title set into title levels while maintaining the title text and / or the title semantics unchanged, wherein the title levels include at least two levels; and, upon determining that a first edit distance between the text of the first output title and the text of the first original title is less than or equal to a preset distance threshold, replaces the preset initial level of the first original title in the second parsed data with the title level corresponding to the first output title, thereby obtaining updated second parsed data, wherein the first original title is any title in the first title set, and the first output title is any title in the second title set; and determines the first parsed data based on the updated second parsed data. The number of titles included in the second title set is equal to or different from the number of titles included in the first title set.
[0073] As an example, the preset distance threshold may be T, where T satisfies the formula: T=0.2*max[len(title1), len(title2)], where len(title1) and len(title2) are the first original title and the first output title, respectively.
[0074] Understandably, if the large language model identifies the first original title as unsuitable as a title or as content that can be used as a title with corresponding modifications, the large language model may add, delete, or modify the first original title. To ensure model accuracy, the electronic device should adopt the first output title corresponding to the first original title output by the large language model if the changes are minor, and reject it if the changes are significant.
[0075] By adopting the method provided in the embodiment of the present application, the level of the first output title output by the large language model is determined only when the first edit distance is less than or equal to the preset distance threshold, thereby improving the accuracy of title level classification.
[0076] In some possible implementations, the electronic device may select an input strategy for the large language model based on the amount of the second parsed data, the amount of preset prompt words, and a preset maximum data size, where the preset maximum data size is the maximum amount of input data that the large language model can accommodate. For example, if the amount of the second parsed data and the amount of the preset prompt words exceed the maximum data size, the electronic device inputs the first set of titles and the preset prompt words into the large language model to obtain a second set of titles as output by the large language model. This approach simplifies and refines the data in the large language model, improving its processing speed.
[0077] Alternatively, when the data volume of the second parsed data and the data volume of the preset prompt words are less than or equal to the maximum data volume, the electronic device may also input the second parsed data and the preset prompt words into the large language model. In this case, the content output by the large language model may be the above-mentioned second title set, or it may be the third parsed data, which not only includes the title elements after the level update, but also includes the data content corresponding to the title elements (including one or more of the subtitles, text, pictures, tables, and links). In this way, the data input into the large language model is a complete data source, which helps the large language model to extract the title level based on the semantic information in the text other than the title, thereby improving the accuracy of the title level structure output by the large language model.
[0078] In some possible implementations, if the first edit distance is determined to be less than a preset distance threshold, the electronic device may not only replace the preset initial level of the first original title in the second parsed data with the title level corresponding to the first output title, but may also replace the title text of the first original title in the second parsed data with the title text corresponding to the first output title. In this manner, when the first edit distance is less than the preset distance threshold, adopting the slightly modified title text of the first output title output by the large language model can improve the accuracy of the title text in the first parsed data.
[0079] In some possible implementations, when it is determined that the first edit distance between the first original title and any first output title in the second title set is greater than a preset distance threshold, the first original title is downgraded to the main text. Exemplarily, the first preset title (e.g., the '###' mark) before the title text of the first original title is deleted. In this way, when the large language model makes too large a change to the original title, the original title is downgraded to the main text, which can reduce the influence of the title level of the original title on the title hierarchy structure in the first parsed data, especially when the initial level of the original title is high, the accuracy of the title hierarchy relationship in the first parsed data can be greatly improved.
[0080] In some other possible implementations, when it is determined that the first edit distance between the first original title and any first output title in the second title set is greater than a preset distance threshold, the title level of the first original title is set to a target level, which is lower than the level of any title in the second title set. Exemplarily, the lowest title level in the second title set is a level three title (e.g., '###'), then the target level is a level four title (e.g., '####'). In this way, when the large language model makes too much change to the first original title, the title level of the first original title is set to a target level that is lower than the level of any other title. This can retain the title attributes of the first original title, avoid the title text of the first original title and the data content (text, charts, links, etc.) corresponding to the first original title being classified as belonging to the previous original title (e.g., the second original title), resulting in too much data associated with the second original title, and the problem of too much data content displayed based on the search for the second original title during data block vectorization retrieval.
[0081] In some possible implementations, a second preset marker is present between elements in the first parsed data that are adjacent to each other on the page. This second preset marker is used to distinguish between elements in the adjacent pages. This approach preserves as many document data features as possible in the first parsed data, further effectively preserving the contextual semantic relationships between pages, and helping to improve the accuracy of vectorized retrieval based on the first parsed data.
[0082] As an example, the second preset mark includes at least two line breaks.
[0083] In some possible implementations, after determining the first parsed data, the electronic device may further perform single-slide multi-title processing on the first parsed data to obtain updated first parsed data. The single-slide multi-title processing includes: the electronic device determines, based on a second preset tag, N title elements contained in the elements of the same presentation page after the first parsed data; when N is greater than or equal to 2, retaining the target title with the highest title level among the N title elements of the first parsed data, and setting the other titles among the N titles except the target title as the body elements of the target title. If the target title is not the first element of the page to which it belongs, the target title is moved to the first line of the page to which it belongs.
[0084] As an example, Figure 4 After the first parsed data shown in FIG. 1 is processed into a single slide with multiple titles, the updated first parsed data obtained can be as follows: Figure 5 shown.
[0085] Generally, the data on the same presentation page is strongly correlated, and when a document is divided into chunks, the data content of different titles is generally divided into different chunks. If a presentation page contains multiple title elements, the data on the same presentation page that is strongly correlated will be divided into different chunks, resulting in fragmentation. However, by adopting the method provided in the embodiment of the present application, for the case of multiple titles in a single slide in the first parsed data, the highest-level target title is retained, and other titles are downgraded to the body elements of the target title. This can enhance the semantic relevance of the data content on the same presentation page and improve the fragmentation problem of document chunking when performing vectorized retrieval based on the first parsed data.
[0086] In some possible implementations, the electronic device obtaining the elements contained in the at least two demonstration pages may also include: when the attribute of the second text box containing second text content in the first page is a header attribute, adding a third preset mark to the second text content, and the third preset mark is used to indicate that the second text content belongs to the header content; when the attribute of the third text box containing third text content in the first page is a footer attribute, adding a fourth preset mark to the third text content, and the fourth preset mark is used to indicate that the third text content belongs to the footer content.
[0087] In some possible implementations, after the electronic device determines the first parsed data containing the target structure data, the above method also includes: when the electronic device determines that the elements of each page of the demonstration page contain the target header content and the target header content is not recognized as a title, the electronic device deletes all the target header content originally contained in the first parsed data, adds the target header content at the end of the first parsed data, and sets the target header content to a title element belonging to the target level, and the target level is lower than the title level of other formal title elements contained in the first parsed data; when it is determined that the elements of each page of the demonstration page contain the target footer content and the target footer content is not recognized as a title, the electronic device deletes all the target footer content originally contained in the first parsed data, adds the target footer content at the end of the first parsed data, and sets the target footer content to a title element belonging to the target level.
[0088] In some possible implementations, after the electronic device determines the first parsed data containing the target structure data, the above method also includes: when the electronic device determines that the elements of each page of the demonstration page contain the target image content (indicating that the target image is likely to be a logo), deleting all the target image content originally contained in the first parsed data, adding the target image content at the end of the first parsed data, and setting the target image content to a title element belonging to the target level.
[0089] Generally, in the page-by-page parsing mode, the document is divided into blocks and pages, and there is a problem of multiple encoding for unimportant repeated elements contained in each page. However, the document parsing method provided by the present application only retains one record at the end of the document for unimportant repeated elements (headers, footers, logos). On the one hand, it can reduce the data content redundancy problem of unimportant repeated elements in the parsed data, thereby improving the problem of repeated elements being encoded multiple times in vectorized retrieval; on the other hand, by setting the header, footer, and logo elements to a target level lower than the title level of other formal title elements, the semantic association between them and other title elements is retained, which can better protect the integrity and originality of the first parsed data.
[0090] It is understandable that the embodiment of the present application uses an electronic device as an example to illustrate the execution subject of the document parsing method provided by the present application, and the electronic device can also be understood as the document parsing device shown in the embodiment of the present application. In the embodiment of the present application, the electronic device can be a microprocessor or computer for executing program code, etc., and the electronic devices that can be used to execute the method provided by the embodiment of the present application all fall within the protection scope of the embodiment of the present application, and the present application does not impose any restrictions. Exemplarily, the electronic device can be a desktop computer, a portable notebook, a mobile terminal, a 32-bit microprocessor or a 64-bit microprocessor, etc., and the embodiment of the present application does not limit this.
[0091] Example 6:
[0092] The following combination Figure 6 , introduces another document parsing method provided by the embodiment of this application. Figure 6 As shown, the method includes:
[0093] S601, receiving a PPTX format document.
[0094] Specifically, the electronic device receives a presentation document in PPTX format to be parsed.
[0095] S602, obtain markdown file 1.
[0096] Specifically, the electronic device extracts elements from the presentation to be parsed and converts the elements in the presentation to be parsed into a markdown format, thereby obtaining a markdown file 1. The elements in the markdown file 1 correspond sequentially to the elements included in the presentation page of the presentation to be parsed in the visual presentation order. The markdown file 1 can be understood as the second parsed data in the above embodiment.
[0097] Specifically, the electronic device can use a depth-first traversal algorithm to process the elements in each slide, extract all elements in the order of visual presentation, identify title elements and body elements, and convert the format of text content, pictures, tables, and link objects based on preset conversion rules to obtain the above-mentioned markdown file 1.
[0098] Exemplarily, after extracting the text box element, the electronic device performs classification judgment on the text box, and identifying the title element and the body element includes:
[0099] i) When the text content in the text box meets a preset title condition, the text content is determined to be a title element.
[0100] ii) If the text content in the text box does not meet the preset title condition, the text content is determined to be a body element.
[0101] Exemplarily, the electronic device performing format conversion based on a preset conversion rule includes:
[0102] i) Add '###' mark to the title element.
[0103] ii) The text element remains the original text.
[0104] iii) Convert images, tables, and link elements to the corresponding markdown format.
[0105] For the markdown image format, markdown table format, and markdown link format, please refer to the relevant description in the above embodiment and will not be described in detail here.
[0106] iiiiii) Add two line breaks (\n\n) immediately between slides.
[0107] The electronic device determines the markdown file 1 based on the extracted elements and preset conversion rules.
[0108] For the description of the preset title conditions and the preset conversion rules, please refer to the relevant description in the above embodiment, which will not be described in detail here.
[0109] S603: Determine a first title set.
[0110] Specifically, the electronic device determines the first title set based on the title element with the '###' mark added in the markdown1 file.
[0111] S604: Determine a second title set including three-level structure titles.
[0112] Specifically, the electronic device may input the first title set and preset prompt words into the large language model, divide the title elements in the first title set into three levels of title levels, and obtain the second title set output by the large language model.
[0113] Exemplarily, the large language model processing flow includes:
[0114] The electronic device selects an input strategy based on the token length corresponding to markdown file 1:
[0115] i) When the token length is less than or equal to the preset maximum length, input markdown file 1.
[0116] ii) When the token length is greater than the preset maximum length, the first title set is input.
[0117] The electronic device is configured with preset prompt words, which include:
[0118] i) A first title set, which includes titles marked with '###'.
[0119] ii) Keep the title content unchanged, and / or keep the title semantics unchanged.
[0120] iii) Divide the titles in the first title set into a three-level title structure.
[0121] The output requirements of electronic devices for constructing large language models include:
[0122] i) Strictly follow the three-level title system (#, ##, ###)
[0123] ii) The inclusion of text content is prohibited.
[0124] iii) Keep the original semantics unchanged.
[0125] It should be noted that the specific preset prompt words and output requirements can be determined based on the specific design, and this document does not limit this. For example, when the data volume of the markdown file 1 is less than or equal to 80% of the maximum data volume that can be input by the large language model, the electronic device inputs the markdown file 1 and the preset prompt words into the large language model, and obtains the third parsed data output by the large language model. The third parsed data not only includes the title after the level is updated, but also includes the corresponding text content. Among them, the above-mentioned preset prompt words also include keeping the original semantics and / or original text unchanged, and the above-mentioned output requirements do not include the requirement of prohibiting the inclusion of text content.
[0126] S605: Match and merge titles to obtain markdown file 2.
[0127] Specifically, the electronic device matches the original title with the title output by the model using a greedy approach from top to bottom, for example, by calculating the edit distance based on Levenshtein.
[0128] Exemplarily, title matching and fusion include the following steps:
[0129] i) Calculate the edit distance d between the first original title and the first output title.
[0130] ii) Determine whether d is less than or equal to a preset distance threshold T. For the description of T, reference may be made to the relevant description in the above embodiment and will not be elaborated here.
[0131] iii) If it is determined that d is less than or equal to T, the first original title in markdown file 1 is replaced with the first output title. If it is determined that the edit distance d between the first original title and any first output title in the second title set is greater than T, the title level of the first original title in markdown file 1 is set to text. Based on this, the markdown file 1 with the updated title level is obtained, that is, markdown file 2. In this embodiment of the present application, the first parsed data in the above embodiment can be this markdown file 2.
[0132] S606, processing multiple titles in a single slide to obtain markdown file 3.
[0133] Specifically, the electronic device determines N title elements included in the elements of the same presentation page in the markdown file 2 based on the information of adding two line breaks between adjacent slides in S602.
[0134] When N is greater than or equal to 2, i) the electronic device retains the target title with the highest title level among the N title elements of markdown file 2; ii) the electronic device sets the other titles among the N title elements except the target title as the body elements of the target title. Among them, if the target title is not the first element of the page to which it belongs, the target title is moved to the first line of the page to which it belongs. Based on this, the updated markdown file 2, that is, markdown file 3, is obtained. In the embodiment of the present application, the first parsed data in the above embodiment can be the markdown file 3.
[0135] S607, output markdown file 3.
[0136] In this embodiment of the present application, markdown file 3 includes target structure data, which includes a data structure relationship between a reference title and a reference sub-element. The reference sub-element includes at least one of the following elements, located after the reference title and before the reference element in the visual presentation order: a subtitle, body text, an image, a table, and a link. The subtitle has a lower heading level than the reference title, and the reference sub-element includes elements belonging to a different presentation page than the reference title.
[0137] As an example, the electronic device performs document segmentation based on the target structure data in the markdown file 3, including: the electronic device groups the sub-elements belonging to the same first and third-level titles into associated blocks (for example, the same block, or two different blocks, and there is an associated identifier belonging to the same third-level title between the two different blocks); the electronic device groups the text, pictures, tables, links and other data belonging to the same first and second-level titles into associated blocks, and establishes an affiliated relationship with the third-level sub-titles; the electronic device groups the text, pictures, tables, links and other data belonging to the same first and second-level titles into associated blocks, and establishes an affiliated relationship with the third-level sub-titles and the second-level sub-titles respectively. When the electronic device performs vector encoding based on chunks, it may not encode the first picture and the first link sub-element corresponding to the first title. After completing the vectorized encoding, after receiving the vectorized search request for the first title, when displaying the search results, all sub-elements corresponding to the first title are displayed, including the text content such as the text, sub-titles, tables and the first picture and the first link sub-element corresponding to the first title, as well as the first picture and the first link sub-element.
[0138] By adopting the document parsing method provided in the embodiment of the present application, on the one hand, based on the levels of different title elements, the sub-element data belonging to the same title in different pages are associated with the title, which can effectively retain the semantic information between pages in the presentation. On the other hand, title fusion is performed only when the matching conditions (edit distance is less than or equal to the preset distance threshold) between the original title and the output title are met, thereby improving the accuracy of title level classification. On the other hand, for a single slide with multiple titles, the highest level title is retained and other titles are downgraded to the text elements of the target title, which can enhance the semantic relevance of the data content in the same presentation page.
[0139] In some possible implementations, after extracting the text box element, the above-mentioned electronic device performs classification judgment on the text box to identify the title element and the body element. Among them, when the text content in the text box does not meet the preset title condition, the text content is determined to be a body element. Specifically, it can also be: when the text content in the text box does not meet the preset title condition, if the attribute of the text box is a header attribute, then the text content is determined to belong to the header element; when the text content in the text box does not meet the preset title condition, if the attribute of the text box is a footer attribute, then the text content is determined to belong to the footer element; when the text content in the text box does not meet the preset title condition and the text box attribute is neither a header attribute nor a footer attribute, then the text content is determined to belong to the body element. The above-mentioned preset conversion rules also include: adding a third preset mark for the page to the header element; adding a fourth preset mark for the footer element. Before outputting the markdown file 3, the electronic device can also perform header and footer processing on the markdown file 3.
[0140] As an example, the electronic device performs header and footer processing on the markdown file 3, including: deleting all target header contents originally contained in the markdown file 3, adding a copy of the target header contents at the end of the markdown file 3, and setting the target header contents added at the end of the text to a title element belonging to a target level, which is lower than the title level of other formal title elements contained in the markdown file 3, for example, the target level is level four. When it is determined that the elements of each presentation page contain target footer contents, deleting all target footer contents originally contained in the markdown file 3, adding the target footer contents at the end of the markdown file 3, and setting the target footer contents added at the end of the text to a title element belonging to the target level.
[0141] In some possible implementations, before outputting the markdown file 3, the electronic device may further perform duplicate image deduplication processing on the markdown file 3. Specifically, when it is determined that the elements of each presentation page contain the target image content, all target image content originally contained in the markdown file 3 is deleted, and the target image content is added to the end of the first parsed data, and the target image content added to the end of the text is set as a title element belonging to the target level.
[0142] By adopting this approach, the data content of unimportant repeated elements is reduced, and the semantic association between repeated elements and other title elements is retained, thereby better ensuring the integrity and originality of the first parsed data.
[0143] It should be noted that the document parsing method provided in this application is not limited to application in the data parsing scenario of PPTX presentations, but can also be applied to any document with unclear chapter and paragraph structure, and this article does not limit this.
[0144] The present application also provides a document parsing device, comprising a Figure 1 and Figure 6 A unit that implements any of the document parsing methods in .
[0145] For example, please refer to Figure 7 , the document parsing device may include:
[0146] A first acquiring unit 701 is configured to acquire a presentation document to be parsed, wherein the presentation document to be parsed includes at least two presentation pages;
[0147] The first determining unit 702 is configured to determine first parsed data including target structure data based on the elements included in the at least two presentation pages, where the target structure data includes a data structure relationship between a reference title and a reference sub-element.
[0148] Please refer to Figure 8 , the document parsing device may further include:
[0149] A second acquiring unit 703 is configured to acquire elements included in the at least two presentation pages in a visual presentation order;
[0150] The second acquiring unit 703 specifically includes:
[0151] The first acquiring subunit 7031 is configured to acquire the first text content contained in the first text box in the first page;
[0152] The first determination subunit 7032 is used to determine whether the first text content meets the preset title condition; if it is determined that the first text content meets the preset title condition, add a first preset tag to the first text content, and the first preset tag is used to indicate that the first text content belongs to the title element; if it is determined that the first text content does not meet the preset title condition, determine that the first text content is the text data content corresponding to the first title.
[0153] The above-mentioned second acquisition unit 703 also includes a second acquisition sub-unit 7033, which is used to obtain the first image contained in the first page, convert the first image into a lightweight markup language markdown syntax, and determine that the first image is the image data content corresponding to the second title.
[0154] The first determining unit 702 specifically includes:
[0155] A generating subunit 7021 is configured to generate second parsed data based on the elements included in the at least two presentation pages;
[0156] A second determining subunit 7022 is configured to determine a first title set based on the second parsed data;
[0157] A third determining subunit 7023 is configured to obtain a second title set based on the first title set, preset prompt words, and a large language model;
[0158] The first processing sub-unit 7024 is configured to, if it is determined that the first edit distance between the text of the first output title and the text of the first original title is less than or equal to a preset distance threshold, replace the preset initial level of the first original title in the second parsed data with the title level corresponding to the first output title, thereby obtaining updated second parsed data;
[0159] The fourth determining subunit 7025 is configured to determine the first parsed data based on the updated second parsed data.
[0160] The document parsing device further includes:
[0161] A second determining unit 704 is configured to determine N title elements included in the elements of the same presentation page based on the second preset mark;
[0162] The first updating unit 705 is used to retain the target title with the highest title level among the N title elements when N is greater than or equal to 2, and set the other titles among the N titles except the target title as the body elements of the target title, thereby updating the first parsed data.
[0163] The second obtaining unit 703 specifically includes:
[0164] The second processing sub-unit 7034 is configured to add a third preset mark to the second text content when the attribute of the second text box containing the second text content on the first page is a header attribute;
[0165] The third processing sub-unit 7035 is configured to add a fourth preset mark to the third text content when the attribute of the third text box containing the third text content on the first page is a footer attribute;
[0166] The document parsing device further includes:
[0167] The second updating unit 706 is configured to, if it is determined that all elements of each presentation page include target header content and the target header content is not recognized as a title, delete all target header content originally included in the first parsed data, add the target header content to the end of the first parsed data, and set the target header content as a title element belonging to a target level;
[0168] The third updating unit 707 is used to delete all the target footer contents originally contained in the first parsed data when it is determined that the elements of each demonstration page contain target footer contents and the target footer contents are not identified as titles, add the target footer contents at the end of the first parsed data, and set the target footer contents to be title elements belonging to the target level.
[0169] For the description of the reference title, reference sub-element, first original title, first output title, preset distance threshold, preset initial level, preset prompt word, first title set, second parsed data, first page, first preset mark, second preset mark, third preset mark, fourth preset mark, preset title condition, target level, first title, and second title, please refer to the relevant description in the above embodiments and will not be described in detail here.
[0170] In the embodiments of the present application, any implementation method mentioned in the method embodiments is also applicable to the document parsing device provided in the present application. The specific execution steps can be found in the description of the aforementioned method embodiments and will not be described in detail here.
[0171] The embodiment of the present application also provides a document parsing device, including a processor, wherein the processor is configured to execute the following Figure 1 or Figure 6 Any document parsing method in .
[0172] Please refer to Figure 9 , is a structural diagram of another document parsing device provided in an embodiment of the present application, such as Figure 9 As shown, the document parsing device 900 may include: at least one processor 901, such as a CPU, at least one communication interface 903, a memory 904, and at least one communication bus 902. The communication bus 902 is used to realize the connection and communication between these components. The communication interface 903 may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface or a Bluetooth interface, etc.). The memory 904 may be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk memory. The memory 904 may optionally also be at least one storage device located away from the aforementioned processor 901. As Figure 9 As shown, the memory 904 as a computer storage medium may include an operating system, a network communication module, and program instructions.
[0173] exist Figure 9 In the document parsing device 900 shown, the processor 901 may be configured to load program instructions stored in the memory 904 and specifically perform the following operations:
[0174] Obtaining a presentation to be parsed, wherein the presentation to be parsed includes at least two presentation pages;
[0175] Based on the elements contained in the at least two demonstration pages, first parsed data containing target structure data is determined, and the target structure data includes a data structure relationship between a reference title and a reference sub-element; wherein, the reference title is a title element contained in the at least two demonstration pages, and the sub-elements of the reference title are located after the reference title in the visual presentation order, and the sub-elements include at least one of the following elements: sub-title, text, picture, table, link, the title level of the sub-title is lower than that of the reference title, and the reference sub-element contains elements belonging to a different demonstration page from the reference title.
[0176] It should be noted that the specific execution process can be found in the specific description of the above method embodiment and will not be described in detail here.
[0177] The specific execution steps can be found in the description of the aforementioned method embodiment and will not be described in detail here.
[0178] An embodiment of the present application also provides a computer storage medium, which can store multiple instructions, and the instructions are suitable for being loaded by a processor and executing the document parsing method provided by the embodiment of the present application. The specific execution process can be found in the specific description of the method embodiment shown above, and will not be described in detail here.
[0179] An embodiment of the present application further provides a computer program product comprising instructions, which, when executed on an electronic device, enables the electronic device to execute the method steps of the method embodiment shown above.
[0180] An embodiment of the present application also provides a chip module, including a transceiver component and a chip, wherein the chip is used to execute the method steps of the method embodiment shown above.
[0181] It is understood that the document parsing device, computer storage medium, computer program, computer program product, and chip provided above are all used to execute the method shown in any implementation of the corresponding aspects of the embodiments of this application. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the corresponding methods and will not be described in detail here.
[0182] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it includes the processes of the embodiments of the above-mentioned methods.
[0183] The at least one (item) involved in this application indicates one (item) or more (items). More than one (item) refers to two (items) or more than two (items). "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. In addition, it should be understood that although the terms first, second, etc. may be used to describe each object in this application, these objects should not be limited to these terms. These terms are only used to distinguish each object from each other.
[0184] As mentioned above, the terms "includes" and "having" and any variations thereof, are intended to cover a non-exclusive inclusion.
[0185] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A document parsing method, characterized in that: The method comprises: Obtaining a presentation to be parsed, wherein the presentation to be parsed includes at least two presentation pages; Determining first parsed data including target structure data based on elements included in the at least two presentation pages, the target structure data including a data structure relationship between a reference title and a reference sub-element; The reference title is a title element contained in the at least two demonstration pages, and the sub-elements of the reference title are located after the reference title in the visual presentation order. The sub-elements include at least one of the following elements: sub-title, text, picture, table, link. The title level of the sub-title is lower than that of the reference title, and the reference sub-element contains elements belonging to a different demonstration page from the reference title.
2. The method according to claim 1, wherein Before determining the first parsed data containing the target structure data, the method further includes: Obtaining elements included in the at least two demonstration pages in a visual presentation order; The obtaining of the elements contained in the at least two demonstration pages includes: Obtaining first text content contained in a first text box on a first page, and determining whether the first text content meets a preset title condition, wherein the first page is any one of the at least two demonstration pages; If it is determined that the first text content meets the preset title condition, adding a first preset tag to the first text content, where the first preset tag is used to indicate that the first text content belongs to a title element; When it is determined that the first text content does not meet the preset title condition, the first text content is determined to be the main text data content corresponding to the first title, the first title is the title that is closest to the first text content and precedes the first text content in the visual presentation order, the page to which the first title belongs is the first page, or, the page to which the first title belongs is before the first page in the visual presentation order.
3. The method according to claim 2, wherein The preset title condition includes that the font attribute of the text contains the title feature, and the text length is less than or equal to a preset length threshold.
4. The method according to claim 2 or 3, wherein: The step of obtaining the elements included in the at least two demonstration pages further includes: Obtain a first image contained in a first page, convert the first image into a lightweight markup language markdown syntax, and determine that the first image is image data content corresponding to a second title, the second title is a title that is closest to the first image and precedes the first image in a visual presentation order, and the page to which the second title belongs is the first page, or, in a visual presentation order, the page to which the second title belongs is before the first page.
5. The method according to any one of claims 1 to 4, characterized in that The determining, based on the elements included in the at least two presentation pages, first parsed data including target structure data comprises: generating second parsed data based on elements included in the at least two demonstration pages, wherein the elements in the second parsed data correspond sequentially to the elements included in the at least two demonstration pages in a visual presentation order, and the title level of any title element included in the second parsed data is a preset initial level; determining a first title set based on the second parsed data, the first title set being a title set included in the second parsed data; Obtaining a second title set based on the first title set, a preset prompt word, and a large language model, wherein the preset prompt word is used to instruct the large language model to classify the titles in the first title set into title levels while maintaining the semantic meaning of the titles, wherein the title levels include at least two levels; If it is determined that a first edit distance between the text of the first output title and the text of the first original title is less than or equal to a preset distance threshold, replacing a preset initial level of the first original title in the second parsed data with a title level corresponding to the first output title, thereby obtaining updated second parsed data, wherein the first original title is any one title in the first title set, and the first output title is any one title in the second title set; The first parsed data is determined based on the updated second parsed data.
6. The method according to any one of claims 1 to 5, characterized in that There is a second preset mark between the elements of the adjacent pages in the first parsed data, and the second preset mark is used to distinguish different elements of the adjacent pages.
7. The method according to any one of claims 1 to 6, wherein: There is a second preset marker between elements of the adjacent pages in the first parsed data, and the second preset marker is used to distinguish different elements of the adjacent pages. After determining the first parsed data containing the target structure data, the method further includes: Determining N title elements included in the elements of the same presentation page based on the second preset mark; When N is greater than or equal to 2, retain the target title with the highest title level among the N title elements, and set the other titles among the N titles except the target title as the body elements of the target title, and update the first parsed data.
8. The method according to any one of claims 1 to 7, wherein: The obtaining of the elements contained in the at least two demonstration pages includes: When the attribute of the second text box in the first page is a header attribute, adding a third preset mark to the second text content of the second text box, wherein the third preset mark is used to indicate that the second text content belongs to the header content; When the attribute of the third text box in the first page is a footer attribute, adding a fourth preset mark to the third text content of the third text box, wherein the fourth preset mark is used to indicate that the third text content belongs to the footer content; After determining the first parsed data including the target structure data, the method further includes: If it is determined that all elements of each presentation page contain target header content and the target header content is not recognized as a title, deleting all target header content originally contained in the first parsed data, adding the target header content to the end of the first parsed data, and setting the target header content as a title element of a target level, the target level being lower than the title level of other formal title elements contained in the first parsed data; When it is determined that the elements of each demonstration page include target footer content and the target footer content is not identified as a title, all the target footer content originally included in the first parsed data is deleted, and the target footer content is added to the end of the first parsed data, and the target footer content is set to a title element belonging to the target level.
9. A document parsing device, characterized in that: Comprising means for performing the method according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store a computer program. When the computer program is executed, the method according to any one of claims 1 to 8 is performed.