Analyzing and processing method and system for chapter titles of unstructured document
By combining embedded directory information, layout analysis model and text presentation rules, the accuracy and robustness problems of chapter title recognition in unstructured documents are solved, and more efficient chapter title parsing and restoration are achieved.
Patent Information
- Application Number
- CN202510707308.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When dealing with complex and changeable unstructured documents, existing document structure parsing technologies have insufficient accuracy, poor robustness and limited applicability in chapter title recognition.
By combining embedded catalog information parsing, layout analysis models and text presentation rules, multiple information sources are integrated to identify chapter titles, a pre-set chapter title fusion strategy is used to select a list of candidate chapter titles, and the chapter hierarchical structure is restored based on the title attribute information.
The accuracy and robustness of chapter title parsing are improved, making it applicable to more types of unstructured documents and overcoming the limitations of a single method.
Smart Images

Figure CN120654684A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a method and system for analyzing and processing chapter titles of unstructured documents. Background Art
[0002] In today's world, unstructured documents, as a common information carrier, contain vast amounts of knowledge and data. However, their diverse formats and lack of a unified structure make it difficult for machines to automatically process and understand their content. Section headings in unstructured documents are crucial for understanding document structure, enabling rapid navigation, and information retrieval. Accurately and efficiently parsing and restoring the section heading structure from various unstructured documents is crucial for increasing their usefulness.
[0003] At present, existing document structure parsing technologies often face challenges such as insufficient accuracy, poor robustness, and limited applicability when processing complex and changeable unstructured documents. Especially when identifying chapter titles, due to the actual complex and diverse document formats and content, a single parsing method is difficult to meet the actual parsing needs.
[0004] Therefore, how to achieve more accurate chapter title parsing and restoration for different types of unstructured documents is an urgent problem to be solved. Summary of the Invention
[0005] In view of this, the main purpose of the present invention is to provide a method and system for analyzing and processing chapter titles of unstructured documents, by fusing the information sources obtained by parsing the embedded directory information, the layout analysis model and the text presentation rules, and using them for parsing and restoring chapter titles, thereby effectively solving the problems of insufficient accuracy, poor robustness and limited scope of application in recognizing chapter titles of unstructured documents.
[0006] In a first aspect, an embodiment of the present application provides a method for analyzing and processing chapter titles of unstructured documents, comprising the following steps: For unstructured documents to be processed, title recognition is performed using embedded directory information parsing, layout analysis model and text presentation rules to obtain a corresponding list of candidate chapter titles; Selecting the corresponding candidate chapter title list based on a preset chapter title fusion strategy, and determining the target chapter title list using the results in the corresponding candidate chapter title list; The chapter hierarchical structure of the unstructured document is restored according to the title attribute information in the target chapter title list.
[0007] In some embodiments, for the unstructured document to be processed, title recognition is performed using embedded directory information parsing, layout analysis model, and text presentation rules to obtain a corresponding candidate chapter title list, including: Checking whether the unstructured document to be processed has embedded directory information, and after checking the embedded directory information, extracting chapter title text, title level and position information to obtain a candidate chapter title list based on the directory; Inputting the unstructured document to be processed into a layout analysis model, and obtaining a model-based candidate chapter title list based on text, image and layout information in the unstructured document to be processed; The text and text attributes in the unstructured document to be processed are extracted, and a rule-based candidate section title list is obtained from the text and text attributes using text presentation rules.
[0008] In some embodiments, extracting text and text attributes from the unstructured document to be processed, and identifying a rule-based candidate section title list from the text and text attributes using text presentation rules, includes: Obtaining the unstructured document to be processed; Determine the specific type of the unstructured document to be processed and select a corresponding tool to extract text and text attributes; A rule-based candidate chapter title list is obtained from the text and text attributes using pre-set text presentation rules.
[0009] In some embodiments, the text presentation rules are set based on the format characteristics and text character characteristics of the chapter title.
[0010] In some embodiments, when determining the specific type of the unstructured document to be processed and selecting the corresponding tool to extract text and text attributes, if the unstructured document to be processed is a scanned type, an OCR recognition tool is used for extraction; if the unstructured document to be processed is a text type, a PDF extraction tool is used for extraction.
[0011] In some embodiments, the chapter title fusion strategy includes: When the embedded directory information exists and is valid, the directory-based candidate chapter title list in the candidate chapter title list is used as the basic result, and can be supplemented by the model-based candidate chapter title list and / or the rule-based candidate chapter title list; When the embedded catalog information does not meet the application standard, the target chapter title list can be determined based on the results of the model-based candidate chapter title list and / or the rule-based candidate chapter title list.
[0012] In some embodiments, the chapter title fusion strategy further includes: When there is a conflict between the model-based candidate chapter title list and the rule-based candidate chapter title list, a preset conflict resolution logic is used to perform conflict resolution.
[0013] In some embodiments, when restoring the chapter hierarchical structure of the unstructured document based on the title attribute information in the target chapter title list, the hierarchical relationship between the chapter titles is determined based on the text content, document presentation order and format attributes of the target chapter title list, and a tree-like chapter structure is constructed to restore the chapter hierarchical structure of the unstructured document.
[0014] In a second aspect, an embodiment of the present application provides a system for analyzing and processing chapter titles of unstructured documents, including: The title recognition module is configured to perform title recognition on the unstructured document to be processed by using embedded directory information parsing, layout analysis model and text presentation rules to obtain a corresponding candidate chapter title list; an information fusion module configured to select the corresponding candidate chapter title list based on a preset chapter title fusion strategy, and determine a target chapter title list using the results in the corresponding candidate chapter title list; The structure restoration module is configured to restore the chapter hierarchical structure of the unstructured document according to the title attribute information in the target chapter title list.
[0015] In some embodiments, the title identification module includes: An embedded directory parsing module is configured to check whether the unstructured document to be processed has embedded directory information, and after checking the embedded directory information, extract the chapter title text, title level and position information to obtain a candidate chapter title list based on the directory; a model reasoning module configured to input the unstructured document to be processed into a layout analysis model, and obtain a model-based candidate chapter title list based on text, image and layout information in the unstructured document to be processed; The rule matching module is configured to extract text and text attributes from the unstructured document to be processed, and use text presentation rules to identify and obtain a rule-based candidate chapter title list from the text and text attributes.
[0016] Technical effects of the present invention: This application improves the accuracy of chapter title parsing by integrating multiple information sources. Specifically, it includes: for the unstructured documents to be processed, using embedded directory information parsing, layout analysis model and text presentation rules to perform title recognition to obtain a corresponding candidate chapter title list; selecting a corresponding candidate chapter title list based on a pre-set chapter title fusion strategy, and using the results in the corresponding candidate chapter title list to determine the target chapter title list; restoring the chapter hierarchical structure of the unstructured document based on the title attribute information in the target chapter title list. It can be seen that by combining the three different types of information sources identified based on embedded directory information, text presentation rules and layout analysis models, and using multi-source evidence for cross-validation and supplementation, the problem of missed recognition or misidentification that may exist in a single method can be effectively solved, thereby improving the overall accuracy of chapter title parsing and restoration.
[0017] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the disclosure and, together with the description, serve to explain the principles of the disclosure.
[0019] Figure 1 A flowchart illustrating a method for analyzing and processing chapter titles of unstructured documents according to an embodiment of the present application is shown; Figure 2 A flowchart illustrating a method for analyzing and processing section titles of unstructured documents according to another embodiment of the present application is shown; Figure 3 A schematic block diagram of a system for analyzing and processing chapter titles of unstructured documents according to an embodiment of the present application is shown; Figure 4 The figure shows a schematic block diagram of the structure of a title recognition module in a system for analyzing and processing chapter titles of unstructured documents according to an embodiment of the present application. DETAILED DESCRIPTION
[0020] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.
[0021] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.
[0022] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.
[0023] For unstructured documents, due to the lack of a predefined data model or organizational structure, identifying section headings is more complex than for structured documents, and the accuracy and robustness of the recognition results are relatively poor. However, section headings in documents often serve as key clues for understanding document structure and enabling fast navigation. Therefore, accurately and efficiently parsing section headings from unstructured documents and improving their usefulness is a key research issue in the industry.
[0024] like Figure 1 As shown, in a first aspect, an embodiment of the present application provides a method for analyzing and processing chapter titles of unstructured documents, which runs on a computer system and includes the following steps: S100: for the unstructured document to be processed, title recognition is performed using embedded directory information parsing, layout analysis model and text presentation rules to obtain a corresponding candidate chapter title list; In this step, after obtaining the unstructured document to be processed, check whether there is embedded directory information, i.e., ToC information, and parse it if it exists. Extract the metadata of the chapter title from it as the initial, more credible candidate title set, i.e., the candidate chapter title list based on the directory.
[0025] In addition, the system uses layout analysis models and text rendering rules to identify titles in unstructured documents. Specifically, after inputting the acquired unstructured document into the layout analysis model, the model infers the location of the title based on the document's content elements, generating a list of candidate titles based on model recognition. Predefined text rendering rules are then used to match the unstructured document, analyzing the pattern of the title text within the document to generate a rule-based list of candidate chapter titles.
[0026] It should be noted that most unstructured documents can be converted to PDF for further processing, but not all PDF files contain ToC metadata that complies with the specifications. For example, many PDF files generated by scanning or simple conversion lack this structured data. At the same time, for unstructured documents in non-PDF formats, such as scans of Word documents, plain text, and image formats, or documents with incorrect or incomplete ToC information, it is difficult to parse out the ToC information. That is, the use of embedded directory parsing to identify chapter titles in unstructured documents depends on the existence, accuracy, and completeness of the ToC. For unstructured documents that do not have ToC information or have poor ToC information quality, it is difficult to accurately identify chapter titles using this method. Therefore, in this step, on the basis of using embedded directory parsing to perform title recognition, layout analysis models and text presentation rules are also used for title recognition to solve the problems of insufficient accuracy, poor robustness, and limited scope of application of a single method in the parsing and restoration of chapter titles in unstructured documents.
[0027] S200, selecting a corresponding candidate chapter title list based on a pre-set chapter title fusion strategy, and determining a target chapter title list using the results in the corresponding candidate chapter title list; In this step, the chapter title fusion strategy is used to select a list of candidate chapter titles for each specific situation to generate the final chapter title list. This cross-validates and supplements information sources of different natures based on the actual situation, reducing the potential for missed or misidentified results from a single approach.
[0028] Specifically, for parsing chapter titles using embedded directory information, by checking and parsing the embedded directory information that meets the specifications, a list of candidate chapter titles based on the directory can be obtained, that is, a ToC candidate title list. For layout analysis models built using machine learning or deep learning technology, after training on documents with the title position and content marked, chapter titles can be identified from unstructured documents to obtain a model-based list of candidate chapter titles, that is, a model-recognized candidate title list. For chapter title recognition in unstructured documents using text presentation rules, a rule-based list of candidate chapter titles is obtained by matching the format and text features of the title with predefined text presentation rules, that is, a rule-recognized candidate title list.
[0029] When determining the final list of chapter titles, the corresponding candidate chapter title list can be selected according to the actual situation, and the results in the corresponding list can be used as the basis to summarize all the adopted titles, thereby obtaining the final list. That is, this step can meet more actual situations, and can adopt any one of the ToC candidate title list, model recognition candidate title list or rule recognition candidate title list to generate the final list, or select at least two of the lists to generate the final list. When a certain information source is unavailable or ineffective, such as when there is no embedded directory information to be checked, or there is an error or missing in the embedded directory information, thereby resulting in the inability to obtain or difficulty in obtaining an accurate candidate chapter title list based on the directory, the results in the rule-based candidate chapter title list and / or the model-based candidate chapter title list can be used to determine the final target chapter title list. Alternatively, when only ToC information exists and is valid, and there is no recognition in the model and the rules, the results in the ToC candidate title list can be directly adopted. In other words, compared to a single method, in this step, the corresponding results can be selected according to the specific circumstances of each candidate title list, further ensuring the generation of the final target candidate chapter title list.
[0030] S300: Restore the chapter hierarchical structure of the unstructured document according to the title attribute information in the target chapter title list.
[0031] In this step, the hierarchical relationship between chapter titles in the unstructured document can be inferred through the title attribute information of the target chapter title list, thereby restoring the chapter hierarchical structure of the unstructured document. It should be noted that title recognition is the basis for restoring the document structure. The more accurate the result of title recognition is, the more accurate the restored document hierarchical structure will be. This step utilizes the target chapter title list obtained in step S200. This list combines three different types of information sources from ToC, rules, and models, and uses multi-source evidence for cross-validation and supplementation, thereby reducing the missed recognition and misrecognition that may be caused by a single method, thereby greatly improving the overall accuracy of chapter title parsing. The chapter hierarchical structure restored based on this is more consistent with the chapter structure of the original unstructured document.
[0032] In some embodiments, for the unstructured document to be processed, title recognition is performed using embedded directory information parsing, layout analysis model, and text presentation rules to obtain a corresponding candidate chapter title list, including: Checking whether the unstructured document to be processed has embedded directory information, and after checking the embedded directory information, extracting the chapter title text, title level and position information to obtain a candidate chapter title list based on the directory; Inputting the unstructured document to be processed into the layout analysis model, and obtaining a model-based candidate chapter title list according to text, image and layout information in the unstructured document to be processed; The text and text attributes in the unstructured document to be processed are extracted, and a rule-based candidate chapter title list is obtained from the text and text attributes using text presentation rules.
[0033] In this embodiment, when using embedded directory information for title recognition, first use a parsing tool such as pymupdf to check whether there is embedded ToC information in the document. If so, the ToC information is parsed, the ToC information is read, and the chapter title text, title hierarchy and position information, such as page numbers, are extracted. This is used as the initial, highly reliable candidate title set, that is, a directory-based candidate chapter title list containing text, hierarchy and position.
[0034] When using the layout analysis model for title recognition, the unstructured document to be processed is input into a layout analysis model that has been trained or specifically trained, such as the LayoutLMv3 model. By training on documents with the title position and content marked, the model learns to distinguish between titles and non-title content from the visual features or text features of the text blocks. Among them, visual features include position, shape and size, and text features include words, characters and n-grams. The trained layout analysis model can combine the text, image and layout information in the unstructured document to infer the location of the title, thereby obtaining a model-based candidate section title list containing text, position and confidence. It should be noted that in addition to the layout analysis model, this embodiment can also use a graph neural network-based model to construct the relationship between document elements for title recognition, thereby obtaining a model-based candidate section title list.
[0035] When using text rendering rules for title recognition, title information is identified from the text and text attributes extracted from unstructured documents through a set of predefined text rendering rules, that is, potential section titles are identified by matching predefined rules, thereby obtaining a rule-based candidate section title list.
[0036] It should be noted that in the process of parsing the chapter titles of unstructured documents, at least two candidate chapter title lists are cross-validated and supplemented, based on the rule-based candidate chapter title list, the directory-based candidate chapter title list, and the model-based candidate chapter title list, to obtain the final target chapter title list. When a certain information source is unavailable or ineffective, other information sources can still provide effective clues, making this embodiment capable of handling unstructured documents of various types and qualities, and more robust and adaptable.
[0037] In some of the embodiments, text and text attributes are extracted from an unstructured document to be processed, and a rule-based list of candidate chapter titles is obtained from the text and text attributes using text rendering rules, including: obtaining an unstructured document to be processed; determining the specific type of the unstructured document to be processed, and selecting corresponding tools to extract text and text attributes; and using pre-set text rendering rules to identify a rule-based list of candidate chapter titles from the text and text attributes.
[0038] In this embodiment, when using text presentation rules to identify chapter titles of unstructured documents to be processed, it is first necessary to determine the specific type of the unstructured document, and use different tools to extract plain text and text attributes for different types of documents. Among them, text attributes include the font, size, position, and whether it is centered. After the text and text attributes are extracted, the pre-set text presentation rules are used to identify chapter titles from the text and text attributes. That is, this embodiment can select appropriate tools to extract text and attributes based on the differences in the structure and content storage methods of unstructured documents to meet different processing requirements, thereby obtaining chapter title information more accurately. And combined with the pre-set rules, it is possible to identify which text content is the chapter title, thereby obtaining a more accurate list of candidate chapters, providing an accurate basis for the subsequent final chapter title list.
[0039] In some of these embodiments, text presentation rules are set based on format features and text character features of the chapter title.
[0040] In this embodiment, the setting of text presentation rules is based on the common format features of the title and the text character features, including the text conforming to a specific pattern and the text format conforming to a specific pattern. For example, the text conforms to a specific pattern, such as "Chapter X", "Part X", "XY", and the chapter title usually has specific words and phrases, is relatively short in length, and is relatively concise in expression. By setting rules based on these character features, the recognition of the chapter title can be refined; the text format conforms to the pattern of the feature, such as "font size larger than a specific threshold of the surrounding text", "bold font", "specific indentation or centering format", by setting text presentation rules based on format features, the chapter title can be accurately distinguished from the text content, thereby obtaining a rule-based list of candidate chapter titles.
[0041] In some embodiments, when determining the specific type of the unstructured document to be processed and selecting the corresponding tool to extract text and text attributes, if the unstructured document to be processed is a scanned type, an OCR recognition tool is used for extraction; if the unstructured document to be processed is a text type, a PDF extraction tool is used for extraction.
[0042] In this embodiment, Figure 2 As shown, different tools are used to extract text for different types of unstructured documents. Specifically, after obtaining the unstructured document, it is first determined whether the document is an editable text type or a scanned type. Among them, for editable text-type documents, PDF extraction tools such as pdfplumber or PyMuPDF are used to extract them as plain text; for scanned documents, OCR recognition tools are used to extract them as plain text. It can be seen that the analysis and processing method of this embodiment is not limited to unstructured documents of a specific format, such as PDF documents with ToC, but is also applicable to PDF documents and Word documents without ToC, and can even be combined with OCR technology to be applied to scanned image documents, thereby expanding the scope of application to meet more practical needs.
[0043] In some embodiments, the chapter title fusion strategy includes: When the embedded directory information exists and is valid, the directory-based candidate chapter title list in the candidate chapter title list is used as the basic result, and can be supplemented by the model-based candidate chapter title list and / or the rule-based candidate chapter title list; When the embedded bibliography information does not meet the application standard, the target chapter title list can be determined based on the results of the model-based candidate chapter title list and / or the rule-based candidate chapter title list.
[0044] In this embodiment, a chapter title fusion strategy is adopted to fuse the embedded directory information, the layout analysis model, and the text presentation rule identification to determine the final chapter title list. Specifically, when setting the chapter title fusion strategy, it is first determined whether the ToC information in the unstructured document exists and is valid. If so, the result of the candidate chapter title list based on the directory obtained by parsing the ToC information is preferentially adopted. At the same time, the result of the candidate chapter title list based on the directory is supplemented by combining the specific circumstances of the candidate chapter title list based on the rule and / or the candidate chapter title list based on the model.
[0045] The reason is that if the ToC information exists and is accurate, it is relatively direct and accurate to read it directly and extract the chapter titles and their hierarchical relationships defined therein. In one scenario, if the unstructured document being identified has a problem with irregular document format, that is, the rule-based candidate chapter title list may be unavailable or ineffective, the directory-based candidate chapter title list and the rule-based candidate chapter title list can be combined to determine the final chapter title list, so that when the heuristic rules are limited by format interference and the document rule matching fails, the remaining two candidate chapter title lists can still be used to identify the chapter titles, thereby ensuring the accuracy of the results; in another scenario, if there is a layout for which the model has not been trained, that is, the model's generalization ability may be insufficient, and the model-based candidate chapter title list may be unavailable or ineffective, the directory-based candidate chapter title list and the rule-based candidate chapter title list can be combined to determine the final chapter title list, so that when the model's generalization ability is insufficient, the remaining two candidate chapter title lists can still be used to identify the chapter titles, thereby preventing recognition failures caused by the failure of a single method.
[0046] If it is determined that the ToC information in the unstructured document does not meet the application standard, that is, the ToC information does not exist, the ToC information is invalid, or the ToC information is partially missing, then the results of the rule-based candidate chapter title list and / or the model-based candidate chapter title list are selected according to different situations to determine the final target chapter title list, overcoming the complete reliance on the ToC information embedded in the unstructured document and effectively parsing the title when the ToC is missing or inaccurate. It can be seen that the fusion strategy of this embodiment effectively overcomes the limitations of a single information source, integrating three different types of information sources: document-embedded directory information, text presentation rules, and layout analysis models, to achieve chapter title parsing and subsequent restoration.
[0047] In some embodiments, the chapter title fusion strategy further includes: When there is a conflict between the model-based candidate chapter title list and the rule-based candidate chapter title list, a preset conflict resolution logic is used to perform conflict resolution.
[0048] In this embodiment, when conflicts arise within the candidate title list generated based on rule matching and model prediction, a pre-defined conflict resolution logic is employed to determine the target chapter title list. Specifically, this conflict resolution logic includes prioritization, contextual checks, and comparisons of similar title formats and model confidence. This multi-faceted conflict analysis and judgment reduces errors caused by single-factor judgments, resulting in more reliable results.
[0049] The following is a detailed explanation of the chapter title fusion strategy and conflict resolution logic.
[0050] 1. Input the candidate chapter title list based on the table of contents (ToC), the candidate chapter title list based on the model, and the candidate chapter title list based on the rule; The above list is specifically: (1) T_toc: ToC candidate title list [{text, level, page, source: 'ToC', id: 'toc_id'}, ...]; In the above content, text is the title text content, level is the title level, page is the page number corresponding to the title, source: 'ToC' identifies the data source as a table of contents, which is used to distinguish title information from other sources, and id: 'toc_id' is a unique identifier for each title; (2) T_model: model recognition candidate title list [{text, location_bbox, confidence, source: 'model', id: 'model_id'}, ...]; In the above content, text is the candidate title text content recognized by the model, location_bbox is the location coordinates of the title in the document, confidence is the model's confidence in the recognition result, source: 'model' indicates that the data source is model recognition, and id: 'model_id' is the unique identifier for each recognition result; (3) T_rule: rule identification candidate title list [{text, location_info, format_features, source: 'rule', id: 'rule_id'}, ...]; In the above content, text is the candidate title text content identified by the rule, location_info is the location information of the title in the document, format_features is the format features of the title, source: 'rule' identifies the data source as rule matching, and id: 'rule_id' is the unique identifier of each recognition result; It should also be noted that location_info includes page number, bbox (the border coordinates of the title on the page) and line number, and format_feature includes font size, boldness, centering and numbering mode.
[0051] 2. Preprocessing and alignment Specifically include: (1) Text normalization: Clean the text of all candidate titles, such as removing extra spaces, standardizing punctuation, and converting specific characters, to facilitate subsequent comparison; (2) Spatial alignment: The purpose is to determine whether candidate titles from different sources refer to the same physical text block in the document; Specifically, the spatial overlap between candidate titles from different sources is calculated, such as the bounding box-based IoU, and a threshold is set. For example, if IoU>0.8, it is considered to point to the same potential title above this threshold; a "fused candidate cluster" is created for each independent potential title at each physical location.
[0052] 3. Fusion decision engine (for each "fusion candidate cluster" or independent candidate title) When the ToC information exists and is valid: (1) Unique to ToC: There is a ToC candidate for a certain position, but there is no corresponding identification in the model or rules. In this case, the ToC candidate is adopted, provided that its page number is valid and the text is reasonable.
[0053] (2) The ToC corresponds to a single other source: that is, there is a ToC candidate at a certain position, and it corresponds to any one of the models or rules; When the ToC corresponds to the model but not to the rules, in this case, when making a specific ruling, the hierarchy of the ToC will be given priority. For the texts in the two documents, if the texts are highly similar, such as "Introduction" vs. "1. Introduction", and the model confidence is high, the text with a more complete model or numbered text can be selected; if the texts of the two documents are very different, the text corresponding to the ToC will be selected first. In other words, the ToC is the main focus, and the model is used to confirm or fine-tune the text; When the ToC matches the rules but not the model, the ToC hierarchy takes precedence when making a specific ruling. For text within the ToC, if the two texts are highly similar and the formatting features matched by the rules, such as numbering and font, match the ToC hierarchy, the rule will be selected to identify the more standardized or numbered text in the candidate title list. In other words, the ToC is the primary focus, and the rules are used to confirm or fine-tune the text or numbering.
[0054] (3) The ToC corresponds to both sources: that is, there is a ToC at a certain location, and it corresponds to both the model and the rules; When the three corresponding normalized texts and levels (explicit or inferred) are consistent, they are directly adopted, and the credibility is the highest at this time; When there are inconsistencies, the corresponding conflicts should be handled. When making specific decisions, for the hierarchy, the hierarchical information of the ToC is usually the highest priority. However, if the model and the rules consistently point to a different hierarchy than the ToC, and there is strong evidence, such as clear numbering, then it is necessary to evaluate the possibility of revising the ToC hierarchy, or mark it for manual review. For the text, if the three texts are different but have similar semantics, the numbered, more complete, and formatted text provided by the model or rule that is consistent with the title of the same level of the document is preferred. At this time, the model confidence is an important reference; if the model and the rule text are consistent, but different from the ToC, for example, the ToC lacks numbering, then the model / rule text is tended to be adopted, the hierarchy refers to the ToC, or is revised according to the numbering.
[0055] (4) The model and / or rules identify headings not included in the ToC In this case, it will be considered as a potential supplementary title in the specific ruling. Specifically, if the following conditions are met: its number can be logically embedded in the existing ToC structure, the format characteristics are consistent with the confirmed titles at the same level, and the source is credible (such as high model confidence and high rule strength), then it can be added as a supplementary title.
[0056] When ToC information is missing, invalid, or incomplete: (1) Neither the model nor the rule is recognized: that is, if neither the model nor the rule is recognized at a certain position, the position is deemed to have no title.
[0057] (2) Single-source identification: that is, only model identification without corresponding rules, or only rule identification without corresponding models; When only model identification is performed and there is no corresponding rule, the model's results will be adopted in the specific ruling, provided that the model confidence is high, such as >0.85, and its inferred numbering and format are consistent with other identified titles in the document (if any); When only rules are identified and there is no corresponding model, the result of the rules is adopted in the specific judgment, provided that strong rules such as "References" and "Chapter X" or specificity rules are matched and their format and numbering are consistent with the context.
[0058] (3) Both the model and the rules recognize the candidate title: that is, both the model and the rules recognize the candidate title (in the same or similar position) When the two results are consistent, if the normalized text and the inferred hierarchy through numbering / formatting are highly consistent, then it is adopted with high confidence; When there is a conflict between the two results, the following logic is used for execution: Rule numbers take precedence, i.e., if the rule identifies a clear numbering pattern, such as "3.1", and the model has no numbering or is irregular, the rule's text and inferred hierarchy will be adopted first, provided that the model also has a high confidence score at that position, i.e., the model confirms that this is likely to be a title; Models with high confidence are prioritized. That is, if the model confidence is very high, such as >0.95, and the rule only matches weak features, such as "centered but with a font size close to the main text", the model's result will be prioritized. Contextual coherence: Select the candidate whose numbering / content fits better into the contextual heading sequence.
[0059] 4. Final list generation and redundancy removal (1) Collection: Summarize the titles adopted in all steps; (2) Redundancy removal: remove completely duplicate titles based on normalized text and precise physical location, such as starting line number or closely overlapping bboxes; (3) Sorting: Sort the finalized title list according to the physical order of the titles in the document, such as page number or position within the page, to obtain the final list.
[0060] In some of the embodiments, when restoring the chapter hierarchy of an unstructured document based on the title attribute information in the target chapter title list, the hierarchical relationship between the chapter titles is determined based on the text content, document presentation order and format attributes of the target chapter title list, and a tree-like chapter structure is constructed to restore the chapter hierarchy of the unstructured document.
[0061] In this embodiment, the hierarchical relationship between titles is inferred based on the finalized list of chapter titles and their positions in the document. Specifically, based on the textual content of the titles, such as the numbers "1.", "1.1," "1.1.1," their order of appearance in the document, and possible information such as indentation and font size, the hierarchical relationship between titles is determined, a tree-like chapter structure is constructed, the title information is reconstructed, and a chapter structure that conforms to the original document's intent is output.
[0062] It can be seen from this that the analysis and processing method of the present application, by integrating multiple information sources, namely, the candidate chapter title catalog obtained by means of embedded directory information, text presentation rules and layout analysis models, is used for parsing and restoring chapter titles, combining ToC parsing, rule matching, model reasoning, information fusion and structure restoration to realize the analysis and processing of chapter titles of unstructured documents, overcome the limitations of a single information source, and improve the accuracy, robustness and universality of parsing for different types of unstructured documents.
[0063] like Figure 3 As shown, in a second aspect, an embodiment of the present application provides a system 10 for analyzing and processing chapter titles of unstructured documents, which is used to execute the above method, including: The title recognition module 100 is configured to perform title recognition on the unstructured document to be processed by using embedded directory information parsing, layout analysis model and text presentation rules to obtain a corresponding candidate chapter title list; The information fusion module 200 is configured to select a corresponding candidate chapter title list based on a preset chapter title fusion strategy, and determine a target chapter title list using the results in the corresponding candidate chapter title list; The structure restoration module 300 is configured to restore the chapter hierarchical structure of the unstructured document according to the title attribute information in the target chapter title list.
[0064] like Figure 4 As shown, in some embodiments, the title identification module includes: The embedded directory parsing module 110 is configured to check whether the unstructured document to be processed has embedded directory information, and after checking the embedded directory information, extract the chapter title text, title level and position information to obtain a candidate chapter title list based on the directory; The model reasoning module 120 is configured to input the unstructured document to be processed into the layout analysis model, and obtain a model-based candidate section title list based on the text, image and layout information in the unstructured document to be processed; The rule matching module 130 is configured to extract text and text attributes from the unstructured document to be processed, and identify a rule-based candidate chapter list from the text and text attributes using text presentation rules.
[0065] On the third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a storage medium and a computer program. The computer program is stored in the storage medium. When the computer program is executed by the processor, it implements any of the above-mentioned methods for analyzing and processing chapter titles of unstructured documents. For specific examples, please refer to the examples described in the above-mentioned embodiments and optional implementation methods, and this embodiment will not be repeated here.
[0066] Throughout this specification, references to terms such as "one embodiment," "example," or "specific example" indicate that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0067] Unless otherwise defined, the technical terms or scientific terms involved in this application should have the usual meanings understood by people with ordinary skills in the technical field to which this application belongs. The words "one", "the" and similar words involved in this application do not indicate quantity restrictions and can indicate the singular or plural. The terms "include", "comprise", "have" and any variations thereof involved in this application are intended to cover non-exclusive inclusions. The words "connect", "connected", "coupled" and similar words involved in this application are not limited to physical or mechanical connections, but include electrical connections, whether direct or indirect. The "multiple" involved in this application refers to two or more, and "and / or" describes the association relationship of associated objects, indicating that three relationships can exist. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. The terms "first", "second", "third" and the like involved in this application are merely to distinguish similar objects and do not represent a specific ordering of objects.
[0068] The above embodiments merely illustrate several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make several modifications or improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A method for analyzing and processing chapter titles of unstructured documents, characterized in that: The steps include: For unstructured documents to be processed, title recognition is performed using embedded directory information parsing, layout analysis model and text presentation rules to obtain a corresponding list of candidate chapter titles; Selecting the corresponding candidate chapter title list based on a preset chapter title fusion strategy, and determining the target chapter title list using the results in the corresponding candidate chapter title list; The chapter hierarchical structure of the unstructured document is restored according to the title attribute information in the target chapter title list.
2. The method for analyzing and processing chapter titles of unstructured documents according to claim 1, characterized in that: For the unstructured document to be processed, title recognition is performed by using embedded directory information parsing, layout analysis model and text presentation rules to obtain a corresponding candidate chapter title list, including: Checking whether the unstructured document to be processed has embedded directory information, and after checking the embedded directory information, extracting chapter title text, title level and position information to obtain a candidate chapter title list based on the directory; Inputting the unstructured document to be processed into a layout analysis model, and obtaining a model-based candidate chapter title list based on text, image and layout information in the unstructured document to be processed; The text and text attributes in the unstructured document to be processed are extracted, and a rule-based candidate section title list is obtained from the text and text attributes using text presentation rules.
3. The method for analyzing and processing chapter titles of unstructured documents according to claim 2, characterized in that: The step of extracting text and text attributes from the unstructured document to be processed and identifying a rule-based candidate chapter title list from the text and text attributes using text presentation rules includes: Obtaining the unstructured document to be processed; Determine the specific type of the unstructured document to be processed and select a corresponding tool to extract text and text attributes; A rule-based candidate chapter title list is obtained from the text and text attributes using pre-set text presentation rules.
4. The method for analyzing and processing chapter titles of unstructured documents according to claim 3, characterized in that: The text presentation rules are set based on the format features and text character features of the chapter title.
5. The method for analyzing and processing chapter titles of unstructured documents according to claim 3, characterized in that: When determining the specific type of the unstructured document to be processed and selecting the corresponding tool to extract text and text attributes, if the unstructured document to be processed is a scanned type, an OCR recognition tool is used for extraction; if the unstructured document to be processed is a text type, a PDF extraction tool is used for extraction.
6. The method for analyzing and processing chapter titles of unstructured documents according to claim 1, characterized in that: The chapter title fusion strategy includes: When the embedded directory information exists and is valid, the directory-based candidate chapter title list in the candidate chapter title list is used as the basic result, and can be supplemented by the model-based candidate chapter title list and / or the rule-based candidate chapter title list; When the embedded catalog information does not meet the application standard, the target chapter title list can be determined based on the results of the model-based candidate chapter title list and / or the rule-based candidate chapter title list.
7. The method for analyzing and processing chapter titles of unstructured documents according to claim 6, characterized in that: The chapter title fusion strategy also includes: When there is a conflict between the model-based candidate chapter title list and the rule-based candidate chapter title list, a preset conflict resolution logic is used to perform conflict resolution.
8. The method for analyzing and processing chapter titles of unstructured documents according to any one of claims 1 to 7, characterized in that: When restoring the chapter hierarchical structure of the unstructured document based on the title attribute information in the target chapter title list, the hierarchical relationship between the chapter titles is determined based on the text content, document presentation order and format attributes of the target chapter title list, a tree-like chapter structure is constructed, and the chapter hierarchical structure of the unstructured document is restored.
9. A system for analyzing and processing chapter titles of unstructured documents, characterized in that: include: The title recognition module is configured to perform title recognition on the unstructured document to be processed by using embedded directory information parsing, layout analysis model and text presentation rules to obtain a corresponding candidate chapter title list; an information fusion module configured to select the corresponding candidate chapter title list based on a preset chapter title fusion strategy, and determine a target chapter title list using the results in the corresponding candidate chapter title list; The structure restoration module is configured to restore the chapter hierarchical structure of the unstructured document according to the title attribute information in the target chapter title list.
10. The unstructured document chapter title analysis and processing system according to claim 9, characterized in that: The title recognition module includes: An embedded directory parsing module is configured to check whether the unstructured document to be processed has embedded directory information, and after checking the embedded directory information, extract the chapter title text, title level and position information to obtain a candidate chapter title list based on the directory; a model reasoning module configured to input the unstructured document to be processed into a layout analysis model, and obtain a model-based candidate chapter title list based on text, image and layout information in the unstructured document to be processed; The rule matching module is configured to extract text and text attributes from the unstructured document to be processed, and use text presentation rules to identify and obtain a rule-based candidate chapter title list from the text and text attributes.