High-performance document title hierarchy recognition system, method and equipment and medium
Through deep learning technology, pre-trained title classification model and title hierarchical model are used to solve the problems of title loss and hierarchical confusion in electronic documents, efficient document processing and structured data extraction are achieved, and the readability and retrievalability of the document are improved.
Patent Information
- Application Number
- CN202510179353.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art often faces the problems of title loss and title hierarchy confusion during the conversion of electronic documents or photocopy documents, resulting in inefficient readability, retrievalability and structured data extraction of documents.
Using deep learning methods, the pre-trained title classification model and title hierarchical model are used to automatically learn the characteristics and hierarchical relationships of the title from a large amount of data, and directly process the original document without cumbersome rule formulation and data processing.
The document processing process is simplified, the work efficiency is improved, and the complexity, low universality and low accuracy of rule recognition methods are avoided. The title classification model and title hierarchical model have certain generalization capabilities and can accurately judge unusual title forms.
Smart Images

Figure CN120126164A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence and data processing, and particularly to a high-performance document title hierarchy recognition system, method, device, and medium. Background Art
[0002] In the current process of document processing and information conversion, electronic documents or documents converted from photocopied documents generally face problems such as title loss and chaotic title hierarchies. These problems greatly affect the readability, retrievability, and the extraction efficiency of structured data of the documents.
[0003] First, when converting PDF documents formed through scanning and OCR (Optical Character Recognition) technology into structured data, there is often a lack of effective title classification and hierarchy division. For example, during the conversion of some PDF documents, due to the limitations of format recognition technology, all line contents may be misidentified as titles of the same level, resulting in the confusion of title information and the loss of the hierarchical structure. In this case, even if the title information in the original document is clear, the converted document may have phenomena such as titles like "Preface", "Abstract", "Summary", etc. being mixed with the body content and unable to be correctly distinguished.
[0004] Second, some original documents do not clearly mark titles and title levels during writing or typesetting. After these documents are converted into electronic format, due to the lack of clear hierarchical structure information, subsequent processing and utilization become particularly difficult. For example, some technical documents or regulatory documents may only use indentation or line breaks to distinguish different paragraphs and chapters without clear title markings, which makes it difficult to accurately divide and extract the content of each part when converting into structured data.
[0005] In response to the above problems, traditional rule-based methods have many limitations in title recognition and hierarchy division. On the one hand, rule formulation requires high technical requirements and professional knowledge, and conflicts are prone to occur between rules, resulting in the difficulty of continuously improving the accuracy of title recognition. On the other hand, rule-based methods often lack universality. Different document categories and formats may require the development of different extraction rules, which not only increases the time cost, but also in practical applications, due to the diversity and complexity of document types, the adaptability and flexibility of the rules are greatly restricted. Summary of the Invention
[0006] This application provides a high-performance document title hierarchy recognition system, method, device, and medium for solving the problems of title loss and chaotic title hierarchies existing in existing electronic documents or documents converted from photocopied documents.
[0007] In a first aspect, the present application provides a high-performance document title hierarchy recognition system, which includes:
[0008] A document receiving unit for obtaining a document to be recognized;
[0009] A title classification unit for tagging each title sequence in the document to be recognized with a title label through a pre-trained title classification model, so as to obtain a first text sequence carrying the title label;
[0010] A title hierarchy recognition unit for determining the title hierarchy corresponding to each title sequence based on the first text sequence through a pre-trained title hierarchy model, and tagging each title sequence with a title hierarchy label in the first text sequence, so as to obtain a second text sequence carrying the title hierarchy label;
[0011] A document output unit for, for each title sequence, determining the text style corresponding to the title sequence from the text styles corresponding to each pre-configured title hierarchy according to the title hierarchy label corresponding to the title sequence in the second text sequence; and organizing the second text sequence into a document with a title hierarchy according to the text styles corresponding to each title sequence and the pre-configured body text style.
[0012] In a second aspect, the present application further provides a high-performance document title hierarchy recognition method, which includes:
[0013] Obtaining a document to be recognized;
[0014] Tagging each title sequence in the document to be recognized with a title label through a pre-trained title classification model, so as to obtain a fourth text sequence carrying the title label;
[0015] Determining the title hierarchy corresponding to each title sequence based on the fourth text sequence through a pre-trained title hierarchy model, and tagging each title sequence with a title hierarchy label in the fourth text sequence, so as to obtain a fifth text sequence carrying the title hierarchy label;
[0016] For each title sequence, determining the text style corresponding to the title sequence from the text styles corresponding to each pre-configured title hierarchy according to the title hierarchy label corresponding to the title sequence in the fifth text sequence;
[0017] Organizing the fifth text sequence into a document with a title hierarchy according to the text styles corresponding to each title sequence and the pre-configured body text style.
[0018] In a third aspect, the present application provides a computer device, which includes a processor for implementing the steps of the high-performance document title hierarchy recognition method as described above when executing a computer program stored in a memory.
[0019] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the steps of the high-performance document title hierarchy recognition method as described above.
[0020] The beneficial effects of the present application are as follows:
[0021] The high-performance document title hierarchy recognition system adopts a deep learning method. Through a pre-trained title classification model and a title hierarchy model, it automatically learns the features and hierarchical relationships of titles from a large amount of data. The system can directly process the original document without the need for cumbersome rule-making and data processing. This simplifies the document processing process, improves work efficiency, and avoids many problems existing in the rule recognition method, such as complex extraction and processing procedures, low universality, and low accuracy. Moreover, both the title classification model and the title hierarchy model have a certain generalization ability, and even when encountering uncommon title forms, they can make accurate judgments based on the knowledge they have learned. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0023] Figure 1 It is a schematic structural diagram of a high-performance document title hierarchy recognition system provided by an embodiment of the present application;
[0024] Figure 2 It is a schematic process diagram of a high-performance document title hierarchy recognition provided by an embodiment of the present application;
[0025] Figure 3 It is a schematic structural diagram of a computer device provided by an alternative embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] In order to make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the present application in detail with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0027] In order to effectively identify the titles and title levels in electronic documents or documents converted from photocopied documents, the present application provides a high-performance document title level identification system, method, device and medium.
[0028] Example 1:
[0029] The present application provides a high-performance document title level identification system. Figure 1 FIG. is a schematic structural diagram of a high-performance document title level identification system provided by an embodiment of the present application. The system includes:
[0030] A document receiving unit 11, configured to obtain a document to be identified;
[0031] A title classification unit 12, configured to label each title sequence in the document to be identified with a title label through a pre-trained title classification model, so as to obtain a first text sequence carrying the title label;
[0032] A title level identification unit 13, configured to determine the title level corresponding to each title sequence based on the first text sequence through a pre-trained title level model, and label each title sequence with a title level label in the first text sequence, so as to obtain a second text sequence carrying the title level label;
[0033] A document output unit 14, configured to, for each title sequence, determine the text style corresponding to the title sequence from the text styles corresponding to each pre-configured title level according to the title level label corresponding to the title sequence in the second text sequence; and organize the second text sequence into a document with title levels according to the text styles corresponding to each title sequence and the pre-configured body text style.
[0034] The purpose of this high-performance document title level identification system is to accurately identify the title levels in the document and apply the identification results to the organization and output of the document, so as to improve the readability and standardization of the document. The system mainly consists of four parts: a document receiving unit 11, a title classification unit 12, a title level identification unit 13, and a document output unit 14. The following is an explanation of each unit:
[0035] I. Document receiving unit 11
[0036] The main function of the document receiving unit 11 is to obtain the document to be recognized. In practical applications, the document to be recognized can come from various data sources. For example, users can directly upload the documents stored locally through the system's user interface, and the formats of these documents can include but are not limited to common document formats such as.doc,.docx,.pdf,.txt, etc. The system can be compatible with different formats of documents through the corresponding file reading interface and convert them into an internal data format that the system can process.
[0037] In addition, the document to be recognized can also be obtained from network resources. The system supports downloading documents from a specified network address through network protocols (such as HTTP, FTP, etc.).
[0038] In a possible implementation manner, after obtaining the document to be recognized, the system will perform preliminary preprocessing on the document to be recognized to simplify the subsequent processing flow. Among them, the preprocessing includes but is not limited to: removing redundant characters and blank lines in the document.
[0039] II. Title classification unit 12
[0040] The title classification unit 12 classifies each title sequence in the document to be recognized through a pre-trained title classification model and assigns corresponding title labels, thereby obtaining a first text sequence carrying title labels. Among them, the title classification model can adopt a deep learning model, such as a convolutional neural network (CNN) or a recurrent neural network (RNN), and be trained in combination with natural language processing technology. For example, for a technical document containing multiple chapters and sections, the label classification model can accurately identify each chapter title and section title and assign different labels to them, such as "first-level title", "second-level title", etc. After being processed by the title classification unit 12, each title sequence in the document to be recognized is assigned a title label, forming a first text sequence carrying title labels.
[0041] In one example, the title classification unit 12 is specifically used for:
[0042] Segment the document to be recognized into sentences to generate a sentence sequence of the document to be recognized;
[0043] Classify each sentence sequence through the title classification model, identify the sentence sequences classified as title sequences; in each sentence sequence, assign title labels to the sentence sequences classified as title sequences to obtain the first text sequence.
[0044] After the title classification unit 12 obtains the document to be recognized, the document to be recognized can be segmented into sentences to generate a sentence sequence of the document to be recognized. During the sentence segmentation process, the title classification unit 12 can determine the boundaries of sentences based on punctuation marks (such as full stops, question marks, exclamation marks, etc.), line breaks, and specific language rules. For example, for Chinese documents, segmentation will be performed in combination with Chinese grammar structures and punctuation marks; for English documents, processing will be based on English punctuation and grammar rules. This can ensure that the document is accurately segmented into individual sentence sequences, providing a basis for subsequent title classification work.
[0045] Then, using the pre-trained title classification model, classify each generated sentence sequence. After the title classification model classifies each sentence sequence, it will identify the sentence sequences classified as title sequences. The title classification model can assign title labels to these sentence sequences classified as title sequences in each sentence sequence. For example, if the sentence sequence "Chapter 1 Overview of Artificial Intelligence" is recognized as a title by the model, the system will assign it the label "first-level title"; after "1.1 Definition of Artificial Intelligence" is recognized as a title, it will be assigned the label "second-level title". After this series of operations, the first text sequence with title labels is obtained.
[0046] In order to accurately identify the title sequences in the document to be recognized, in this application, it is necessary to pre-construct a title classification sample set to train the original title classification model through this title classification sample set. Exemplarily, collect a large number of documents with pre-labeled titles, and record these documents as the first text samples. All the first text samples together constitute the title classification sample set. These first text samples cover various types of documents, such as academic papers, news reports, technical manuals, business reports, etc. The purpose is to enable the title classification model to fully learn the characteristics of titles in documents of different fields and genres, including semantic, grammatical, and structural characteristics. By learning from the title classification sample set, the title classification model can possess a powerful judgment ability to accurately determine whether a text sequence is a title and assign the corresponding title label to it.
[0047] In one example, when annotating the first text sample, only the titles can be annotated, or both the titles and the text can be annotated. The following will separately describe the training process of the label classification model provided in this application for the two annotation situations:
[0048] Situation 1: Only annotate the title labels corresponding to each title in the first text sample.
[0049] For ease of understanding, the following is an annotation example for this situation two. In this annotation example, the format for annotating each sentence sequence in the first text sample is: {label} title sequence. For example:
[0050] {1}Chapter 1 General Provisions
[0051] {1}Section 1 General Requirements
[0052] {1}(1) Military Ships
[0053] The significant wave heights corresponding to the 5% guarantee rate are divided into three levels, A, B, and C, from high to low...
[0054] {1}(2) Fishing Vessels
[0055] Classification of navigation areas and sections...
[0056] {1}(3) Sports Boats
[0057] Classification of navigation areas and sections...
[0058] {1}(4) Yachts
[0059] New ship: It refers to a newly built ship for which a construction contract is signed on or after the effective date of this Code...
[0060] In this example, the label "1" represents that the sentence sequence is a title. This annotation method clearly distinguishes the sentence sequences corresponding to the titles, providing clear supervision information for model training.
[0061] During the training process, any first text sample is randomly selected from the title classification sample set. Each title in this sample has been pre-annotated with a title label, and the representation of the title label has diversity. It can be a number, for example, using "1" to represent that the text sequence is a title; it can also be a string, such as "title", etc. The specific representation method can be flexibly determined according to actual needs, and this application does not make specific limitations on this.
[0062] Input the first text sample into the original title classification model. Among them, the original title classification model can adopt a deep learning model, such as a convolutional neural network (CNN) or a recurrent neural network (RNN), etc. Through the original title classification model, the first text sample is processed to calculate the probabilities of each predicted title sequence in the first text sample. Specifically, the original title classification model will extract features from each sentence sequence in the first text sample, and then through its internal neural network structure, map these features to different title categories, and output the probability that each sequence belongs to each title category. Based on the probabilities of each predicted title sequence and each title label, the loss value is determined to train the original title classification model according to this loss value.
[0063] Case 2: Annotate the title labels corresponding to each title in the first text sample and the body labels corresponding to each body sequence.
[0064] For ease of understanding, the following is an annotation example for this Case 2. In this annotation example, the format for annotating each sentence sequence in the first text sample is: {tag} The first line content of the sentence sequence. For example:
[0065] {1}Chapter 1 General Provisions
[0066] {1}Section 1 General Provisions
[0067] {1}(1) Military ships
[0068] {0}The significant wave heights corresponding to the 5% guarantee rate are divided into three levels, A, B, and C, from high to low
[0069] {1}(2) Fishing vessels
[0070] {0}Classification of navigation areas and sections
[0071] {1}(3) Sports boats
[0072] {0}Classification of navigation areas and sections
[0073] {1}(4) Yachts
[0074] {0}New ship: It refers to a newly built ship for which the construction contract is signed on or after the effective date of this Code.
[0075] In this example, the tag "1" represents that the sentence sequence is a title, and the tag "0" represents that the sentence sequence is the main text. This annotation method clearly distinguishes between titles and the main text, providing clear supervision information for model training.
[0076] After inputting the first text sample into the original title classification model, the original title classification model not only calculates the probabilities of each predicted title sequence but also calculates the probabilities of each predicted main text sequence. Based on the probabilities of each predicted title sequence and their corresponding title tags, and at the same time combining the probabilities of each predicted main text sequence and their corresponding main text tags, the loss value is comprehensively determined. Then, based on this loss value, the original title classification model is trained, and by continuously adjusting the model parameters, the original title classification model can accurately distinguish between titles and the main text simultaneously, and finally, a trained title classification model is obtained.
[0077] Regardless of which of the above annotation cases is adopted, for each first text sample, the corresponding operations above need to be performed to obtain the loss value of the first text sample. Finally, based on the sum of the loss values of all first text samples in the current iteration, further parameter adjustment is performed on the title classification model in the current iteration. When the preset convergence condition is met, the training of the title classification model is completed.
[0078] Among them, meeting the preset convergence condition can be that the sum of the loss values determined based on each first text sample is less than a pre-configured loss threshold, or the determined loss value has been in a downward trend and tends to flatten out, or the number of iterations for training the title classification model reaches the set maximum number of iterations, etc. In specific implementations, it can be flexibly set and will not be specifically limited here.
[0079] As a possible implementation manner, when training the title classification model, the first text samples can be divided into training samples and test samples. First, the title classification model is trained based on the training samples, and then the reliability of the above-mentioned trained title classification model is verified based on the test samples.
[0080] III. Title Hierarchy Recognition Unit 13
[0081] After obtaining the first text sequence, the title hierarchy recognition unit 13 can, based on the first text sequence, use a pre-trained title hierarchy model to determine the title hierarchy corresponding to each title sequence. Among them, the title hierarchy model can also adopt a deep learning model, and by learning a large number of document data with clear title hierarchy annotations, it can master the relationships and characteristics between different title hierarchies.
[0082] Exemplarily, after obtaining the first text sequence, the title hierarchy recognition unit 13 can input the first text sequence into a pre-trained title hierarchy model. Through the processing of the input first text sequence by the title hierarchy model, combined with the context information of each title sequence, the hierarchy of each title sequence is determined, and a title hierarchy label is assigned to each title sequence in the first text sequence, so as to obtain a second text sequence carrying the title hierarchy label.
[0083] In a document processing system, accurately identifying the title hierarchy is crucial for clearly presenting the document structure and improving the readability of the document. To achieve this goal, it is necessary to effectively train the title hierarchy model. In this application, in order to train the original title hierarchy model, first, a representative title hierarchy sample set needs to be constructed. A large number of documents in different fields and different formats are collected and annotated to obtain text sequence samples carrying title labels, that is, in any text sequence sample, each title sequence carries a title label. Then, for each text sequence sample, the hierarchy of each title sequence in the text sequence sample is annotated so that each title sequence corresponds to a title hierarchy label (such as "first-level title", "second-level title", "third-level title", etc.).
[0084] In a possible implementation manner, these text sequence samples can cover various possible title hierarchy structures and text styles to ensure that the title hierarchy model can learn comprehensive title hierarchy features.
[0085] For example, the format for annotating the title hierarchy tags for any text sequence sample is: {label} Title features in the document. For example:
[0086] {1} Chapter 1 General Provisions
[0087] {2} Section 1 General Provisions
[0088] {3} 1.1.1 Scope of Application
[0089] {4} (1) Military ships
[0090] {4} (2) Fishing boats
[0091] {4} (3) Sports boats
[0092] {4} (4) Yachts
[0093] {1} Chapter 2 Hull Structure
[0094] {2} Section 1 Deck
[0095] {3} 2.1.1 General Definitions
[0096] {4} (1) Length between perpendiculars at full load.
[0097] Among them, the numbers in "{}" represent the levels of the titles.
[0098] Exemplarily, for any text sequence sample in the title hierarchy sample set, the selected text sequence sample is input into the original title hierarchy model. The original title hierarchy model processes the first text sample, extracts features, and maps these features to different title hierarchy categories through a neural network structure, outputting the probability distribution of each title sequence belonging to each title hierarchy category (denoted as the predicted title hierarchy probability distribution). For example, for a title sequence, the model may output that it has a 70% probability of being a first-level title, a 20% probability of being a second-level title, and a 10% probability of being a third-level title. Based on the predicted title hierarchy probability distributions corresponding to each title sequence and their corresponding title hierarchy tags respectively, the loss value is determined. According to this loss value, an optimization algorithm (such as stochastic gradient descent, Adam optimizer, etc.) is used to train the original title hierarchy model, adjusting the parameters of the model to gradually reduce the loss value.
[0099] Since several text sequence samples are obtained in the above embodiments, for each text sequence sample, corresponding operations are performed to obtain the loss value of the sample. In the current iteration, the loss values of all text sequence samples are added together to obtain the total loss value. Based on this total loss value, further parameter adjustment is performed on the title hierarchy model of the current iteration. For example, the title hierarchy recognition model is constructed by continuously adjusting the LSTM (Long Short-Term Memory) layer and the CRF (Conditional Random Field) layer of the original title hierarchy model. CRF is introduced into the original title hierarchy model to solve the problem of weak document context dependence in the model. When the preset convergence condition is met, it is considered that the training of the title hierarchy model is completed.
[0100] Among them, meeting the preset convergence condition can be that the sum of the loss values determined based on each text sequence sample is less than the pre-configured loss threshold, or the sum of the determined loss values has been in a downward trend and tends to be flat, or the number of iterations for training the title hierarchy model reaches the set maximum number of iterations, etc. In specific implementations, it can be flexibly set and will not be specifically limited here.
[0101] As a possible implementation manner, when training the title hierarchy model, the text sequence samples can be divided into training samples and test samples. First, the title hierarchy model is trained based on the training samples, and then the reliability of the above-mentioned trained title hierarchy model is verified based on the test samples.
[0102] In a possible implementation manner, the title classification model and the title hierarchy model are jointly trained in the following way:
[0103] Obtain any second text sample in the comprehensive text sample set; wherein, each title in the second text sample corresponds to a title hierarchy label.
[0104] Through the original title classification model, based on the second text sample, title labels are assigned to each predicted title sequence in the second text sample to obtain a third text sequence carrying the title labels.
[0105] Through the original title hierarchy model, based on the third text sequence, determine the title hierarchy probability distribution corresponding to each predicted title sequence.
[0106] Based on the title hierarchy probability distribution corresponding to each predicted title sequence and the title hierarchy labels corresponding to each title, jointly train the original title classification model and the original title hierarchy model to obtain the trained title classification model and title hierarchy model.
[0107] In a document processing system, accurate title classification and title hierarchy recognition are crucial for clearly presenting the document structure and enhancing document readability. To improve the accuracy and coordination of these two tasks, the title classification model and the title hierarchy model can be jointly trained.
[0108] To conduct joint training, a comprehensive text sample set needs to be constructed. Collect a large number of documents from different fields (such as academic, business, news, etc.) and different formats (such as Word documents, PDF documents, plain text, etc.). Select the second text samples from these documents and annotate each title in them, assigning corresponding title hierarchy labels to each title, such as "first-level title", "second-level title", "third-level title", etc. These second text samples should be widely representative to ensure that the models (including the title classification model and the title hierarchy model) can learn the characteristics and hierarchical relationships of titles in various types of documents.
[0109] Exemplarily, randomly select any second text sample from the constructed comprehensive text sample set. Each title in this sample has been pre-annotated with a title hierarchy label, which can be represented by numbers (such as "1" for the first-level title, "2" for the second-level title, etc.) or other appropriate identifiers. The specific representation method can be flexibly determined according to actual needs. Input the selected second text sample into the original title classification model. The original title classification model processes the second text sample and assigns title labels to each predicted title sequence in it, thus obtaining a third text sequence carrying title labels. Input the obtained third text sequence into the original title hierarchy model. The original title hierarchy model processes each predicted title sequence based on the third text sequence and determines the corresponding title hierarchy probability distribution for each of them. For example, for a predicted title sequence, the original title hierarchy model may output that it has a 70% probability of being a first-level title, a 20% probability of being a second-level title, and a 10% probability of being a third-level title. Based on the title hierarchy probability distribution corresponding to each predicted title sequence and the title hierarchy label corresponding to each title, determine the joint loss value. This joint loss value comprehensively reflects the degree of difference between the prediction results of the title classification model and the title hierarchy model and the true labels. According to this joint loss value, jointly train the original title classification model and the original title hierarchy model.
[0110] For each second text sample, perform corresponding operations to obtain the combined loss value of the sample. In the current iteration, add up the combined loss values of all second text samples to obtain the total combined loss value. Based on this total combined loss value, further adjust the parameters of the title classification model and the title hierarchy model in the current iteration. During the training process, adjust the parameters of both models simultaneously to gradually reduce the combined loss value. Through multiple iterations of training, the two models can cooperate with each other and learn together to improve the accuracy of title classification and title hierarchy recognition. When the preset convergence condition is met, it is considered that the joint training of the title classification model and the title hierarchy model is completed.
[0111] Among them, meeting the preset convergence condition can be that the sum of the combined loss values determined based on each second text sample is less than the pre-configured loss threshold, or the sum of the determined combined loss values has been in a downward trend and tends to be flat, or the number of iterations for jointly training the two models reaches the set maximum number of iterations, etc. In the specific implementation process, the convergence condition can be flexibly set according to the actual situation.
[0112] As a possible implementation method, in order to ensure that the jointly trained title classification model and title hierarchy model have good generalization ability, during training, the second text samples can be divided into training samples and test samples. First, jointly train the title classification model and the title hierarchy model based on the training samples, allowing the two models to learn the title features and hierarchy relationships in the training samples, and continuously adjust the parameters to reduce the combined loss value, so that the models gradually adapt to the training data. Then, verify the reliability of the two jointly trained models based on the test samples. Input the test samples into the two models, and calculate evaluation metrics such as accuracy, recall rate, F1 value, etc. according to the prediction results and true labels of the models. Use these evaluation metrics to measure the performance of the two models on unseen data, and ensure that they can be accurately applied to actual document processing tasks, providing reliable title classification and title hierarchy recognition results for subsequent document sorting and presentation. After the above training and verification process, finally obtain two models with good performance and working together, namely the trained title classification model and title hierarchy model.
[0113] IV. Document Output Unit 14
[0114] The document output unit 14 determines the text style corresponding to each title sequence according to the title hierarchy labels corresponding to the title sequences in the second text sequence from the text styles respectively corresponding to each pre-configured title hierarchy. Among them, the pre-configured text styles can include attributes such as font, font size, color, bold, italic, etc. For example, the first-level title can be set to a larger font size, bold font, and a specific color to highlight its importance; the second-level title can be distinguished by a slightly smaller font size and a different color.
[0115] Meanwhile, the system is also pre-configured with a body text style to standardize the display format of the body part of the document. The document output unit 14 sorts the second text sequence into a document with title levels according to the text styles corresponding to each title sequence and the body text style respectively. The finally generated document can be output in various formats, such as.docx,.pdf, etc., which is convenient for users to view and use.
[0116] The beneficial effects of this application are as follows:
[0117] This high-performance document title level recognition system adopts a deep learning method. Through a pre-trained title classification model and a title level model, it automatically learns the features and level relationships of titles from a large amount of data. The system can directly process the original document without cumbersome rule formulation and data processing. This simplifies the document processing process, improves work efficiency, and avoids many problems existing in the rule recognition method, such as complex extraction and processing procedures, low universality, and low accuracy. Moreover, both the title classification model and the title level model have a certain generalization ability. Even when encountering uncommon title forms, they can make accurate judgments based on the knowledge they have learned.
[0118] Embodiment 2:
[0119] Based on the same inventive concept, this application also provides a high-performance document title level recognition method. Figure 2 FIG. is a schematic diagram of the process of high-performance document title level recognition provided by an embodiment of this application. The process includes:
[0120] S201: Obtain the document to be recognized.
[0121] S202: Use the pre-trained title classification model to label each title sequence in the document to be recognized with a title label, so as to obtain a fourth text sequence carrying the title label.
[0122] S203: Use the pre-trained title level model to determine the title level corresponding to each title sequence based on the fourth text sequence, and label each title sequence with a title level label in the fourth text sequence, so as to obtain a fifth text sequence carrying the title level label.
[0123] S204: For each title sequence, determine the text style corresponding to the title sequence from the text styles corresponding to each title level pre-configured according to the title level label corresponding to the title sequence in the fifth text sequence.
[0124] S205: Sort the fifth text sequence into a document with title levels according to the text styles corresponding to each title sequence and the pre-configured body text style.
[0125] The high-performance document title level recognition method provided by this application is applied to a computer device, which can be an intelligent terminal, such as a computer, a robot, etc., or a server, such as an application server, a business server, etc. This high-performance document title level recognition method aims to efficiently and accurately recognize the title level of a document and generate a document with a clear title level structure.
[0126] First, obtain the document to be recognized. Among them, the document to be recognized can be obtained through the following multiple ways:
[0127] 1. Obtain from local storage: The user can select the required document from storage devices such as a local hard disk and a USB flash drive through the file selection interface provided by the computer device. The computer device can recognize common document formats, such as Microsoft Word documents (.doc,.docx), PDF documents (.pdf), plain text files (.txt), etc. For documents in different formats, the computer device will use corresponding parsers for processing and convert them into a unified internal text representation form for subsequent title classification and level recognition operations.
[0128] 2. Obtain from network resources: The computer device supports obtaining the document to be recognized from the network. The user can input the network link of the document (such as a URL with HTTP or HTTPS protocol), and the computer device will download the document through a network request. During the download process, the computer device will perform error handling to ensure the complete acquisition of the document.
[0129] 3. Obtain from other data sources: It is also possible to obtain the document to be recognized from data sources such as a database and cloud storage. The computer device will perform data query and extraction operations according to the interface specifications of different data sources and load the document into the computer device for processing.
[0130] Input the document to be recognized into the trained title classification model. The title classification model will analyze each sentence sequence in the document to be recognized, extract its features (such as word vectors, syntactic structures, etc.), and judge whether the sentence sequence is a title according to the learned patterns. If it is judged that the sentence sequence is a title, a corresponding title label will be assigned to it. After being processed by the label classification model, each title sequence in the document to be recognized is assigned a title label, thus forming a fourth text sequence carrying the title label.
[0131] Input the fourth text sequence with title tags into the trained title hierarchy model. The title hierarchy model comprehensively considers factors such as the content of the title sequence, context information, and its relationship with other titles to determine the title hierarchy corresponding to each title sequence. Then, add the corresponding title hierarchy tags to each title sequence in the fourth text sequence to obtain the fifth text sequence with title hierarchy tags.
[0132] Five pre-configures text styles corresponding to each title hierarchy, and these styles can include attributes such as font, font size, color, bold, italic, underline, etc. For example, the fourth-level title may use a relatively large font size (such as No. 3 font), bold font, and a specific color (such as red); the fifth-level title uses a slightly smaller font size (such as No. 4 font), bold font, and a different color (such as blue), etc. After obtaining the fifth text sequence with title hierarchy tags, five will traverse each title sequence therein and look up the corresponding text style from the pre-configured text style library according to its title hierarchy tag.
[0133] Five also pre-configures the text style for the main body text to standardize the display format of the main body part of the document. After determining the text style of each title sequence, five will organize the fifth text sequence according to these styles. Specifically, for each title sequence, apply the corresponding text style for formatting; for the main body part, apply the pre-configured text style for the main body for formatting. Then, combine the processed title sequences and the main body content in the order of the original document to generate four documents with a clear title hierarchy structure. The finally generated documents can be saved in common document formats such as.docx,.pdf, etc., for convenient viewing and use by users.
[0134] Embodiment 3:
[0135] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of a computer device provided by an alternative embodiment of the present application, as Figure 3As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting the components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if needed, multiple processors and / or multiple buses can be used with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as an array of servers, a set of blade servers, or a multi-processor system). Figure 3 Take one processor 10 as an example in
[0136] The processor 10 can be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 can further include a hardware chip. The above hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above programmable logic device can be a complex programmable logic device, a field programmable gate array, a generic array logic, or any combination thereof.
[0137] Among them, the memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiments.
[0138] The memory 20 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device presented by a kind of landing page of a small program, etc. In addition, the memory 20 can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 can optionally include a memory remotely set relative to the processor 10, and these remote memories can be connected to the computer device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0139] The memory 20 can include a volatile memory, such as a random access memory; the memory can also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state drive; the memory 20 can also include a combination of the above types of memories.
[0140] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30, and the output device 40 may be connected through a bus or other means. Figure 3 Taking the connection through the bus as an example.
[0141] The input device 30 can receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED), and a tactile feedback device (e.g., a vibration motor), etc. The above display device includes but is not limited to a liquid crystal display, a light-emitting diode, a display, and a plasma display. In some alternative embodiments, the display device may be a touch screen.
[0142] Embodiment 4:
[0143] Based on the above embodiments, an embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program executable by a processor. When the program runs on the processor, the processor is caused to perform the following steps when executing:
[0144] Obtain a document to be recognized;
[0145] Through a pre-trained title classification model, label each title sequence in the document to be recognized with a title label to obtain a fourth text sequence carrying the title label;
[0146] Through a pre-trained title hierarchy model, based on the fourth text sequence, determine the title hierarchy corresponding to each title sequence, and label each title sequence with a title hierarchy label in the fourth text sequence to obtain a fifth text sequence carrying the title hierarchy label;
[0147] For each title sequence, according to the title hierarchy label corresponding to the title sequence in the fifth text sequence, determine the text style corresponding to the title sequence from the text styles corresponding to each pre-configured title hierarchy;
[0148] Arrange the fifth text sequence into a document with a title hierarchy according to the text styles corresponding to each title sequence and the pre-configured body text style.
[0149] Since the principle of the computer-readable storage medium for solving problems is similar to that of the high-performance document title hierarchy recognition method, the implementation of the above computer-readable storage medium can refer to the embodiments of the method, and the repeated parts will not be elaborated.
[0150] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to cover these changes and modifications.
Claims
1. A high-performance document title level recognition system, characterized in that: The system comprises: A document receiving unit, used for acquiring a document to be identified; A title classification unit, used for labeling each title sequence in the to-be-identified document with a title label by using a pre-trained title classification model, so as to obtain a first text sequence carrying the title label; a title level identification unit, configured to determine the title levels corresponding to the respective title sequences based on the first text sequence by using a pre-trained title level model, and to label the respective title sequences with title level labels in the first text sequence to obtain a second text sequence carrying the title level labels; The document output unit is used to determine, for each title sequence, the text style corresponding to the title sequence from the pre-configured text styles corresponding to each title level according to the title level label corresponding to the title sequence in the second text sequence; and organize the second text sequence into a document with a title level according to the text styles corresponding to each title sequence and the pre-configured body text style.
2. The system according to claim 1, characterized in that The title classification unit is specifically used for: Segmenting the document to be identified according to sentences to generate a sentence sequence of the document to be identified; The title classification model is used to classify each of the sentence sequences, and sentence sequences classified as title sequences are identified; in each of the sentence sequences, the sentence sequences classified as title sequences are tagged with title tags to obtain the first text sequence.
3. The system according to claim 1, characterized in that The title classification model is trained in the following way: Obtain any first text sample in the title classification sample set; wherein each title in the first text sample corresponds to a title tag; Obtaining the probability of each predicted title sequence in the first text sample based on the first text sample through the original title classification model; Based on the probabilities of the predicted title sequences and the title tags, the original title classification model is trained to obtain a trained title classification model.
4. The system according to claim 3, characterized in that If each text sequence in the first text sample corresponds to a text label, then the probability of each predicted title sequence and the probability of each predicted text sequence in the first text sample are obtained based on the first text sample through the original title classification model; based on the probability of each predicted title sequence and the title labels, and the probability of each predicted text sequence and the text labels, the original title classification model is trained to obtain a trained title classification model.
5. The system according to claim 1, wherein: The title level model is trained as follows: Obtain any text sequence sample in the title level sample set; wherein each title sequence in the text sequence sample carries a title tag, and any title sequence corresponds to a title level tag; Obtaining predicted title level probability distributions corresponding to each title sequence in the text sequence sample based on the text sequence sample through the original title level model; Based on the predicted title level probability distributions corresponding to the respective title sequences and the respective corresponding title level labels, the original title level model is trained to obtain a trained title level model.
6. The system according to claim 1, wherein: The title classification model and the title hierarchy model are jointly trained in the following manner: Obtain any second text sample in the comprehensive text sample set; wherein each title in the second text sample corresponds to a title level label; Using the original title classification model, based on the second text sample, label each predicted title sequence in the second text sample with a title label to obtain a third text sequence carrying the title label; Determine, by using the original title hierarchy model, the title hierarchy probability distribution corresponding to each predicted title sequence based on the third text sequence; Based on the title level probability distributions corresponding to the predicted title sequences and the title level labels corresponding to the titles, the original title classification model and the original title level model are jointly trained to obtain a trained title classification model and title level model.
7. A high-performance document title level identification method, characterized in that: The method comprises: Get the document to be identified; Using a pre-trained title classification model, labeling each title sequence in the to-be-identified document with a title label to obtain a fourth text sequence carrying the title label; Determining the title levels corresponding to the title sequences based on the fourth text sequence by using a pre-trained title level model, and marking the title levels of the title sequences in the fourth text sequence to obtain a fifth text sequence carrying the title level labels; For each title sequence, according to the title level label corresponding to the title sequence in the fifth text sequence, determine the text style corresponding to the title sequence from the pre-configured text styles corresponding to each title level; According to the text styles corresponding to the title sequences and the pre-configured body text style, the fifth text sequence is organized into a document with a title level.
8. A computer device, characterized in that: The computer device comprises a processor, and the processor is used to implement the steps of the high-performance document title hierarchy identification method as claimed in claim 7 when executing a computer program stored in a memory.
9. A computer-readable storage medium, characterized in that: The device stores a computer program, which, when executed by a processor, implements the steps of high-performance document title hierarchy identification as described in claim 7.
Citation Information
Cited By
Multi-type text hierarchical directory construction method, device and equipment based on large model
CN121166840A