Method and device for identifying title in document
By extracting features from the initial paragraphs and text lines in the document in semantic, spatial and visual dimensions, possible title paragraphs are screened out, solving the problem of the existing technology that is unable to recognize non-standard structured document titles, and achieving fast and accurate multi-line title recognition.
Patent Information
- Application Number
- CN202410342740.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-25
- Publication Date
- 2025-09-26
AI Technical Summary
The existing technology cannot effectively recognize the titles in documents with non-standard structures, resulting in the inability to accurately recognize the titles.
By extracting features of semantic, spatial and visual dimensions from the initial paragraphs and text lines in the document, possible title paragraphs are screened out, and whether they are titles is determined based on the feature processing results.
It achieves fast and accurate recognition of titles in documents, especially supports the recognition of multi-line titles, and improves recognition performance and accuracy.
Smart Images

Figure CN120706416A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of document recognition, and in particular to a method and device for recognizing titles in documents. Background Art
[0002] Whether it's a flow document or a layout document, the concept of a title is part of its structure. However, many documents are not created according to the standard structure, resulting in the title being placed at the same level as the main text, making it difficult to identify the title. Therefore, an effective solution is urgently needed to solve this problem. Summary of the Invention
[0003] In view of the problems existing in the prior art, embodiments of the present invention provide a method and apparatus for identifying titles in a document.
[0004] The present invention provides a method for identifying titles in a document, comprising:
[0005] Identifying at least one initial paragraph in a target document and at least one text line in each initial paragraph;
[0006] performing feature extraction on each of the text lines with respect to at least one dimension among a semantic dimension, a spatial dimension, and a visual dimension to obtain a feature extraction result for each of the text lines;
[0007] Filtering a target paragraph from the at least one initial paragraph according to each of the feature extraction results;
[0008] Processing the feature extraction results of each text line in the target paragraph to obtain a feature processing result of the target paragraph;
[0009] According to the feature processing result, it is determined whether the target paragraph in the target document is a title.
[0010] According to a method for identifying titles in a document provided by the present invention, the feature extraction of each text line is performed for at least one dimension among a semantic dimension, a spatial dimension, and a visual dimension to obtain a feature extraction result for each text line, including:
[0011] Perform at least one of the following feature extraction processes on each of the text lines:
[0012] Extracting features from each text line from a semantic dimension to obtain semantic feature extraction results for each text line;
[0013] Extracting features of each text line from a spatial dimension to obtain spatial feature extraction results of each text line;
[0014] Extract features from each of the text lines from a visual dimension to obtain visual feature extraction results for each of the text lines;
[0015] For each of the text lines, at least one of the semantic feature extraction result, the spatial feature extraction result, and the visual feature extraction result of the text line is processed to obtain a feature extraction result of the text line.
[0016] According to a method for identifying titles in documents provided by the present invention, the semantic feature extraction results include sequence number feature extraction results, special title feature extraction results, and common title feature extraction results;
[0017] The extracting features of each text line from a semantic dimension to obtain a semantic feature extraction result of each text line includes:
[0018] For each of the text lines, perform the following steps:
[0019] Performing serial number recognition on the beginning of the text line; if the serial number is recognized, determining that the serial number feature extraction result of the text line has the serial number feature; if the serial number is not recognized, determining that the serial number feature extraction result of the text line does not have the serial number feature;
[0020] Matching the text line with a preset set of special title symbols; if the match is successful, determining that the special title feature extraction result of the text line is that the special title symbol exists; if the match fails, determining that the special title feature extraction result of the text line is that the special title symbol does not exist;
[0021] The text line is matched with a preset set of common title symbols; if the match is successful, the common title feature extraction result of the text line is determined to be the presence of common title symbols; if the match fails, the common title feature extraction result of the text line is determined to be the absence of the common title symbols.
[0022] According to a method for identifying titles in a document provided by the present invention, the spatial feature extraction result includes a content feature extraction result, a first font size feature extraction result, and a second font size feature extraction result;
[0023] The extracting features of each text line from a spatial dimension to obtain a spatial feature extraction result of each text line includes:
[0024] For each of the text lines, perform the following steps:
[0025] If the content of the text line is all text content, determining that the content feature extraction result of the text line has text content features; if the content of the text line contains non-text content, determining that the content feature extraction result of the text line does not have the text content features;
[0026] Determining whether a first average font size corresponding to the text line is greater than a second average font size, where the second average font size is the average font size of all text in the page to which the text line belongs; if so, determining that the first font size feature extraction result of the text line has the font size feature; if not, determining that the first font size feature extraction result of the text line does not have the font size feature;
[0027] In the absence of a first designated text line, determining that the second font size feature extraction result of the text line is a first value, and the first designated text line is the next text line of the text line; in the presence of the first designated text line, determining whether the first average font size corresponding to the text line is greater than the third average font size of the first designated text line; if so, determining that the second font size feature extraction result of the text line is a second value; if not, determining that the second font size feature extraction result of the text line is a third value; the first value, the second value and the third value are all different.
[0028] According to a method for identifying titles in a document provided by the present invention, the visual feature extraction results include font feature extraction results, modification feature extraction results, and interval feature extraction results;
[0029] The extracting features of each text line from a visual dimension to obtain a visual feature extraction result of each text line includes:
[0030] Performing font recognition on each text content in the text line to obtain a font feature extraction result of the text line;
[0031] Identifying modified content on the text line, the modified content being used to modify the text content in the text line; if the modified content is identified, determining that a modified feature extraction result of the text line has a modified feature; if the modified content is not identified, determining that a modified feature extraction result of the text line does not have the modified feature;
[0032] Perform associated interval matching on the text line to obtain an interval feature extraction result of the text line.
[0033] According to a method for identifying titles in a document provided by the present invention, the font feature extraction result includes a first bold feature extraction result, a second bold feature extraction result, and a same-font feature extraction result;
[0034] The performing font recognition on each text content in the text line to obtain a font feature extraction result of the text line includes:
[0035] Determine whether each text content in the text line is in a bold font; if so, determine that the first bold feature extraction result of the text line has all the bold features; if not, determine that the first bold feature extraction result of the text line does not have all the bold features;
[0036] Determine whether there is text content in a bold font in the text line; if so, determine that the second bold feature extraction result of the text line has a partial font bold feature; if not, determine that the second bold feature extraction result of the text line does not have the partial font bold feature;
[0037] Determine whether the fonts of each text content in the text line are the same; if so, determine that the font same feature extraction result of the text line has the font same feature; if not, determine that the font same feature extraction result of the text line does not have the font same feature.
[0038] According to a method for identifying titles in documents provided by the present invention, the interval feature extraction results include font size interval feature extraction results and character number interval feature extraction results;
[0039] The performing associated interval matching on the text line to obtain the interval feature extraction result of the text line includes:
[0040] Determining a first target font size interval from at least one preset initial font size interval based on a first average font size corresponding to the text line; determining a font size interval feature extraction result of the text line as a first serial number, the first serial number representing a ranking of the first target font size interval in each of the initial font size intervals;
[0041] Based on the first number of characters contained in the text line, a first target character number interval is determined from at least one preset initial character number interval; the character number interval feature extraction result of the text line is determined as a second serial number, and the second serial number represents the ranking of the first target character number interval in each of the initial character number intervals.
[0042] According to a method for identifying titles in a document provided by the present invention, screening out a target paragraph from the at least one initial paragraph based on each of the feature extraction results includes:
[0043] For each of these initial paragraphs, perform the following steps:
[0044] When the number of text lines in the initial paragraph is a first preset value, determining the initial paragraph as a target paragraph;
[0045] When the number of text lines in the initial paragraph is not the first preset value, determine whether the feature extraction results of each text line in the initial paragraph meet the set merging conditions; if so, determine the initial paragraph as the target paragraph.
[0046] According to a method for identifying titles in a document provided by the present invention, determining whether the feature extraction results of each text line in the initial paragraph meet a set merging condition includes:
[0047] Obtaining the alignment and first average font size of each text line in the initial paragraph;
[0048] determining a first absolute value of a difference between first average font sizes corresponding to adjacent text lines in the initial paragraph based on each of the first average font sizes;
[0049] It is determined whether the feature extraction results and the alignment methods of the text lines in the initial paragraph, as well as the first absolute values, all meet the set merging conditions.
[0050] According to a method for identifying titles in a document provided by the present invention, the feature extraction results include a content feature extraction result, a first bold feature extraction result, a second bold feature extraction result, a font-same feature extraction result, a modification feature extraction result, a first font size feature extraction result, and a sequence number feature extraction result;
[0051] The feature extraction results of each text line in the initial paragraph meet the set merging conditions, including:
[0052] The content feature extraction results of each of the text lines in the initial paragraph are the same, the first bold feature extraction results are the same, the second bold feature extraction results are the same, the font same feature extraction results are the same, the modification feature extraction results are the same, and the first font size feature extraction results are the same; and
[0053] The result of the sequence number feature extraction of the first text line in the initial paragraph is that it has the sequence number feature, and the results of the sequence number feature extraction of each specific text line in the initial paragraph do not have the sequence number feature, and the specific text lines are the text lines in the initial paragraph other than the first text line;
[0054] Each of the first absolute values corresponding to the initial paragraphs meets a set merging condition, including:
[0055] Each of the first absolute values corresponding to the initial paragraph is smaller than the set font size difference;
[0056] The alignment of each text line in the initial paragraph meets the set merging conditions, including:
[0057] The alignment of each text line in the initial paragraph is the same.
[0058] According to a method for identifying titles in a document provided by the present invention, obtaining the alignment of each text line in the initial paragraph includes:
[0059] For each of the text lines in the initial paragraph, perform the following steps:
[0060] Obtaining the bounding rectangle and line height of the text line, and obtaining the visual area of the page to which the text line belongs;
[0061] Calculating a left distance between a left border of the circumscribed rectangular frame and a left border of the visualization area, and calculating a right distance between a right border of the circumscribed rectangular frame and a right border of the visualization area;
[0062] If the left spacing and the right spacing meet the center alignment condition, then the alignment of the text line is determined to be center alignment, and the center alignment condition is that the second absolute value of the difference between the left spacing and the right spacing is greater than the line height, and both the left spacing and the right spacing are greater than the product of the line height and a set value;
[0063] If the left spacing and the right spacing meet a right alignment condition, determining that the alignment of the text line is right alignment, the right alignment condition is that the left spacing is smaller than the right spacing, and the right spacing is smaller than the product;
[0064] If the left spacing and the right spacing do not satisfy the center alignment condition and do not satisfy the right alignment condition, then the alignment of the text line is determined to be left alignment.
[0065] According to a method for identifying titles in a document provided by the present invention, processing the feature extraction results of each text line in the target paragraph to obtain the feature processing results of the target paragraph includes:
[0066] When the number of text lines in the target paragraph is a second preset value, using the feature extraction results of the text lines in the target paragraph as the feature processing results of the target paragraph;
[0067] When the number of text lines in the target paragraph is not the second preset value, the feature extraction results of the text lines in the target paragraph are merged according to the set merging strategy to obtain the feature processing result of the target paragraph.
[0068] According to a method for identifying titles in a document provided by the present invention, the feature extraction results include content feature extraction results, first bold feature extraction results, second bold feature extraction results, font same feature extraction results, sequence number feature extraction results, special title feature extraction results, common title feature extraction results, modification feature extraction results, first font size feature extraction results, second font size feature extraction results, font size interval feature extraction results, and character number interval feature extraction results;
[0069] Merging the feature extraction results of each text line in the target paragraph according to a set merging strategy to obtain the feature processing result of the target paragraph includes:
[0070] The content feature extraction result, the first bold feature extraction result, the second bold feature extraction result, the same font feature extraction result, the modified feature extraction result, and the first font size feature extraction result of the first text line in the target paragraph are respectively used as the content feature extraction result, the first bold feature extraction result, the second bold feature extraction result, the same font feature extraction result, the modified feature extraction result, and the first font size feature extraction result of the target paragraph;
[0071] Performing an OR operation on the sequence number feature extraction results of each text line in the target paragraph to obtain the sequence number feature extraction result of the target paragraph; performing an OR operation on the special title feature extraction results of each text line in the target paragraph to obtain the special title feature extraction result of the target paragraph; performing an OR operation on the common title feature extraction results of each text line in the target paragraph to obtain the common title feature extraction result of the target paragraph;
[0072] In the absence of a second designated text line, determining that the second font size feature extraction result of the target paragraph is a first value, and the second designated text line is the next text line of the target paragraph; in the presence of the second designated text line, performing an average calculation on the first average font sizes corresponding to each text line in the target paragraph to obtain a target average font size, and determining whether the target average font size is greater than a third average font size of the second designated text line; if so, determining that the second font size feature extraction result of the target paragraph is a second value; if not, determining that the second font size feature extraction result of the target paragraph is a third value; the first value, the second value, and the third value are all different;
[0073] determining a second target font size interval from at least one preset initial font size interval based on the fourth average font size of the target paragraph; determining a font size interval feature extraction result of the target paragraph as a third serial number, the third serial number representing a ranking of the second target font size interval in each of the initial font size intervals;
[0074] Determining a second target character number interval from at least one preset initial character number interval based on a second number of characters included in the target paragraph; determining a character number interval feature extraction result of the text line as a fourth serial number, the fourth serial number representing a ranking of the second target character number interval in each of the initial character number intervals;
[0075] The feature processing result of the target paragraph includes the content feature extraction result of the target paragraph, the first bold feature extraction result, the second bold feature extraction result, the font same feature extraction result, the serial number feature extraction result, the special title feature extraction result, the common title feature extraction result, the modification feature extraction result, the first font size feature extraction result, the second font size feature extraction result, the font size interval feature extraction result and the character number interval feature extraction result.
[0076] According to a method for identifying titles in a document provided by the present invention, determining whether the target paragraph in the target document is a title based on the feature processing result includes:
[0077] Determine whether the target paragraph in the target document is a title based on the feature processing result and the set title condition; or
[0078] The feature processing result is input into a trained title recognition model to determine whether the target paragraph in the target document is a title.
[0079] According to a method for identifying titles in a document provided by the present invention, the feature processing results include content feature extraction results, first bold feature extraction results, second bold feature extraction results, font same feature extraction results, sequence number feature extraction results, special title feature extraction results, common title feature extraction results, modification feature extraction results, first font size feature extraction results, second font size feature extraction results, font size range feature extraction results, and character number range feature extraction results;
[0080] The title setting conditions include:
[0081] The content feature extraction result, the first bold feature extraction result, the second bold feature extraction result, the same font feature extraction result, the serial number feature extraction result, the special title feature extraction result, the common title feature extraction result, the modification feature extraction result, and the first font size feature extraction result are all characterized by having corresponding features; and
[0082] The second font size feature extraction result is a second value; and
[0083] The font size interval feature extraction result is greater than a first threshold; and
[0084] The character number interval feature extraction result is less than a second threshold.
[0085] The present invention also provides a device for identifying titles in a document, comprising:
[0086] a recognition module configured to recognize at least one initial paragraph in a target document and at least one text line in each initial paragraph;
[0087] an extraction module configured to perform feature extraction on each of the text lines with respect to at least one of a semantic dimension, a spatial dimension, and a visual dimension, to obtain a feature extraction result for each of the text lines;
[0088] a screening module configured to screen out a target paragraph from the at least one initial paragraph based on the feature extraction results;
[0089] a processing module configured to process the feature extraction results of each text line in the target paragraph to obtain a feature processing result of the target paragraph;
[0090] The determination module is configured to determine whether the target paragraph in the target document is a title according to the feature processing result.
[0091] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for identifying titles in a document as described above is implemented.
[0092] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above methods for identifying titles in a document.
[0093] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above methods for identifying titles in a document.
[0094] The method and device for identifying titles in documents provided by the present invention identify at least one initial paragraph in a target document and at least one text line in each of the initial paragraphs; perform feature extraction on each of the text lines with respect to at least one dimension among semantic dimension, spatial dimension and visual dimension to obtain feature extraction results for each of the text lines; screen out a target paragraph from the at least one initial paragraph based on each of the feature extraction results; process the feature extraction results of each of the text lines in the target paragraph to obtain feature processing results for the target paragraph; determine whether the target paragraph in the target document is a title based on the feature processing results; the method and device can quickly and accurately identify titles in documents, have good performance, extremely high accuracy, and support multi-line title recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0095] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0096] Figure 1 1 is a flow chart of a method for identifying titles in a document provided by the present invention;
[0097] Figure 2 Schematic diagram of the effect of the method for identifying titles in documents provided by the present invention;
[0098] Figure 3 It is a schematic structural diagram of the device for identifying titles in documents provided by the present invention;
[0099] Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0100] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0101] In order to facilitate a clearer understanding of the embodiments of the present invention, some relevant knowledge is first introduced as follows.
[0102] A title is a short phrase that summarizes your content and grabs your reader's attention.
[0103] Unicode is a character encoding standard used to represent characters, symbols, and punctuation marks in various languages. It provides consistent encoding for all texts worldwide to resolve confusion and compatibility issues caused by different character sets and encoding schemes.
[0104] Features: In machine learning and data analysis, features are attributes or metrics used to describe a data sample. Features can be individual dimensions, attributes, or combinations of attributes. In machine learning models, features are used for training and prediction. The model makes predictions by learning the relationship between input features and output labels.
[0105] A font is the visual representation of a set of characters, including their appearance, style, and arrangement. Fonts have many properties that can be adjusted and configured, including but not limited to: Font Family, Font Style, Font Size, and Font Weight.
[0106] A font family is a category or classification of fonts, such as Songti, Arial, and Times New Roman. Font styles include Regular, Italic, and Bold. Font size determines the display size of characters. Font weight specifies the thickness of characters.
[0107] TextObject is a basic element in a document and belongs to the text type element. It contains information such as Unicode (Unicode), CharCode (character code), Font, and BBox (bounding rectangle).
[0108] The prior art provides four methods for identifying titles in documents.
[0109] Method 1: After converting a PDF (Portable Document Format) or Word document into HTML (Hypertext Markup Language), the plain text is then matched against several possible title patterns to identify the title. This method is only suitable for documents containing plain text, and due to its limited matching pattern, it can only recognize the most basic titles, resulting in poor recognition results.
[0110] Method 2: Use deep learning to obtain row coordinate information and use this information as a machine learning classification feature to determine whether it is a title. This method has poor performance and poor recognition effect.
[0111] Method 3: Determine whether it is a title and the title level based on whether a serial number appears in the document. However, the recognition effect is poor based on the serial number alone, and multi-line title recognition is not supported.
[0112] Method 4: Identify titles and title hierarchies using a rule-based and machine learning approach. However, the rule setting is relatively simple, targeting only serial number formats. Machine learning requires first extracting word vector features and relies purely on semantic features. This method relies solely on text content and is relatively simple.
[0113] The following combination Figure 1-Figure 4 The method and apparatus for identifying titles in documents of the present invention are described.
[0114] Figure 1 This is a flow chart of the method for identifying titles in documents provided by the present invention. Figure 1 As shown, it includes steps 101 to 105, wherein:
[0115] Step 101: Identify at least one initial paragraph in a target document and at least one text line in each initial paragraph.
[0116] First of all, it should be noted that the execution subject of the present invention can be any electronic device that can recognize the title in the document, for example, it can be any one of a smart phone, a smart watch, a desktop computer, a laptop computer, etc.
[0117] In addition, the method for identifying titles in documents provided by the present invention can be applied to scenarios such as document conversion, document rearrangement, document content retrieval, document structure tree generation, document directory generation, and document bookmark generation.
[0118] Specifically, the target document refers to the document whose title needs to be identified. It can be a layout document, such as a PDF document, an OFD (Open Fixed-layout Document) document, a streaming document, such as a Word document, a TXT document, a presentation document, such as a slide document, or a picture document. The present invention does not impose any restrictions on this.
[0119] A text line is a line in a document used to display content. The number of text lines is the number of text lines included. The set threshold is a preset number of text lines, such as 3 or 4 lines.
[0120] The initial paragraph refers to a paragraph in the target document whose number of text lines is less than a set threshold, that is, the number of text lines of the initial paragraph is less than the set threshold. Exemplarily, the initial paragraph is a paragraph whose number of text lines is less than 4.
[0121] In practical applications, before determining the area to be identified in the target document, the target document must be acquired first. There are many methods for acquiring the target document, which are not limited in the present invention.
[0122] Exemplarily, a user uploads a target document through a document upload page, and accordingly, the execution entity obtains the target document.
[0123] Exemplarily, the execution subject receives a document acquisition instruction or a title identification instruction, and accordingly, the execution subject acquires the target document from the storage area pointed to by the document acquisition instruction or the title identification instruction.
[0124] After the target document is acquired, each initial paragraph and each text line in the target document is further identified.
[0125] Because titles are relatively short, they occupy relatively few lines of text. Therefore, paragraphs with too many lines of text are not considered titles. That is, paragraphs with a number of lines greater than or equal to a set threshold are not considered titles, and the lines within them are not considered paragraphs. To reduce data processing, paragraphs with a number of lines greater than or equal to the set threshold and the lines within them are not recognized. Instead, initial paragraphs with a number of lines less than the set threshold and the lines within them are recognized.
[0126] In addition, text lines can be identified by a line identifier: if it is a layout document, it is necessary to determine it based on the position of the TextObject in the document; if it is a streaming document, no special identification is required. The initial paragraph can be identified by a paragraph identifier: if it is a layout document, it can be determined based on the method for identifying paragraphs in a layout document; if it is a streaming document, no special operation is required. Among them, the method for identifying paragraphs in a layout document includes: determining the area to be identified in the target layout document, the area to be identified contains at least one text line; performing text line association detection on each text line to obtain an association detection result of whether the area to be identified meets the preset line association condition; performing alignment detection on each text line to obtain an alignment detection result of whether the area to be identified meets the preset alignment condition; performing color detection on the color range of the text object in each text line to obtain a color detection result of whether the area to be identified meets the preset color condition; and determining whether the area to be identified in the target layout document is a paragraph based on the association detection result, alignment detection result and color detection result. It should be noted that the above process is only an example, and any method can be used to identify paragraphs and / or text lines of the target document in step 101 of this application.
[0127] It should be noted that when identifying the initial paragraphs and text lines, they are identified from areas in the target document that are not tables and are not in headers or footers.
[0128] Step 102: performing feature extraction on each of the text lines with respect to at least one dimension among the semantic dimension, the spatial dimension, and the visual dimension to obtain a feature extraction result for each of the text lines.
[0129] Specifically, the semantic dimension refers to the level of content expressed by the linguistic form of the title. The spatial dimension refers to the level of spatial distribution of the title in the document. The visual dimension refers to the level of the title's presentation in the document.
[0130] In practical applications, for each text line, feature extraction can be performed on the text line along at least one of the semantic, spatial, and visual dimensions to obtain sub-feature extraction results for at least one dimension. When feature extraction is performed on the text line along a single dimension, the sub-feature extraction results for that dimension can be used as the feature extraction result for the text line. When feature extraction is performed on the text line along multiple dimensions, the sub-feature extraction results for each dimension are fused or concatenated to obtain the feature extraction result for the text line.
[0131] Step 103: Filter out a target paragraph from the at least one initial paragraph according to the feature extraction results.
[0132] Specifically, the target paragraph refers to a paragraph that may be a title.
[0133] In practical applications, for each initial paragraph, whether the initial paragraph is a target paragraph is identified based on the feature extraction results of each text line in the initial paragraph.
[0134] Exemplarily, the feature extraction results of each text line in the initial paragraph are compared with the set merging conditions. If both meet the requirements, the initial paragraph is determined to be the target paragraph. If there is a feature extraction result that does not meet the set merging conditions, the initial paragraph is not the target paragraph.
[0135] Exemplarily, the feature extraction results of each text line in the initial paragraph are input into a pre-trained target paragraph recognition model for recognition, thereby determining whether the initial paragraph is a target paragraph.
[0136] Step 104: Process the feature extraction results of each text line in the target paragraph to obtain a feature processing result of the target paragraph.
[0137] Specifically, the feature processing result refers to the feature extraction result corresponding to the target paragraph.
[0138] In practical applications, for each target paragraph, the feature extraction results of each text line in the target paragraph may be processed, and the processed feature extraction results are the feature processing results of the target paragraph.
[0139] Step 105: Determine whether the target paragraph in the target document is a title based on the feature processing result.
[0140] In practical applications, after obtaining the feature processing results of each target paragraph, the feature processing results can be parsed to determine whether the target document is a title.
[0141] For example, the feature processing result can be compared with the set title condition to determine whether the feature processing result meets the set title condition: if it meets the condition, the target paragraph is a title; if it does not meet the condition, the target paragraph is not a title.
[0142] For example, the feature processing result may be input into a trained title recognition model to determine whether the target paragraph is a title.
[0143] After identifying the titles in the document, you can generate a document structure tree based on the title content, generate document bookmarks based on the title, generate a document directory based on the title, etc.
[0144] The method for identifying titles in documents provided by the present invention comprises the following steps: identifying at least one initial paragraph in a target document and at least one text line in each of the initial paragraphs, wherein the number of text lines in the initial paragraphs is less than a set threshold; performing feature extraction on each of the text lines for at least one dimension among a semantic dimension, a spatial dimension, and a visual dimension to obtain a feature extraction result for each of the text lines; screening out a target paragraph from the at least one initial paragraph based on each of the feature extraction results; processing the feature extraction results of each of the text lines in the target paragraph to obtain a feature processing result for the target paragraph; determining whether the target paragraph in the target document is a title based on the feature processing result; and being able to quickly and accurately identify titles in documents with good performance, extremely high accuracy, and support for multi-line title recognition.
[0145] In one or more optional embodiments of the present invention, the feature extraction is performed on each text line for at least one dimension among the semantic dimension, the spatial dimension, and the visual dimension to obtain the feature extraction result of each text line. The feature extraction is performed on each text line to obtain the feature extraction result of each text line. The specific implementation process can be as follows:
[0146] Perform at least one of the following feature extraction processes on each of the text lines:
[0147] Extracting features from each text line from a semantic dimension to obtain semantic feature extraction results for each text line;
[0148] Extracting features of each text line from a spatial dimension to obtain spatial feature extraction results of each text line;
[0149] Extract features from each of the text lines from a visual dimension to obtain visual feature extraction results for each of the text lines;
[0150] For each of the text lines, at least one of the semantic feature extraction result, the spatial feature extraction result, and the visual feature extraction result of the text line is processed to obtain a feature extraction result of the text line.
[0151] Specifically, the semantic feature extraction result is the result of feature extraction from the semantic dimension; the spatial feature extraction result is the result of feature extraction from the spatial dimension; and the visual feature extraction result is the result of feature extraction from the visual dimension.
[0152] In practical applications, for each text line, feature extraction can be performed on the text line from at least one dimension among semantic dimension, spatial dimension and visual dimension to obtain sub-feature extraction results under at least one dimension, that is, at least one of the following can be performed on the text line: feature extraction is performed on the text line from the semantic dimension to obtain the semantic feature extraction result of the text line, feature extraction is performed on the text line from the spatial dimension to obtain the spatial feature extraction result of the text line, and feature extraction is performed on the text line from the visual dimension to obtain the visual feature extraction result of the text line.
[0153] It should be noted that when extracting features from a text line from multiple dimensions, the features can be extracted from each dimension simultaneously, or in a certain order. The certain order can be a movement order, i.e., the order in which the text lines move on the display interface when reading the target document. The present invention does not impose any limitation on this.
[0154] Furthermore, at least one of the semantic feature extraction results, spatial feature extraction results, and visual feature extraction results of the text line is processed to obtain a feature extraction result for the text line. When feature extraction is performed on the text line from a single dimension, the sub-feature extraction results for that dimension can be used as the feature extraction result for the text line. When feature extraction is performed on the text line from multiple dimensions, the sub-feature extraction results for each dimension are concatenated according to a certain concatenation rule to obtain the feature extraction result for the text line.
[0155] According to the above method, each text line is traversed to obtain the feature extraction results of each text line.
[0156] In this way, feature extraction is performed from different dimensions, and the results of feature extraction in each dimension are spliced together to obtain the feature extraction results of the text line. This can ensure the comprehensiveness, reliability and accuracy of the feature extraction results, thereby improving the accuracy of title recognition.
[0157] In one or more optional embodiments of the present invention, the semantic feature extraction results include a sequence number feature extraction result, a special title feature extraction result, and a common title feature extraction result; and the feature extraction of each text line from a semantic dimension to obtain the semantic feature extraction result of each text line may be specifically implemented as follows:
[0158] For each of the text lines, perform the following steps:
[0159] Performing serial number recognition on the beginning of the text line; if the serial number is recognized, determining that the serial number feature extraction result of the text line has the serial number feature; if the serial number is not recognized, determining that the serial number feature extraction result of the text line does not have the serial number feature;
[0160] Matching the text line with a preset set of special title symbols; if the match is successful, determining that the special title feature extraction result of the text line is that the special title symbol exists; if the match fails, determining that the special title feature extraction result of the text line is that the special title symbol does not exist;
[0161] The text line is matched with a preset set of common title symbols; if the match is successful, the common title feature extraction result of the text line is determined to be the presence of common title symbols; if the match fails, the common title feature extraction result of the text line is determined to be the absence of the common title symbols.
[0162] Specifically, the sequence number feature extraction result is used to characterize whether the text line has a sequence number feature, where the sequence number feature is that the beginning of the text line contains a sequence number.
[0163] The special title feature extraction results are used to characterize whether there are uncommon title symbols in the text line, namely special title symbols, where special title symbols are characters in the special title character set, and the special title character set is a collection of symbols and characters that cannot appear in the title, such as "@#¥%…&" and so on.
[0164] The common title feature extraction results are used to characterize whether there are common title symbols in the text line, where common title symbols are characters in the common title character set, and the common title character set is a collection of symbols and characters that may appear in the title, such as ".()", ", etc.
[0165] In practical applications, when determining the result of sequence number feature extraction, the beginning of a text line, i.e., the beginning, can be subjected to sequence number recognition in the semantic dimension to determine the sequence number feature extraction result of the text line: if there is a sequence number at the beginning of the text line, the sequence number feature extraction result of the text line is determined to have sequence number features; if there is no sequence number at the beginning of the text line, the sequence number feature extraction result of the text line is determined to not have sequence number features. The sequence number here is a text that represents the sequence number, such as "1.", "2.1", "(1)", "One," "Chapter One," and "A." In this way, the accuracy and reliability of the sequence number feature extraction result can be improved.
[0166] For example, if there is a serial number text or serial number at the beginning of the Unicode text in the current text line, feature[serial number] = True, that is, the serial number feature extraction result has a serial number feature; otherwise, feature[serial number] = False, that is, the serial number feature extraction result does not have a serial number feature.
[0167] In addition, whether it is a streaming document or a layout document, a preset feature set such as "1.", "2.1", "(1)", "One," "Chapter One," and "A," is obtained by querying whether the Unicode at the beginning of the text line exists, and the length N of the serial number text or serial number is obtained.
[0168] When determining the special title feature extraction result, the text line can be matched with a preset special title symbol set, that is, whether the text line contains characters in the special title symbol set. If so, the special title feature extraction result for the text line is determined to be the presence of special title symbols; if not, the special title feature extraction result for the text line is determined to be the absence of special title symbols. This can improve the accuracy and reliability of the special title feature extraction result.
[0169] For example, if the current text line or starting from the Nth character contains a Unicode character in the preset character set 1 (special title symbol set), then feature[special title] = True, indicating that the special title feature extraction result for the line contains a special title symbol. Otherwise, feature[special title] = False, indicating that the special title feature extraction result does not contain a special title symbol. Unicode characters in the preset character set 1 are symbols that are not typically found in titles.
[0170] When determining the common title feature extraction results, the text line can be matched against a preset set of common title symbols. Specifically, a determination is made as to whether the text line contains characters from the set. If so, the common title feature extraction result for the text line is determined to be the presence of common title symbols; if not, the common title feature extraction result for the text line is determined to be the absence of common title symbols. This improves the accuracy and reliability of the common title feature extraction results.
[0171] For example, if there is a Unicode character that may be in the preset character set 2 (common title symbol set) in the current text line or starting from the Nth character, then feature[common title] = True, i.e., the common title feature extraction result is that common title symbols exist; otherwise, feature[common title] = False, i.e., the common title feature extraction result is that common title symbols do not exist. The Unicode characters in the preset character set 2 are symbols that may appear in common titles.
[0172] In one or more optional embodiments of the present specification, the spatial feature extraction result includes a content feature extraction result, a first font size feature extraction result, and a second font size feature extraction result; and the feature extraction of each text line from a spatial dimension to obtain the spatial feature extraction result of each text line may be specifically implemented as follows:
[0173] For each of the text lines, perform the following steps:
[0174] If the content of the text line is all text content, determining that the content feature extraction result of the text line has text content features; if the content of the text line contains non-text content, determining that the content feature extraction result of the text line does not have the text content features;
[0175] Determining whether a first average font size corresponding to the text line is greater than a second average font size, where the second average font size is the average font size of all text in the page to which the text line belongs; if so, determining that the first font size feature extraction result of the text line has the font size feature; if not, determining that the first font size feature extraction result of the text line does not have the font size feature;
[0176] In the absence of a first designated text line, determining that the second font size feature extraction result of the text line is a first value, and the first designated text line is the next text line of the text line; in the presence of the first designated text line, determining whether the first average font size corresponding to the text line is greater than the third average font size of the first designated text line; if so, determining that the second font size feature extraction result of the text line is a second value; if not, determining that the second font size feature extraction result of the text line is a third value; the first value, the second value and the third value are all different.
[0177] Specifically, the content feature extraction result is used to characterize whether a text line has text content features, where the text content feature indicates that all content in the text line is text. Text content refers to text; non-text content refers to elements such as images and / or vector paths.
[0178] The first font size feature extraction result is used to characterize whether the text line has a font size feature, where the font size feature is that the average font size of the current text line is greater than the average font size of the current page.
[0179] The second font size feature extraction result is used to indicate whether a text line has a next text line and whether the average font size of the current text line is greater than the average font size of the next text line. The first value indicates that there is no text line after the current text line, that is, there is no next text line; the second value indicates that the current text line has a next text line and the average font size of the current text line is greater than the average font size of the next text line; and the second value indicates that the current text line has a next text line and the average font size of the current text line is less than or equal to the average font size of the next text line.
[0180] In practical applications, when determining the content feature extraction results, it is possible to determine whether the content in the text line is all text content. If so, the content feature extraction result for the text line is determined to have text content features; if not, the content feature extraction result for the text line is determined to not have text content features. This can improve the accuracy and reliability of the content feature extraction results.
[0181] For example, if all the content in the current text line is text (text content) and there is no non-text content such as pictures or vector paths, then feature[text content] = True, that is, the content feature extraction result has text content features; otherwise, feature[text content] = False, that is, the content feature extraction result does not have text content features.
[0182] When determining the first font size feature extraction result, the average font size of all text in the text line can be calculated to obtain a first average font size, and the average font size of all text in the page associated with the text line can be calculated to obtain a second average font size. Furthermore, a determination is made as to whether the first average font size is greater than the second average font size. If so, the first font size feature extraction result for the text line is determined to have a font size feature; if not, the first font size feature extraction result for the text line is determined to not have a font size feature. This improves the accuracy and reliability of the first font size feature extraction result.
[0183] For example, if the average font size of all text in the current text line (first average font size) is greater than the average font size of all text in the entire page (second average font size), then feature[first font size] = True, that is, the first font size feature extraction result has font size features; otherwise, feature[first font size] = False, that is, the first font size feature extraction result does not have font size features.
[0184] The average font size (average font size) in the layout document is: the average value of the FontSize used by each Unicode in all TextObjects.
[0185] For example, a TextObject may contain multiple Unicode characters. For example, there are two TextObjects in a row. The first TextObject contains 3 characters with a font size of 10, and the second TextObject contains 7 characters with a font size of 15. The average font size is (3*10+7*15) / (3+7), not (10+15) / 2.
[0186] When determining the second font size feature extraction result, it is possible to first determine whether the current text line has a next text line, namely, a first designated text line. If not, the second font size feature extraction result of the text line is determined to be the first value. If so, the first average font size corresponding to the text line and the third average font size of the first designated text line are obtained. Furthermore, it is determined whether the first average font size is greater than the third average font size. If so, the second font size feature extraction result of the text line is determined to be the second value. If not, the second font size feature extraction result of the text line is determined to be the third value. In this way, the accuracy and reliability of the second font size feature extraction result can be improved.
[0187] For example, if the average font size of all text in the current text line (the first specified text line) is greater than the average font size of all text in the next line (the first specified text line) (the third average font size), then feature[second font size] = 2 (second numerical value), that is, the second font size feature extraction result of the text line is determined to be the second numerical value; otherwise, feature[second font size] = 1 (third numerical value), that is, the second font size feature extraction result of the text line is determined to be the third numerical value; if there is no next line or there is no text in the next line, then feature[second font size] = 0 (first numerical value), that is, the second font size feature extraction result of the text line is determined to be the first numerical value.
[0188] In one or more optional embodiments of the present invention, the visual feature extraction results include font feature extraction results, modification feature extraction results, and interval feature extraction results; and the feature extraction of each text line from a visual dimension to obtain the visual feature extraction results of each text line may be specifically implemented as follows:
[0189] Performing font recognition on each text content in the text line to obtain a font feature extraction result of the text line;
[0190] Identifying modified content on the text line, the modified content being used to modify the text content in the text line; if the modified content is identified, determining that a modified feature extraction result of the text line has a modified feature; if the modified content is not identified, determining that a modified feature extraction result of the text line does not have the modified feature;
[0191] Perform associated interval matching on the text line to obtain an interval feature extraction result of the text line.
[0192] Specifically, the font feature extraction result is used to characterize whether the text line has font features, wherein the font features are used to characterize features corresponding to the font of the text in the text line.
[0193] The modification feature extraction result is used to characterize whether the text line has a modification feature, where the modification feature is the presence of modification content in the current text line, such as underline, strikethrough, etc.
[0194] The interval feature extraction result is used to represent the interval feature corresponding to or associated with the text line, wherein the interval feature represents the serial number corresponding to or associated with the text line.
[0195] In practical applications, when determining the font feature extraction result, font recognition can be performed on each text content in the text line, and the font feature extraction result of the text line can be determined based on the font recognition result. In this way, the efficiency of determining the font feature extraction result can be improved.
[0196] When determining the modified feature extraction result, the modified content of the text line can be identified, that is, whether the text line contains modified content. If modified content is present, the modified feature extraction result of the text line is determined to have the modified feature. If no modified content is present, the modified feature extraction result of the text line is determined to not have the modified feature. In this way, the accuracy and reliability of the modified feature extraction result can be improved.
[0197] For example, if there is modifying content that modifies the text content in the current text line, feature[modification]=True, that is, the modification feature extraction result has the modification feature; otherwise, feature[modification]=False, that is, the modification feature extraction result does not have the modification feature.
[0198] When determining the interval feature extraction result, the text line may be matched with associated intervals, and the interval feature extraction result of the text line may be determined based on the result of the associated interval matching. In this way, the efficiency of determining the interval feature extraction result may be improved.
[0199] In one or more optional embodiments of the present invention, the font feature extraction result includes a first bold feature extraction result, a second bold feature extraction result, and a font same feature extraction result; the font recognition is performed on each text content in the text line to obtain the font feature extraction result of the text line. The specific implementation process can be as follows:
[0200] Determine whether each text content in the text line is in a bold font; if so, determine that the first bold feature extraction result of the text line has all the bold features; if not, determine that the first bold feature extraction result of the text line does not have all the bold features;
[0201] Determine whether there is text content in a bold font in the text line; if so, determine that the second bold feature extraction result of the text line has a partial font bold feature; if not, determine that the second bold feature extraction result of the text line does not have the partial font bold feature;
[0202] Determine whether the fonts of each text content in the text line are the same; if so, determine that the font same feature extraction result of the text line has the font same feature; if not, determine that the font same feature extraction result of the text line does not have the font same feature.
[0203] Specifically, the first bold feature extraction result is used to indicate whether the text line has a full bold feature, where the full bold feature indicates that the text content in the text line is all bold fonts. Bold fonts refer to fonts that have a bold attribute, have bold characteristics (such as boldface), have a Bold attribute of True, or have a FontWeight greater than a bold threshold (such as 700).
[0204] The second bold feature extraction result is used to indicate whether the text line has a partial font bold feature, where the partial font bold feature is the presence of text content in a bold font in the current text line.
[0205] The result of the font same feature extraction is used to characterize whether a text line has the font same feature, wherein the font same feature means that the text contents in the text line are all in the same font.
[0206] In practical applications, when determining the first bold feature extraction result, it is possible to determine whether each text content (text element) in the text line is in a bold font. If so, the first bold feature extraction result for the text line is determined to have all bold features. If not, the first bold feature extraction result for the text line is determined to not have all bold features. This can improve the accuracy and reliability of the first bold feature extraction result.
[0207] For example, if all text elements (text content) in the current text line have the bold attribute, then feature[all fonts bold] = True, that is, the first bold feature extraction result is to have all fonts bold features; otherwise, feature[all fonts bold] = False, that is, the first bold feature extraction result is to not have all fonts bold features.
[0208] For streaming documents, you can directly obtain the bold attribute (i.e., the attribute value) of the font to determine whether it is a bold font. For layout documents, you need to query the font used in the TextObject to see if it has bold properties (such as black font, which has bold properties), or whether the Bold property is True, or whether the FontWeight is greater than a given threshold (e.g., 700).
[0209] When determining the second bold feature extraction result, it can be determined whether the text line contains text content in a bold font. If so, the second bold feature extraction result for the text line is determined to have the partial bold feature. If not, the second bold feature extraction result for the text line is determined to not have the partial bold feature. In this way, the accuracy and reliability of the second bold feature extraction result can be improved.
[0210] For example, if there is a text element (text content) with a bold attribute in the current text line, that is, whether there is text content in a bold font in the text line, then feature[partial font bold] = True, that is, the second bold feature extraction result is that it has a partial font bold feature; otherwise, feature[partial font bold] = False, that is, the second bold feature extraction result is that it does not have a partial font bold feature.
[0211] When determining the result of the font identity feature extraction, it is possible to determine whether the fonts of all text contents in the text line are the same. If they are, the font identity feature extraction result of the text line is determined to have the font identity feature. If not, the font identity feature extraction result of the text line is determined to not have the font identity feature. This can improve the accuracy and reliability of the font identity feature extraction result.
[0212] For example, if all text elements (text content) in the current text line are written and drawn using the same font, then feature[same font] = True, that is, the result of the font same feature extraction is the font same feature; otherwise, feature[same font] = False, that is, the result of the font same feature extraction is not the font same feature.
[0213] In one or more optional embodiments of the present invention, the interval feature extraction result includes a font size interval feature extraction result and a character number interval feature extraction result; and the associated interval matching is performed on the text line to obtain the interval feature extraction result of the text line. The specific implementation process may be as follows:
[0214] Determining a first target font size interval from at least one preset initial font size interval based on a first average font size corresponding to the text line; determining a font size interval feature extraction result of the text line as a first serial number, the first serial number representing a ranking of the first target font size interval in each of the initial font size intervals;
[0215] Based on the first number of characters contained in the text line, a first target character number interval is determined from at least one preset initial character number interval; the character number interval feature extraction result of the text line is determined as a second serial number, and the second serial number represents the ranking of the first target character number interval in each of the initial character number intervals.
[0216] Specifically, the font size interval feature extraction result is used to represent the first target font size interval corresponding to the text line. A font size interval refers to the interval corresponding to a preset font size range, which is obtained by dividing the minimum value Min and the maximum value Max of the font size into intervals n. For example, the interval is divided into intervals with a minimum value of 1, a maximum value of 25, and an interval of 1.
[0217] The character count interval feature extraction result is used to represent the first target character count interval corresponding to the text line. A character count interval is the interval corresponding to a preset character count range. This is obtained by dividing the minimum and maximum character counts in the title into intervals m. For example, the minimum and maximum character counts are 1, 50, and the interval is 2.
[0218] When determining the font size interval feature extraction results, several initial font size intervals can be set and a first average font size corresponding to the text line can be obtained. The initial font size interval to which the first average font size belongs is then determined as the first target font size interval. The font size interval feature extraction result for the text line is then determined as the sequence number of the first target font size interval within each of the initial font size intervals, i.e., the first sequence number. This improves the accuracy and reliability of the font size interval feature extraction results.
[0219] For example, the interval is divided into 24 initial font size intervals, i.e., interval 0 to interval 23, with a minimum value of 1 and a maximum value of 25. If the average font size of all text in the current text line (the first average font size) falls within the i-th initial font size range (the first target font size range), then feature[font size range] = i.
[0220] It should be noted that if the first average font size is smaller than the minimum font size, the first target font size zone is the 0th initial font size zone; if it is larger than the maximum font size, the first target font size zone is the last initial font size zone.
[0221] When determining the character number interval feature extraction results, you can first set several initial character number intervals and obtain the first number of characters contained in the text line, where the first number of characters can be the number of all characters contained in the text line, or the number of all characters in the text line except the serial number, that is, the number of all characters in the text line minus N, where N is the length of the serial number text.
[0222] Furthermore, the initial character count interval to which the first character count belongs is determined as the first target character count interval, and the character count interval feature extraction result of the text line is determined as the serial number of the first target character count interval within each initial character count interval, i.e., the second serial number. In this way, the accuracy and reliability of the character count interval feature extraction result can be improved.
[0223] For example, the minimum font size is 1, the maximum font size is 50, and the interval is 2, resulting in 24 initial character number intervals, namely interval 0 to interval 23. If the result of subtracting N from the number of characters in the current text line falls within the jth initial character number interval, then feature[character number interval] = j.
[0224] It should be noted that if the first number of characters is less than the minimum number of characters, the first target character number interval is the 0th initial character number interval; if the first number of characters is greater than the maximum number of characters, the first target character number interval is the last initial character number interval.
[0225] In one or more optional embodiments of the present invention, screening out a target paragraph from each of the initial paragraphs based on the at least one feature extraction result includes:
[0226] For each of these initial paragraphs, perform the following steps:
[0227] When the number of text lines in the initial paragraph is a first preset value, determining the initial paragraph as a target paragraph;
[0228] When the number of text lines in the initial paragraph is not the first preset value, determine whether the feature extraction results of each text line in the initial paragraph meet the set merging conditions; if so, determine the initial paragraph as the target paragraph.
[0229] In practical applications, the initial paragraph containing the first preset number of text lines can be directly used as the target paragraph.
[0230] For an initial paragraph containing text lines whose number is not the first preset value, it is necessary to determine whether the initial paragraph is a target paragraph based on whether the text lines in the initial paragraph can be merged: if the feature extraction results of the text lines in the initial paragraph meet the set merging conditions, that is, the text lines can be merged, then the initial paragraph is the target paragraph; if the feature extraction results of the text lines in the initial paragraph do not meet the set merging conditions, that is, the text lines cannot be merged, then the initial paragraph is not the target paragraph.
[0231] In this way, by combining the number of text lines and feature extraction results to filter target paragraphs, the integrity of the filtered target paragraphs can be guaranteed. Furthermore, by filtering target paragraphs based on the set merging conditions, the accuracy of the screening can be guaranteed.
[0232] It should be noted that if the number of text lines in the initial paragraph is not the first preset value, it is necessary to determine whether the initial paragraph is the target paragraph based on whether the text lines in the initial paragraph can be merged. The premise for merging text lines is that there are multiple text lines; if there is only one text line, merging is not required. Therefore, the first preset value here is one. Accordingly, an initial paragraph containing one text line can be directly used as the target paragraph; for an initial paragraph containing more than one text line, it is necessary to determine whether the initial paragraph is the target paragraph based on whether the text lines in the initial paragraph can be merged.
[0233] In one or more optional embodiments of the present invention, determining whether the feature extraction results of each text line in the initial paragraph meet the set merging conditions includes:
[0234] Obtaining the alignment and first average font size of each text line in the initial paragraph;
[0235] determining a first absolute value of a difference between first average font sizes corresponding to adjacent text lines in the initial paragraph based on each of the first average font sizes;
[0236] It is determined whether the feature extraction results and the alignment methods of the text lines in the initial paragraph, as well as the first absolute values, all meet the set merging conditions.
[0237] In practical applications, the alignment of each text line in the initial paragraph and the first average font size corresponding to each text line can be obtained. Then, based on the first average font size corresponding to each text line, the difference between the first average font sizes corresponding to two adjacent text lines in the initial paragraph is calculated, and the absolute value of the difference is determined, thereby obtaining the first absolute value.
[0238] Furthermore, it is determined whether the feature extraction results, alignment, and first absolute value of each text line in the initial paragraph meet the set merging conditions. If so, the initial paragraph is the target paragraph; if not, the initial paragraph is not the target paragraph. This can further improve the reliability and accuracy of determining the target paragraph.
[0239] It should be noted that the feature extraction results include content feature extraction results, first bold feature extraction results, second bold feature extraction results, font same feature extraction results, modification feature extraction results, first font size feature extraction results and sequence number feature extraction results;
[0240] The feature extraction results of each text line in the initial paragraph meet the set merging conditions, including:
[0241] The content feature extraction results of each of the text lines in the initial paragraph are the same, the first bold feature extraction results are the same, the second bold feature extraction results are the same, the font same feature extraction results are the same, the modification feature extraction results are the same, and the first font size feature extraction results are the same; and
[0242] The result of the sequence number feature extraction of the first text line in the initial paragraph is that it has the sequence number feature, and the results of the sequence number feature extraction of each specific text line in the initial paragraph do not have the sequence number feature, and the specific text lines are the text lines in the initial paragraph other than the first text line;
[0243] Each of the first absolute values corresponding to the initial paragraphs meets a set merging condition, including:
[0244] Each of the first absolute values corresponding to the initial paragraph is smaller than the set font size difference;
[0245] The alignment of each text line in the initial paragraph meets the set merging conditions, including:
[0246] The alignment of each text line in the initial paragraph is the same.
[0247] The font size difference is set to a preset font size difference, such as 0.1.
[0248] For example, the merging conditions may include the following (1)-(4):
[0249] (1) The features [text content], [all bold], [partial bold], [same font], [modification], and [first font size] of each text line in the initial paragraph are all equal.
[0250] (2) The value of feature[serial number] can be True only when it is the first line. If it is not the first line, all values are False. That is, the feature[serial number] of the first text line in the initial paragraph is True, and the feature[serial number] of a specific text line is False.
[0251] (3) The first absolute value of the first average font size difference between adjacent text lines in the initial paragraph does not exceed 0.1 (the set font size difference).
[0252] (4) Adjacent lines of text in the initial paragraph are aligned in the same way.
[0253] In one or more optional embodiments of the present invention, the obtaining of the alignment of each text line in the initial paragraph may be specifically implemented as follows:
[0254] For each of the text lines in the initial paragraph, perform the following steps:
[0255] Obtaining the bounding rectangle and line height of the text line, and obtaining the visual area of the page to which the text line belongs;
[0256] Calculating a left distance between a left border of the circumscribed rectangular frame and a left border of the visualization area, and calculating a right distance between a right border of the circumscribed rectangular frame and a right border of the visualization area;
[0257] If the left spacing and the right spacing satisfy the centered alignment condition, then determine that the alignment of the text line is centered alignment. The centered alignment condition is that the second absolute value of the difference between the left spacing and the right spacing is greater than the line height, and both the left spacing and the right spacing are greater than the product of the line height and a set value;
[0258] If the left spacing and the right spacing satisfy the right alignment condition, then determine that the alignment of the text line is right alignment. The right alignment condition is that the left spacing is less than the right spacing, and the right spacing is less than the product;
[0259] If the left spacing and the right spacing do not satisfy the centered alignment condition and do not satisfy the right alignment condition, then determine that the alignment of the text line is left alignment.
[0260] In practical applications, the union of the bounding rectangles of all elements (text content) within the text line is the bounding rectangle of the text line, denoted as LineBBox, and the line height of the text line is denoted as height. Denote the visual area of the current page (the page to which the text line belongs) as PageBBox.
[0261] Then, subtract the left boundary PageBBox.left of the visual area from the left boundary LineBBox.left of the bounding rectangle to obtain the left spacing leftLength, that is, leftLength = LineBBox.left - PageBBox.left; subtract the right boundary LineBBox.right of the bounding rectangle from the right boundary PageBBox.right of the visual area to obtain the right spacing rightLength, that is, rightLength = PageBBox.right - LineBBox.right.
[0262] If |leftLength - rightLength| < height and leftLength > set value (such as 2.5) * height, and rightLength > set value (such as 2.5) * height, then the alignment of the text line is centered alignment.
[0263] If leftLength < rightLength and rightLength > set value (such as 2.5) * height, then the alignment of the text line is right alignment.
[0264] In all remaining cases, the alignment of the text line is left alignment.
[0265] In this way, the accuracy and reliability of determining the alignment can be improved.
[0266] In one or more optional embodiments of the present invention, the feature extraction results of each text line in the target paragraph are processed to obtain the feature processing result of the target paragraph. The specific implementation process may be as follows:
[0267] When the number of text lines in the target paragraph is a second preset value, using the feature extraction results of the text lines in the target paragraph as the feature processing results of the target paragraph;
[0268] When the number of text lines in the target paragraph is not the second preset value, the feature extraction results of the text lines in the target paragraph are merged according to the set merging strategy to obtain the feature processing result of the target paragraph.
[0269] In actual applications, for a target paragraph that only contains a second preset number of text lines, the feature extraction results of the text lines can be directly used as the feature processing results of the target paragraph.
[0270] For a target paragraph containing text lines whose number is not the second preset value, the feature extraction results of the text lines in the target paragraph need to be merged according to the set merging strategy to obtain the feature processing result of the target paragraph.
[0271] In this way, the feature processing result of the target paragraph is determined by combining the number of text lines and the feature extraction result, which can reduce the number of merged feature extraction results to a certain extent, thereby reducing the amount of data processing.
[0272] It should be noted that, since the number of text lines in the initial paragraph is not the second preset value, the feature extraction results of each text line in the target paragraph need to be merged according to the set merging strategy to obtain the feature processing result of the target paragraph. The premise for merging the feature extraction results is that there are multiple feature extraction results. For the case of only one feature extraction result, merging is not required, and one feature extraction result corresponds to one text line. Therefore, the second preset value here is one. Accordingly, for an initial paragraph containing one text line, the feature extraction result of the text line can be directly used as the feature processing result of the target paragraph; for an initial paragraph containing more than one text line, the feature extraction results of each text line in the target paragraph need to be merged according to the set merging strategy to obtain the feature processing result of the target paragraph.
[0273] In one or more optional embodiments of the present invention, the feature extraction results include content feature extraction results, first bold feature extraction results, second bold feature extraction results, font same feature extraction results, sequence number feature extraction results, special title feature extraction results, common title feature extraction results, modification feature extraction results, first font size feature extraction results, second font size feature extraction results, font size interval feature extraction results, and character number interval feature extraction results;
[0274] Merging the feature extraction results of each text line in the target paragraph according to a set merging strategy to obtain the feature processing result of the target paragraph includes:
[0275] The content feature extraction result, the first bold feature extraction result, the second bold feature extraction result, the same font feature extraction result, the modified feature extraction result, and the first font size feature extraction result of the first text line in the target paragraph are respectively used as the content feature extraction result, the first bold feature extraction result, the second bold feature extraction result, the same font feature extraction result, the modified feature extraction result, and the first font size feature extraction result of the target paragraph;
[0276] Performing an OR operation on the sequence number feature extraction results of each text line in the target paragraph to obtain the sequence number feature extraction result of the target paragraph; performing an OR operation on the special title feature extraction results of each text line in the target paragraph to obtain the special title feature extraction result of the target paragraph; performing an OR operation on the common title feature extraction results of each text line in the target paragraph to obtain the common title feature extraction result of the target paragraph;
[0277] In the absence of a second designated text line, determining that the second font size feature extraction result of the target paragraph is a first value, and the second designated text line is the next text line of the target paragraph; in the presence of the second designated text line, performing an average calculation on the first average font sizes corresponding to each text line in the target paragraph to obtain a target average font size, and determining whether the target average font size is greater than a third average font size of the second designated text line; if so, determining that the second font size feature extraction result of the target paragraph is a second value; if not, determining that the second font size feature extraction result of the target paragraph is a third value; the first value, the second value, and the third value are all different;
[0278] determining a second target font size interval from at least one preset initial font size interval based on the fourth average font size of the target paragraph; determining a font size interval feature extraction result of the target paragraph as a third serial number, the third serial number representing a ranking of the second target font size interval in each of the initial font size intervals;
[0279] Determining a second target character number interval from at least one preset initial character number interval based on a second number of characters included in the target paragraph; determining a character number interval feature extraction result of the text line as a fourth serial number, the fourth serial number representing a ranking of the second target character number interval in each of the initial character number intervals;
[0280] The feature processing result of the target paragraph includes the content feature extraction result of the target paragraph, the first bold feature extraction result, the second bold feature extraction result, the font same feature extraction result, the serial number feature extraction result, the special title feature extraction result, the common title feature extraction result, the modification feature extraction result, the first font size feature extraction result, the second font size feature extraction result, the font size interval feature extraction result and the character number interval feature extraction result.
[0281] In practical applications, the features [text content], [all bold], [part bold], [same font], [modification], and [first font size] of the first text line in the target paragraph are used as the features [text content], [all bold], [part bold], [same font], [modification], and [first font size] of the target paragraph, respectively.
[0282] Perform an OR operation on the feature[serial number] of each text line in the target paragraph to obtain the feature[serial number] of the target paragraph; perform an OR operation on the feature[special title] of each text line in the target paragraph to obtain the feature[special title] of the target paragraph; perform an OR operation on the feature[common title] of each text line in the target paragraph to obtain the feature[common title] of the target paragraph.
[0283] According to the first average font size of each text line in the target paragraph, the feature [second font size] of the target paragraph is obtained. That is, the average font size of consecutive lines is obtained by averaging the average font sizes of multiple lines, and the feature [second font size] of the target paragraph is recalculated.
[0284] Specifically, first determine whether there is a next text line in the target paragraph, that is, the second designated text line. If so, determine that the second font size feature extraction result of the target paragraph is the first value, that is, feature[second font size] of the target paragraph = 0 (first value); if not, calculate the average of the first average font sizes corresponding to each text line in the target paragraph to obtain the target average font size, and then determine whether the target average font size is greater than the third average font size of the second designated text line; if greater, determine that the second font size feature extraction result of the target paragraph is the second value, that is, feature[second font size] of the target paragraph = 2 (second value); if not greater, determine that the second font size feature extraction result of the target paragraph is the third value, that is, feature[second font size] of the target paragraph = 1 (first value).
[0285] The fourth average font size of the target paragraph is calculated, or the fourth average font size of the target paragraph is calculated based on the first average font size corresponding to each text line in the target paragraph. Then, the initial font size range to which the fourth average font size belongs is determined as the second target font size range, and feature [font size range] of the target paragraph is determined as the third sequence number of the second target font size range.
[0286] The number of second characters included in the target paragraph is counted, and then the initial character number interval to which the second character number belongs is determined as the second target character number interval, and feature[character number interval] of the target paragraph is determined as the fourth serial number of the second target character number interval.
[0287] In this way, the accuracy of merging can be improved, thereby improving the accuracy and reliability of feature processing results.
[0288] In one or more optional embodiments of the present invention, determining whether the target paragraph in the target document is a title according to the feature processing result includes:
[0289] Determine whether the target paragraph in the target document is a title based on the feature processing result and the set title condition; or
[0290] The feature processing result is input into a trained title recognition model to determine whether the target paragraph in the target document is a title.
[0291] In practical applications, a title template can be constructed. Specifically, title conditions are set, and then the feature processing results of the target paragraph are compared with the set title conditions. If the feature processing results meet the set title conditions, the target paragraph is considered a title. If the feature processing results do not meet the set title conditions, the target paragraph is not a title. This can improve the accuracy of title recognition.
[0292] A title recognition model can also be pre-trained, for example, by training an initial model based on sample feature processing results and labels (such as the isHeading flag) to obtain a title recognition model. The initial model can be a machine learning classifier, including but not limited to logistic regression, support vector machine, neural network, decision tree, random forest, Bayesian classifier, etc.
[0293] After obtaining the title recognition model, the feature processing results of the target paragraph can be input into the trained title recognition model to determine whether the target paragraph is a title. In this way, title recognition can be performed in large quantities, improving the efficiency of title recognition.
[0294] It should be noted that the feature processing results include content feature extraction results, first bold feature extraction results, second bold feature extraction results, font same feature extraction results, sequence number feature extraction results, special title feature extraction results, common title feature extraction results, modification feature extraction results, first font size feature extraction results, second font size feature extraction results, font size interval feature extraction results, and character number interval feature extraction results;
[0295] The title setting conditions include:
[0296] The content feature extraction result, the first bold feature extraction result, the second bold feature extraction result, the same font feature extraction result, the serial number feature extraction result, the special title feature extraction result, the common title feature extraction result, the modification feature extraction result, and the first font size feature extraction result are all characterized by having corresponding features; and
[0297] The second font size feature extraction result is a second value; and
[0298] The font size interval feature extraction result is greater than a first threshold; and
[0299] The character number interval feature extraction result is less than a second threshold.
[0300] Specifically, setting the title conditions includes the following (1)-(4).
[0301] (1) The target paragraph’s features [text content], [all bold], [partial bold], [same font], [serial number], [special title], [common title], [modification], and [first font size] are all True.
[0302] (2) The feature[second font size] of the target paragraph is the second value, such as 2.
[0303] (3) The feature [font size interval] of the target paragraph is greater than a first threshold, such as 10.
[0304] (4) The feature [character number interval] of the target paragraph is less than a second threshold, such as 10.
[0305] For example, the method of the title in the document is as follows:
[0306] 1. Preconditions:
[0307] Line identifier. If it is a layout document, it needs to be determined based on the position of the TextObject in the document. If it is a flow document, no special identification is required.
[0308] Paragraph recognizer. If the document is a layout document, it needs to be determined according to the method of identifying paragraphs in layout documents. If the document is a flow document, no special operation is required.
[0309] 2. Get all text lines that are not in tables, headers, or footers, and obtain a set of candidate header lines.
[0310] 3. Based on the following rules, extract the features of the text lines in each title candidate line set. Then each text line can correspond to a feature vector feature.
[0311] (1) Determine the content feature extraction result: If all the content in the current text line is text and there is no picture, vector path, etc., then feature[0] = True, otherwise it is False.
[0312] (2) Determine the first bold feature extraction result: If all text contents in the current text line have the bold attribute, then feature[1] = True, otherwise it is False. For layout documents, it is necessary to query whether the font used in the TextObject is a bold font, or whether the Bold attribute is True, or whether the FontWeight is greater than a given threshold (e.g., 700).
[0313] (3) Determine the second bold feature extraction result: If there is a text element with a bold attribute in the current text line, then feature[2] = True, otherwise it is False.
[0314] (4) Determine the result of the font-identical feature extraction: If all text elements in the current text line are written and drawn using the same font, then feature[3] = True, otherwise it is False.
[0315] (5) Determine the extraction result of serial number feature: If there is a serial number feature at the beginning of the text Unicode of the current text line, then feature[4] = True, otherwise it is False. Whether it is a streaming or layout document, it is to query whether there are preset feature sets such as "1.", "2.1", "(1)", "一、", "第一章" etc. at the beginning of the Unicode of the text line, and obtain the length N of the serial number text.
[0316] (6) Determine the extraction result of special title feature: If there are Unicodes belonging to the possible preset character set 1 starting from the Nth character of the text in the current text line, then feature[5] = True, otherwise it is False. The Unicodes in the preset character set 1 are symbols that usually do not appear in titles. For example, "@#¥%…&" etc.
[0317] (7) Determine the extraction result of common title feature: If there are Unicodes belonging to the possible preset character set 2 starting from the Nth character of the text in the current text line, then feature[6] = True, otherwise it is False. The Unicodes in the preset character set 2 are symbols that may appear in titles usually. For example, ".()、" etc.
[0318] (8) Determine the extraction result of modification feature: If there is text modification content in the current text line, such as underlines, strikethroughs, etc., then feature[7] = True, otherwise it is False.
[0319] (9) Determine the extraction result of the first font size feature: If the average font size of all the text in the current text line is greater than the average font size of all the text on the entire page, then feature[8] = True, otherwise it is False.
[0320] (10) Determine the extraction result of the second font size feature: If the average font size of all the text in the current text line is greater than the average font size of all the text in the next line, then feature[9] = 2, otherwise it is 1. If there is no next line or the next line has no text, then it is 0.
[0321] (11) Determine the extraction result of font size interval feature: Divide the interval with a minimum value of 1, a maximum value of 25, and an interval of 1. If the average font size of all the text in the current text line falls into the i-th interval, then feature
[10] = i. Among them, if the font size is less than the minimum value, it is the 0th interval. If it is greater than the maximum value, it is the last interval.
[0322] (12) Determine the result of character number interval feature extraction: The interval is divided into intervals with a minimum value of 1, a maximum value of 50, and an interval of 2. If the result after subtracting N from the number of characters in the current text line falls in the i-th interval, then feature
[11] = i. If the number is less than the minimum value, it is the 0th interval; if it is greater than the maximum value, it is the last interval.
[0323] 4. Traverse all paragraphs in the current document. If the number of text lines in a paragraph is greater than 1 and less than a given threshold (MaxLine, which can be set to 4 based on experience), then perform a merge feature rule. Once the merge condition is met, it is considered that all lines in the current paragraph may belong to the same title, that is, there may be a multi-line title, and they can be merged according to the merge rule. The merged multiple lines can obtain characteristics similar to those of a single line, and can be subsequently identified and processed in the same way as a single line.
[0324] (1) Definition: Assume that the feature of line i in the current paragraph is curFeature, and the feature of line i+1 is nextFeature. The total number of lines in the current paragraph is numLine.
[0325] (2) Merge conditions: The 0th, 1st, 2nd, 3rd, 7th, and 8th features of curFeature and nextFeature are all equal; the 4th feature can be True only when it is the first row, and is False when it is not the first row; the absolute value of the average font size difference between the i-th row and the i+1-th row does not exceed 0.1; the alignment of the i-th row and the i+1-th row is the same.
[0326] (3) Merging rules: The 0th, 1st, 2nd, 3rd, 7th and 8th features are respectively taken as the feature results of the first line; the 4th, 5th and 6th features are respectively taken as the result of the OR of the corresponding bits of multiple lines of features; the 10th and 11th feature share ratios are taken as the result of recalculating the Unicode splicing of all lines of text in the current paragraph; the average font size of consecutive lines is obtained by averaging the average font sizes of multiple lines, and the 9th feature is recalculated and updated.
[0327] (4) Alignment calculation:
[0328] A. The bounding rectangle of all elements in the current row is recorded as LineBBox, the visual area of the current PDF page is recorded as PageBBox, and the row height is height;
[0329] B. leftLength=LineBBox.left-PageBBox.left;
[0330] C. rightLength=PageBBox.right-LineBBox.right;
[0331] D. If |leftLength - rightLength| < height and leftLength > 2.5 * height and rightLength > 2.5 * height, then it is center alignment;
[0332] E. If leftLength < rightLength and rightLength > 2.5 * height, then it is right alignment;
[0333] F. In all other cases, it is left alignment.
[0334] 5. After obtaining all the merged feature results, a determination can be made. The determination method can be pattern matching or machine learning. The application scenarios of the two are different. The former uses pattern recognition and requires artificial design of the pattern of the title features. It is usually used when there are strict requirements for the accuracy rate of title detection, but the recall rate can be relaxed. The latter uses the machine learning mode of big data and requires a suitable data set and a large amount of data. It is usually used when there are requirements for both the accuracy rate and the recall rate, and a clean training data set can be constructed. The following are examples of using the two methods
[0335] (1) Construct the title pattern: When the 0th to 8th bits in feature are all True, the value of the 9th bit is 2, the value of the 10th bit is greater than 10, and the value of the 11th bit is less than 10, then it is considered that the (multi-) line corresponding to this feature is a title.
[0336] [[ID=^{}15]](2) Give an additional isHeading flag bit to indicate whether the current (multi-) line is a title. Extract a large number of line features and label each line feature. If it is a title, then isHeading = True, otherwise it is False. Then use a machine learning classifier, including but not limited to logistic regression, support vector machine, neural network, decision tree, random forest, Bayesian classifier, etc., to learn and predict the input data.
[0337] See Figure 2 , Figure 2 is a schematic diagram of the effect of the method for identifying titles in a document provided by the present invention: Through the method for identifying titles in a document provided by the present invention, not only can titles occupying a single text line be identified, such as Title 1 formed by Text Line 1, but also titles occupying multiple text lines can be identified, such as Title 2 formed by Text Line 3 and Text Line 4. Among them, Text Line 2 is not a title, which improves the accuracy rate of code segment identification.
[0338] This embodiment extracts features from three dimensions: semantic, spatial, and visual, which can improve the accuracy of title recognition. It also provides a method for matching and fusing multiple lines of features to support the recognition of multiple lines of titles. In addition, it also provides a method for determining whether it is a title based on pattern recognition, as well as a machine learning discrimination method.
[0339] The following describes the device for recognizing titles in a document provided by the present invention. The device for recognizing titles in a document described below and the method for recognizing titles in a document described above can be referenced to each other.
[0340] Figure 3 Schematic diagram of the structure of the device for identifying titles in documents provided by the present invention. Figure 3 As shown, the apparatus 300 for identifying titles in a document includes: an identification module 301, an extraction module 302, a screening module 303, a processing module 304, and a determination module 305, wherein:
[0341] The recognition module 301 is configured to recognize at least one initial paragraph in a target document and at least one text line in each initial paragraph;
[0342] An extraction module 302 is configured to perform feature extraction on each text line with respect to at least one of a semantic dimension, a spatial dimension, and a visual dimension, to obtain a feature extraction result for each text line;
[0343] A screening module 303 is configured to screen out a target paragraph from the at least one initial paragraph based on the feature extraction results;
[0344] The processing module 304 is configured to process the feature extraction results of each text line in the target paragraph to obtain a feature processing result of the target paragraph;
[0345] The determination module 305 is configured to determine whether the target paragraph in the target document is a title according to the feature processing result.
[0346] The device for identifying titles in documents provided by the present invention identifies at least one initial paragraph in a target document and at least one text line in each of the initial paragraphs, wherein the number of text lines in the initial paragraphs is less than a set threshold; performs feature extraction on each of the text lines with respect to at least one dimension among a semantic dimension, a spatial dimension, and a visual dimension to obtain a feature extraction result for each of the text lines; screens out a target paragraph from the at least one initial paragraph based on each of the feature extraction results; processes the feature extraction results of each of the text lines in the target paragraph to obtain a feature processing result for the target paragraph; determines whether the target paragraph in the target document is a title based on the feature processing result; and can quickly and accurately identify titles in documents, has good performance, extremely high accuracy, and supports multi-line title recognition.
[0347] According to one or more optional embodiments of the present invention, the extraction module 302 is further configured to:
[0348] Perform at least one of the following feature extraction processes on each of the text lines:
[0349] Extracting features from each text line from a semantic dimension to obtain semantic feature extraction results for each text line;
[0350] Extracting features of each text line from a spatial dimension to obtain spatial feature extraction results of each text line;
[0351] Extract features from each of the text lines from a visual dimension to obtain visual feature extraction results for each of the text lines;
[0352] For each of the text lines, at least one of the semantic feature extraction result, the spatial feature extraction result, and the visual feature extraction result of the text line is processed to obtain a feature extraction result of the text line.
[0353] According to one or more optional embodiments of the present invention, the semantic feature extraction result includes a sequence number feature extraction result, a special title feature extraction result, and a common title feature extraction result;
[0354] The extraction module 302 is further configured to:
[0355] For each of the text lines, perform the following steps:
[0356] Performing serial number recognition on the beginning of the text line; if the serial number is recognized, determining that the serial number feature extraction result of the text line has the serial number feature; if the serial number is not recognized, determining that the serial number feature extraction result of the text line does not have the serial number feature;
[0357] Matching the text line with a preset set of special title symbols; if the match is successful, determining that the special title feature extraction result of the text line is that the special title symbol exists; if the match fails, determining that the special title feature extraction result of the text line is that the special title symbol does not exist;
[0358] The text line is matched with a preset set of common title symbols; if the match is successful, the common title feature extraction result of the text line is determined to be the presence of common title symbols; if the match fails, the common title feature extraction result of the text line is determined to be the absence of the common title symbols.
[0359] According to one or more optional embodiments of the present invention, the spatial feature extraction result includes a content feature extraction result, a first font size feature extraction result, and a second font size feature extraction result;
[0360] The extraction module 302 is further configured to:
[0361] For each of the text lines, perform the following steps:
[0362] If the content of the text line is all text content, determining that the content feature extraction result of the text line has text content features; if the content of the text line contains non-text content, determining that the content feature extraction result of the text line does not have the text content features;
[0363] Determining whether a first average font size corresponding to the text line is greater than a second average font size, where the second average font size is the average font size of all text in the page to which the text line belongs; if so, determining that the first font size feature extraction result of the text line has the font size feature; if not, determining that the first font size feature extraction result of the text line does not have the font size feature;
[0364] In the absence of a first designated text line, determining that the second font size feature extraction result of the text line is a first value, and the first designated text line is the next text line of the text line; in the presence of the first designated text line, determining whether the first average font size corresponding to the text line is greater than the third average font size of the first designated text line; if so, determining that the second font size feature extraction result of the text line is a second value; if not, determining that the second font size feature extraction result of the text line is a third value; the first value, the second value and the third value are all different.
[0365] According to one or more optional embodiments of the present invention, the visual feature extraction result includes a font feature extraction result, a modification feature extraction result, and an interval feature extraction result;
[0366] The extraction module 302 is further configured to:
[0367] Performing font recognition on each text content in the text line to obtain a font feature extraction result of the text line;
[0368] Identifying modified content on the text line, the modified content being used to modify the text content in the text line; if the modified content is identified, determining that a modified feature extraction result of the text line has a modified feature; if the modified content is not identified, determining that a modified feature extraction result of the text line does not have the modified feature;
[0369] Perform associated interval matching on the text line to obtain an interval feature extraction result of the text line.
[0370] According to one or more optional embodiments of the present invention, the font feature extraction result includes a first bold feature extraction result, a second bold feature extraction result, and a font same feature extraction result;
[0371] The extraction module 302 is further configured to:
[0372] Determine whether each text content in the text line is in a bold font; if so, determine that the first bold feature extraction result of the text line has all the bold features; if not, determine that the first bold feature extraction result of the text line does not have all the bold features;
[0373] Determine whether there is text content in a bold font in the text line; if so, determine that the second bold feature extraction result of the text line has a partial font bold feature; if not, determine that the second bold feature extraction result of the text line does not have the partial font bold feature;
[0374] Determine whether the fonts of each text content in the text line are the same; if so, determine that the font same feature extraction result of the text line has the font same feature; if not, determine that the font same feature extraction result of the text line does not have the font same feature.
[0375] According to one or more optional embodiments of the present invention, the interval feature extraction result includes a font size interval feature extraction result and a character number interval feature extraction result;
[0376] The extraction module 302 is further configured to:
[0377] Determining a first target font size interval from at least one preset initial font size interval based on a first average font size corresponding to the text line; determining a font size interval feature extraction result of the text line as a first serial number, the first serial number representing a ranking of the first target font size interval in each of the initial font size intervals;
[0378] Based on the first number of characters contained in the text line, a first target character number interval is determined from at least one preset initial character number interval; the character number interval feature extraction result of the text line is determined as a second serial number, and the second serial number represents the ranking of the first target character number interval in each of the initial character number intervals.
[0379] According to one or more optional embodiments of the present invention, the screening module 303 is further configured to:
[0380] For each of these initial paragraphs, perform the following steps:
[0381] When the number of text lines in the initial paragraph is a first preset value, determining the initial paragraph as a target paragraph;
[0382] When the number of text lines in the initial paragraph is not the first preset value, determine whether the feature extraction results of each text line in the initial paragraph meet the set merging conditions; if so, determine the initial paragraph as the target paragraph.
[0383] According to one or more optional embodiments of the present invention, the screening module 303 is further configured to:
[0384] Obtaining the alignment and first average font size of each text line in the initial paragraph;
[0385] determining a first absolute value of a difference between first average font sizes corresponding to adjacent text lines in the initial paragraph based on each of the first average font sizes;
[0386] It is determined whether the feature extraction results and the alignment methods of the text lines in the initial paragraph, as well as the first absolute values, all meet the set merging conditions.
[0387] According to one or more optional embodiments of the present invention, the feature extraction result includes a content feature extraction result, a first bold feature extraction result, a second bold feature extraction result, a font same feature extraction result, a modification feature extraction result, a first font size feature extraction result, and a sequence number feature extraction result;
[0388] The feature extraction results of each of the text lines in the initial paragraph meet the requirements, including:
[0389] The content feature extraction results of each of the text lines in the initial paragraph are the same, the first bold feature extraction results are the same, the second bold feature extraction results are the same, the font same feature extraction results are the same, the modification feature extraction results are the same, and the first font size feature extraction results are the same; and
[0390] The result of the sequence number feature extraction of the first text line in the initial paragraph is that it has the sequence number feature, and the results of the sequence number feature extraction of each specific text line in the initial paragraph do not have the sequence number feature, and the specific text lines are the text lines in the initial paragraph other than the first text line;
[0391] Each of the first absolute values corresponding to the initial paragraphs meets a set merging condition, including:
[0392] Each of the first absolute values corresponding to the initial paragraph is smaller than the set font size difference;
[0393] The alignment of each text line in the initial paragraph meets the set merging conditions, including:
[0394] The alignment of each text line in the initial paragraph is the same.
[0395] According to one or more optional embodiments of the present invention, the screening module 303 is further configured to:
[0396] For each of the text lines in the initial paragraph, perform the following steps:
[0397] Obtaining the bounding rectangle and line height of the text line, and obtaining the visual area of the page to which the text line belongs;
[0398] Calculating a left distance between a left border of the circumscribed rectangular frame and a left border of the visualization area, and calculating a right distance between a right border of the circumscribed rectangular frame and a right border of the visualization area;
[0399] If the left spacing and the right spacing meet the center alignment condition, then the alignment of the text line is determined to be center alignment, and the center alignment condition is that the second absolute value of the difference between the left spacing and the right spacing is greater than the line height, and both the left spacing and the right spacing are greater than the product of the line height and a set value;
[0400] If the left spacing and the right spacing meet a right alignment condition, determining that the alignment of the text line is right alignment, the right alignment condition is that the left spacing is smaller than the right spacing, and the right spacing is smaller than the product;
[0401] If the left spacing and the right spacing do not satisfy the center alignment condition and do not satisfy the right alignment condition, then the alignment of the text line is determined to be left alignment.
[0402] According to one or more optional embodiments of the present invention, the processing module 304 is further configured to:
[0403] When the number of text lines in the target paragraph is a second preset value, using the feature extraction results of the text lines in the target paragraph as the feature processing results of the target paragraph;
[0404] When the number of text lines in the target paragraph is not the second preset value, the feature extraction results of the text lines in the target paragraph are merged according to the set merging strategy to obtain the feature processing result of the target paragraph.
[0405] According to one or more optional embodiments of the present invention, the feature extraction results include content feature extraction results, first bold feature extraction results, second bold feature extraction results, font same feature extraction results, sequence number feature extraction results, special title feature extraction results, common title feature extraction results, modification feature extraction results, first font size feature extraction results, second font size feature extraction results, font size interval feature extraction results, and character number interval feature extraction results;
[0406] The processing module 304 is further configured to:
[0407] The content feature extraction result, the first bold feature extraction result, the second bold feature extraction result, the same font feature extraction result, the modified feature extraction result, and the first font size feature extraction result of the first text line in the target paragraph are respectively used as the content feature extraction result, the first bold feature extraction result, the second bold feature extraction result, the same font feature extraction result, the modified feature extraction result, and the first font size feature extraction result of the target paragraph;
[0408] Performing an OR operation on the sequence number feature extraction results of each text line in the target paragraph to obtain the sequence number feature extraction result of the target paragraph; performing an OR operation on the special title feature extraction results of each text line in the target paragraph to obtain the special title feature extraction result of the target paragraph; performing an OR operation on the common title feature extraction results of each text line in the target paragraph to obtain the common title feature extraction result of the target paragraph;
[0409] In the absence of a second designated text line, determining that the second font size feature extraction result of the target paragraph is a first value, and the second designated text line is the next text line of the target paragraph; in the presence of the second designated text line, performing an average calculation on the first average font sizes corresponding to each text line in the target paragraph to obtain a target average font size, and determining whether the target average font size is greater than a third average font size of the second designated text line; if so, determining that the second font size feature extraction result of the target paragraph is a second value; if not, determining that the second font size feature extraction result of the target paragraph is a third value; the first value, the second value, and the third value are all different;
[0410] determining a second target font size interval from at least one preset initial font size interval based on the fourth average font size of the target paragraph; determining a font size interval feature extraction result of the target paragraph as a third serial number, the third serial number representing a ranking of the second target font size interval in each of the initial font size intervals;
[0411] Determining a second target character number interval from at least one preset initial character number interval based on a second number of characters included in the target paragraph; determining a character number interval feature extraction result of the text line as a fourth serial number, the fourth serial number representing a ranking of the second target character number interval in each of the initial character number intervals;
[0412] The feature processing result of the target paragraph includes the content feature extraction result of the target paragraph, the first bold feature extraction result, the second bold feature extraction result, the font same feature extraction result, the serial number feature extraction result, the special title feature extraction result, the common title feature extraction result, the modification feature extraction result, the first font size feature extraction result, the second font size feature extraction result, the font size interval feature extraction result and the character number interval feature extraction result.
[0413] According to one or more optional embodiments of the present invention, the determining module 305 is further configured to:
[0414] Determine whether the target paragraph in the target document is a title based on the feature processing result and the set title condition; or
[0415] The feature processing result is input into a trained title recognition model to determine whether the target paragraph in the target document is a title.
[0416] According to one or more optional embodiments of the present invention, the feature processing results include content feature extraction results, first bold feature extraction results, second bold feature extraction results, font same feature extraction results, sequence number feature extraction results, special title feature extraction results, common title feature extraction results, modification feature extraction results, first font size feature extraction results, second font size feature extraction results, font size interval feature extraction results, and character number interval feature extraction results;
[0417] The title setting conditions include:
[0418] The content feature extraction result, the first bold feature extraction result, the second bold feature extraction result, the same font feature extraction result, the serial number feature extraction result, the special title feature extraction result, the common title feature extraction result, the modification feature extraction result, and the first font size feature extraction result are all characterized by having corresponding features; and
[0419] The second font size feature extraction result is a second value; and
[0420] The font size interval feature extraction result is greater than a first threshold; and
[0421] The character number interval feature extraction result is less than a second threshold.
[0422] Figure 4 An example of a physical structure diagram of an electronic device is shown below. Figure 4 As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communication interface 420, and the memory 430 communicate with each other via the communication bus 440. The processor 410 may call the logic instructions in the memory 430 to execute a method for identifying titles in a document, the method comprising: identifying at least one initial paragraph in a target document and at least one text line in each of the initial paragraphs; performing feature extraction on each of the text lines for at least one dimension among semantic dimension, spatial dimension, and visual dimension to obtain feature extraction results for each of the text lines; selecting a target paragraph from each of the initial paragraphs based on the at least one feature extraction result; processing the feature extraction results of each of the text lines in the target paragraph to obtain a feature processing result for the target paragraph; and determining whether the target paragraph in the target document is a title based on the feature processing result.
[0423] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0424] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the method for identifying titles in documents provided by the above methods, which includes: identifying at least one initial paragraph in the target document and at least one text line in each of the initial paragraphs; performing feature extraction on each of the text lines for at least one dimension among semantic dimension, spatial dimension and visual dimension to obtain feature extraction results for each of the text lines; based on each of the feature extraction results, screening out a target paragraph from the at least one initial paragraph; processing the feature extraction results of each of the text lines in the target paragraph to obtain a feature processing result of the target paragraph; and determining whether the target paragraph in the target document is a title based on the feature processing result.
[0425] On the other hand, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the method for identifying titles in documents provided by the above-mentioned methods, the method comprising: identifying at least one initial paragraph in a target document and at least one text line in each of the initial paragraphs; performing feature extraction on each of the text lines for at least one dimension among semantic dimension, spatial dimension and visual dimension to obtain feature extraction results for each of the text lines; screening out a target paragraph from the at least one initial paragraph based on each of the feature extraction results; processing the feature extraction results of each of the text lines in the target paragraph to obtain a feature processing result of the target paragraph; and determining whether the target paragraph in the target document is a title based on the feature processing result.
[0426] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0427] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0428] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for identifying titles in a document, characterized in that: include: Identifying at least one initial paragraph in a target document and at least one text line in each initial paragraph; performing feature extraction on each of the text lines with respect to at least one dimension among a semantic dimension, a spatial dimension, and a visual dimension to obtain a feature extraction result for each of the text lines; Filtering a target paragraph from the at least one initial paragraph according to each of the feature extraction results; Processing the feature extraction results of each text line in the target paragraph to obtain a feature processing result of the target paragraph; According to the feature processing result, it is determined whether the target paragraph in the target document is a title.
2. The method for identifying titles in a document according to claim 1, wherein: The step of performing feature extraction on each of the text lines with respect to at least one of the semantic dimension, the spatial dimension, and the visual dimension to obtain a feature extraction result for each of the text lines includes: Perform at least one of the following feature extraction processes on each of the text lines: Extracting features from each text line from a semantic dimension to obtain semantic feature extraction results for each text line; Extracting features of each text line from a spatial dimension to obtain spatial feature extraction results of each text line; Extract features from each of the text lines from a visual dimension to obtain visual feature extraction results for each of the text lines; For each of the text lines, at least one of the semantic feature extraction result, the spatial feature extraction result, and the visual feature extraction result of the text line is processed to obtain a feature extraction result of the text line.
3. The method for identifying titles in a document according to claim 2, wherein: The semantic feature extraction results include sequence number feature extraction results, special title feature extraction results and common title feature extraction results; The extracting features of each text line from a semantic dimension to obtain a semantic feature extraction result of each text line includes: For each of the text lines, perform the following steps: Performing serial number recognition on the beginning of the text line; if the serial number is recognized, determining that the serial number feature extraction result of the text line has the serial number feature; if the serial number is not recognized, determining that the serial number feature extraction result of the text line does not have the serial number feature; Matching the text line with a preset set of special title symbols; if the match is successful, determining that the special title feature extraction result of the text line is that the special title symbol exists; if the match fails, determining that the special title feature extraction result of the text line is that the special title symbol does not exist; The text line is matched with a preset set of common title symbols; if the match is successful, the common title feature extraction result of the text line is determined to be the presence of common title symbols; if the match fails, the common title feature extraction result of the text line is determined to be the absence of the common title symbols.
4. The method for identifying titles in a document according to claim 2, wherein: The spatial feature extraction result includes a content feature extraction result, a first font size feature extraction result, and a second font size feature extraction result; The extracting features of each text line from a spatial dimension to obtain a spatial feature extraction result of each text line includes: For each of the text lines, perform the following steps: If the content of the text line is all text content, determining that the content feature extraction result of the text line has text content features; if the content of the text line contains non-text content, determining that the content feature extraction result of the text line does not have the text content features; Determining whether a first average font size corresponding to the text line is greater than a second average font size, where the second average font size is the average font size of all text in the page to which the text line belongs; if so, determining that the first font size feature extraction result of the text line has the font size feature; if not, determining that the first font size feature extraction result of the text line does not have the font size feature; In the absence of a first designated text line, determining that the second font size feature extraction result of the text line is a first value, and the first designated text line is the next text line of the text line; in the presence of the first designated text line, determining whether the first average font size corresponding to the text line is greater than the third average font size of the first designated text line; if so, determining that the second font size feature extraction result of the text line is a second value; if not, determining that the second font size feature extraction result of the text line is a third value; the first value, the second value and the third value are all different.
5. The method for identifying titles in a document according to claim 2, wherein: The visual feature extraction results include font feature extraction results, modification feature extraction results and interval feature extraction results; The extracting features of each text line from a visual dimension to obtain a visual feature extraction result of each text line includes: Performing font recognition on each text content in the text line to obtain a font feature extraction result of the text line; Identifying modified content on the text line, the modified content being used to modify the text content in the text line; if the modified content is identified, determining that a modified feature extraction result of the text line has a modified feature; if the modified content is not identified, determining that a modified feature extraction result of the text line does not have the modified feature; Perform associated interval matching on the text line to obtain an interval feature extraction result of the text line.
6. The method for identifying titles in a document according to claim 5, characterized in that: The font feature extraction results include a first bold feature extraction result, a second bold feature extraction result, and a font-same feature extraction result; The performing font recognition on each text content in the text line to obtain a font feature extraction result of the text line includes: Determine whether each text content in the text line is in a bold font; if so, determine that the first bold feature extraction result of the text line has all the bold features; if not, determine that the first bold feature extraction result of the text line does not have all the bold features; Determine whether there is text content in a bold font in the text line; if so, determine that the second bold feature extraction result of the text line has a partial font bold feature; if not, determine that the second bold feature extraction result of the text line does not have the partial font bold feature; Determine whether the fonts of each text content in the text line are the same; if so, determine that the font same feature extraction result of the text line has the font same feature; if not, determine that the font same feature extraction result of the text line does not have the font same feature.
7. The method for identifying titles in a document according to claim 5, wherein: The interval feature extraction results include font size interval feature extraction results and character number interval feature extraction results; The performing associated interval matching on the text line to obtain the interval feature extraction result of the text line includes: Determining a first target font size interval from at least one preset initial font size interval based on a first average font size corresponding to the text line; determining a font size interval feature extraction result of the text line as a first serial number, the first serial number representing a ranking of the first target font size interval in each of the initial font size intervals; Based on the first number of characters contained in the text line, a first target character number interval is determined from at least one preset initial character number interval; the character number interval feature extraction result of the text line is determined as a second serial number, and the second serial number represents the ranking of the first target character number interval in each of the initial character number intervals.
8. The method for identifying titles in a document according to any one of claims 1 to 7, characterized in that: The step of selecting a target paragraph from the at least one initial paragraph based on the feature extraction results includes: For each of these initial paragraphs, perform the following steps: When the number of text lines in the initial paragraph is a first preset value, determining the initial paragraph as a target paragraph; When the number of text lines in the initial paragraph is not the first preset value, determine whether the feature extraction results of each text line in the initial paragraph meet the set merging conditions; if so, determine the initial paragraph as the target paragraph.
9. The method for identifying titles in a document according to claim 8, wherein: The determining whether the feature extraction results of each text line in the initial paragraph meet the set merging conditions includes: Obtaining the alignment and first average font size of each text line in the initial paragraph; determining a first absolute value of a difference between first average font sizes corresponding to adjacent text lines in the initial paragraph based on each of the first average font sizes; It is determined whether the feature extraction results and the alignment methods of the text lines in the initial paragraph, as well as the first absolute values, all meet the set merging conditions.
10. The method for identifying titles in a document according to claim 9, wherein: The feature extraction results include content feature extraction results, first bold feature extraction results, second bold feature extraction results, font same feature extraction results, modification feature extraction results, first font size feature extraction results and sequence number feature extraction results; The feature extraction results of each text line in the initial paragraph meet the set merging conditions, including: The content feature extraction results of each of the text lines in the initial paragraph are the same, the first bold feature extraction results are the same, the second bold feature extraction results are the same, the font same feature extraction results are the same, the modification feature extraction results are the same, and the first font size feature extraction results are the same; and The result of the sequence number feature extraction of the first text line in the initial paragraph is that it has the sequence number feature, and the results of the sequence number feature extraction of each specific text line in the initial paragraph do not have the sequence number feature, and the specific text lines are the text lines in the initial paragraph other than the first text line; Each of the first absolute values corresponding to the initial paragraphs meets a set merging condition, including: Each of the first absolute values corresponding to the initial paragraph is smaller than the set font size difference; The alignment of each text line in the initial paragraph meets the set merging conditions, including: The alignment of each text line in the initial paragraph is the same.
11. The method for identifying titles in a document according to claim 9, wherein: The obtaining of the alignment of each text line in the initial paragraph includes: For each of the text lines in the initial paragraph, perform the following steps: Obtaining the bounding rectangle and line height of the text line, and obtaining the visual area of the page to which the text line belongs; Calculating a left distance between a left border of the circumscribed rectangular frame and a left border of the visualization area, and calculating a right distance between a right border of the circumscribed rectangular frame and a right border of the visualization area; If the left spacing and the right spacing meet the center alignment condition, then the alignment of the text line is determined to be center alignment, and the center alignment condition is that the second absolute value of the difference between the left spacing and the right spacing is greater than the line height, and both the left spacing and the right spacing are greater than the product of the line height and a set value; If the left spacing and the right spacing meet a right alignment condition, determining that the alignment of the text line is right alignment, the right alignment condition is that the left spacing is smaller than the right spacing, and the right spacing is smaller than the product; If the left spacing and the right spacing do not satisfy the center alignment condition and do not satisfy the right alignment condition, then the alignment of the text line is determined to be left alignment.
12. The method for identifying titles in a document according to any one of claims 1 to 7, characterized in that: The step of processing the feature extraction results of each text line in the target paragraph to obtain the feature processing results of the target paragraph includes: When the number of text lines in the target paragraph is a second preset value, using the feature extraction results of the text lines in the target paragraph as the feature processing results of the target paragraph; When the number of text lines in the target paragraph is not the second preset value, the feature extraction results of the text lines in the target paragraph are merged according to the set merging strategy to obtain the feature processing result of the target paragraph.
13. The method for identifying titles in a document according to claim 12, wherein: The feature extraction results include content feature extraction results, first bold feature extraction results, second bold feature extraction results, font same feature extraction results, sequence number feature extraction results, special title feature extraction results, common title feature extraction results, modification feature extraction results, first font size feature extraction results, second font size feature extraction results, font size interval feature extraction results, and character number interval feature extraction results; Merging the feature extraction results of each text line in the target paragraph according to a set merging strategy to obtain the feature processing result of the target paragraph includes: The content feature extraction result, the first bold feature extraction result, the second bold feature extraction result, the same font feature extraction result, the modified feature extraction result, and the first font size feature extraction result of the first text line in the target paragraph are respectively used as the content feature extraction result, the first bold feature extraction result, the second bold feature extraction result, the same font feature extraction result, the modified feature extraction result, and the first font size feature extraction result of the target paragraph; Performing an OR operation on the sequence number feature extraction results of each text line in the target paragraph to obtain the sequence number feature extraction result of the target paragraph; performing an OR operation on the special title feature extraction results of each text line in the target paragraph to obtain the special title feature extraction result of the target paragraph; performing an OR operation on the common title feature extraction results of each text line in the target paragraph to obtain the common title feature extraction result of the target paragraph; In the absence of a second designated text line, determining that the second font size feature extraction result of the target paragraph is a first value, and the second designated text line is the next text line of the target paragraph; in the presence of the second designated text line, performing an average calculation on the first average font sizes corresponding to each text line in the target paragraph to obtain a target average font size, and determining whether the target average font size is greater than a third average font size of the second designated text line; if so, determining that the second font size feature extraction result of the target paragraph is a second value; if not, determining that the second font size feature extraction result of the target paragraph is a third value; the first value, the second value, and the third value are all different; determining a second target font size interval from at least one preset initial font size interval based on the fourth average font size of the target paragraph; determining a font size interval feature extraction result of the target paragraph as a third serial number, the third serial number representing a ranking of the second target font size interval in each of the initial font size intervals; Determining a second target character number interval from at least one preset initial character number interval based on a second number of characters included in the target paragraph; determining a character number interval feature extraction result of the text line as a fourth serial number, the fourth serial number representing a ranking of the second target character number interval in each of the initial character number intervals; The feature processing result of the target paragraph includes the content feature extraction result of the target paragraph, the first bold feature extraction result, the second bold feature extraction result, the font same feature extraction result, the serial number feature extraction result, the special title feature extraction result, the common title feature extraction result, the modification feature extraction result, the first font size feature extraction result, the second font size feature extraction result, the font size interval feature extraction result and the character number interval feature extraction result.
14. The method for identifying titles in a document according to any one of claims 1 to 7, characterized in that: The determining, based on the feature processing result, whether the target paragraph in the target document is a title includes: Determine whether the target paragraph in the target document is a title based on the feature processing result and the set title condition; or The feature processing result is input into a trained title recognition model to determine whether the target paragraph in the target document is a title.
15. The method for identifying titles in a document according to claim 14, wherein: The feature processing results include content feature extraction results, first bold feature extraction results, second bold feature extraction results, font same feature extraction results, sequence number feature extraction results, special title feature extraction results, common title feature extraction results, modification feature extraction results, first font size feature extraction results, second font size feature extraction results, font size interval feature extraction results, and character number interval feature extraction results; The title setting conditions include: The content feature extraction result, the first bold feature extraction result, the second bold feature extraction result, the same font feature extraction result, the serial number feature extraction result, the special title feature extraction result, the common title feature extraction result, the modification feature extraction result, and the first font size feature extraction result are all characterized by having corresponding features; and The second font size feature extraction result is a second value; and The font size interval feature extraction result is greater than a first threshold; and The character number interval feature extraction result is less than a second threshold.
16. A device for identifying titles in a document, characterized in that: include: a recognition module configured to recognize at least one initial paragraph in a target document and at least one text line in each initial paragraph; an extraction module configured to perform feature extraction on each of the text lines with respect to at least one of a semantic dimension, a spatial dimension, and a visual dimension, to obtain a feature extraction result for each of the text lines; a screening module configured to screen out a target paragraph from the at least one initial paragraph based on the feature extraction results; a processing module configured to process the feature extraction results of each text line in the target paragraph to obtain a feature processing result of the target paragraph; The determination module is configured to determine whether the target paragraph in the target document is a title according to the feature processing result.