Document element extraction method, related device, equipment and storage medium

By identifying and classifying text lines in official documents, and modifying and extracting them in combination with the expression specifications of element categories, the problem of inaccurate extraction of complex layout documents is solved, and higher extraction accuracy and reliability are achieved.

CN120164227APending Publication Date: 2025-06-17HEFEI IFLY DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510141838.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing official document element extraction technology has the problem of inaccurate extraction when dealing with complex layout official documents, resulting in insufficient accuracy.

Method used

By identifying text lines in official documents, predicting the first category of text lines based on multimodal features, and modifying the first category through the expression specification of element categories to obtain the second category. Then, the starting and ending characters are determined based on the expression specifications of the second category, the corresponding element content is extracted, and structured data is finally extracted from the official document.

Benefits of technology

It improves the accuracy of official document factor extraction, enhances the category prediction ability in complex layout situations, and improves the reliability of factor categories through rule correction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164227A_ABST
    Figure CN120164227A_ABST
Patent Text Reader

Abstract

The invention discloses an official document element extraction method, a related device, equipment and a storage medium, and the official document element extraction method comprises the steps: recognizing each text line in a target official document; predicting a first category of the text line based on the multi-modal features of the text line; correcting the first category of the text lines based on the expression specifications of the plurality of element categories to obtain a second category of the text lines; wherein the first category and the second category are both one of a plurality of element categories; determining a starting character belonging to a second category in the text line based on an expression specification of the second category to which the text line belongs, and traversing from the starting character to an ending character of the second category in the target official document as element content of the second category; and extracting structured data from the target official document based on the element content of the second category. According to the scheme, the document element extraction accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of natural language processing, and particularly to a method for extracting official document elements and related devices, equipment, and storage media. Background Art

[0002] The official document element extraction technology aims to automatically extract key information from official documents, such as content like title, issuing agency, issuing date, attachments, text, etc.

[0003] Currently, existing official document element extraction technologies usually rely on OCR (Optical Character Recognition) technology and NLP (Natural Language Processing) technology. Although these technologies can improve the automation level to a certain extent, there are still obvious limitations when dealing with official documents with complex layouts, resulting in problems such as inaccurate extraction. In view of this, how to improve the accuracy of official document element extraction has become an urgent problem to be solved. Summary of the Invention

[0004] The main technical problem to be solved by this application is to provide a method for extracting official document elements and related devices, equipment, and storage media, which can improve the accuracy of official document element extraction.

[0005] To solve the above technical problem, in the first aspect of this application, a method for extracting official document elements is provided, including: identifying each text line in the target official document; predicting the first category of the text line based on the multi-modal features of the text line; correcting the first category of the text line based on the expression specifications of several element categories to obtain the second category of the text line; where the first category and the second category are both one of several element categories; determining the starting character belonging to the second category in the text line based on the expression specification of the second category to which the text line belongs, and traversing in the target official document from the starting character until the ending character of the second category as the element content of the second category; extracting structured data from the target official document based on the element content of the second category.

[0006] To solve the above technical problems, a second aspect of the present application provides a document element extraction device, including: an identification module, a classification module, a correction module, an extraction module, and an integration module. The identification module is used to identify each text line in the target document; the classification module is used to predict the first category of the text line based on the multi-modal features of the text line; the correction module is used to correct the first category of the text line based on the expression specifications of several element categories to obtain the second category of the text line; wherein the first category and the second category are both one of several element categories; the extraction module is used to determine the starting character belonging to the second category in the text line based on the expression specification of the second category to which the text line belongs, and traverse from the starting character in the target document until the ending character of the second category as the element content of the second category; the integration module is used to extract structured data from the target document based on the element content of the second category.

[0007] To solve the above technical problems, a third aspect of the present application provides an electronic device, at least including a memory and a processor coupled to each other. The memory stores at least program instructions, and the processor is used to execute the program instructions to implement the document element extraction method in the first aspect above.

[0008] To solve the above technical problems, a fourth aspect of the present application provides a computer-readable storage medium storing program instructions that can be run by a processor. The program instructions are used to implement the document element extraction method in the first aspect above.

[0009] In the above solution, each text line in the target document is identified, and the first category of the text line is predicted based on the multi-modal features of the text line. Then, the first category of the text line is corrected based on the expression specifications of several element categories to obtain the second category of the text line. Both the first category and the second category are one of several element categories. Furthermore, based on the expression specification of the second category to which the text line belongs, the starting character belonging to the second category in the text line is determined, and traversed from the starting character in the target document until the ending character of the second category as the element content of the second category, so as to extract structured data from the target document based on the element content of the second category. Therefore, on the one hand, taking the text line as the basic unit of category prediction can greatly improve the recall rate and reduce missed detections, and by integrating the multi-modal features of the text line for category prediction, it helps to improve the accuracy of category prediction for text lines even when facing documents with complex layouts. On the other hand, after predicting the category through the multi-modal features of the text line, further combining the expression specifications of the element categories can accurately correct the predicted category by combining rules, so that the element categories with clear rules are more reliably identified. Therefore, the accuracy of document element extraction can be improved. Description of the Drawings

[0010] Figure 1It is a schematic flowchart of an embodiment of the method for extracting document elements in this application;

[0011] Figure 2 It is a schematic diagram of an embodiment of extracting the element content of the second category in this application;

[0012] Figure 3 It is a schematic framework diagram of an embodiment of the document element extraction device in this application;

[0013] Figure 4 It is a schematic framework diagram of an embodiment of the electronic device in this application;

[0014] Figure 5 It is a schematic framework diagram of an embodiment of the computer-readable storage medium in this application. Detailed implementation manners

[0015] The following will combine the accompanying drawings of the specification to elaborate in detail on the solutions of the embodiments of this application.

[0016] In the following description, specific details such as specific system architectures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand this application.

[0017] The terms "system" and "network" are often used interchangeably in this article. The term "and / or" in this article merely describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the fragment " / " in this article generally represents an "or" relationship between the preceding and following associated objects. In addition, "multiple" in this article means two or more than two.

[0018] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of an embodiment of the method for extracting document elements in this application. Specifically, it may include the following steps:

[0019] Step S11: Identify each text line in the target document.

[0020] In one implementation scenario, the target document may include, but is not limited to: letters, notices, decisions, announcements, circulars, motions, reports, requests for instructions, replies, etc. The specific type of the target document is not limited here, and no further examples will be given here.

[0021] In an implementation scenario, the target official document may include, but is not limited to, PDF documents, images, word documents, etc. The document format of the target official document is not limited herein. In addition, as a possible example, a detection model of deep learning such as YOLO can be used to identify each text line in the target official document. For the specific process, reference can be made to the technical details of detection models such as YOLO, which will not be elaborated herein. As another possible example, OCR can also be used to identify the target official document to obtain each text line in the target official document, such as the specific content of each text line and the position information (e.g., coordinate information, etc.) in the target official document. For the specific process, reference can be made to the technical details of OCR, which will not be elaborated herein.

[0022] Step S12: Based on the multi-modal features of the text line, predict the first category of the text line.

[0023] It should be noted that the multi-modal features of the text line may include at least two of content features, layout features, position features, visual features, and page number features. For the sake of convenience of description, the target official document can be denoted as D, which contains a total of K pages, and the i-th page can be denoted as p i ∈{1,…,K}, and the page images contained therein can be denoted as I = {I1, I2, …, I K}, and the i-th text line can be denoted as u i , and its recognized text can be denoted as s i , and its detection area (such as a bounding box) in the page image can be denoted as b i =(x min , y min , x max , y max ), where x min , y min respectively represent the minimum coordinate values of the detection area on the X-axis and Y-axis, and x max , y max respectively represent the maximum coordinate values of the detection area on the X-axis and Y-axis. In addition, the content feature can be denoted as x text , the layout feature can be denoted as x layout , the position feature can be denoted as x pos , the visual feature can be denoted as x vis , and the page number feature can be denoted as x page . As a possible example, in the case where the multi-modal features include content features, the content feature can be obtained by encoding the recognized text of the text line by a sentence-level encoder and then performing a linear mapping. Taking the i-th text line as an example, its content feature can be expressed as:

[0024]

[0025] In the above formula (1), SentenceBERT represents a sentence-level encoder, and LinearProj represents a linear mapping. As another possible example, in the case where the multimodal features include layout features, the layout features can be obtained by fusing the first layout information and the second layout information of the text lines after they are respectively embedded and characterized. The first layout information may include the extreme coordinates of the text line on the horizontal axis and the width of the text line, and the second layout information may include the extreme coordinates of the text line on the vertical axis and the height of the text line. Still taking the i-th text line as an example, its layout features can be expressed as:

[0026]

[0027] In the above formula (2), Emb represents the embedded features, and Concat represents the concatenation operation. As yet another possible example, in the case where the multimodal features include position features, the position features can be obtained by embedding and characterizing the line number encoding of the text line in the target document. Still taking the i-th text line as an example, its position features can be expressed as:

[0028]

[0029] In the above formula (3), i is the line number encoding, and Emb1D represents the embedded representation, specifically a one-dimensional embedded representation. As yet another possible example, in the case where the multimodal features include visual features, the visual features can be extracted from the feature map of the target image in the detection region of the text line in the target image, and the target image can be the page image where the text line is located in the target document. Still taking the i-th text line as an example, its visual features can be expressed as:

[0030]

[0031] In the above formula (4), I pi represents the page image where the i-th text line is located in the target document, ResNet(I pi ) represents the feature map extracted from the page image I pi using ResNet, b i represents the detection region of the i-th text line in the page image I oi , and RoIAlign represents the local representation of the detection region extracted from the feature map, that is, the visual features of the i-th text line. As yet another possible example, in the case where the multimodal features include page number features, the page number features can be obtained by embedding and characterizing the page number encoding of the text line in the target document. Still taking the i-th text line as an example, its page number features can be expressed as:

[0032]

[0033] In the above formula (5), p i represents the page number encoding of the i-th text line in the target official document, and EmbPage represents the embedding representation of the page number encoding.

[0034] In one implementation scenario, as a possible implementation method, after obtaining the multi-modal features of the text line, the multi-modal features of the text line can be classified and predicted through a classifier to obtain the first category of the text line. It should be noted that the classifier may include, but is not limited to, a fully connected layer, softmax, etc., and the specific structure of the classifier is not limited herein. In addition, when the classifier processes the multi-modal features, it can first obtain the probability values of the text line belonging to several element categories respectively, and then select the element category with the maximum probability value as the first category of the text line. The several element categories may at least include, but are not limited to, the issuing agency, title (such as first-level title, second-level title, third-level title, etc.), text body, footer, etc., and the several element categories are not limited herein.

[0035] In another implementation scenario, different from the foregoing implementation method, as another possible implementation method, after obtaining the multi-modal features of the text line, layer normalization can also be performed based on the multi-modal features to obtain the text line features. Still taking the multi-modal features including content features, layout features, position features, visual features, and page number features as an example, the text line feature x of the i-th text line i can be expressed as:

[0036]

[0037] In the above formula (6), represents the content feature of the i-th text line, represents the layout feature of the i-th text line, represents the position feature of the i-th text line, represents the visual feature of the i-th text line, represents the page number feature of the i-th text line. In addition, LN represents layer normalization. On this basis, encoding can be performed based on the text line features to obtain the line-level encoding features. Still taking the i-th text line as an example, the line-level encoding feature can be expressed as:

[0038]

[0039] In the above formula (7), BiEncoder represents a bidirectional encoder, which can be specifically based on the Transformer architecture, but is not limited thereto. The specific structure of the encoder for encoding text line features is not defined herein. After obtaining the line-level encoded features, prediction can continue based on the line-level encoded features to obtain the first category of the text line. Still taking the i-th text line as an example, the first category C of the text line i can be expressed as:

[0040]

[0041] In the above formulas (8) and (9), LinearProj represents a linear mapping, represents the probability values that the i-th text line belongs to each element category, and argmax represents taking the maximum value. For specific details, reference can be made to the relevant descriptions of the aforementioned classifier, which will not be elaborated herein. It should be noted that, as a possible example, the first category of the text line can be predicted by a text line classification model, and the text line classification model can include a feature extraction network for extracting multi-modal features and a classifier for classifying and predicting the multi-modal features. Then, before using the text line classification model to classify and predict the text line, the network parameters of the text line classification model can be gradient-optimized through a loss function such as FocalLoss. Exemplarily, when using FocalLoss, the training loss of the text line classification model can be expressed as:

[0042]

[0043] In the above formula (10), L represents the total number of text lines in the sample official document used to train the text line classification model. For specific technical details of loss functions such as FocalLoss, reference can be made thereto, which will not be elaborated herein.

[0044] Of course, the above examples are only several possible examples of classifying and predicting through multi-modal features to obtain the first category of the text line. Other possible ways of classifying and predicting are not defined herein, and no further examples will be given one by one.

[0045] Step S13: Correct the first category of the text line based on the expression specifications of several element categories to obtain the second category of the text line.

[0046] In the embodiments of the present disclosure, both the first category and the second category are each one of several element categories. In addition, the expression norms of element categories are used to characterize how the element categories should be written in official documents. Taking the element category of "second-level heading" as an example, its expression norm can be that "the first character is a capital Chinese numeral in parentheses" (e.g., "(I) XXX"); or, taking the element category of "third-level heading" as an example, its expression norm can be that "the first character is an Arabic numeral" (e.g., "1. XXX"). Of course, the above examples are only several possible examples of the expression norms in the actual application process, and other situations are not limited here, nor are the expression norms of other element categories exemplified one by one.

[0047] In an implementation scenario, as a possible implementation method, for any text line, the element content of the text line can be verified first using the expression norm of the first category. If the element content of the text line conforms to the expression norm of the first category, the first category of the text line can be maintained and used as the second category (that is, in this case, the second category is the same as the first category); conversely, if the element content of the text line does not conform to the expression norm of the first category, the element content of the text line can be continuously verified in turn using the expression norms of other element categories until the element content of the text line conforms to the expression norm of a certain element category, and then that element category can be used as the second category of the text line (that is, in this case, the second category is different from the first category). As a possible example, if the first category of the text line "(I) XXX" is "second-level heading", the expression norm of the first category of "second-level heading", that is, "the first character is a capital Chinese numeral in parentheses", can be used to verify the element content "(I) XXX" of the foregoing text line. Since the element content of the text line conforms to the expression norm of the first category, the first category of the text line, "second-level heading", can be maintained and used as the second category of the text line (that is, the second category of the text line is still "second-level heading"). As another possible example, if the first category of the text line "1. XXX" is "second-level heading", the expression norm of the first category of "second-level heading", that is, "the first character is a capital Chinese numeral in parentheses", can be used to verify the element content "1. XXX" of the foregoing text line. Since the element content of the text line does not conform to the expression norm of the first category of "second-level heading", the element content of the text line can be verified using the expression norms of other element categories respectively. When the expression norm of the element category of "third-level heading", that is, "the first character is an Arabic numeral", is used to verify the element content "1. XXX" of the text line, since the element content of the text line conforms to the expression norm of the element category of "third-level heading", the element category of "third-level heading" can be corrected as the second category of the text line (that is, the second category of the text line is "third-level heading"). Of course, the above examples are only several possible examples in the actual application process, and other possible situations are not exemplified one by one here.

[0048] In another implementation scenario, as another possible implementation method, different from the foregoing implementation method, for any text line, the expression specifications of each element category can be used in sequence to verify the element content of the text line. If the verification fails, continue to the next element category; if the verification passes, the element category can be used as the second category of the text line.

[0049] It should be noted that the above implementation methods are only several possible implementation examples for correcting the first category of the text line based on the expression specifications of the element category, and do not limit other possible implementation methods accordingly. Other possible implementation methods are not listed one by one here.

[0050] Step S14: Based on the expression specification of the second category to which the text line belongs, determine the starting character of the second category in the text line, and traverse from the starting character in the target official document until the ending character of the second category, which is used as the element content of the second category.

[0051] In one implementation scenario, after correcting to obtain the second category of the text line, the starting character of the second category in the text can be determined based on the expression specification of the second category to which the text line belongs. Exemplarily, for a text line belonging to the second category of "second-level title", according to the expression specification of "second-level title", the starting character can be determined to be a bracket; or, for a text line belonging to the second category of "third-level title", according to the expression specification of "third-level title", the starting character can be determined to be an Arabic numeral. Of course, the above examples are only several possible situations for determining the starting character in the actual application process, and do not limit other possible situations accordingly. Other possible situations are not listed one by one here.

[0052] In one implementation scenario, after determining the starting character, it is possible to traverse from the starting character in the target official document until the ending character of the second category, which can be used as the element content of the second category. It should be noted that the ending character of the second category is determined by the expression specification of the second category. Taking the element category of "title" as an example, its ending character can include but is not limited to a period, a semicolon, the starting character of the title (that is, when the starting character of a certain title is detected again after traversing from the starting character of the title, it can be considered that the current title ends here), etc. Other possible situations are not listed one by one here. Please refer to Figure 2 , Figure 2 is a schematic diagram of an embodiment for extracting the element content of the second category of this application. As Figure 2As shown, the text enclosed by the black dashed box is the recognized text line. Among them, the second category of the text line "XXX Group" is "Issuing Agency", the second category of the text line "XXXXX Office" is "Issuing Agency", the second category of the text line "XXXXX Business Unit" is "Issuing Agency", the second category of the text line "(1) First, it is necessary to XXXXXXXXXXXXXX" is "Second-level Heading", and so on. The second category of the text line "XXXXXXXXXXXX. Specifically, XXXXX" is "Main Text", the second category of the text line "(2) Second, it is necessary to XXXXXXXXXXXX." is "Third-level Heading", the second category of the text line "(3) Third, it is necessary to XXXXXXXXXXXX." is "Third-level Heading", the second category of the text line "(3) Third, it is necessary to XXXXXXXXXXXX." is "Third-level Heading", the second category of the text line "(4) Fourth, it is necessary to XXXXXXXXXXXX." is "Third-level Heading", the second category of the text line "XXX Group" is "Issuing Agency", and the second category of the text line "CC: XXX General Manager's Office Date of Issue: X year X month X day" is "Footer". For the text line "(1) First, it is necessary to XXXXXXXXXXXXXX", according to the expression specification of the second category "Second-level Heading", the starting character (i.e., the parenthesis) can be determined, and starting from the starting character of this text line, traverse in the target official document shown in Figure 2 until the ending character of the second category (such as a period), then an element content of the second category "Second-level Heading" "(1) First, it is necessary to XXXXXXXXXXXXXX XXXXXXXXXXXX." can be extracted (as shown by the blue dashed box in Figure 2 ). By analogy, the element content of the second category "Main Text" "Specifically, XXXX", the element content of another second category "Second-level Heading" "(2) Second, it is necessary to XXXXXXXXXXXX.", the element content of another second category "Second-level Heading" "(3) Third, it is necessary to XXXXXXXXXXXX.", the element content of another second category "Second-level Heading" "(4) Fourth, it is necessary to XXXXXXXXXXXX.", the element content of another second category "Issuing Agency" "XXX Group", and the element content of the second category "Footer" "CC: XXX General Manager's Office Date of Issue: X year X month X day" can be continuously extracted. Of course, the above example is only one possible example of extracting the element content of the second category, and it does not limit other possible situations, and no more examples will be given here.

[0053] Step S15: Extract structured data from the target official document based on the element content of the second category.

[0054] In an implementation scenario, after extracting the element content of the second category, structured data representing the element extraction of the target official document can be formed. Still taking the Figure 2 target official document shown as an example, the following structured data can be extracted:

[0055] Table 1 Structured data of the target official document after element extraction

[0056]

[0057] Of course, the above representation is only one possible form of structured data, and other possible forms will not be exemplified one by one here. For example, structured data can also be represented in formats such as Json and XML, which are not limited here. It should be noted that in Table 1 above, "Text" represents the element content, and "box" represents the position information of the element content (such as the upper left coordinate and the lower right coordinate of the area where the element content is located).

[0058] In an implementation scenario, after extracting the structured data, it is also possible to select the element content of the second category as the target category and the text lines of which are continuous as the target content, and the target category includes at least the issuing agency. Then, based on the position information of the text lines to which each target content belongs, the reading order of each target content in the structured data is adjusted. The above method can improve the reading fluency of the structured data by modulating the reading order in the structured data according to the position information.

[0059] In a specific implementation scenario, as mentioned above, the position information can include the abscissa and the ordinate. The coordinate system can take the upper left vertex of the target official document as the coordinate origin, the horizontal extension direction of the text as the horizontal axis, and the vertical extension direction of the text as the vertical axis. Then, when adjusting the reading order based on the position information, specifically, the reading order of each target content in the structured data can be first adjusted based on the abscissa of the text line to which each target content belongs. For example, the smaller the abscissa, the earlier the reading order of the target content, and vice versa, the larger the abscissa, the later the reading order of the target content. On this basis, for the target content that still cannot distinguish the reading order after the first adjustment, the reading order can be secondarily adjusted based on the ordinate of the text line to which it belongs. For example, the smaller the ordinate, the earlier the reading order, and vice versa, the larger the ordinate, the later the reading order. The above method can make the adjusted reading order more in line with the reading habit by first adjusting and then secondarily adjusting through the abscissa and the ordinate.

[0060] In a specific implementation scenario, please refer to Table 1 and Figure 2, taking the target category as "issuing agency" as an example, the element contents of the second category "issuing agency" and with continuous text lines include: "XXXXX Office", "XXX Group", "XXXXX Business Unit", and their original reading order is also like this. Then, the above element contents can be used as the target content first. Since the abscissa of the text line where the target content "XXX Group" is located is the smallest, and the abscissas of the text lines where the other two are located cannot be distinguished in size, the first adjustment can be made first so that the reading order of the target content "XXX Group" is the foremost among the three. Then, based on the ordinate, the second adjustment is made to the remaining two. Since the ordinate of the text line where the target content "XXXXX Office" is located is smaller, its reading order can be placed before the target content "XXXXX Business Unit". That is to say, the adjusted reading order is: "XXX Group", "XXXXX Office", "XXXXX Business Unit". Of course, the above example is just a possible example of adjusting the reading order in the actual application process, and other possible situations will not be exemplified one by one here.

[0061] In an implementation scenario, after extracting the structured data, it is also possible to select the element contents of the same second category and with continuous text lines for splicing to obtain the spliced content of the second category, and then split the spliced content of the second category based on the expression specification of the second category to obtain the element contents of different element units under the second category. Please refer to Figure 2The element contents that are in the same second category and have consecutive text lines in Table 1 include: "(2) Second, key XXXXXXXXXXX.", "(3) Third, key XXXXXXXXXXX.", "(4) Fourth, key XXXXXXXXXXX.". Therefore, the above three can be concatenated first to obtain the concatenated content "(2) Second, key XXXXXXXXXXX. (3) Third, key XXXXXXXXXXX. (4) Fourth, key XXXXXXXXXXX.", and then the above concatenated content can be split based on the expression specification of the second category "Second Title" (specifically, refer to the relevant description above) to obtain the element contents of different element units under the "Second-level Title" of the second category: "(2) Second, key XXXXXXXXXXX.", "(3) Third, key XXXXXXXXXXX.", "(4) Fourth, key XXXXXXXXXXX.". Of course, the above example is only one possible example in the actual application process, and other possible situations will not be listed one by one here. In the above method, the element contents that are in the same second category and have consecutive text lines are selected for concatenation to obtain the concatenated content of the second category, and then the concatenated content of the second category is split based on the expression specification of the second category to obtain the element contents of different element units under the second category. Therefore, even when the concatenated content contains multiple elements of the same type, the accuracy of the finally extracted element content can be further improved.

[0062] In an implementation scenario, after extracting the structured data, multi-element matching can also be performed on the element content of a single text line based on regular expressions to obtain a matching result. Thus, in response to the matching result indicating that it contains multiple element units, the character separation positions of the multiple element units can be obtained, and based on the character separation positions, the element content of the text line can be divided. Moreover, for the element content after the division of a single text line, the spatial position of the element content in the target official document can be determined according to the division ratio. It should be noted that regular expressions can be used to detect, for example, whether there are spaces, etc., so as to achieve multi-element matching of the element content of a single text line through regular expressions. Please refer to Figure 2 and Table 1, Figure 2When performing multi-element matching on the element content of the last text line in " Figure 2 as shown by the green dashed box in

[0063] In an implementation scenario, after extracting the structured data, the document format specifications of the target official document can also be obtained, and then based on the document format specifications, various element contents of the second category are filtered respectively to obtain new structured data of the target official document. It should be noted that the document format specifications of the target official document may include, but are not limited to: the issuing agency can only appear in the center of the article, if an attachment is detected, the content subsequent to the endnote does not need to be recognized, the date of issue is generally located at the end of the article, the content outside the typesetting area does not need to be recognized, etc., and no more examples of the document format specifications are given here. Please continue to refer to Figure 2 and Table 1. For the element content "For XXX Group" of the second second category "issuing agency" in Table 1 (which is actually Figure 2The content in the seal shown), since its position (as shown in box: {x91, y91, x92, y92} in Table 1) is not the central position of the article, the content of this element can be screened and filtered. Of course, the above examples are only several possible examples of screening and filtering element content in the actual application process, and other possible situations will not be listed one by one here. In the above manner, the writing specification of the target official document is obtained, and then based on the writing specification, the content of various elements of the second category is screened and filtered respectively to obtain new structured data of the target official document, which can minimize the misidentification of elements.

[0064] In the above solution, each text line in the target official document is identified, and based on the multi-modal features of the text line, the first category of the text line is predicted. Then, based on the expression specifications of several element categories, the first category of the text line is corrected to obtain the second category of the text line. Both the first category and the second category are one of several element categories. Furthermore, based on the expression specification of the second category to which the text line belongs, the starting character belonging to the second category in the text line is determined, and in the target official document, it is traversed from the starting character until the ending character of the second category, which is used as the element content of the second category. In order to extract structured data from the target official document based on the element content of the second category, on the one hand, taking the text line as the basic unit of category prediction can greatly improve the recall rate and reduce missed detections. And by integrating the multi-modal features of the text line for category prediction, it helps to improve the accuracy of category prediction for text lines even when facing official documents with complex layouts. On the other hand, after predicting the category through the multi-modal features of the text line, further combining the expression specifications of element categories can accurately correct the predicted category by combining rules, so that the element categories with clear rules can be more reliably identified. Therefore, the accuracy of official document element extraction can be improved.

[0065] Please refer to Figure 3 , Figure 3 is a schematic framework diagram of an embodiment of the official document element extraction device of the present application. The official document element extraction device 30 includes: an identification module 31, a classification module 32, a correction module 33, an extraction module 34, and an integration module 35. The identification module 31 is used to identify each text line in the target official document; the classification module 32 is used to predict the first category of the text line based on the multi-modal features of the text line; the correction module 33 is used to correct the first category of the text line based on the expression specifications of several element categories to obtain the second category of the text line. Wherein, both the first category and the second category are one of several element categories; the extraction module 34 is used to determine the starting character belonging to the second category in the text line based on the expression specification of the second category to which the text line belongs, and in the target official document, it is traversed from the starting character until the ending character of the second category, which is used as the element content of the second category; the integration module 35 is used to extract structured data from the target official document based on the element content of the second category.

[0066] In the above solution, the official document element extraction device 30 identifies each text line in the target official document, predicts the first category of the text line based on the multi-modal features of the text line, and then corrects the first category of the text line based on the expression specifications of several element categories to obtain the second category of the text line. Both the first category and the second category are one of several element categories. Furthermore, based on the expression specification of the second category to which the text line belongs, the starting character belonging to the second category in the text line is determined, and the target official document is traversed from the starting character until the ending character of the second category, which is used as the element content of the second category. Based on the element content of the second category, structured data is extracted from the target official document. Therefore, on the one hand, taking the text line as the basic unit of category prediction can greatly improve the recall rate and reduce missed detections. And by integrating the multi-modal features of the text line for category prediction, it helps to improve the accuracy of category prediction for text lines even when facing official documents with complex layouts. On the other hand, after predicting the category through the multi-modal features of the text line, further combining the expression specifications of the element categories can accurately correct the predicted category according to the rules, so that the element categories with clear rules can be more reliably recognized. Therefore, the accuracy of official document element extraction can be improved.

[0067] In some disclosed embodiments, the classification module 32 includes a layer normalization sub-module for performing layer normalization based on the multi-modal features to obtain text line features; the classification module 32 includes a feature encoding sub-module for encoding based on the text line features to obtain row-level encoded features; the classification module 32 includes a classification prediction sub-module for predicting based on the row-level encoded features to obtain the first category.

[0068] In some disclosed embodiments, when the multi-modal features include content features, the content features are obtained by encoding the recognized text of the text line by a sentence-level encoder and then performing a linear mapping.

[0069] In some disclosed embodiments, when the multi-modal features include layout features, the layout features are obtained by respectively performing embedding representation on the first layout information and the second layout information of the text line and then fusing them. The first layout information includes the extreme coordinates of the text line on the horizontal axis and the width of the text line, and the second layout information includes the extreme coordinates of the text line on the vertical axis and the height of the text line.

[0070] In some disclosed embodiments, when the multi-modal features include position features, the position features are obtained by performing embedding representation on the line number encoding of the text line in the target official document.

[0071] In some disclosed embodiments, when the multimodal features include visual features, the visual features are extracted from the feature map of the target image in the detection area of the target image through text lines, where the target image is the page image where the text lines are located in the target official document.

[0072] In some disclosed embodiments, when the multimodal features include page number features, the page number features are embedded and characterized by the page numbers where the text lines are located in the target official document.

[0073] In some disclosed embodiments, the official document element extraction device 30 includes a content selection module for selecting the element content with the second category being the target category and the text lines to which it belongs being continuous as the target content; wherein, the target category at least includes the issuing agency; the official document element extraction device 30 includes an order adjustment module for adjusting the reading order of each target content in the structured data based on the position information of the text lines to which each target content belongs.

[0074] In some disclosed embodiments, the position information includes the abscissa and the ordinate. The official document element extraction device 30 includes a first adjustment module for making a first adjustment to the reading order of each target content in the structured data based on the abscissa of the text lines to which each target content belongs; the official document element extraction device 30 includes a second adjustment module for making a second adjustment to the reading order of the target content that still cannot be distinguished after the first adjustment based on the ordinate of the text lines to which it belongs.

[0075] In some disclosed embodiments, the official document element extraction device 30 includes a content splicing module for selecting and splicing the element content with the same second category and the text lines to which it belongs being continuous to obtain the spliced content of the second category; the official document element extraction device 30 includes a content splitting module for splitting the spliced content of the second category based on the expression specification of the second category to obtain the element content of different element units under the second category.

[0076] In some disclosed embodiments, the official document element extraction device 30 includes a regular matching module for performing multi-element matching on the element content of a single text line based on a regular expression to obtain a matching result; the official document element extraction device 30 includes an element division module for, in response to the matching result being characterized as including multiple element units, obtaining the character separation positions of the multiple element units, and based on the character separation positions, dividing the element content of the text line, and for the element content after division of a single text line, determining the spatial position of the element content in the target official document according to the division ratio.

[0077] In some disclosed embodiments, the official document element extraction device 30 includes a specification acquisition module for acquiring the writing specification of the target official document; the official document element extraction device 30 includes a content filtering module for screening and filtering various second-category element contents based on the writing specification to obtain new structured data of the target official document.

[0078] In some disclosed embodiments, several element categories at least include the issuing agency, title, body text, and footer; and / or, the ending character of the second category is determined by the expression specification of the second category; and / or, the multimodal features include at least two of content features, layout features, position features, visual features, and page number features.

[0079] Please refer to Figure 4 , Figure 4 is a framework schematic diagram of an embodiment of the electronic device of the present application. The electronic device 40 at least includes a memory 41 and a processor 42 that are coupled to each other. At least program instructions are stored in the memory 41, and the processor 42 is configured to execute the program instructions to implement the steps in any of the above-mentioned official document element extraction method embodiments. Specifically, reference can be made to the foregoing disclosed embodiments, which will not be elaborated here. As a possible example, the electronic device 40 may include, but is not limited to, a smart phone, a tablet computer, a server, etc. The specific type of the electronic device 40 is not limited here.

[0080] Specifically, the processor 42 is configured to control itself and the memory 41 to implement the steps in any of the above-mentioned official document element extraction method embodiments. The processor 42 may also be referred to as a CPU (Central Processing Unit). The processor 42 may be an integrated circuit chip with signal processing capabilities. The processor 42 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 42 may be implemented jointly by integrated circuit chips.

[0081] In the above solution, the electronic device 40 identifies each text line in the target official document, predicts the first category of the text line based on the multi-modal features of the text line, and then corrects the first category of the text line based on the expression specifications of several element categories to obtain the second category of the text line. Both the first category and the second category are one of several element categories. Furthermore, based on the expression specification of the second category to which the text line belongs, the starting character belonging to the second category in the text line is determined, and the target official document is traversed from the starting character until the ending character of the second category, which is used as the element content of the second category. Based on the element content of the second category, structured data is extracted from the target official document. Therefore, on the one hand, taking the text line as the basic unit of category prediction can greatly improve the recall rate and reduce missed detections. And by integrating the multi-modal features of the text line for category prediction, it helps to improve the accuracy of category prediction for text lines even when facing official documents with complex layouts. On the other hand, after predicting the category through the multi-modal features of the text line and further combining the expression specifications of the element categories, it can accurately correct the predicted category according to the rules, making the element categories with clear rules more reliably identified. Therefore, the accuracy of official document element extraction can be improved.

[0082] Please refer to Figure 5 , Figure 5 FIG. is a schematic framework diagram of an embodiment of the computer-readable storage medium 50 of the present application. The computer-readable storage medium 50 stores program instructions 51 that can be run by a processor. The program instructions 51 are used to implement the steps in any of the above-mentioned embodiments of the official document element extraction method.

[0083] In the above solution, the computer-readable storage medium 50 identifies each text line in the target official document, predicts the first category of the text line based on the multi-modal features of the text line, and then corrects the first category of the text line based on the expression specifications of several element categories to obtain the second category of the text line. Both the first category and the second category are one of several element categories. Furthermore, based on the expression specification of the second category to which the text line belongs, the starting character belonging to the second category in the text line is determined, and the target official document is traversed from the starting character until the ending character of the second category, which is used as the element content of the second category. Based on the element content of the second category, structured data is extracted from the target official document. Therefore, on the one hand, taking the text line as the basic unit of category prediction can greatly improve the recall rate and reduce missed detections. And by integrating the multi-modal features of the text line for category prediction, it helps to improve the accuracy of category prediction for text lines even when facing official documents with complex layouts. On the other hand, after predicting the category through the multi-modal features of the text line and further combining the expression specifications of the element categories, it can accurately correct the predicted category according to the rules, making the element categories with clear rules more reliably identified. Therefore, the accuracy of official document element extraction can be improved.

[0084] In some embodiments, the functions or modules included in the apparatus provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0085] The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to each other. For the sake of brevity, they will not be repeated in this article.

[0086] In several embodiments provided in the present application, it should be understood that the disclosed methods and apparatuses can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0087] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0088] In addition, in each embodiment of the present application, the various functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0089] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0090] If the technical solution of this application involves personal information, before the product applying this technical solution processes personal information, it has clearly informed the personal information processing rules and obtained the personal's autonomous consent. If the technical solution of this application involves sensitive personal information, before the product applying this technical solution processes sensitive personal information, it has obtained the personal's separate consent and at the same time meets the requirement of "express consent". For example, at a personal information collection device such as a camera, a clear and prominent sign is set to inform that the personal information collection range has been entered and personal information will be collected. If an individual voluntarily enters the collection range, it is regarded as consenting to the collection of their personal information; or on the personal information processing device, when the personal information processing rules are informed by obvious signs / information, personal authorization is obtained through pop-up messages or by asking the individual to upload their personal information themselves; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.

Claims

1. A method for extracting elements of official documents, characterized in that: include: Identify individual text lines in the target document; Predicting a first category of the text line based on the multimodal features of the text line; Modifying the first category of the text line based on expression specifications of several element categories to obtain a second category of the text line; wherein the first category and the second category are both one of the several element categories; Based on the expression specification of the second category to which the text line belongs, determining the starting character of the text line that belongs to the second category, and traversing from the starting character to the ending character of the second category in the target document as the element content of the second category; Based on the element content of the second category, structured data is extracted from the target document.

2. The method according to claim 1, characterized in that The predicting, based on the multimodal features of the text line, a first category of the text line includes: Performing layer normalization based on the multimodal features to obtain text line features; Encoding is performed based on the text line features to obtain line-level encoding features; Prediction is performed based on the row-level coding features to obtain the first category.

3. The method according to claim 2, characterized in that In the case where the multimodal feature includes a content feature, the content feature is obtained by encoding the recognition text of the text line by a sentence-level encoder and then linearly mapping it; And / or, in the case where the multimodal feature includes a layout feature, the layout feature is obtained by respectively embedding and representing the first layout information and the second layout information of the text line and then fusing them, and the first layout information includes the extreme value coordinates of the text line on the horizontal axis and the width of the text line, and the second layout information includes the extreme value coordinates of the text line on the vertical axis and the height of the text line; And / or, in the case where the multimodal feature includes a position feature, the position feature is obtained by embedding the line number code of the text line in the target document; And / or, in the case where the multimodal feature includes a visual feature, the visual feature is obtained by extracting the detection area of ​​the text line in the target image from a feature map of the target image, and the target image is an image of a page where the text line is located in the target document; And / or, in the case where the multimodal feature includes a page number feature, the page number feature is obtained by embedding the page number code of the text line in the target document.

4. The method according to claim 1, characterized in that: After extracting structured data from the target document based on the element content of the second category, the method further includes: Selecting the element contents whose second category is a target category and which belong to the continuous text lines as target contents; wherein the target category at least includes an issuing institution; Based on the position information of the text line to which each of the target contents belongs, the reading order of each of the target contents in the structured data is adjusted.

5. The method according to claim 4, characterized in that The position information includes a horizontal coordinate and a vertical coordinate, and adjusting the reading order of each target content of the structured data based on the position information of the text line to which each target content belongs respectively includes: Based on the horizontal coordinates of the text lines to which the target contents respectively belong, a first adjustment is made to the reading order of the target contents of the structured data; For the target content whose reading order is still indistinguishable after the first adjustment, a second adjustment is made to the reading order based on the vertical coordinate of the corresponding text line.

6. The method according to claim 1, characterized in that After extracting structured data from the target document based on the element content of the second category, the method further includes: Selecting and splicing element contents of the same second category and belonging to continuous text lines to obtain spliced ​​contents of the second category; The spliced ​​content of the second category is split based on the expression specification of the second category to obtain element contents of different element units under the second category.

7. The method according to claim 1, characterized in that After extracting structured data from the target document based on the element content of the second category, the method further includes: Performing multi-element matching on the element content of a single text line based on regular expressions to obtain a matching result; In response to the matching result being characterized as containing multiple element units, the character separation positions of the multiple element units are obtained, and based on the character separation positions, the element content of the text line is divided. As for the element content after the single text line is divided, the spatial position of the element content in the target document is determined according to the division ratio.

8. The method according to claim 1, characterized in that After extracting structured data from the target document based on the element content of the second category, the method further includes: Obtaining the writing specifications of the target official document; Based on the writing specifications, the various elements of the second category are screened and filtered to obtain new structured data of the target document.

9. The method according to any one of claims 1 to 8, characterized in that: The several element categories at least include the issuing agency, title, text, and colophon; and / or, the end character of the second category is determined by the expression specification of the second category; And / or, the multimodal feature includes at least two of a content feature, a layout feature, a position feature, a visual feature, and a page number feature.

10. A document element extraction device, characterized in that: include: A recognition module, used to recognize individual text lines in the target document; A classification module, configured to predict a first category of the text line based on the multimodal features of the text line; A correction module, configured to correct the first category of the text line based on expression specifications of several element categories to obtain a second category of the text line; wherein the first category and the second category are both one of the several element categories; an extraction module, configured to determine, based on the expression specification of the second category to which the text line belongs, a starting character in the text line that belongs to the second category, and traverse from the starting character to the ending character of the second category in the target document as element content of the second category; An integration module is used to extract structured data from the target document based on the element content of the second category.

11. An electronic device, characterized in that: The method comprises at least a memory and a processor coupled to each other, wherein the memory at least stores program instructions, and the processor is used to execute the program instructions to implement the method for extracting official document elements as described in any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that: Program instructions that can be executed by a processor are stored, and the program instructions are used to implement the document element extraction method described in any one of claims 1 to 9.