Text information extraction method and device, computer equipment and storage medium

By identifying and analyzing text information line by line named entity, combined with multiple column processing steps, the problems of information loss and incorrect classification in the existing text information extraction methods are solved, and more accurate and efficient text information extraction is achieved.

CN120047963APending Publication Date: 2025-05-27PING AN HEALTH INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411763469.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-02
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing text information extraction methods can easily lead to information loss and mixing when processing multiple columns of information in a row, or incorrectly classify the major category names into detailed project names, resulting in errors in overall data extraction.

Method used

By identifying the line-by-line named entity, obtain text information, parse the target line text, confirm whether each text element is a title element, determine the number of tables and data columns of the data table based on the same number of elements and the number of text elements, divide the text information, and perform multiple column processing to obtain the data table.

Benefits of technology

It effectively solves the problem of multiple columns of information loss, improves the accuracy and efficiency of text information extraction, and significantly reduces changes to existing processing systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047963A_ABST
    Figure CN120047963A_ABST
Patent Text Reader

Abstract

The invention discloses a text information extraction method and device, computer equipment and a storage medium. A to-be-extracted text image is obtained, named entity recognition is performed on the text image line by line, and text information in the text image is obtained. A target line text in the text information is analyzed, multiple text elements contained in the target line text are obtained, and whether each text element is a title element or not is determined. And if each text element is the title element, obtaining the same element number of each text element in the plurality of text elements. And determining the table number of the data tables after the text information extraction and the data column number in each data table according to the number of the same elements and the number of the text elements. The text information is divided according to the table number of the data table, and multiple pieces of sub-text information are obtained. And extracting each piece of sub-text information according to the table number, the data column number and the line number of the text information, obtaining a data table corresponding to each piece of sub-text information, and completing text information extraction of the text image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image recognition technology, and in particular to a text information extraction method, device, computer equipment and storage medium. Background Art

[0002] Extracting text information from images can quickly extract text information from images and summarize and organize data. It has a wide range of application needs in scenarios with large amounts of text images, such as medical care, finance, and education.

[0003] At present, the automatic extraction of text information adopts the combination of optical character recognition (OCR) and text entity extraction. However, in the current text image extraction, it is assumed that one line includes one detail, and the number of details in the actual text image is uncertain. Because in the existing extraction processing method, for each entity type, only one line is extracted, resulting in partial information loss and mixing when there are multiple columns of information in a line, or the general category name is misclassified as the detailed item name, resulting in the overall data being extracted incorrectly, which in turn leads to the unsatisfactory effect of the existing text extraction method. Summary of the invention

[0004] The present application provides a text information extraction method, apparatus, computer equipment and storage medium, which aims to solve the problem that in the existing extraction processing method, for each entity type, only one row is extracted, resulting in partial information loss and mixing when there are multiple columns of information in a row, or the general category name is incorrectly classified as a detailed item name, resulting in the overall data being extracted incorrectly, and thus causing the existing text extraction method to be less than ideal. The method provided compensates for the problem of multi-column information loss in the current conventional claim list extraction, and by prepending the multi-column processing steps, it minimizes the changes to the current existing processing system.

[0005] In a first aspect, the present application provides a text information extraction method, comprising:

[0006] Obtain the text image to be extracted, perform named entity recognition on the text image line by line, and obtain the text information in the text image;

[0007] Parse the target line text in the text information, obtain multiple text elements contained in the target line text, and confirm whether each text element is a title element; wherein the target line text is located in the first line of the text information according to the recognition order;

[0008] If each text element is a title element, get the number of identical elements of each text element in multiple text elements;

[0009] Determine the number of tables in the data table after text information extraction and the number of data columns in each data table according to the number of identical elements and the number of text elements;

[0010] Divide the text information according to the number of tables in the data table to obtain multiple sub-text information;

[0011] Each sub-text information is extracted according to the number of tables, the number of data columns and the number of rows of text information, and the data table corresponding to each sub-text information is obtained to complete the text information extraction of the text image.

[0012] In some embodiments, before obtaining the data table corresponding to each sub-text information, it also includes: obtaining the title element type corresponding to the text element of the target row; the title element type includes identification elements and detail elements; obtaining the target column element corresponding to the detail element from the multiple column elements extracted from each sub-text information; wherein each column element includes multiple text elements arranged row by row; constructing a detail dictionary according to each target column element, and determining whether there is an empty element in the detail dictionary; if there is an empty element in the detail dictionary, determining the text element corresponding to the empty element according to the text elements in the detail dictionary in the same row as the empty element, and completing the update of the detail dictionary; updating the target column element according to the updated detail dictionary.

[0013] Exemplarily, before constructing a detailed dictionary based on each target column element, the method includes: inputting the text element of the target row corresponding to each target column element into a preset title element prediction model, and the title element prediction model outputs a title prediction result corresponding to each text element; the title prediction result includes any one of an identification element and a detail element; if the title prediction result includes multiple detail elements containing preset keywords, obtaining the first text position information of the text element containing the preset keywords; confirming whether the text element containing the preset keywords is an identification element based on the first text position information; if the text element containing the preset keywords is an identification element, confirming that the text element is a target text element, and removing the target column element corresponding to the target text element from multiple target column elements.

[0014] Exemplarily, text information is divided according to the number of tables in a data table to obtain multiple sub-text information, including: obtaining the second text position information of each text element in a target row; dividing the multiple text elements in the target row into text element groups according to the number of tables, the second text position information of each text element and the type of title element corresponding to each text element; wherein the number of identification elements and detail elements in each text element group is the same; determining the division position for dividing the text information according to the second text position information corresponding to the text elements in each text element group; dividing the text information according to the division position to obtain multiple sub-text information.

[0015] In some embodiments, each sub-text information is extracted according to the number of tables, the number of data columns and the number of rows of text information to obtain a data table corresponding to each sub-text information, including: generating multiple blank tables according to the number of tables, the number of data columns and the number of rows of text information; the blank table includes a title row and multiple content rows; filling the text elements of the target row into the title row in sequence; determining the number of extractions for the sub-text information according to the number of data columns; extracting the text elements of each row of the sub-text information except the title row one by one according to the number of extractions to obtain multiple extracted elements; adding each extracted element to the content row of the blank table according to the number of extractions to complete the filling of the blank table and obtain a data table corresponding to each sub-text information.

[0016] In some embodiments, the text image to be extracted is stored in a preset insurance claims system; confirming whether each text element is a title element includes: obtaining multiple preset titles corresponding to the insurance claims system; matching each text element with multiple preset titles, and calculating the similarity between the text element and each preset title; if at least one similarity among the multiple similarities is greater than a preset threshold, confirming that the text element is a title element.

[0017] Exemplarily, before determining the number of tables in the data table after text information extraction and the number of data columns in each data table based on the number of identical elements and the number of text elements, it also includes: if each text element is not a title element, calculating the matching degree of the text element with each preset title based on a self-attention mechanism; determining the title element corresponding to each text element based on multiple matching degrees; and obtaining the number of identical elements of the title element corresponding to each text element among the title elements corresponding to multiple text elements.

[0018] In a second aspect, the present application provides a text information extraction device, comprising:

[0019] An image acquisition module is used to acquire the text image to be extracted, perform named entity recognition on the text image line by line, and acquire the text information in the text image;

[0020] A text parsing module, used to parse a target line of text in the text information, obtain multiple text elements contained in the target line of text, and confirm whether each text element is a title element; wherein the target line of text is located in the first line of the text information according to the recognition order;

[0021] An element acquisition module, used for acquiring the number of identical elements of each text element in multiple text elements if each text element is a title element;

[0022] A column number acquisition module is used to determine the number of tables in the data table after the text information is extracted and the number of data columns in each data table according to the number of identical elements and the number of text elements;

[0023] An information acquisition module, used to divide the text information according to the number of tables in the data table, and acquire multiple sub-text information;

[0024] The text extraction module is used to extract each sub-text information according to the number of tables, the number of data columns and the number of rows of text information, obtain the data table corresponding to each sub-text information, and complete the text information extraction of the text image.

[0025] In a third aspect, the present application further provides a computer device, including:

[0026] Memory and processor;

[0027] The memory is used to store computer programs;

[0028] The processor is used to execute the computer program and implement the steps of the text information extraction method described in the first aspect when executing the computer program.

[0029] In a fourth aspect, the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor implements the steps of the text information extraction method as described in the first aspect above.

[0030] The present application discloses a method, device, computer equipment and storage medium for extracting text information. The method first obtains a text image to be extracted, performs named entity recognition on the text image line by line, and obtains text information in the text image. Then, the target line text in the text information is parsed, multiple text elements contained in the target line text are obtained, and it is confirmed whether each text element is a title element; wherein the target line text is located in the first line of the text information in the recognition order. Then, if each text element is a title element, the number of identical elements of each text element in multiple text elements is obtained. Further, the number of tables in the data table after the text information is extracted and the number of data columns in each data table are determined according to the number of identical elements and the number of text elements. Furthermore, the text information is divided according to the number of tables in the data table to obtain multiple sub-text information. Finally, each sub-text information is extracted according to the number of tables, the number of data columns and the number of rows of the text information, and the data table corresponding to each sub-text information is obtained to complete the text information extraction of the text image.

[0031] Furthermore, the provided method obtains the number of identical elements of each text element in multiple text elements to determine the number of tables in the data table and the number of data columns in each data table after the text information is extracted, by obtaining, when each text element is a title element. The method can then accurately divide the text information according to the number of tables in the data table, obtain multiple sub-text information, and then extract each sub-text information according to the number of tables, the number of data columns and the number of rows of text information. This can make up for the problem of multi-column information loss in the current conventional claims list extraction, and by pre-pending multi-column processing steps, minimize changes to the current existing processing system.

[0032] At the same time, adding a multi-column information extraction method can significantly improve the automation of information extraction of expense detail lists in financial systems, such as claim materials, and greatly reduce the difficulty and time of manual input (from more than 1 hour for a list to less than 10 minutes, and from the original manual input to the correction of a small number of typos). In general, the provided method completes the extraction of multi-column text information, greatly shortens the customer's claim time, improves the claim speed and quick claim ratio, and improves user satisfaction in the claim process.

[0033] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0035] Figure 1 is a schematic flow chart of the steps of a text information extraction method provided in an embodiment of the present application;

[0036] Figure 2 is a schematic flow chart of the steps of a method for updating a target column element provided by an embodiment of the present application;

[0037] Figure 3 is a schematic flow chart of the steps of a data table acquisition method provided in an embodiment of the present application;

[0038] Figure 4 It is a structural schematic diagram of a text information extraction device provided in an embodiment of the present application;

[0039] Figure 5 It is a schematic block diagram of the structure of a computer device provided in one embodiment of the present application.

[0040] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. DETAILED DESCRIPTION

[0041] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0042] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may also be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions.

[0043] It should be understood that, in order to facilitate the clear description of the technical solutions of the embodiments of the present invention, in the embodiments of the present invention, the words "first", "second", etc. are used to distinguish the same items or similar items with substantially the same functions and effects. Those skilled in the art can understand that the words "first", "second", etc. do not limit the quantity and execution order, and the words "first", "second", etc. do not necessarily limit the difference.

[0044] It should be understood that the terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit the application. As used in this application specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.

[0045] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0046] In conjunction with the accompanying drawings, some embodiments of the present application are described in detail below. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.

[0047] Extracting text information from images can quickly extract text information from images and summarize the data. It has a wide range of application needs in scenarios with large amounts of text images, such as medical care, finance, and education.

[0048] At present, the automatic extraction of text information adopts the combination of optical character recognition (OCR) and text entity extraction. However, in the current text image extraction, it is assumed that one line includes one detail, and the number of details in the actual text image is uncertain. Because in the existing extraction processing method, for each entity type, only one line is extracted, resulting in partial information loss and mixing when there are multiple columns of information in a line, or the general category name is misclassified as the detailed item name, resulting in the overall data being extracted incorrectly, which in turn leads to the unsatisfactory effect of the existing text extraction method.

[0049] To solve the above problems, please refer to Figure 1 , Figure 1 The figure is a schematic flow chart of a text information extraction method provided by an embodiment of the present application. The text information extraction method can be implemented by a computer device, which can be deployed on a single server or a server cluster. It can also be deployed on a handheld terminal, a laptop computer, a wearable device or a robot, etc.

[0050] The provided method can be applied in the financial field to extract information such as expense details in claim materials, and the provided method can quickly extract text from a large amount of materials. At the same time, this application can also be applied to other scenarios such as medical care and education that require a large amount of text extraction, so the embodiments of this application do not limit the application scenarios of the provided method.

[0051] Text information extraction is a technology that extracts specific information from text images. Text images are images composed of some specific units, such as sentences, paragraphs, and chapters. Text information is an image composed of some small specific units, such as characters, words, phrases, sentences, paragraphs, or a combination of these specific units. Extracting noun phrases, names, place names, etc. from text images are all text information extraction. Of course, the information extracted by text information extraction technology can be various types of information.

[0052] To solve the above problems, please refer to Figure 1 Specifically, Figure 1 As shown, the provided method includes steps S101 to S106. The details are as follows:

[0053] S101. Obtain a text image to be extracted, perform named entity recognition on the text image line by line, and obtain text information in the text image.

[0054] Specifically, when a server equipped with the method provided in an embodiment of the present application has a text image to be extracted in the application scenario (such as a claim list in a financial scenario that requires text extraction), it can complete the preliminary recognition of the text image by first recognizing the text image line by line.

[0055] In some embodiments, performing named entity recognition on a text image further includes first performing optical character recognition (OCR) on the text image line by line, and performing named entity recognition (NER) on the result of the optical character recognition to obtain text information in the text image. The provided method is then combined with the OCR+NER method to complete rapid recognition of text images.

[0056] In some embodiments, before performing named entity recognition on the text image line by line, the text image is further corrected to ensure that the difference between the rotation angle of each text image and 0° is less than a preset angle (e.g., 5°). For example, the rotation angle of the text image is determined by identifying whether a preset number of characters are rotated.

[0057] S102. Parse the target line text in the text information, obtain multiple text elements contained in the target line text, and confirm whether each text element is a title element; wherein the target line text is located in the first line of the text information according to the recognition order.

[0058] Specifically, the server parses the target line text in the text information, such as if the target line text is located in the first line from the top of the text information, and then determines whether the target line text located in the first line of the text information corresponds to the title, so as to ensure that the provided method can accurately extract the content of the target line text.

[0059] S103. If each text element is a title element, obtain the number of identical elements of each text element in multiple text elements.

[0060] Specifically, if all text elements in the target row text are title elements (such as XXX name, XXX category, etc.), by obtaining the number of identical elements of each text element in multiple text elements, it can be determined that the target row includes multiple text elements of the same category, that is, the server will think that there are corresponding numbers of tables to be extracted in the text image at this time. The above method can greatly improve the accuracy of extraction.

[0061] S104. Determine the number of tables in the data table after text information extraction and the number of data columns in each data table according to the number of identical elements and the number of text elements.

[0062] Specifically, the server can determine the number of identical elements based on the number of identical elements and the number of text elements, and then determine the number of data columns in each data table, thereby determining the extraction result that should be obtained in this extraction and clarifying the extraction target.

[0063] In some embodiments, the text image to be extracted is stored in a preset insurance claims system; confirming whether each text element is a title element includes: obtaining multiple preset titles corresponding to the insurance claims system; matching each text element with multiple preset titles, and calculating the similarity between the text element and each preset title; if at least one similarity among the multiple similarities is greater than a preset threshold, confirming that the text element is a title element.

[0064] By determining that a corresponding scene preset title and a general standard preset title are combined to form a corresponding plurality of preset titles according to the characteristics of the application of the provided method, the text similarity between the text element and the text corresponding to each preset title can be calculated respectively, and then when at least one similarity is greater than a preset threshold (such as 0.8, which can be arbitrarily set according to requirements), it is possible to accurately determine that the text element is a title element and the type of the corresponding title element.

[0065] Exemplarily, before determining the number of tables in the data table after text information extraction and the number of data columns in each data table based on the number of identical elements and the number of text elements, it also includes: if each text element is not a title element, calculating the matching degree of the text element with each preset title based on a self-attention mechanism; determining the title element corresponding to each text element based on multiple matching degrees; and obtaining the number of identical elements of the title element corresponding to each text element among the title elements corresponding to multiple text elements.

[0066] When the text element is not a title element, the provided method can analyze the title corresponding to the current text element. For example, if the text element is 10 yuan, it can be analyzed that its corresponding title is expenses, etc. The provided method combines the self-attention mechanism, such as by calculating the Softmax function value corresponding to the text element and each preset title, it can extract the correlation between the text element and the title element, and then quickly determine the title corresponding to the text element in each column when the text image does not contain a title.

[0067] S105. Divide the text information according to the number of tables in the data table to obtain a plurality of sub-text information.

[0068] Specifically, after determining the number of tables in the data table, the server can equally divide the text information based on the number of tables to split the complete text information into multiple sub-text information. At the same time, combined with the distribution pattern of the text elements of the target row, the position of dividing the text information can be more accurately determined.

[0069] S106. Extract each sub-text information according to the number of tables, the number of data columns and the number of rows of text information, obtain the data table corresponding to each sub-text information, and complete the text information extraction of the text image.

[0070] Specifically, after dividing and obtaining multiple sub-text information, the server can extract text from each sub-text information according to the number of tables, the number of data columns and the number of rows of text information, thereby filling and obtaining multiple data tables. The provided method can accurately extract text images containing multiple columns of text (the present application does not limit the number of columns of text images).

[0071] In some embodiments, Figure 2 As shown, before obtaining the data table corresponding to each sub-text information, steps S107a to S107e are also provided.

[0072] like Figure 2 As shown, the provided target column element updating method includes the following steps:

[0073] Step S107a. Obtain the title element type corresponding to the text element of the target row; the title element type includes an identification element and a detail element;

[0074] Step S107b. Obtaining a target column element corresponding to a detail element from a plurality of column elements extracted from each sub-text information; wherein each column element includes a plurality of text elements arranged row by row;

[0075] Step S107c. Construct a detailed dictionary based on each target column element, and determine whether there is an empty element in the detailed dictionary;

[0076] Step S107d. If there is an empty element in the detail dictionary, determine the text element corresponding to the empty element according to the text element in the same row as the empty element in the detail dictionary, and complete the update of the detail dictionary;

[0077] Step S107e. Update the target column element according to the updated detail dictionary.

[0078] Through the above steps, the text elements of the target row can be further subdivided, identification elements such as product / drug / course name, number, etc. Detail elements such as fees, number of times used, etc. This application is aimed at the method of multi-column detail elements (such as XX yuan, XX branches, etc. need to be split into 2 columns) by constructing a detail dictionary to determine whether the empty element is extracted (such as XX yuan is merged into the same column as the number, the actual yuan should be extracted separately from the amount). Determine the text element corresponding to the empty element based on the text elements in the same row as the empty element in the detail dictionary, such as determining the information corresponding to the blank space based on the adjacent information (left and right or up and down), and also make corresponding corrections to the positions of multiple text elements extracted at the same time. Therefore, this application can ensure that the details of each column and each row can be accurately extracted.

[0079] Exemplarily, before constructing a detailed dictionary based on each target column element, the method includes: inputting the text element of the target row corresponding to each target column element into a preset title element prediction model, and the title element prediction model outputs a title prediction result corresponding to each text element; the title prediction result includes any one of an identification element and a detail element; if the title prediction result includes multiple detail elements containing preset keywords, obtaining the first text position information of the text element containing the preset keywords; confirming whether the text element containing the preset keywords is an identification element based on the first text position information; if the text element containing the preset keywords is an identification element, confirming that the text element is a target text element, and removing the target column element corresponding to the target text element from multiple target column elements.

[0080] To avoid misidentification, such as the situation where a major expense category is mistakenly identified as an expense detail, the above method can prevent the same preset keyword from appearing multiple times in the detail category, thereby ensuring the accuracy of identification.

[0081] Exemplarily, text information is divided according to the number of tables in a data table to obtain multiple sub-text information, including: obtaining the second text position information of each text element in a target row; dividing the multiple text elements in the target row into text element groups according to the number of tables, the second text position information of each text element and the type of title element corresponding to each text element; wherein the number of identification elements and detail elements in each text element group is the same; determining the division position for dividing the text information according to the second text position information corresponding to the text elements in each text element group; dividing the text information according to the division position to obtain multiple sub-text information.

[0082] Combined with the second text position information of the target text element and the title element type corresponding to each text element, the multiple text elements of the target row are divided into a table of text element groups. For example, when the title element type corresponding to the text element is identification class, identification class, detail class, detail class, identification class..., then it is considered that [text element corresponding to the identification class, text element corresponding to the identification class, text element corresponding to the detail class, text element corresponding to the detail class] is a text element group, and then the division position can be determined according to the position information of the head and tail text elements of each text element group, ensuring that multiple sub-text elements can be accurately divided.

[0083] In some embodiments, Figure 3 Step S105 also includes steps S105a to S105d.

[0084] like Figure 3 As shown, the provided data table acquisition method includes the following steps:

[0085] S105a generates multiple blank tables according to the number of tables, the number of data columns and the number of rows of text information; the blank table includes a header row and multiple content rows;

[0086] S105b. Fill the text elements of the target row into the title row in sequence;

[0087] S105c determines the number of extractions of sub-text information according to the number of columns of data, and extracts text elements of each row of sub-text information except the title row one by one according to the number of extractions to obtain multiple extraction elements;

[0088] S105d. Add each extracted element to the content row of the blank table according to the extraction order, complete the filling of the blank table, and obtain the data table corresponding to each sub-text information.

[0089] Furthermore, through the provided method, the creation and filling of a blank data table can be completed accurately, and different table styles can also be generated based on the type of the corresponding title row. Furthermore, the provided method can quickly realize the extraction of multi-column text information.

[0090] See also Figure 4 As shown, Figure 4 1 is a schematic diagram of the structure of a text information extraction device 200 provided in an embodiment of the present application. The text information extraction device 200 is used to execute the steps of the text information extraction method shown in the above embodiments. The text information extraction device 200 can be a single server or a server cluster, or the text information extraction device 200 can be a terminal, which can be a handheld terminal, a laptop computer, a wearable device or a robot, etc.

[0091] like Figure 4As shown, the text information extraction device 200 includes:

[0092] The image acquisition module 201 is used to acquire the text image to be extracted, perform named entity recognition on the text image line by line, and acquire the text information in the text image;

[0093] The text parsing module 202 is used to parse the target line text in the text information, obtain multiple text elements contained in the target line text, and confirm whether each text element is a title element; wherein the target line text is located in the first line of the text information according to the recognition order;

[0094] The element acquisition module 203 is used to acquire the number of identical elements of each text element in multiple text elements if each text element is a title element;

[0095] A column number acquisition module 204, used to determine the number of tables in the data table after the text information is extracted and the number of data columns in each data table according to the number of identical elements and the number of text elements;

[0096] The information acquisition module 205 is used to divide the text information according to the number of tables in the data table to obtain a plurality of sub-text information;

[0097] The text extraction module 206 is used to extract each sub-text information according to the number of tables, the number of data columns and the number of rows of text information, obtain the data table corresponding to each sub-text information, and complete the text information extraction of the text image.

[0098] It should be noted that technical personnel in the relevant field can clearly understand that, for the convenience and conciseness of description, the specific working process of the text information extraction device and each module described above can refer to the corresponding process in the text information extraction method embodiments described in the above embodiments, and will not be repeated here.

[0099] The above-mentioned text information extraction method can be implemented in the form of a computer program. Figure 4 Run on the device shown.

[0100] See also Figure 5 , Figure 5 1 is a schematic block diagram of the structure of a computer device provided in an embodiment of the present application. The computer device includes a processor, a memory and a network interface connected via a device bus, wherein the memory may include a storage medium and an internal memory.

[0101] The storage medium can store an operating device and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any text information extraction method.

[0102] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.

[0103] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any text information extraction method.

[0104] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the scheme of the present application, and does not constitute a limitation on the terminal to which the scheme of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0105] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0106] In one embodiment, the processor is used to run a computer program stored in the memory to implement the following steps:

[0107] Obtain the text image to be extracted, perform named entity recognition on the text image line by line, and obtain the text information in the text image;

[0108] Parsing the target line text in the text information, obtaining multiple text elements contained in the target line text, and confirming whether each text element is a title element; wherein the target line text is located in the first line of the text information according to the recognition order;

[0109] If each text element is a title element, get the number of identical elements of each text element in multiple text elements;

[0110] Determine the number of tables in the data table after text information extraction and the number of data columns in each data table according to the number of identical elements and the number of text elements;

[0111] Divide the text information according to the number of tables in the data table to obtain multiple sub-text information;

[0112] Each sub-text information is extracted according to the number of tables, the number of data columns and the number of rows of text information, and the data table corresponding to each sub-text information is obtained to complete the text information extraction of the text image.

[0113] In some embodiments, before obtaining the data table corresponding to each sub-text information, it also includes: obtaining the title element type corresponding to the text element of the target row; the title element type includes identification elements and detail elements; obtaining the target column element corresponding to the detail element from the multiple column elements extracted from each sub-text information; wherein each column element includes multiple text elements arranged row by row; constructing a detail dictionary according to each target column element, and determining whether there is an empty element in the detail dictionary; if there is an empty element in the detail dictionary, determining the text element corresponding to the empty element according to the text elements in the detail dictionary in the same row as the empty element, and completing the update of the detail dictionary; updating the target column element according to the updated detail dictionary.

[0114] Exemplarily, before constructing a detailed dictionary based on each target column element, the method includes: inputting the text element of the target row corresponding to each target column element into a preset title element prediction model, and the title element prediction model outputs a title prediction result corresponding to each text element; the title prediction result includes any one of an identification element and a detail element; if the title prediction result includes multiple detail elements containing preset keywords, obtaining the first text position information of the text element containing the preset keywords; confirming whether the text element containing the preset keywords is an identification element based on the first text position information; if the text element containing the preset keywords is an identification element, confirming that the text element is a target text element, and removing the target column element corresponding to the target text element from multiple target column elements.

[0115] Exemplarily, text information is divided according to the number of tables in a data table to obtain multiple sub-text information, including: obtaining the second text position information of each text element in a target row; dividing the multiple text elements in the target row into text element groups according to the number of tables, the second text position information of each text element and the type of title element corresponding to each text element; wherein the number of identification elements and detail elements in each text element group is the same; determining the division position for dividing the text information according to the second text position information corresponding to the text elements in each text element group; dividing the text information according to the division position to obtain multiple sub-text information.

[0116] In some embodiments, each sub-text information is extracted according to the number of tables, the number of data columns and the number of rows of text information to obtain a data table corresponding to each sub-text information, including: generating multiple blank tables according to the number of tables, the number of data columns and the number of rows of text information; the blank table includes a title row and multiple content rows; filling the text elements of the target row into the title row in sequence; determining the number of extractions for the sub-text information according to the number of data columns; extracting the text elements of each row of the sub-text information except the title row one by one according to the number of extractions to obtain multiple extracted elements; adding each extracted element to the content row of the blank table according to the number of extractions to complete the filling of the blank table and obtain a data table corresponding to each sub-text information.

[0117] In some embodiments, the text image to be extracted is stored in a preset insurance claims system; confirming whether each text element is a title element includes: obtaining multiple preset titles corresponding to the insurance claims system; matching each text element with multiple preset titles, and calculating the similarity between the text element and each preset title; if at least one similarity among the multiple similarities is greater than a preset threshold, confirming that the text element is a title element.

[0118] Exemplarily, before determining the number of tables in the data table after text information extraction and the number of data columns in each data table based on the number of identical elements and the number of text elements, it also includes: if each text element is not a title element, calculating the matching degree of the text element with each preset title based on a self-attention mechanism; determining the title element corresponding to each text element based on multiple matching degrees; and obtaining the number of identical elements of the title element corresponding to each text element among the title elements corresponding to multiple text elements.

[0119] A computer-readable storage medium is also provided in an embodiment of the present application, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and the processor executes the program instructions to implement the steps of the text information extraction method provided in the above embodiments of the present application.

[0120] The computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiment, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart memory card (SmartMedia Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc., equipped on the computer device.

[0121] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.

Claims

1. A text information extraction method, characterized in that: include: Acquire a text image to be extracted, perform named entity recognition on the text image line by line, and acquire text information in the text image; Parsing the target line text in the text information, obtaining multiple text elements contained in the target line text, and confirming whether each of the text elements is a title element; wherein the target line text is located in the first line of the text information according to the recognition order; If each of the text elements is the title element, obtain the number of identical elements of each text element in the multiple text elements; Determine the number of tables in the data table after text information extraction and the number of data columns in each of the data tables according to the number of identical elements and the number of text elements; Dividing the text information according to the number of tables in the data table to obtain a plurality of sub-text information; Each sub-text information is extracted according to the number of tables, the number of data columns and the number of rows of the text information, and the data table corresponding to each sub-text information is obtained to complete the text information extraction of the text image.

2. The method according to claim 1, characterized in that Before obtaining the data table corresponding to each sub-text information, the method further includes: Acquire the title element type corresponding to the text element of the target row; the title element type includes an identification element and a detail element; Obtaining a target column element corresponding to the detail element from a plurality of column elements extracted from each sub-text information; wherein each of the column elements includes a plurality of text elements arranged row by row; Construct a detailed dictionary according to each target column element, and determine whether there is an empty element in the detailed dictionary; If there is an empty element in the detailed dictionary, determine the text element corresponding to the empty element according to the text element in the same row as the empty element in the detailed dictionary, and complete the update of the detailed dictionary; The target column element is updated according to the updated detail dictionary.

3. The method according to claim 2, characterized in that Before constructing the detailed dictionary according to each of the target column elements, the method includes: Inputting the text element of the target row corresponding to each target column element into a preset title element prediction model, the title element prediction model outputs a title prediction result corresponding to each of the text elements; the title prediction result includes any one of the identification element and the detail element; If the title prediction result includes a plurality of detail elements containing preset keywords, obtaining first text position information of the text element containing the preset keywords; Determining whether the text element containing the preset keyword is the identification element according to the first text position information; If the text element containing the preset keyword is the identification element, the text element is confirmed to be a target text element, and the target column element corresponding to the target text element is removed from the plurality of target column elements.

4. The method according to claim 2, characterized in that: The step of dividing the text information according to the number of tables in the data table to obtain a plurality of sub-text information includes: Acquire second text position information of each text element of the target row; Divide the plurality of text elements of the target row into text element groups of the same number as the table according to the number of tables, the second text position information of each text element and the title element type corresponding to each text element; wherein the number of the identification elements and the detail elements in each text element group is the same; Determining a division position for dividing the text information according to second text position information corresponding to the text elements in each text element group; The text information is divided according to the division position to obtain a plurality of sub-text information.

5. The method according to claim 1, characterized in that: The extracting each sub-text information according to the number of tables, the number of data columns and the number of rows of the text information to obtain a data table corresponding to each sub-text information includes: Generate multiple blank tables according to the number of tables, the number of data columns and the number of rows of text information; the blank tables include a title row and multiple content rows; Fill the text elements of the target row into the title row in sequence; Determining the number of times to extract the subtext information according to the number of data columns; Extracting text elements of each line of the sub-text information except the title line one by one according to the number of extractions to obtain a plurality of extracted elements; Each extracted element is added to the content row of the blank table according to the extraction order, the blank table is filled, and the data table corresponding to each sub-text information is obtained.

6. The method according to claim 1, characterized in that The text image to be extracted is stored in a preset insurance claim system; The step of confirming whether each of the text elements is a title element comprises: Obtaining multiple preset titles corresponding to the insurance claims system; Matching each text element with a plurality of the preset titles, and calculating the similarity between the text element and each preset title; If at least one of the multiple similarities is greater than a preset threshold, the text element is confirmed to be the title element.

7. The method according to claim 6, characterized in that Before determining the number of tables in the data table after text information extraction and the number of data columns in each data table according to the number of identical elements and the number of text elements, the method further includes: If each of the text elements is not the title element, calculating the matching degree between the text element and each of the preset titles based on the self-attention mechanism; Determine the title element corresponding to each of the text elements according to multiple matching degrees; Gets the number of identical elements of the title element corresponding to each text element among the title elements corresponding to multiple text elements.

8. A text information extraction device, characterized in that: include: An image acquisition module is used to acquire a text image to be extracted, perform named entity recognition on the text image line by line, and acquire text information in the text image; A text parsing module, used to parse the target line text in the text information, obtain multiple text elements contained in the target line text, and confirm whether each of the text elements is a title element; wherein the target line text is located in the first line of the text information according to the recognition order; An element acquisition module, used for acquiring the number of identical elements of each text element in a plurality of text elements if each of the text elements is the title element; A column number acquisition module, used to determine the number of tables in the data table after the text information is extracted and the number of data columns in each of the data tables according to the number of identical elements and the number of text elements; An information acquisition module, used for dividing the text information according to the number of tables in the data table to acquire a plurality of sub-text information; The text extraction module is used to extract each sub-text information according to the number of tables, the number of data columns and the number of rows of the text information, obtain the data table corresponding to each sub-text information, and complete the text information extraction of the text image.

9. A computer device, characterized in that: The computer device includes a memory and a processor; The memory is used to store computer programs; The processor is configured to execute the computer program and implement the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, causes the processor to implement the method according to any one of claims 1 to 7.